Avatar signaling in scene description

The proposed format extension for MPEG-I scene descriptions addresses the limitation of lacking avatar representation by signaling avatar types and providing semantic descriptions, enabling effective reconstruction and animation in extended reality applications.

JP2026510678APending Publication Date: 2026-04-10INTERDIGITALCE PATENT HLDG SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
INTERDIGITALCE PATENT HLDG SAS
Filing Date
2024-03-15
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Current MPEG-I scene description frameworks lack the ability to provide visual and animation information about avatar assets, limiting their representation and interaction in extended reality applications.

Method used

Introduce a format extension that signals the type of avatar representation under an avatar node, using attributes like 'MPEG_node_avatar' to specify the avatar type, enabling reconstruction and animation based on MPEG-I standards and other formats, and providing high-level semantic descriptions for avatars.

Benefits of technology

Enables accurate reconstruction and animation of avatars in extended reality scenes, facilitating interoperability and efficient network transmission by providing high-level semantic descriptions and allowing transcoding between different avatar formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026510678000001_ABST
    Figure 2026510678000001_ABST
Patent Text Reader

Abstract

One implementation proposes providing signaling to inform applications about what to expect from avatar nodes. The proposed format representation is reconstructed upon given some user input and mapped onto a 3D reference avatar model, enabling a shared geometric and feature representation for any humanoid character. For example, the "Format" and "Model" attributes provide additional information about the avatar, such as the avatar format, high-level semantics, and 3D media linked to the geometry or texture. In the case of a 3D model, the avatar type and skeletal structure can also be described. This common basic representation facilitates animation and representation in different forms or contexts. Such high-level semantic descriptions are costly in terms of network transmission and also enable easy transcoding between different formats.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This embodiment generally relates to the representation and interaction of digital humans within a 3D-designed virtual scene, and more specifically, to identifying an avatar used beneath an avatar node in a scene description.

Background Art

[0002] Extended reality (XR) is a technology that enables an interactive experience in which the real-world environment and / or video content is enhanced by virtual content that can be defined across multiple sensory modalities including vision, hearing, touch, etc. During the runtime of an application, virtual content (e.g., 3D content or audio / video files) is rendered in real time in a way that matches the user context (environment, viewpoint, device, etc.). A scene graph (e.g., as proposed by Khronos / glTF (Graphics Language Transmission Format), and its extensions defined in the MPEG scene description format or Apple / USDZ) is a possible means for representing the content to be rendered. These combine, on the one hand, a declarative description of the scene structure that links real-world environment objects and virtual objects, and on the other hand, the binary representation of the virtual content. A scene description framework ensures that the time-domain media and the corresponding associated virtual content are available at any point during the rendering of the application. The scene description can also carry scene-level data that describes how a user can interact with scene objects at runtime for an immersive XR experience.

Summary of the Invention

[0003] According to one embodiment, a method is presented which includes: obtaining a format used to represent an avatar from a description for an augmented reality scene; obtaining at least one type of avatar from the description; and reconstructing the avatar based on the format and at least one type of avatar.

[0004] In another embodiment, a method is presented which includes: indicating a format used to represent an avatar in a description for an augmented reality scene; indicating at least one type of avatar in the description; generating data representing the avatar based on the format and at least one type of avatar; and indicating data representing the avatar in the description.

[0005] According to another embodiment, a device is presented comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to obtain a format used to represent an avatar from a description for an augmented reality scene, obtain at least one type of avatar from the description, and reconstruct an avatar based on the format and the at least one type of avatar.

[0006] According to another embodiment, a device is presented comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to indicate a format used to represent an avatar in a description for an augmented reality scene, to indicate at least one type of avatar in the description, to generate data representing an avatar based on the format and the at least one type of avatar, and to indicate data representing an avatar in the description.

[0007] One or more embodiments also provide a computer program that, when executed by one or more processors, includes instructions causing one or more processors to carry out a method according to any of the embodiments described herein. One or more of these embodiments also provide a computer-readable storage medium storing instructions for processing a scene description according to the method described herein.

[0008] One or more embodiments also provide a computer-readable storage medium storing scene descriptions generated according to the methods described above. One or more embodiments also provide methods and apparatus for transmitting or receiving scene descriptions generated according to the methods described herein. [Brief explanation of the drawing]

[0009] [Figure 1] This shows an exemplary architecture for an XR processing engine. [Figure 2] This shows an example of the syntax for a data stream that encodes an augmented reality scene description. [Figure 3] An illustrative graph of an augmented reality scene description is shown. [Figure 4] An example of an augmented reality scene description is shown. [Figure 5] This document illustrates current solutions for representing avatars using MPEG-I Scene Description (SD). [Figure 6] This document illustrates which information is encoded in an avatar encoder according to one embodiment, and provides examples of different possible renderings. [Figure 7] An example of avatar representation for network transmission and mobile rendering according to one embodiment is provided. [Figure 8] One embodiment illustrates a semantic mapping for animation control between an MPEG reference model and a VRM model. [Figure 9]One embodiment illustrates a semantic blend shape for face mapping for animation control between an MPEG reference model and a VRM model. [Figure 10A] One embodiment illustrates semantic mapping and motion transfer using the MPEG reference model. [Figure 10B] One embodiment illustrates semantic mapping and motion transfer using the MPEG reference model. [Figure 11A] Rig parameters are shown as an example. [Figure 11B] Rig parameters are shown as an example. [Modes for carrying out the invention]

[0010] Various XR applications can be applied to different contexts and real or virtual environments. For example, in industrial XR applications, when a reference object (engine part B) is detected in the real environment by a camera equipped on a head-mounted display device, a virtual 3D content item (e.g., engine part A) is displayed. The 3D content item is positioned in the real world at a position and scale defined relative to the detected reference object.

[0011] For example, in an XR application for interior design, a 3D model of furniture is displayed when a given image from a catalog is detected within the input camera view. The 3D content is positioned in the real world at a defined position and scale relative to the detected reference image. In another application, several audio files may start playing when the user enters an area close to a church (which is rendered in reality or virtually in the augmented reality environment). In yet another example, an advertising jingle file may play when the user sees a given can of soda in the real environment. In an outdoor game application, various virtual characters may appear depending on the semantics of the landscape observed by the user. For example, since bird characters are suitable for trees, if the sensors of the XR device detect a real object described by the semantic label "tree," birds flying around the tree can be added. In a companion application implemented with smart glasses, a car noise may be emitted in the user's headset when a car is detected within the user camera's field of view to warn the user of a potential hazard. Furthermore, the sound may be spatialized to ensure the sound reaches from the direction in which the car was detected.

[0012] XR applications can also extend video content rather than the real environment. The video is displayed on a rendering device, and when a timed event is detected within the video, a virtual object described in a node tree is overlaid. In such a scenario, the node tree contains only descriptions of the virtual object.

[0013] Figure 1 shows an exemplary architecture of an XR processing engine 130 that may be configured to implement the methods described herein. Devices according to the architecture of Figure 1 are linked with other devices via their bus 131 and / or via I / O interface 136.

[0014] Device 130 comprises the following elements, which are linked to each other by a data and address bus 131: - A microprocessor 132 (or CPU), for example, a DSP (or Digital Signal Processor), - A ROM (or Read Only Memory) 133, - A RAM (or Random Access Memory) 134, - A storage interface 135, - An I / O interface 136 for receiving data to be transmitted from an application, and - A power supply (not shown in FIG. 1), for example, a battery.

[0015] According to an example, the power supply is external to the device. In each of the above memories, the word "register" as used herein may correspond to a small area (a few bits) or a very large area (e.g., an entire program or a large amount of received or decoded data). The ROM 133 includes at least programs and parameters. The ROM 133 may store algorithms and instructions for implementing the technology according to this principle. When switched on, the CPU 132 uploads the program into the RAM and executes the corresponding instructions.

[0016] The RAM 134 includes a program executed by the CPU 132 and uploaded after the device 130 is switched on in the register, input data in the register, intermediate data of different states of the method in the register, and other variables used for executing the method in the register.

[0017] Device 130 is linked, via, for example, bus 131, to a set of sensors 137 and a set of rendering devices 138. The sensors 137 may be, for example, cameras, microphones, temperature sensors, inertial measurement devices, GPS, humidity measurement sensors, IR or UV light sensors, or wind sensors. The rendering devices 138 may be, for example, displays, speakers, vibrators, heaters, fans, etc.

[0018] According to an example, device 130 is configured to implement a method according to the present principle and belongs to a set including: - Mobile devices, - Communication devices, - Gaming devices, - Tablets (or tablet computers), - Laptops, - Still cameras, - Video cameras.

[0019] In XR applications, the scene description is used to combine an explicit and easily analyzable description of the scene structure with some binary representations of media content. FIG. 2 shows an example of the syntax of a data stream encoding an extended reality scene description. FIG. 2 shows an exemplary structure 210 of an XR scene description. The structure exists within a container that organizes the stream into individual syntax elements. This structure may include a header portion 220 that is a set of data common to all syntax elements of the stream. For example, the header portion includes a part of the metadata regarding the syntax elements and describes their respective properties and roles. The structure also includes a payload that includes elements of syntax 230 and elements of syntax 240. The syntax element 230 comprises data representing media content items described at nodes of a scene graph related to virtual elements. Images, meshes, and other raw data may be compressed according to a compression method. The elements of syntax 240 are part of the payload of the data stream and include data encoding the scene description as described according to the present principle.

[0020] Figure 3 shows an exemplary graph 310 of an augmented reality scene description. In this example, the scene graph may include descriptions of real objects, e.g., a "plane horizontal plane" (which could be a table or a road), and descriptions of virtual objects 312, e.g., the animation of a car. The scene description is organized as an array of nodes. Nodes can link to child nodes to form a scene structure 311. Nodes can carry descriptions of real objects (e.g., semantic descriptions) or virtual objects. In the example in Figure 3, node 301 describes a virtual camera located within a 3D volume of the XR application. Node 302 describes a virtual car and includes an index of the car's representation, e.g., an index in an array of 3D meshes. Node 303 is a child of node 302 and includes a description of one of the car's wheels. Similarly, it includes an index to the wheel's 3D mesh. Since the scale, position, and orientation of objects are described in the scene nodes, the same 3D mesh may be used for several objects in the 3D scene. Scene graph 310 also includes nodes, which describe the spatial relationships between real and virtual objects.

[0021] In time-based media streaming, the scene description itself can evolve over time, providing relevant virtual content for each sequence in the media stream. For example, for advertising purposes, a virtual bottle could be displayed on a table during a video sequence showing people sitting around it. This type of behavior can be achieved by relying on a framework defined in the scene description of an MPEG media document.

[0022] Currently, the MPEG-I scene description framework uses "behavior" data to extend time-evolving scene descriptions, providing a description of how users can interact with scene objects at runtime for immersive XR experiences. These behaviors relate to predefined virtual objects that enable runtime interactivity for user-specific XR experiences. These behaviors are also time-evolving and are updated through existing scene description update mechanisms.

[0023] Figure 4 shows an example of an augmented reality scene description that includes behavioral data, described at the node level and stored at the scene level at runtime for an immersive XR experience, describing how the user can interact with scene objects. When an XR application starts, media content items (e.g., meshes of virtual objects visible to the camera) are loaded, rendered, and buffered so that they are displayed when triggered. For example, when a sensor detects a plane in the real environment, the application displays the buffered media content item as described in the relevant scene node. Timing is managed by the application according to the timing of features detected in the real environment and animations. Some nodes in the scene graph may not contain descriptions and only act as parents of child nodes. Figure 4 shows the relationship between behaviors included in the scene-level scene description and the nodes that are components of the scene graph. Behavior 410 relates to a predefined virtual object that enables runtime interactivity for a user-specific XR experience. Behavior 410 is also time-evolving and is updated through a scene description update mechanism.

[0024] The behavior is, - Trigger 420 defines the conditions that must be met for it to be activated. - Trigger control parameters that define logical operations between defined triggers, - Action 430 to be processed when the trigger is activated, - Action control parameters that define the order in which related actions are executed. - A priority number that allows selection of the highest priority behavior when several behaviors conflict simultaneously on the same virtual object. -Includes an optional interrupt action that specifies how to terminate this behavior if it is no longer defined in a newly received scene update, for example, the behavior is no longer defined if the related object does not belong to the new scene, or if the behavior is no longer relevant to this current media (e.g., audio or video) sequence.

[0025] Behavior 410 is performed at the scene level. Triggers are linked to nodes and their child nodes. In the example in Figure 4, trigger 1 is linked to nodes 1, 2, and 8. Since node 31 is a child of node 1, trigger 1 is linked to node 31. Trigger 1 is also linked to node 14 as a child of node 8. Trigger 2 is linked to node 1. In fact, the same node may be linked to several triggers. Trigger n is linked to nodes 5, 6, and 7. A behavior may include several triggers. For example, the first behavior may be activated by trigger 1 AND trigger 2, where AND is the trigger control parameter for the first behavior. A behavior may have several actions. For example, the first behavior may perform action m first, then action 1, where "first, then" is the action control parameter for the first behavior. The second behavior may be activated by trigger n, and for example, perform action 1 first, then action 2.

[0026] Different formats can be used to represent the node tree. For example, an MPEG-I scene description framework using the Khronos glTF extension mechanism may be used for the node tree. In this example, an interactive extension may be applied at the glTF scene level and is called MPEG_scene_interactivity. The corresponding semantics are provided in Table 1, where "M" in the "Usage" column indicates that the field is required in the XR scene description format, and "O" indicates that the field is optional.

[0027] [Table 1]

[0028] Figure 5 illustrates the current solution for representing avatars in MPEG-I Scene Description (SD). In particular, Figure 5 illustrates the glTF file structure. The entry point is a "scene" node (530) containing a "node" node (535) which can be either a "camera" (510), a "light" (565), or a "mesh" (540). MPEG-I SD defines a new boolean attribute "MPEG_node_avatar" (535) that indicates whether a node corresponds to an avatar node, as illustrated in Table 2. The purpose of this boolean attribute is to define which nodes define the user representation.

[0029] Today, this approach is limited because it does not provide any visual or animation information about avatar assets. Such information must be explicitly defined on the following nodes: Appearance is defined as a "Material" (570), which includes "Texture" (590), "Technique" (575), "Program" (580), and "Shader" (585), referenced directly by either a "Source" (595) or an "Image" (598). Animation is defined in an "Accessor" (545), which includes "Animation" (520) and "Skin" (525). All of these are displayed using the "Buffer View" (550) and "Buffer" node information (555). "MPEG_media" (560) can redefine several existing attributes, such as geometry, appearance, sound, tactile, and / or animation, by referencing external data.

[0030] [Table 2]

[0031] Using Booleans to represent avatars is restricted because it does not provide any visual or animation information about the avatar assets. Avatar components must accurately represent different types of avatar representations that have different methods of initialization, manipulation, or representation. MPEG-I SD presents a reference avatar ("MPEG_reference_avatar") that includes several attributes, such as the eyes, jaw, face, body, or geometry for detailed semantic descriptions of body parts. These attributes (eyes, jaw, face, body, and semantics) can generally represent the human body and are therefore not limited to a single avatar representation. In practice, other avatar representations may include additional details to their own representation, but are still compatible with the presented reference avatar. Therefore, a unique avatar representation in MPEG-I SD is insufficient when dealing with the reference avatar or any other avatar representation.

[0032] This specification aims to signal what type of avatar representation an avatar node (e.g., "MPEG_node_avatar") refers to, so that applications can query the correct assets and construct a valid semantic domain.

[0033] Avatar signaling in MPEG-I scene description In this specification, the proposed format conforms to the glTF format and is compatible with recent MPEG-I SD efforts to extend glTF using the MPEG-I SD extension. However, its meaning and use are general and can be coded in any other format (e.g., XML, USD).

[0034] In current MPEG-I, while the geometry of an avatar is defined, the semantics (e.g., high-level descriptions of body parts such as which vertices represent the left or right hand, head or legs, or high-level descriptions of the avatar type such as whether it is humanoid or not) are not defined, meaning that the client / engine doesn't necessarily know how to process the avatar. Therefore, we propose providing signaling to inform applications about what to expect from the avatar node (e.g., "MPEG_node_avatar").

[0035] In one embodiment, the avatar should comply with the following requirements: 1. An avatar is represented by a node specified for the avatar. For example, "MPEG_node_avatar" is used to specify an MPEG-I avatar node, and any other character representation is not eligible as "avatar" under the MPEG-I SD standard. 2. Reconstruction and animation respect the primitives supported in scene descriptions for MPEG and / or other standards and animation formats. 3. Avatars may or may not be explicitly represented by a “format” with a URI for each 3D asset. 4. When an avatar sets a "format," the default is "MPEG," which is fully supported by the MPEG-I SD standard for reconstruction and animation. 5. Any other “format” representation must include 3D mesh media to allow reconstruction by user applications.

[0036] The proposed format representation is reconstructed upon receiving some user input and mapped to a 3D avatar model, e.g., an "MPEG_reference_avatar" containing a full-body-based mesh representation, skeletal and skin weights, facial blend shapes and landmarks, and additional geometry including eyes, jaw, teeth, and tongue. This enables shared geometry and feature representations for any humanoid character (naturally, allowing appropriate deformations to adapt the shape). This common base representation facilitates animation and representation of different forms or contexts.

[0037] This "format" extension provides a signal to identify the avatar used under the avatar node ("MPEG_node_avatar"). Considering the proposed extension properties, the client will know which avatar to rebuild / render. The default representation is the reference avatar ("MPEG_avatar_reference"), which is fully supported under this proposed extension.

[0038] If a node in the scene graph is an avatar node, the format described in Table 3 is taken into consideration; otherwise, it is ignored.

[0039] Table 3 describes the newly proposed attributes to support the "isAvatar" signaling in Table 2. The "Format" and "Model" attributes provide additional information about the avatar, such as the avatar format, high-level semantics, and 3D media linked to geometry or textures.

[0040] [Table 3]

[0041] Table 4 lists the 3D formats that can be used to represent avatars.

[0042] [Table 4]

[0043] Table 5 illustrates the semantic description of the property "Model." This introduces the semantics and media of the avatar model.

[0044] [Table 5]

[0045] Table 6 lists examples of different avatar types.

[0046] [Table 6]

[0047] Table 7 illustrates an avatar skeletal structure semantic description, including the naming of joints, 3D positions, bone structure, influences between bone and geometry, and high-level semantic control points for creating different levels of manipulation.

[0048] [Table 7]

[0049] Media attributes are defined by "MPEG_media.media.alternative" in the MPEG_media extension of the ISO / IEC 23090-14 scene description specification, which includes "uri" and "mimeType" used to refer to any media input format. Media is an array of MPEG_media.

[0050] Example of a glTF schema The next glTF will instantiate "MPEG_node_avatar" in clients that support this extension, and otherwise revert to the standard "MPEG_node" without semantic extensions.

[0051] The following example illustrates the glTF format using MPEG extensions. This example demonstrates the “MPEG_node_avatar” node extension to support the basic representation of avatars in MPEG SD. In this “MPEG_node_avatar” node, the model is instantiated using “format=1” which references an MPEG reference avatar. The “type” semantically identifies the avatar and its mobility; for example, “terrestrial” has the ability to walk and run and is manipulable (“articulated”). This provides two levels of resolution (“low resolution” and “high resolution”) and an example (not complete) of the “skeleton” object structure.

[0052] [Table 8]

[0053] Use case The purpose of this use case is to demonstrate the new features provided by additional signaling. The proposed method facilitates the creation of new codec formats tailored to avatar media formats. Figure 6 illustrates a novel avatar coded representation (using the glTF example above) for lightweight network transmission and decoding of different avatar representations and contexts, according to one embodiment.

[0054] In particular, Figure 6 demonstrates which information is encoded in the "avatar encoder" module (610) and the different possible renderings. As shown in Figure 6, the following information is encoded in this example: -format - Model ○Name ○Type ○ Resolution ○Skeleton ■Semantics ■ Joints ■Bone ■bone_influence ■control points ○ Media

[0055] In this example, the decoder (620) is intended to interpret the transmitted bits in a readable format, and therefore the rendering can take the shape of a human body with a semantic description, a head model with facial joint landmarks, a hair model with control animation points, or the shape of 3D geometry of an animal with a semantic category, as illustrated in Figure 6. The avatar encoder (610) and decoder (620) can be, for example, a glTF encoder and decoder.

[0056] Figure 7 illustrates an avatar encoder representation for lightweight network transmission and mobile rendering via a “resolution” parameter that provides multiple options depending on the application, according to one embodiment.

[0057] In particular, Figure 7 illustrates the usefulness of such a codec for a lightweight mobile decoder that requires information about the resolution of the geometric model. In this example, the mobile decoder (710) would have access to high-level information about the content, for example, if the geometry is high-resolution or low-resolution. This facilitates the selection of a geometric model to load, taking into account the performance capabilities of a given mobile device. Figure 7 shows the possibility of having three levels of resolution, with the number of vertices increasing along with the semantic description. The mobile engine (720) handles loading and displaying the content supported by its hardware specifications, while the mobile decoder (710) instantiates and filters the metadata necessary for lightweight rendering.

[0058] This use case exemplifies the novelty of using this approach in conjunction with the basic avatar MPEG reference model. Following the same codec example shown in Figures 6 and 7, it will have a coded representation of the MPEG reference avatar model.

[0059] Figure 8 illustrates a semantic mapping for animation control between an MPEG reference model and a VRM model according to one embodiment. VRM is a file format for handling human-like 3D avatar (3D model) data for VR applications, and is based on glTF2. The semantics of the VRM itself are illustrated in Table 8.

[0060] [Table 9]

[0061] In particular, Figure 8 illustrates the encoding of the semantic features of an avatar from an MPEG reference model into a glTF file (810). The glTF file is sent to a decoder, which decodes the semantic features of the body parts (820) and transcodes the semantic features from the MPEG model to a different model (VRM). This allows the avatar encoder to transmit a known animated base model, and enables the decoder and engine to map the animated base model to a new proprietary model (VRM model). The mapping is performed on the engine or application side, and its purpose is to correlate semantic body parts between different formats, enabling the exchange of animations as if they were mirror scenarios.

[0062] The following example illustrates the glTF format in Figure 8. In this example, an MPEG avatar is described (i.e., isAvatar=true, format=1). This is based on the MPEG reference model ("Morgan"), which is a multi-jointed humanoid model. The skeleton of this model is described by "semantics", "joints", "bones", and "bone_influence".

[0063] The attribute "semantics" is a list of joint names. The attribute "joint" is a list of three-dimensional vectors, e.g., [0.0, 0.0, 0.0], where each vector represents the x, y, and z coordinates of the corresponding joint. The attribute "bone" is a list of tuples containing joint indices, e.g., (0, 1) represents a pair of joints with indices 0 and 1 that form a single "bone". For simplicity, the "bone_influence" attribute is incomplete and only provides an example of a single vertex, e.g., a vertex with index 0 (the first element of the tuple (0, [0, 1], [0.5, 0.5]) indicates that the indices have bone indices 0 and 1 (the second element of the tuple [0, 1]), each bone is associated with a weight of 0.5, and the third element of the tuple [0.5, 0.5] means that if bone 0 or 1 has a transformation, vertex 0 will be influenced by 0.5 units of the bone transformation matrix). This representation is similar for all vertices of the mesh, and is therefore simplified.

[0064] [Table 10]

[0065] Figure 9 demonstrates another use case using blend shape mapping according to one embodiment. In particular, Figure 9 illustrates semantic blend shapes for face mapping for animation control between an MPEG reference model and a VRM model. Avatar features “blend shape”, “semantics”, and “control_points” are encoded (910) and transmitted. On the client side, a decoder decodes the glTF file (920) and transcodes it into a new format (VRM format) (face mapping). Blend shapes can be seen as deformations, and they generate high-level control for known animations such as blinking or smiling with the left eye. On the engine side, the goal would be mapping from the proposed format to the VRM format.

[0066] Table 9 shows that FACS is an expression shape naming convention defined and used in anatomical structures to classify facial movements. The Morgan shape column is an adaptation of such naming conventions to 3D models of faces. The shapes are separated into left ("_L") and right ("_R") components, and these last components may be further divided into ("_1") and ("_2") sub-components to improve the precision of the naming.

[0067] [Table 11]

[0068] Table 10 lists the VRM blend shape names that can be used by applications to perform mapping between different formats.

[0069] [Table 12]

[0070] Figure 10 illustrates a real-time scenario of semantic mapping and motion transfer using an MPEG reference model according to one embodiment. Input user representations (1001, 1002) are captured from the camera, face parameters (1010) are extracted and used to map MPEG reference face models (1011, 1012) to match the input user representations. The mapped face models are then converted into control_points or blend shapes (rig parameters), as illustrated in Figures 11A and 11B for Figures 10A and 10B, respectively.

[0071] Figures 11A and 11B are visual examples of the blend shape weights used to generate the renderings of Figures 10A and 10B, respectively. The names are semantic descriptions, the bars represent the weights used, for example, the weights are scalar values ​​in the range of 0 to 1, and the last three attributes "yaw," "pitch," and "roll" are global rotation parameters of the face / head model. The rig parameters (1020) are sent to the animation engine (1030, e.g., Unity, Unreal, Blender) to drive the 3D realistic character (1041, 1042). An MPEG reference face model (1011, 1012) is used as an intermediate to facilitate the rig prediction (1010) problem.

[0072] The following example describes the glTF format in Figures 10 and 11. In this example, an MPEG avatar is described (i.e., isAvatar=true, format=1). This is based on the MPEG reference model ("Morgan"), which is a multi-jointed humanoid model. The model is shown as having a resolution ("high resolution") ("MORGAN_FACE_MASK"), which is shown to be a subset of the MPEG reference model. The skeleton of this model is described by "semantics", "joints", "bones", and "bone_influence".

[0073] [Table 13]

[0074] It should be noted that high-level semantic descriptions are less costly from a network transmission perspective and also allow for easy transcoding between different formats. In this example, there are only a total of 10 scalar (floating-point) parameters for transmission over the network; otherwise, the entire mesh would need to be transmitted frame by frame, which could result in longer bitrate transmission times and ultimately loss (depending on the compression parameters).

[0075] Various numerical values ​​are used in this application. Certain values ​​are for illustrative purposes only, and the embodiments described are not limited to these specific values.

[0076] Various methods are described herein, each of which includes one or more steps or actions to achieve the described method. Unless a specific order of steps or actions is required for the normal operation of the method, the order and / or use of any particular steps and / or actions may be modified or combined. Furthermore, terms such as “first,” “second,” etc., may be used in various embodiments to modify elements, components, steps, operations, etc., such as “first decryption” and “second decryption.” The use of such terms does not imply a modified order of operations unless specifically required. Thus, in this example, the first decryption does not need to be performed before the second decryption, and may occur, for example, before, during, or over the overlap period with the second decryption.

[0077] The implementations and embodiments described herein may be implemented, for example, in methods or processes, apparatus, software programs, data streams, or signals. Even when considered only in the context of a single implementation (for example, only as a method), the implementations of the features considered may also be implemented in other forms (for example, apparatus or programs). Apparatus may be implemented, for example, in appropriate hardware, software, and firmware. The method may be performed, for example, in an apparatus, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device, e.g., a processor. Processors also include communication devices, e.g., computers, mobile phones, portable / personal digital assistants ("personal digital assistants, PDAs"), and other devices that facilitate the transmission of information between end users.

[0078] The terms "one embodiment" or "one embodiment," or "one implementation" or "one implementation," and any other variations thereof, mean that the specific features, structures, characteristics, etc., described in relation to the embodiments are included in at least one embodiment. Therefore, the appearance of the phrases "in one embodiment" or "in one embodiment," or "in one implementation" or "in one implementation," and any other variations, found in various places throughout this application, do not necessarily all refer to the same embodiment.

[0079] Additionally, this application may refer to "determining" various types of information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from memory.

[0080] Furthermore, this application may also refer to “accessing” various types of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0081] Additionally, this application may refer to “receiving” various types of information. Receiving is intended to be a broad term, similar to “accessing.” Receiving information may include, for example, accessing information or retrieving information (for example, from memory). Furthermore, “receiving” typically involves, in some way, during operation, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0082] For example, in the cases of "A / B," "A and / or B," and "at least one of A and B," the use of any of the following " / ," "and / or," and "at least one of" should be understood as intended to cover the selection of only the first option (A), only the second option (B), or both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C," such phrasing is intended to cover the selection of only the first option (A), only the second option (B), only the third option (C), only the first and second options (A and B), only the first and third options (A and C), only the second and third options (B and C), or all three options (A, B, and C). This may be extended as many times as there are listed items, as will be obvious to those skilled in the art.

[0083] As will be apparent to those skilled in the art, the implementation can generate a wide variety of signals, for example, that are formatted to carry information that can be stored or transmitted. This information may include, for example, instructions for performing a method, or data generated by one of the implementations described. For example, a signal may be formatted to carry a bitstream of the embodiment described. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a wide variety of different wired or wireless links, as is known. The signal may be stored in a processor-readable medium.

Claims

1. It is a method, From the description for the augmented reality scene, obtain the format used to represent the avatar, Obtain at least one type of avatar from the above description, A method comprising: reconstructing the avatar based on the aforementioned format and the aforementioned at least one type of avatar.

2. It is a method, In describing augmented reality scenes, it is necessary to indicate the format used to represent avatars, The above description must indicate at least one type of avatar, Based on the aforementioned format and the aforementioned at least one type of avatar, generate data representing the avatar, A method comprising, in the above description, showing the data representing the avatar.

3. The method according to claim 1 or 2, wherein the format includes at least one of MPEG, SMPL, VRM, MANO, FAUST, Dynamic FAUST, SMAL, and FLAME.

4. The method according to any one of claims 1 to 3, wherein the at least one type includes at least one of humanoid, non-humanoid, animal, aquatic, aerial, terrestrial, subterranean, and arboreal.

5. The method according to any one of claims 1 to 4, further comprising obtaining the skeletal structure of the avatar.

6. The method according to claim 5, wherein the description for the skeletal structure includes at least joint names, 3D positions of joints, bone structure, influence between bone and geometry, and high-level semantic control points.

7. The method according to any one of claims 1 and 3 to 6, wherein the description further includes information indicating the available resolutions in which the avatar is represented.

8. The method according to claim 7, further comprising selecting a resolution from the available resolutions based on the performance capabilities of the device, wherein the avatar is reconstructed by the device at the selected resolution.

9. The method according to any one of claims 1 and 3 to 8, wherein the description further includes information indicating an array of media.

10. The method according to any one of claims 1 and 3 to 9, further comprising mapping the avatar to represent it in another format.

11. The method according to any one of claims 2 and 3 to 10, further comprising encoding the description.

12. It is a device, One or more processors, The system comprises at least one memory coupled to one or more processors, and the one or more processors From the description for the augmented reality scene, we obtain the format used to represent the avatar. From the above description, obtain at least one type of avatar, A device configured to reconstruct an avatar based on the aforementioned format and the aforementioned at least one type of avatar.

13. It is a device, One or more processors, The system comprises at least one memory coupled to one or more processors, and the one or more processors In describing augmented reality scenes, the format used to represent avatars is shown. The above description indicates at least one type of avatar, Based on the aforementioned format and the aforementioned at least one type of avatar, data representing the avatar is generated. A device configured to show the data representing the avatar in the above description.

14. The apparatus according to claim 12 or 13, wherein the format includes at least one of MPEG, SMPL, VRM, MANO, FAUST, Dynamic FAUST, SMAL, and FLAME.

15. The apparatus according to any one of claims 12 to 14, wherein the at least one type includes at least one of humanoid, non-humanoid, animal, aquatic, aerial, terrestrial, subterranean, and arboreal.

16. The apparatus according to any one of claims 12 to 15, further comprising obtaining the skeletal structure of the avatar.

17. The apparatus according to claim 16, wherein the description for the skeletal structure includes at least joint names, 3D positions of joints, bone structure, influence between bone and geometry, and high-level semantic control points.

18. The apparatus according to any one of claims 12 and 14 to 17, wherein the description further includes information indicating the available resolutions in which the avatar is represented.

19. The one or more processors described above The apparatus according to claim 18, further configured to select a resolution from the available resolutions based on the performance capabilities of the device, wherein the avatar is reconstructed by the device at the selected resolution.

20. The apparatus according to any one of claims 12 and 14 to 19, wherein the description further includes information indicating an array of media.

21. The apparatus according to any one of claims 12 and 14 to 20, further comprising mapping the avatar to represent it in another format.

22. The one or more processors described above The apparatus according to any one of claims 13 and 14 to 21, further configured to encode the above description.

23. A non-temporary computer-readable medium that, when executed by a computer, includes an instruction causing the computer to perform the method according to any one of claims 1 to 11.