Avatar metadata representation
The proposed method and device enhance XR applications by encoding and retrieving detailed avatar parameters, enabling dynamic and socially interactive experiences with improved representation and interaction capabilities.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- INTERDIGITALCE PATENT HLDG SAS
- Filing Date
- 2024-03-18
- Publication Date
- 2026-04-10
AI Technical Summary
Current XR applications lack comprehensive representation and interaction capabilities for avatars, particularly in terms of identity, boundary regions, impairments, abilities, personality, and emotions, limiting immersive experiences and social interactions.
A method and device for encoding and retrieving avatar parameters, including identity, boundary regions, impairments, abilities, personality, and emotions, along with 3D geometry and textures, within augmented reality scenes using extended reality scene descriptions.
Enhances immersive XR experiences by providing detailed avatar representations that facilitate dynamic interactions and social behaviors, accommodating user-specific preferences and disabilities, and ensuring privacy and accessibility.
Smart Images

Figure 2026511175000001_ABST
Abstract
Description
Technical Field
[0001] This embodiment generally relates to digital human representation and interaction within a 3D engineered virtual scene.
Background Art
[0002] Extended Reality (XR) is a technology that enables an interactive experience where the real-world environment and / or video content is enhanced by virtual content that can be defined across multiple sensory modalities including vision, hearing, touch, etc. During the execution of an application, the virtual content (e.g., 3D content or audio / video files) is rendered in real time in a way that matches the user's context (environment, viewpoint, device, etc.).
[0003] Scene graphs (e.g., those proposed by Khronos / glTF (Graphics Language Transmission Format), and its extensions defined in the MPEG scene description format or Apple / USDZ) are a way to represent the content to be rendered. They combine, on the one hand, a declarative description of the scene structure that links objects in the real environment with virtual objects, and on the other hand, the binary representation of the virtual content. The scene description framework ensures that the timed media and the corresponding associated virtual content are available at any time during the rendering of the application. Scene descriptions can also carry scene-level data that describe how a user can interact with scene objects at runtime for immersive XR experiences.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
[0005] According to one embodiment, a method is provided which includes obtaining from a description of an augmented reality scene at least one parameter used to represent an avatar, wherein the at least one parameter includes one or more of identity, boundary regions representing areas of interaction between the avatar and other objects in the scene, impairments of the avatar, abilities of the avatar, personality of the avatar, and emotions of the avatar, and obtaining 3D geometry data and textures associated with the avatar.
[0006] According to another embodiment, a method is provided for describing an augmented reality scene, comprising encoding at least one parameter used to represent an avatar, wherein the at least one parameter includes an identity, a boundary region representing an area of interaction between the avatar and other objects in the scene, the avatar's impairments, the avatar's abilities, the avatar's personality, and the avatar's emotions, and encoding 3D geometric data and textures associated with the avatar.
[0007] According to another embodiment, a device is provided comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to retrieve from a description of an augmented reality scene at least one parameter used to represent an avatar, the at least one parameter including one or more of: identity, boundary regions representing areas of interaction between the avatar and other objects in the scene, impairments of the avatar, abilities of the avatar, personality of the avatar, and emotions of the avatar, and to retrieve 3D geometry data and textures associated with the avatar.
[0008] According to one embodiment, a device is provided comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to encode at least one parameter used to represent an avatar in a description of an augmented reality scene, the at least one parameter including identity, boundary regions representing areas of interaction between the avatar and other objects in the scene, impairments of the avatar, abilities of the avatar, personality of the avatar, and emotions of the avatar, and to encode 3D geometry data and textures associated with the avatar.
[0009] One or more embodiments also provide a computer program that, when executed by one or more processors, includes instructions causing one or more processors to perform a method according to any embodiment described herein. One or more embodiments also provide a computer-readable storage medium storing instructions for processing a scene description according to the method described herein.
[0010] One or more embodiments also provide a computer-readable storage medium storing scene descriptions generated according to the method described above. One or more embodiments also provide a method and apparatus for transmitting or receiving scene descriptions generated according to the method described herein. [Brief explanation of the drawing]
[0011] [Figure 1] This figure shows an exemplary architecture of an XR processing engine. [Figure 2] This figure shows an example of the syntax for a data stream that encodes an augmented reality scene description. [Figure 3] This figure shows an illustrative graph of an augmented reality scene description. [Figure 4] This figure shows an example of an augmented reality scene description. [Figure 5] This shows the architecture of current solutions that utilize synthetic representations. [Figure 6] A pipeline for introducing additional markers according to one embodiment is shown. [Figure 7] This shows the MPEG_node_avatar in MPEG-I SD. [Figure 8] A processing model for handling such metadata according to one embodiment is shown. [Modes for carrying out the invention]
[0012] Various XR applications can be applied to different contexts and real or virtual environments. For example, in industrial XR applications, a virtual 3D content item (e.g., engine part A) is displayed when a reference object (engine part B) is detected in the real environment by a camera attached to a head-mounted display device. The 3D content item is placed in the real world at a defined position and scale relative to the detected reference object.
[0013] For example, in an XR application for interior design, a 3D model of furniture is displayed when a given image from a catalog is detected within the input camera view. The 3D content is placed in the real world at a defined position and scale relative to the detected reference image. In another application, an audio file can start playing when the user enters an area near a church (either real or virtually rendered in an augmented reality environment). In yet another example, an ad jingle file may play when the user sees a given can of soda in the real world. In an outdoor game application, various virtual characters may appear depending on the semantics of the landscape observed by the user. For example, since bird characters are suitable for trees, if the sensors of the XR device detect a real object described by the semantic label "tree," birds flying around a tree can be added. In a companion application implemented with smart glasses, if a car is detected within the user's camera's field of view, the sound of a car may be emitted into the user's headset to warn the user of a potential hazard. Furthermore, the sound may be spatially adjusted to sound as if it is coming from the direction in which the vehicle was detected.
[0014] XR applications can also extend video content rather than the real environment. The video is displayed on a rendering device, and virtual objects described within a node tree are overlaid when timed events are detected within the video. In such a context, the node tree contains only descriptions of virtual objects.
[0015] Figure 1 shows an exemplary architecture of an XR processing engine 130 that may be configured to carry out the methods described herein. A device according to the architecture of Figure 1 is connected to other devices via a bus 131 and / or an I / O interface 136.
[0016] Device 130 comprises the following elements interconnected by data and address bus 131: - A microprocessor 132 (or CPU), such as a DSP (Digital Signal Processor), - A ROM (Read Only Memory) 133, - A RAM (Random Access Memory) 134, - A storage interface 135, - An I / O interface 136 for receiving data transmitted from an application, - A power supply (not shown in FIG. 1), such as a battery.
[0017] According to one example, the power supply is external to the device. In each of the memories mentioned, the term "register" as used herein may correspond to a small (few bits) area or a very large area (e.g., an entire program or a large amount of received or decoded data). ROM 133 includes at least programs and parameters. ROM 133 may store algorithms and instructions for performing techniques based on the principles of the present application. When powered on, CPU 132 uploads the program to RAM and executes the corresponding instructions.
[0018] RAM 134 includes, in registers, a program executed by CPU 132 and uploaded after power-on of device 130, input data in the registers, intermediate data of different states of a method in the registers, and other variables used in the execution of the method in the registers.
[0019] Device 130 is linked, for example via bus 131, to a set of sensors 137 and a set of rendering devices 138. Sensors 137 may be, for example, cameras, microphones, temperature sensors, inertial measurement units, GPS, humidity sensors, infrared or ultraviolet sensors, or wind sensors. Rendering devices 138 may be, for example, displays, speakers, vibrators, heaters, fans, etc.
[0020] According to some examples, device 130 is configured to execute a method based on this principle and belongs to a set including the following: - A mobile device, - A communication device, - A gaming device, - A tablet (or tablet computer), - A laptop, - A still camera, - A video camera.
[0021] In an XR application, a scene description is used to combine an explicit and syntactically easy-to-parse description of the scene structure with some binary representations of media content. FIG. 2 shows an example of the syntax of a data stream encoding an extended reality scene description. FIG. 2 shows an exemplary structure 210 of an XR scene description. This structure is composed of containers that organize the stream into independent elements of syntax. This structure can include a header portion 220 that is a set of data common to all syntax elements of the stream. For example, the header portion can include some metadata regarding syntax elements that describe the nature and role of each syntax element. This structure also includes a payload that includes syntax element 230 and syntax element 240. Syntax element 230 includes data representing media content items described in nodes of a scene graph related to virtual elements. Images, meshes, and other raw data may be compressed according to a compression method. Syntax element 240 is part of the payload of the data stream and includes data encoding a scene description explained according to this principle.
[0022] Figure 3 shows an exemplary graph 310 of an augmented reality scene description. In this example, the scene graph can include descriptions of real objects, such as a “horizontal plane” (which could be a table or a road), and descriptions of virtual objects 312, such as an animated car. The scene description is organized as an array of nodes. Nodes can be linked to child nodes to form a scene structure 311. Nodes can have descriptions of real objects (e.g., semantic descriptions) or virtual objects. In the example in Figure 3, node 301 describes a virtual camera located in the 3D space of the XR application. Node 302 describes a virtual car and includes an index of the car’s representation, such as an index in an array of 3D meshes. Node 303 is a child node of node 302 and includes a description of one of the car’s wheels. Similarly, it includes an index to the wheel’s 3D mesh. Since the scale, position, and orientation of objects are described in the scene nodes, the same 3D mesh can be used for several objects in the 3D scene. The scene graph 310 also includes nodes, which describe the spatial relationships between real and virtual objects.
[0023] In time-based media streaming, the scene description itself can change over time, allowing for the provision of virtual content relevant to each sequence in the media stream. For example, for advertising purposes, a virtual bottle could be displayed on a table during a video sequence showing people sitting around it. This type of behavior can be achieved by relying on a framework defined in the scene description of an MPEG media document.
[0024] Currently, the MPEG-I scene description framework extends time-varying scene descriptions using "behavior" data to provide a description of how users can interact with scene objects at runtime for immersive XR experiences. These behaviors are associated with predefined virtual objects for which runtime interactivity is permitted for user-specific XR experiences. These behaviors also change over time and are updated through existing scene description update mechanisms.
[0025] Figure 4 shows an example of an augmented reality scene description that includes scene-level stored behavior data, describing how a user can interact with node-level described scene objects when running an immersive XR experience. When an XR application starts, media content items (e.g., meshes of virtual objects visible to the camera) are loaded, rendered, and buffered so that they are displayed when triggered. For example, if a sensor detects a plane in the real environment, the application displays the buffered media content item as described in the relevant scene node. Timing is managed by the application according to the timing of features detected in the real environment and animations. Nodes in the scene graph may not contain descriptions and may only act as parents of child nodes. Figure 4 shows the relationship between the behaviors included in the scene-level scene description and the nodes that are components of the scene graph. Behavior 410 relates to a predefined virtual object that enables runtime interactivity for a user-specific XR experience. Behavior 410 also changes over time and is updated via a scene description update mechanism.
[0026] The behavior includes the following elements: - Trigger 420 defines the conditions that must be met for the behavior to be activated. - Trigger control parameters that define logical operations between defined triggers, - Action 430 that is processed when the trigger is activated, - Action control parameters that define the execution order of related actions. - A priority number that allows the selection of the highest priority behavior when several behaviors conflict simultaneously on the same virtual object. - An optional interruption action that specifies how to terminate this behavior if it is no longer defined by a newly received scene update. For example, if the relevant object does not exist in the new scene, or if this behavior is no longer relevant to the current media (audio or video) sequence, the behavior is no longer defined.
[0027] Behavior 410 is executed at the scene level. Triggers are linked to nodes and their child nodes. In the example in Figure 4, trigger 1 is linked to nodes 1, 2, and 8. Since node 31 is a child of node 1, trigger 1 is linked to node 31. Trigger 1 is also linked to node 14 as a child of node 8. Trigger 2 is linked to node 1. In fact, the same node may be linked to several triggers. Trigger n is linked to nodes 5, 6, and 7. A single behavior may contain several triggers. For example, the first behavior may be activated by trigger 1 AND trigger 2, where AND is the trigger control parameter for the first behavior. A single behavior may have several actions. For example, the first behavior may first perform action m and then action 1, where "first and then" is the action control parameter for the first behavior. The second behavior is activated by trigger n, and can, for example, perform action 1 first, then action 2.
[0028] Node trees can be represented using different formats. For example, an MPEG-I scene description framework using the Khronos glTF extension mechanism can be used for a node tree. In this example, the interactivity extension is applied at the glTF scene level and is called "MPEG_scene_interactivity". The corresponding semantics are provided in Table 1, where "M" in the "Usage" column indicates that the field is required in the XR scene description format, and "O" indicates that the field is optional.
[0029] [Table 1]
[0030] Current solutions for avatar representation
[0031] Digital humans are formed through a model-based reconstruction pipeline, which can make assumptions of a subset captured in the form of 3D template models. These assumptions can be generic human body models with skeletal structures, subject-specific models, or statistical shape models. These approaches provide, in an early stage, accurate 3D models that can statistically represent different human body shapes in different postures. These representations are primarily known as synthetic representations and are easily manipulated and fabricated to fit the anatomical structure of a particular body. Synthetic models also facilitate appearance generalization and stylization and can be performed by experts and used for animation and streaming.
[0032] Figure 5 shows an architectural diagram of the current solution for utilizing synthetic representation. The entry point is sensor data, which corresponds to all input data from the user device. This data is divided and processed by different encoding modules, namely, a head encoder (510), a body encoder (520), a hand encoder (530), and a head pose estimator (540). All of this encoded data is streamed over the network to an "Avatar Reconstruction and Animation" module (550), which reconstructs and animates the user's avatar based on an "offline 3D model." This avatar is then passed to any shared space so that other users can view and visualize it. This solution may be limited because it does not take into account important information such as gaze, posture, and skeletal animation. Furthermore, this solution does not transmit non-morphological and non-biomechanical information such as mental state, emotions, or personality.
[0033] As can be seen from Figure 5, the sensor data is limited and does not provide information about the face, eyes, skeletal structure, clothing and accessories, feet, or semantic cues. Therefore, this approach is limited to the use of synthetic models for streaming and interaction with digital humans. Furthermore, this model does not take into account social behavior, time-based animation, and privacy issues common in the real world and social technology world.
[0034] Furthermore, in actual distribution or communication use cases, it is necessary to define the encoded data and 3D models on the receiver (and encoder side). In proprietary systems, all of this data is known, but in open systems, the format of this data needs to be specified. This specification provides a mechanism for providing information and associated encoding formats used for the representation and encoding of user representations (also known as avatars).
[0035] The proposed method describes a humanoid format. However, it can be easily extended to any type of character (e.g., animals, plants).
[0036] Proposed Avatar Representation
[0037] The proposed representation of avatars, intended to be compatible with scene description (SD) content, can be divided into two main areas: - Static representation: Includes metadata (such as IDs) and static representations. - Dynamic representation: Animates and updates static representations.
[0038] The following sections will explain these elements in detail, along with their associated meanings, JSON encoding schemes, and how they can be used.
[0039] In the following description, the proposed format conforms to the glTF format and is compatible with the current MPEG-I Scene Description (SD) efforts to extend glTF with MPEG extensions. However, its meaning and use are general-purpose and can be coded in any other format (e.g., XML, USD).
[0040] Static Representation
[0041] The static representation of an avatar is a description of the exchangeable attributes and characteristics between the user and the 3D digital human (avatar). Tables 2 to 4 list such attributes, dividing them into four main areas. Specifically, "Metadata," "Geometry," "Visual," and "Add-ons" are designated overall areas for characterizing an avatar, and each area is described in more detail.
[0042] [Table 2]
[0043] Metadata
[0044] Identity includes all identity information of an avatar, such as name, gender, age, weight, and health. This identity information is not necessarily the user's actual name or age, but rather information set for the user avatar. If several avatars are used to represent a user, several versions of the user exist. Considering the confidentiality of this data, appropriate mechanisms (e.g., encryption) can be added for protection. Identity information is less sensitive if it is a pseudo-value or fictitious value.
[0045] Boxes: Avatar boundary areas that represent areas of interaction with other avatars or objects in the scene. These can change according to settings defined at the scene level or node level.
[0046] Illustrative use cases in this area include, for example, the following: - A "Social Box" corresponds to an attribute that an avatar defines in its social behavior state. This simply indicates whether an avatar wants to interact with other avatars. In general terms, setting a boundary area around an avatar at a distance of, for example, 1 meter, is called a "Social Box." Within this area, the "Interaction" flag is set and attached to the avatar node. As a result, any other avatar / 3D object that overlaps the "Social Box" will have permission to interact with this avatar. Conversely, the reverse of this setting can be applied to avoid social interaction with other avatars. - The Contact Box shares some similarities with the Social Box in terms of distance-related behavior. However, instead of social attributes, it allows for physical manipulation / interaction with nearby / contacting 3D objects. The Contact Box is the area around the avatar's body parts that signals collisions between the avatar's body parts and 3D objects, or that triggers haptic feedback. - The Restriction box defines an area related to the level of permission an avatar has regarding access to 3D interactive content. This box can be thought of as an access permission feature, such as an email address password. It can restrict or allow privileged access to scene elements. Within a 3D virtual environment, this can be thought of as content that requires a specific identifier to access or interact with, facilitating content creators to create private meeting rooms or restrict user interaction to predefined spaces. - The Experience Box corresponds to a space where users can move freely without restrictions. A use case is a virtual museum experience, where the avatar user is only allowed to move freely within a specific boundary. The area where the "art" is placed cannot be moved or interacted with by the avatar user, but viewing is not restricted. - Parental Box protects children and young adults through parental controls by restricting interaction with permitted content.
[0047] Disability: A description of the avatar's disability, such as impairments in mobility, speech, or any other human sensory impairment, which may represent a human user. It may also include missing body parts of the user (e.g., arms, legs, etc.). It may also include robotic / prosthetic body parts that replace the user's body parts.
[0048] Incorporating information about user disabilities is useful in enabling the engine to process / render appropriate content for the user. For example, if the application is a video conferencing tool, a real user with a speech impairment may not be able to use microphone input, and if the user has a hearing impairment, the application may need to render visual cues and synthesize text from speech.
[0049] Capability: An attribute that describes an avatar's abilities, such as the ability to run, walk, jump, talk, or fly. In the case of Disability, it can also describe the effect the disability has on the ability. The type of ability can be restricted to triggered events that allow only specific actions to be performed; for example, in a meeting room, spectator avatars may only be allowed to use speech, as well as upper body movements such as gestures and head movements. This attribute allows client applications to assign predefined functions to specific time-based events. This attribute is defined at the avatar level and can be combined with other abilities defined at the scene level to restrict or enhance the default avatar abilities. However, this represents what the avatar can do by default, such as a set of predefined animations.
[0050] Personality: Social attributes influence a person's social behavior and affect avatar animations, such as being calm, introverted, extroverted, sociable, or friendly. Personality influences the box area and can affect the realm of social interaction and interaction authority. Therefore, personality attributes, depending on an individual's characteristics, have a significant impact and play a crucial role in social interaction and privacy.
[0051] Emotion: The emotion attribute represents the current state of the avatar, such as happiness, sadness, anger, frustration, etc. Since the type of emotion strongly influences an individual's motor behavior, client applications can infer predefined motor behaviors based on this attribute. For example, unstable emotions manifest as irregular movement patterns.
[0052] Geometry
[0053] Geometry describes the shape and semantics of the entire body. For example, Non-Patent Document 1 describes the skeletal structure and mesh graph connectivity of an avatar. The whole body is separated into two sections, namely the upper and lower body, and each section is further divided into subordinate sections. The body model is represented by a mesh containing vertices, faces, and normals, as detailed in Non-Patent Document 1.
[0054] This semantic attribute is a dictionary that associates each body section with the avatar's vertices / faces. This makes it easier for lower-level information (vertices) to directly access higher-level data (body regions) and generate avatar-to-scene interactions such as contact triggers, as well as avatar-to-avatar interactions.
[0055] Table 3 shows the semantic description of the "Geometry" column in Table 2.
[0056] [Table 3]
[0057] Visuals
[0058] Table 4 shows the semantic description of the "Visual" column in Table 2.
[0059] [Table 4]
[0060] Add-ons
[0061] Table 5 shows the semantic description of the "Add-ons" column in Table 2.
[0062] [Table 5]
[0063] Dynamic Representation
[0064] An avatar is described as a digital representation of a human being. There are two purposes for creating a human avatar: user representation or virtual scene interaction. Human representation in a digital context can be achieved in two different ways: volumetric or synthetic. A volumetric representation is a 3D object that naturally encapsulates the shape and appearance of a human being. Thus, volumetric content accurately represents the anatomical structure of the user's body, such as shape morphology, limb dimensions, height, or a specific wardrobe, as well as the user's appearance, such as eye / hair / skin color, natural skin nuances (freckles), or patterns seen in different clothing. A synthetic representation is a 3D generated or statistically learned 3D object of the human body. Synthetic representations are easier to manipulate and generate to suit specific body structures. Synthetic models also facilitate the generalization and stylization of appearance, which can be performed by experts and used in animation streaming.
[0065] In the previous example shown in Figure 5, when using a synthetic model, the data available from sensor data is limited. Therefore, according to one embodiment, we propose to generate and create a robust avatar that faithfully mimics human anatomy, mechanics, and social interactions by extending the pipeline and introducing additional markers, as shown in Figure 6.
[0066] In the proposed example, the 3D model is augmented with facial landmarks, eye-tracking information, skeletal structure, and body semantics. In this way, the model becomes personalized for the user and capable of responding to interactions within the scene.
[0067] In particular, Figure 6 shows a proposed pipeline extension for processing captured human shape and posture for the purpose of animating and representing a 3D representation of a user, according to one embodiment. With the exception of blocks 610, 620, 621, 630, 635, 645, and 660, the blocks in Figure 6 signal the proposed extension. Sensors (610) are general input devices (camera, controller, or IMU), and head, body, clothing, hand, and foot encoders (620, 621, 622, 623, 624, 625) are processing models that extract specific information from the sensors and feed it to the avatar reconstruction, retargeting, and animation module (650). This process generates personalized avatars that are fed into a shared space (660) for rendering, and each processing model is divided into sub-processing models (631, 632, 633, 634, 636, 637) to process the data from the sensors in more detail. Models (630, 631, 632, 633) relate to head information and aim to extract facial landmarks and position (640), gaze and shape (641), hair shape and posture (642), and jaw shape (643). Models (634, 635, 636) relate to the body and extract skeletal hierarchy, shape, posture, spatial position (644), and the relationship between the semantics of body parts and their respective geometries (644). Similarly, model (623) extracts hand skeletal and posture information (646) from sensor data. Model (624) calculates foot-to-ground contact (647) to synchronize animation. Finally, model (625) includes additional metadata information to facilitate the personalization of the digital avatar.
[0068] Avatar Reproduction Standards
[0069] As mentioned above, avatars can take on many shapes and forms, allowing for the creation of unique avatars that closely resemble the physical and mental characteristics of real humans. Therefore, before instantiating the dynamic 3D mesh of an avatar, its properties must be reconstructed and defined.
[0070] In one embodiment, the avatar should respect the following requirements: 1. Avatars may be reconstructed or animated based on the extensive specifications presented above. 2. Reconstruction and animation respect the primitives supported in the scene description. 3. Avatars are represented by semantic labels that allow for the triggering of interactions between individual body parts. 4. The avatar's shape includes a surrounding area with dynamic dimensions. This facilitates the creation of social barriers between the avatar and the dynamic elements of the scene, triggering interactions with other avatars.
[0071] The proposed avatar is reconstructed upon user input and mapped onto a reference 3D avatar, for example, the 3D avatar "MPEG reference avatar" presented in Non-Patent Document 1, which includes a full-body based mesh representation, skeleton and skinning weights, facial blend shapes and landmarks, and additional geometry including eyes, jaw, teeth, and tongue. This makes it possible to have a common geometric and feature representation for any humanoid character (of course, the shape can be adapted by appropriate deformation). This common representation facilitates animation and mapping of any animation to different forms or contexts.
[0072] Avatar in Scene Description
[0073] By adding the attributes presented above to the scene description, the existing glTF node "MPEG_node_avatar" element is extended.
[0074] The glTF standard allows for the definition of 3D meshes, skeletal structures, and skinning weights, so no extensions are needed to enable the visual appearance of avatars. The proposed extensions indicate what type of avatar the "MPEG node avatar" node refers to (i.e., whether to use "MPEG reference avatar" as a template or something else), add additional information regarding avatar representation, and add interactivity constraints to the avatar and its respective elements.
[0075] Figure 7 shows the contributions to MPEG_node_avatar in MPEG-I SD (Non-Patent Literature 2).
[0076] In particular, Figure 7 shows the glTF file structure. The entry point is a "scene" node (730) containing a "node" node (735) which can be either "camera" (710), "light" (765), or "mesh" (740). MPEG-I SD defines a new boolean attribute, "MPEG_node_avatar" (735), which indicates whether a node corresponds to an avatar node, as shown in Table 6. The purpose of this boolean attribute is to clarify which nodes define the user representation.
[0077] This method is limited because it does not provide visual or animation information about the avatar asset. Such information must be explicitly defined on the following nodes. It appears to be defined as a "material" (770) with "texture" (790), "technique" (775), "program" (780), and "shader" (785), which are directly referenced by either "source" (795) or "image" (798). For animation, it is defined in "animation" (720) and "skin" (725) in "accessor" (745). They are all displayed using "bufferView" (750) and "buffer" node information (755). "MPEG_media" (760) can redefine several existing attributes such as geometry, appearance, sound, haptics, and / or animation by referencing external data.
[0078] [Table 6]
[0079] The inventors propose to use the node "MPEG_node_avatar" to represent the widest possible range of avatar representations, complementing the contributions presented in Non-Patent Literature 2 with a new extension of glTF node elements. This aims to standardize avatar representations in scene descriptions while maintaining previously presented and newly introduced features.
[0080] The metadata includes all of the aforementioned elements that define / identify the avatar (such as identity, age, name, and gender), as shown in Table 7.
[0081] [Table 7]
[0082] Table 8 shows the format description for the "Geometry" property shown in Table 2.
[0083] [Table 8]
[0084] Table 9 shows a semantic list of names available in the avatar model. These names are used in relation to the validity of the mappings shown in Table 8.
[0085] [Table 9]
[0086] Examples of glTF schemes
[0087] The following glTF is an example of instantiating a new avatar attribute in a client that supports "MPEG_node_avatar". Note that a large number of instantiations are possible depending on the application. Here is an example for illustrative purposes.
[0088] The following example demonstrates how to signal the level of detail of the model and texture, along with the semantic definition of the head for vertex indices.
[0089] [Table 10]
[0090] Here, we provide an example to illustrate the capability extension attribute and how it is defined in the glTF format using MPEG extensions. This attribute informs the client about the avatar's capabilities. In this scenario, the avatar is capable of jumping, flying, and running by default. This signal provides semantics to the animation engine, which then generates and processes the following capabilities. The engine's purpose is to ensure that this animation is available to the end user. From a formatting perspective, this signal can be used to evaluate whether the conditions of the 3D environment are possible and to respect the avatar's default capabilities.
[0091] [Table 11]
[0092] Following the same principle as described above, the following examples are merely syntax examples, and interpretation must be handled by the processing / engine model. Here is an example to illustrate the disability extension attribute.
[0093] [Table 12]
[0094] Here is an example to illustrate the emotion extension attribute.
[0095] [Table 13]
[0096] Here is an example to illustrate personality extension attributes.
[0097] [Table 14]
[0098] Here is an example to illustrate the complete metadata extension attributes.
[0099] [Table 15]
[0100] For each of the metadata examples shown above, the user / application has several ways of interpreting the metadata. A common and supported technique is the interaction between scene objects and avatars. The scene will consist of nodes, each representing a 3D mesh. Avatar nodes trigger events when they interact with nodes in the scene, so some node interaction occurs where metadata information stores can be used at the avatar node level. In a practical example, parental attributes are compared to the interactive node, and if the avatar's parental attributes are higher than those present on such a node, the interaction is allowed; otherwise, it is stopped. This behavior is selected by the processing / engine model, and multiple behaviors can be instantiated from this standard format.
[0101] application
[0102] Avatars are represented by 3D geometry that constitutes the majority of the visuals and manipulators for generating realistic content in 3D applications. The problem arises when such user virtual representations become social and allow access to interactions with other virtual users or objects, requiring additional information to facilitate, improve, and restrict such interactions. Therefore, this application provides additional metadata for avatar representations in virtual 3D environments.
[0103] Figure 8 shows a processing model for handling such metadata according to one embodiment. This engine processes all dynamic data (sensor input data) and static data (user metadata input and scene description) to generate unique user avatar reconstructions and animations. In particular, the Human Features Encoder (810) relates to the encoder proposed and illustrated in Figure 6. The engine (835) shows applications such as bots, not limited to Unreal, Unity, Blender, etc., that generate personalized avatars using the decoded (830) content provided by the Human Features Encoder and scene description + avatar metadata (820).
[0104] More specifically, Figure 8 illustrates the processing of dynamic sensor data and user static metadata for generating a personalized user avatar experience. The sensors represent real-world device information that is fed into a human encoder model to extract relevant information (e.g., human posture, movement, and appearance) to generate a personalized avatar. The scene description file and avatar metadata are combined with encoded features to inform the engine of how to render (840), reconstruct (850), animate (860), and interact with other virtual objects (870), i.e., the definition of the personalized avatar.
[0105] Various numerical values are used in this application. Certain values are illustrative, and the embodiments described are not limited to these specific values.
[0106] Various methods are described herein, each of which includes one or more steps or actions to implement the described method. The order and / or use of any particular steps and / or actions can be modified or combined unless a specific order of steps or actions is required for the proper operation of the method. Furthermore, terms such as “first,” “second,” etc., may be used in various embodiments to modify elements, components, steps, operations, etc., such as “first decryption” and “second decryption.” The use of such terms does not imply a modified order of operations unless specifically requested. Therefore, in this example, the first decryption does not need to be performed before the second decryption, but may, for example, be performed before the second decryption, during the second decryption, or during a period overlapping with the second decryption.
[0107] The embodiments and aspects described herein can be implemented, for example, as methods or processes, apparatus, software programs, data streams, or signals. Even if an embodiment of a described feature is described only in the context of a single form (for example, discussed only as a method), it can still be implemented in other forms (for example, apparatus or programs). Apparatus can be implemented, for example, as appropriate hardware, software, and firmware. Methods can be implemented in apparatus such as processors, which generally refer to processing devices, including, for example, computers, microprocessors, integrated circuits, or programmable logic devices. Processors also include communication devices, such as computers, mobile phones, portable / personal digital assistants (PDAs), and other devices that facilitate the communication of information between end users.
[0108] The phrases "one embodiment," "an embodiment," "one implementation," or "an implementation," as well as references to other variations thereof, mean that certain features, structures, characteristics, etc., described in relation to an embodiment are included in at least one embodiment. Therefore, the phrases "in one embodiment," "in one embodiment," "in one implementation," or "in an implementation," appearing in various places throughout this application, as well as any other variations, do not necessarily all refer to the same embodiment.
[0109] Furthermore, this application may refer to "determining" various types of information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from memory.
[0110] Furthermore, this application may refer to "accessing" various types of information. Accessing information may include, for example, one or more of the following: receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0111] Furthermore, this application may refer to "receiving" various types of information. Receiving is intended to be a broad term, similar to "accessing." Receiving information may include, for example, accessing information or retrieving information (for example, from memory) one or more of these. Moreover, "receiving" typically involves, in some way, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information during the operation.
[0112] For example, in the cases of "A / B", "A and / or B", and "at least one of A and B", it should be understood that the use of " / ", "and / or", and "at least one of" is intended to encompass the selection of only the first enumerated option (A), or only the second enumerated option (B), or both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such phrasing is intended to encompass the selection of only the first enumerated option (A), or only the second enumerated option (B), or only the third enumerated option (C), or only the first and second enumerated options (A and B), or only the first and third enumerated options (A and C), or only the second and third enumerated options (B and C), or all three options (A, B, and C). This can be extended to the same number of items listed, as will be obvious to those skilled in the art in this and related fields.
[0113] As will be apparent to those skilled in the art, embodiments can generate a variety of signals formatted to carry information that can be stored or transmitted. This information may include, for example, instructions for performing a method or data generated by one of the embodiments described. For example, a signal may be formatted to carry a bitstream of the embodiment described. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is well known. The signal may be stored in a processor-readable medium.
Claims
1. Obtaining from the description of the augmented reality scene at least one parameter used to represent the avatar, wherein the at least one parameter is - Identity and, - A boundary region representing the area of interaction between the avatar and other objects in the aforementioned scene, - The aforementioned avatar malfunction and, - The abilities of the aforementioned avatar, - The personality of the aforementioned avatar, - The emotions of the aforementioned avatar, Including one or more of the following, Obtaining 3D geometry data and textures associated with the aforementioned avatar, A method that includes this.
2. In describing an augmented reality scene, encoding at least one parameter used to represent an avatar, wherein the at least one parameter is - Identity and, - A boundary region representing the area of interaction between the avatar and other objects in the aforementioned scene, - The aforementioned avatar malfunction and, - The abilities of the aforementioned avatar, - The personality of the aforementioned avatar, - The emotions of the aforementioned avatar, This includes, Encoding the 3D geometry data and textures associated with the aforementioned avatar, A method that includes this.
3. The method according to claim 1 or 2, wherein the boundary region corresponds to any of the following: (1) an attribute defined by the avatar in a social behavior state; (2) a region that allows the avatar to perform physical manipulation or interaction with 3D objects in the scene; (3) a space that restricts user interaction with the avatar; (4) a space in which the avatar is permitted to move; and (5) a region set by parental controls.
4. The method according to any one of claims 1 to 3, wherein the 3D geometry data includes information regarding the level of detail.
5. The method according to any one of claims 1 to 4, wherein the 3D geometry data includes at least one of an eye model and a hair model.
6. The method according to any one of claims 1 to 5, wherein the texture includes information regarding the level of detail.
7. The method according to any one of claims 1 to 6, wherein the representation of the avatar is based on a reference template model.
8. The method according to any one of claims 1 to 7, wherein the at least one parameter further includes an attribute that defines a texture map to be superimposed on the character's skin.
9. Acquiring input data from the sensor, Based on the input from the aforementioned sensors, the avatar is reconstructed and animation is generated. The method according to any one of claims 1 to 8, further comprising:
10. One or more processors, At least one memory connected to the one or more processors, A device comprising, the one or more processors, Obtaining from the description of the augmented reality scene at least one parameter used to represent the avatar, wherein the at least one parameter is - Identity and, - A boundary region representing the area of interaction between the avatar and other objects in the aforementioned scene, - The aforementioned avatar malfunction and, - The abilities of the aforementioned avatar, - The personality of the aforementioned avatar, - The emotions of the aforementioned avatar, Including one or more of the following, Obtaining 3D geometry data and textures associated with the aforementioned avatar, A device configured to perform [a certain action].
11. One or more processors, At least one memory connected to the one or more processors, A device comprising, the one or more processors, In describing an augmented reality scene, encoding at least one parameter used to represent an avatar, wherein the at least one parameter is - Identity and, - A boundary region representing the area of interaction between the avatar and other objects in the aforementioned scene, - The aforementioned avatar malfunction and, - The abilities of the aforementioned avatar, - The personality of the aforementioned avatar, - The emotions of the aforementioned avatar, This includes, Encoding the 3D geometry data and textures associated with the aforementioned avatar, A device configured to perform [a certain action].
12. The apparatus according to claim 10 or 11, wherein the boundary region corresponds to any of the following: (1) an attribute defined by the avatar in a social behavior state; (2) a region that allows the avatar to perform physical manipulation or interaction with 3D objects in the scene; (3) a space that restricts user interaction with the avatar; (4) a space in which the avatar is permitted to move; and (5) a region set by parental controls.
13. The apparatus according to any one of claims 10 to 12, wherein the 3D geometry data includes information regarding the level of detail.
14. The apparatus according to any one of claims 10 to 13, wherein the 3D geometry data includes at least one of an eye model and a hair model.
15. The apparatus according to any one of claims 10 to 14, wherein the texture includes information regarding the level of detail.
16. The apparatus according to any one of claims 10 to 15, wherein the representation of the avatar is based on a reference template model.
17. The apparatus according to any one of claims 10 to 16, wherein the at least one parameter further includes an attribute that defines a texture map to be superimposed on the character's skin.
18. The one or more processors described above are Acquiring input data from the sensor, Based on the input from the aforementioned sensors, the avatar is reconstructed and animation is generated. The apparatus according to any one of claims 10 to 17, configured to further carry out the above.
19. A non-temporary computer-readable medium that, when executed by a computer, includes an instruction causing the computer to perform the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
IEC23090-14