Virtual character metadata representation

By introducing virtual character parameters and 3D geometric data into the extended reality scene description, the problem of insufficient virtual character representation in existing technologies is solved, achieving a more realistic and personalized virtual character representation, and enhancing interactivity and immersion.

CN120958488APending Publication Date: 2025-11-14INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480021738.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-24
Filing Date
2024-03-18
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing methods for representing virtual characters fail to effectively convey key information, such as gaze, body position, and skeletal animation, in extended reality scenarios. Furthermore, they do not consider social behavior, time-based animation, and privacy issues, resulting in limitations in interactivity and immersion.

Method used

By introducing parameters for virtual characters in extended reality scene descriptions, including identity, interaction areas, obstacles, abilities, personality, and emotions, and combining 3D geometric data and textures, the node tree is expanded using the MPEG-I scene description framework and glTF format to achieve richer representations of virtual characters.

Benefits of technology

It enables more realistic and personalized virtual character representations in extended reality scenarios, supports social interaction and privacy protection, and enhances the interactivity and immersion of the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120958488A_ABST
    Figure CN120958488A_ABST
Patent Text Reader

Abstract

In one implementation, additional metadata is presented to normalize social interactions between computer-generated 3D models (e.g., dynamic or static objects) in a virtual environment. For example, a virtual character is characterized using "metadata", "geometry", "visual", and "additional items". In particular, "metadata" may indicate the identity, activity space, obstacle, ability, personality, and emotion of a virtual character; 'Geometry' refers to the shape and semantics of the whole body; the visual sense provides description of a texture map and is used for rendering the material of the virtual character and covering the detail depth of the appearance of the virtual character; the "add-on" may specify the garment and accessories of the virtual character. To generate realistic content in 3D applications, an engine will process all dynamic (sensor input data) and static data (user metadata inputs and scene descriptions) and generate unique user virtual character reconstructions and animations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This embodiment generally relates to the representation and interaction of digital humans within a 3D engineered virtual scene. Background Technology

[0002] Extended Reality (XR) is a technology that enables interactive experiences where real-world environments and / or video content are augmented with virtual content that can be defined across multiple sensory modalities, including visual, auditory, tactile, and other sensory experiences. During application runtime, the virtual content (e.g., 3D content or audio / video files) is rendered in real-time in a manner consistent with the user's context (environment, viewpoint, device, etc.).

[0003] Scene graphs (such as / glTF (Graphics Language Transport Format) proposed by Khronos and its extensions defined in the MPEG scene description format, or Apple / USDZ) are possible ways to represent content to be rendered. They combine a declarative description of the scene structure that links real-world objects and virtual objects with a binary representation of the virtual content. The scene description framework ensures that timing media and corresponding associated virtual content are available at any point during the application's rendering process. Scene descriptions can also carry scene-level data that describes how users interact with scene objects at runtime to achieve an immersive XR experience. Summary of the Invention

[0004] According to one embodiment, a method is proposed, comprising: obtaining at least one parameter for representing a virtual character from a description for an extended reality scene, wherein the at least one parameter includes one or more of the following: identity, a boundary region representing an interaction area between the virtual character and other objects in the scene, obstacles of the virtual character, abilities of the virtual character, personality of the virtual character, and emotions of the virtual character; and obtaining 3D geometric data and textures associated with the virtual character.

[0005] According to another embodiment, a method is proposed, comprising: encoding at least one parameter for representing a virtual character in a description for an extended reality scene, wherein the at least one parameter includes one or more of the following: identity, a boundary region representing an area of ​​interaction between the virtual character and other objects in the scene, obstacles of the virtual character, abilities of the virtual character, personality of the virtual character, and emotions of the virtual character; and encoding 3D geometric data and textures associated with the virtual character.

[0006] According to another embodiment, an apparatus is proposed, comprising: one or more processors; and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to: obtain at least one parameter representing a virtual character from a description for an extended reality scene, wherein the at least one parameter includes one or more of the following: identity, a boundary region representing an area of ​​interaction between the virtual character and other objects in the scene, obstacles of the virtual character, abilities of the virtual character, personality of the virtual character, and emotions of the virtual character; and obtain 3D geometric data and textures associated with the virtual character.

[0007] According to one embodiment, an apparatus is proposed, comprising: one or more processors; and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to: encode at least one parameter for representing a virtual character in a description for an extended reality scene, wherein the at least one parameter includes one or more of the following: identity, a boundary region representing an area of ​​interaction between the virtual character and other objects in the scene, obstacles of the virtual character, abilities of the virtual character, personality of the virtual character, and emotions of the virtual character; and encode 3D geometric data and textures associated with the virtual character.

[0008] One or more embodiments also provide a computer program including instructions that, when executed by one or more processors, cause the one or more processors to perform a method according to any embodiment described herein. One or more embodiments also provide a computer-readable storage medium storing instructions thereon for processing a scenario description according to the method described herein.

[0009] One or more embodiments also provide a computer-readable storage medium storing a scene description generated according to the method described above. One or more embodiments also provide a method and apparatus for transmitting or receiving a scene description generated according to the method described herein. Attached Figure Description

[0010] Figure 1 An exemplary architecture for an XR processing engine is shown.

[0011] Figure 2 This shows a syntax example of encoding a data stream that extends the description of a real-world scene.

[0012] Figure 3 An example diagram illustrating an extended reality scene is shown.

[0013] Figure 4 This example shows an extended reality scene description that includes behavioral data.

[0014] Figure 5 This demonstrates the architecture of current solutions that utilize synthetic representations.

[0015] Figure 6 A pipeline with additional markers introduced according to an embodiment is shown.

[0016] Figure 7 This demonstrates the contribution of the MPEG node avatar to MPEG-I scene description (SD).

[0017] Figure 8 This demonstrates a processing model for processing such metadata according to an embodiment. Detailed Implementation

[0018] Various XR applications can be applied to different contexts and real or virtual environments. For example, in industrial XR applications, when a reference object (part B of an engine) is detected in a real environment by a camera device on a head-mounted display, a virtual 3D content item (e.g., part A of the engine) is displayed. This 3D content item is positioned in the real world, and its position and scale are defined relative to the detected reference object.

[0019] For example, in XR applications for interior design, 3D models of furniture are displayed when a given image from a catalog is detected in the input camera view. The 3D content is positioned in the real world, its location and scale defined relative to the detected reference image. In another application, audio files can be played when a user enters an area near a church (whether real or virtually rendered in an extended reality environment). In yet another example, a short advertising track can be played when a user sees a given can of soda in a real environment. In outdoor gaming applications, various virtual characters can appear, depending on the semantics of the scenery the user is observing. For example, bird characters are suitable for trees, so if the XR device's sensors detect a real object described by the semantic tag "tree," birds can be added flying around the trees. In assistive applications enabled by smart glasses, car noise can be played in the user's headphones when a car is detected within the user's camera's field of view to warn them of potential danger. Furthermore, the sound can be spatialized so that it comes from the direction the car was detected.

[0020] XR applications can also enhance video content rather than the real environment. The video is displayed on a rendering device, and when a timed event is detected in the video, virtual objects described in a node tree are overlaid. In this context, the node tree contains only descriptions of the virtual objects.

[0021] Figure 1 An exemplary architecture of an XR processing engine 130 is shown, which can be configured to implement the methods described herein. Figure 1 The devices in the architecture are linked to other devices via their bus 131 and / or via I / O interface 136.

[0022] Device 130 includes the following elements linked together via data and address bus 131:

[0023] - Microprocessor 132 (or CPU), such as DSP (or digital signal processor);

[0024] -ROM (or read-only memory) 133;

[0025] -RAM (or random access memory) 134;

[0026] - Storage interface 135;

[0027] -I / O interface 136, used to receive data to be transmitted from the application; and

[0028] -Power supply (not in) Figure 1 (as indicated in the text), for example, a battery.

[0029] According to the example, the power supply is external to the device. In each of the mentioned memories, the term "register" as used in the specification can correspond to a small area (a few bits) or a very large area (e.g., the entire program or a large amount of received or decoded data). ROM 133 includes at least one program and parameters. ROM 133 can store algorithms and instructions for performing the technology according to this principle. When powered on, CPU 132 uploads the program to RAM and executes the corresponding instructions.

[0030] RAM 134 contains the program executed by CPU 132 and uploaded after device 130 is turned on, input data in the register, intermediate data in the register at different states of the method, and other variables in the register used to execute the method.

[0031] Device 130 is linked, for example, via bus 131 to a set of sensors 137 and a set of rendering devices 138. Sensors 137 may be, for example, a camera, microphone, temperature sensor, inertial measurement unit, GPS, humidity sensor, infrared or ultraviolet light sensor, or wind sensor. Rendering devices 138 may be, for example, a display, speaker, vibrator, heater, fan, etc.

[0032] According to the example, device 130 is configured to implement the methods according to this principle and belongs to the following set:

[0033] -mobile device;

[0034] - Communication equipment;

[0035] -Gaming devices;

[0036] - Tablet PC (or tablet computer);

[0037] - Laptop;

[0038] - Still image camera;

[0039] -Camera.

[0040] In XR applications, scene descriptions are used to combine an explicit and easily parsed description of the scene structure with some binary representation of the media content. Figure 2 This shows a syntax example of encoding a data stream that extends the description of a real-world scene. Figure 2 An exemplary structure for an XR scene description 210 is shown. This structure consists of containers that organize the stream into individual syntax elements. The structure may include a header section 220, which is a set of data common to each syntax element of the stream. For example, the header section contains metadata about the syntax elements, describing their respective properties and functions. The structure also includes payloads containing syntax elements 230 and 240. Syntax element 230 contains data representing media content items associated with virtual elements described in the scene graph nodes. Images, meshes, and other raw data may have been compressed according to a compression method. Syntax element 240 is part of the data stream payload and contains data encoding the scene description as described according to this principle.

[0041] Figure 3 An example diagram of an extended reality scene description 310 is shown. In this example, the scene diagram may include descriptions of real-world objects (e.g., a “flat horizontal surface”, which could be a table or a road) and descriptions of virtual objects 312 (e.g., an animation of a car). The scene description is organized as an array of nodes. Nodes can be linked to child nodes to form a scene structure 311. Nodes can carry descriptions of real-world objects (e.g., semantic descriptions) or descriptions of virtual objects. Figure 3 In the example, node 301 describes a virtual camera located within the 3D volume of an XR application. Node 302 describes a virtual car and contains an index of the car's representation, such as an index in a 3D mesh array. Node 303 is a child node of node 302 and contains a description of one of the car's wheels. Similarly, it contains an index of the 3D mesh pointing to the wheel. The same 3D mesh can be used for multiple objects in the 3D scene because the scene nodes describe the scale, position, and orientation of the objects. Scene graph 310 also includes nodes that describe the spatial relationships between real and virtual objects.

[0042] In time-based media streaming, the scene description itself can evolve over time to provide relevant virtual content for each sequence of the media stream. For example, for advertising purposes, a virtual bottle could be displayed on a table during a video sequence of people sitting around it. This behavior can be achieved by relying on the framework defined in the "Scene Description for MPEG Media" document.

[0043] Currently, the MPEG-I scene description framework uses "behavioral" data to enhance the scene description that evolves over time and provides a description of how users interact with scene objects at runtime to achieve an immersive XR experience. These behaviors are associated with predefined virtual objects that allow runtime interactivity to achieve user-specific XR experiences. These behaviors also evolve over time and are updated through existing scene description update mechanisms.

[0044] Figure 4 This example illustrates an extended reality scene description containing behavioral data stored at the scene level, describing how a user interacts with scene objects described at the node level at runtime to achieve an immersive XR experience. When an XR application launches, media content items (e.g., a mesh of virtual objects visible from the camera) are loaded, rendered, and buffered to be displayed when triggered. For example, when a sensor detects a planar surface in the real environment, the application displays the buffered media content item as described in the relevant scene node. Timing is managed by the application based on the timing of features detected in the real environment and animations. Nodes in the scene graph may also not contain descriptions and simply act as parent nodes of child nodes. Figure 4 This illustrates the relationships between behaviors contained at the scene level in the scene description and nodes that are components of the scene graph. Behavior 410 is associated with predefined virtual objects that allow runtime interactivity to achieve user-specific XR experiences. Behavior 410 also evolves over time and is updated through the scene description update mechanism.

[0045] An action includes:

[0046] Trigger 420, which defines the conditions that must be met for its activation;

[0047] Trigger control parameters define the logical operations between the defined triggers;

[0048] Action 430: This action must be executed when the trigger is activated.

[0049] Action control parameters define the execution order of related actions;

[0050] Priority numbering enables the selection of the highest priority behavior when multiple behaviors compete for the same virtual object simultaneously.

[0051] An optional interrupt action specifies how to terminate the behavior if it is no longer defined in a newly received scene update; for example, the behavior is no longer defined if the relevant object does not belong to the new scene, or if the behavior is no longer relevant to the current media (e.g., audio or video) sequence.

[0052] Behavior 410 occurs at the scene level. The trigger is linked to the node and its child nodes. Figure 4 In the example, trigger 1 is linked to nodes 1, 2, and 8. Since node 31 is a child of node 1, trigger 1 is linked to node 31. Trigger 1 is also linked to node 14, which is a child of node 8. Trigger 2 is linked to node 1. Indeed, the same node can be linked to multiple triggers. Trigger n is linked to nodes 5, 6, and 7. A line can include multiple triggers. For example, the first line can be activated by trigger 1 and (AND) trigger 2, where AND is the trigger control parameter for the first line. A line can have multiple actions. For example, the first line can execute action m first, then action 1, where "first…then" is the action control parameter for the first line. The second line, for example, can be activated by trigger n and execute action 1 first, then action 2.

[0053] Different formats can be used to represent the node tree. For example, the MPEG-I scene description framework using the Khronos glTF extension mechanism can be used for the node tree. In this example, the interactivity extension can be applied at the glTF scene level and is called MPEG_scene_interactivity. The corresponding semantics are provided in Table 1, where "M" in the "Usage" column indicates that the field is mandatory in the XR scene description format, and "O" indicates that the field is optional.

[0054] Table 1

[0055]

[0056] Current solutions for representing virtual characters

[0057] Digital humans can be morphologically constructed through model-based reconstruction pipelines that make assumptions about a captured subset in the form of 3D template models. These assumptions can be general human models with additional skeletal structures, subject-specific models, or statistical shape models. These methods accurately provide a 3D model as an initialization stage, capable of statistically representing different human shapes in various poses. This representation is often referred to as a synthetic representation and is easily manipulated and fabricated to match specific body anatomy. Synthetic models also facilitate appearance generalization and stylization, which can be performed by professionals and used for animation and streaming.

[0058] The current architecture diagram of the solution utilizing synthetic representation is as follows: Figure 5 As shown. The entry point is sensor data, which corresponds to all input data from the user's device. This data is then broken down and processed by different encoding modules: a head encoder (510), a body encoder (520), a hand encoder (530), and a head pose estimator (540). All this encoded data is then streamed over the network to a "virtual character reconstruction and animation" module (550), which reconstructs and animates the user's virtual character based on an "offline 3D model." The virtual character is then shared in any shared space for other users to display and visualize. This solution may have limitations because it does not consider key information such as gaze, body position, and skeletal animation. Furthermore, it does not convey non-morphological and non-biomechanical information such as psychological state, emotion, or personality.

[0059] from Figure 5 As can be seen, the sensor data is limited and does not provide information about the face, eyes, skeletal structure, clothing and accessories, feet, and semantic queues. Therefore, this approach has limitations when using synthetic models for streaming and interaction with digital humans. Furthermore, the model does not consider social behaviors, time-based animation, and privacy issues common in the real world and the world of social technology.

[0060] Furthermore, in practical distribution or communication use cases, it is necessary to define 3D models of the encoded data and the receiver (and encoder). In proprietary systems, all this data is known, but for open systems, the format of this data needs to be specified. In this document, we provide mechanisms for providing information and related encoding formats used to represent and encode user representations (also known as virtual characters).

[0061] The proposed method is described below for a format suitable for humanoid characters. However, it can be easily extended for any type of character (e.g., animals, plants).

[0062] The proposed virtual character representation

[0063] The proposed virtual character representation, designed to be compatible with scene description (SD) content, can be divided into two main areas:

[0064] -Static representation—It includes metadata (ID, etc.) and static representation.

[0065] -Dynamic representation—It animates and updates the static representation.

[0066] The following details these elements, including their meanings, JSON encoding schemes, and how to use them.

[0067] In the following description, the proposed format follows the glTF format and is compatible with current MPEG-I Scene Description (SD) work that extends glTF through MPEG extensions. However, its meaning and purpose are general and can be encoded in any other format (e.g., XML, USD).

[0068] Static representation

[0069] A static representation of a virtual character is a description of the attributes and characteristics that can be exchanged between the user and the 3D digital human (virtual character). Tables 2 through 4 describe these attributes and categorize them into four main domains. Specifically, “Metadata,” “Geometry,” “Visual,” and “Additional Items” are the overall domains that specify the characterization of the virtual character, each described in more detail.

[0070] Table 2 - Static Information Used for General Representation of Virtual Characters

[0071]

[0072] Metadata

[0073] Identity: Contains all identity information of the virtual character, such as name, gender, age, weight, health status, etc. This identity information is not necessarily the user's real name or age, but rather the name set for the user's virtual character. If multiple virtual characters are used to represent a single user, multiple versions of that user exist. Given the sensitivity of this data, appropriate mechanisms can be added for protection (e.g., encryption). Identity information is less sensitive in the case of spurious or inauthentic values.

[0074] Box: The boundary area of ​​a virtual character, representing the area for interaction with other virtual characters or objects in the scene. It can vary depending on settings defined at the scene or node level.

[0075] Illustrative use cases for this type of area could be as follows:

[0076] - The social box corresponds to an attribute defined by a virtual character in its social behavior state. This simply indicates whether a virtual character wishes to interact with other virtual characters. The common term is setting a boundary area (e.g., 1 meter away) around a virtual character, called the "social box." Within this area, an "interaction" flag is set and attached to the virtual character node. Therefore, any other virtual character / 3D object overlapping the "social box" will gain permission to interact with this virtual character. Converse rules can be applied to avoid social interactions with other virtual characters.

[0077] - The touchbox shares some similarities with the "social box" in terms of distance behavior. However, instead of social attributes, it allows for physical manipulation / interaction with nearby / touching 3D objects. The touchbox will be the area around a body part, used to indicate collisions between virtual character body parts and 3D objects or trigger haptic feedback.

[0078] - The restriction box defines the area associated with the level of permissions a virtual character has in accessing 3D interactive content. This box can be thought of as an access control function, such as a password for an email address. It can restrict or grant privileged access to scene elements. In a 3D virtual environment, it can be seen as content that requires a specific identifier to access or interact with, which helps content creators create private meeting rooms or restrict user interaction to predefined spaces.

[0079] - An experience box corresponds to a space that allows users to move freely without constraints. One use case is a virtual museum experience, where virtual character users are only allowed to move freely within a certain perimeter. The perimeter where the "artwork" is placed prohibits virtual character users from moving or interacting with it, but does not restrict viewing.

[0080] - Parental control boxes protect children and teenagers by limiting interaction to permitted content through parental controls.

[0081] Impairment: A description of a virtual character's impairments, such as mobility, speech, or any other impairments that can represent the human senses of a human user. It may also include missing parts of the user's body (e.g., arms, legs, etc.). It may also include mechanical / prosthetic body parts that replace parts of the user's body.

[0082] Introducing user barriers helps allow the engine to process / render content appropriate for the user. For example, if the application is a video conferencing tool, a real user with a speech impairment will not be able to input using a microphone; if the user has a hearing impairment, the application should render visual cues and synthesize text from speech.

[0083] Abilities: Attributes describing the capabilities of a virtual character, such as running, walking, jumping, speaking, and flying. It can also represent the impact of obstacles on the ability. Ability types can be restricted to triggered events that only allow certain actions; for example, in a meeting room, a bystander virtual character might only be allowed to use voice and upper body movements (such as gestures and head movements). This attribute helps client applications assign predetermined functionality to specific time-based events. This attribute is defined at the virtual character level and can complement other abilities defined at the scene level to limit or enhance the default virtual character abilities. Nevertheless, this represents the actions the virtual character can perform by default, such as predefined animation sets.

[0084] Personality: Social attributes influence a person's social behavior, thus affecting the animation of virtual characters, such as calmness, introversion, extroversion, and friendliness. Personality can influence the box area and affect the social interaction area and interaction permissions. Therefore, personality attributes have a significant impact and role in social interaction and privacy based on individual characteristics.

[0085] Emotion: The emotion attribute describes the current state of the virtual character, such as happiness, sadness, anger, frustration, etc. Emotion type strongly influences an individual's movement behavior, so client applications can infer predetermined movement behaviors based on this attribute. For example, unstable emotions exhibit irregular movement patterns.

[0086] geometry

[0087] Geometry refers to the overall shape and semantics. For example, Q. Avril et al.'s article, "Draft Annex to ISO / IEC 23090-14:2021-MPEG Reference Humanoid Avatar" (ISO / IEC JTC 1 / SC 29 / WG 3m61232, hereinafter referred to as "Avril"), describes the skeletal anatomy and mesh connectivity of the virtual character. The entire body is divided into two parts: the upper body and the lower body, each further divided into subordinate parts. The body model is represented by a mesh containing vertices, faces, and normals, as detailed in Avril's paper.

[0088] Semantic attributes are a dictionary that associates each body segment with a vertex / face of the virtual character. This allows low-level information (vertices) to directly access high-level data (body regions) and generate virtual character-to-scene and virtual character-to-virtual character interactions, such as contact triggers.

[0089] Table 3 shows the semantic description of the "Geometry" column in Table 2.

[0090] Table 3

[0091]

[0092] Visual

[0093] Table 4 shows the semantic description of the "Visual" column in Table 2.

[0094] Table 4

[0095]

[0096] Additional items

[0097] Table 5 shows the semantic description of the "Additional Items" column in Table 2.

[0098] Table 5

[0099]

[0100] Dynamic representation

[0101] Virtual characters are described as digital representations of humans. Creating human virtual characters serves two purposes: user representation or interaction within a virtual scene. In a digital context, human representation can be achieved in two distinct ways: volumetric or synthetic representation. A volumetric representation is a 3D object that naturally encapsulates the shape and appearance of the human body. Therefore, volumetric content accurately represents the user's body anatomy (e.g., shape, limb size, height, or specific clothing) and appearance (e.g., eye / hair / skin color, subtle features of natural skin (freckles), or patterns found in different garments). A synthetic representation is a 3D human object created through 3D hand-modeling or statistical learning. Synthetic representations are easier to manipulate and hand-model to match specific body anatomy. Synthetic models also facilitate appearance generalization and stylization, which can be performed by professionals and used for animation streaming.

[0102] In the past Figure 5 In the example shown, when using a synthetic model, the data utilized from sensor data is limited. Therefore, according to the embodiment, we propose as follows: Figure 6 The extended pipeline shown introduces additional markers to allow the generation and creation of robust virtual characters that closely mimic human anatomy, mechanics, and social interactions.

[0103] In the proposed example, the 3D model is extended with facial landmarks, eye-tracking information, skeletal structure, and body semantics. In this way, the model is personalized for the user and ready for scene interaction.

[0104] Specifically Figure 6 This demonstrates a pipeline extension proposed according to an embodiment for processing captured human body shapes and poses to achieve the goal of animate and represent a user's 3D representation. (Except for boxes 610, 620, 621, 630, 635, 645, and 660) Figure 6The box in the diagram indicates the proposed extension. The sensor (610) is a general-purpose input device (camera, controller, or IMU), and the head, body, clothing, hand, and foot encoders (620, 621, 622, 623, 624, 625) are processing models that extract specific information from the sensors and feed it to the virtual character reconstruction, redirection, and animation module (650). This process generates a personalized virtual character, which is then fed into a shared space (660) for rendering. Each processing model is divided into sub-processing models (631, 632, 633, 634, 636, 637) to specifically process data from the sensors. The models (630, 631, 632, 633) are associated with head information, and their goal is to extract facial landmarks and position (640), eye gaze and shape (641), hair shape and pose (642), and jaw shape (643). Models (634, 635, 636) are body-related, extracting skeletal hierarchy, shape, posture, spatial location (644), and the semantic relationships between body parts and their respective geometries (644). Similarly, model (623) extracts skeletal and posture information of the hand from sensor data (646). Model (624) calculates the contact between the foot and the ground (647) for synchronized animation. Finally, model (625) contains additional metadata information that helps personalize the digital virtual character.

[0105] Virtual Character Reconstruction Standards

[0106] As mentioned above, virtual characters can take on various shapes and forms, allowing for the creation of unique virtual characters that closely resemble the physiological and psychological characteristics of real humans. Therefore, it is necessary to reconstruct and define the attributes before instantiating a dynamic 3D virtual character mesh.

[0107] In one embodiment, the virtual character should meet the following requirements:

[0108] 1. Virtual characters can be reconstructed or animated using the aforementioned extensive specifications.

[0109] 2. Reconstruction and animation processing conform to the primitives supported in the scene description.

[0110] 3. Virtual characters are represented using semantic tags, which allow for interaction triggers for various body parts.

[0111] 4. The shape of a virtual character includes a surrounding area with dynamic dimensions. This facilitates interaction triggers with the scene and other virtual characters, and enables social barriers between virtual characters and dynamic elements of the scene.

[0112] The proposed virtual character is reconstructed based on user input and mapped to a reference 3D virtual character, such as the "MPEG Reference Virtual Character" proposed in Avril. This character includes a full-body mesh representation, skeletal and skinning weights, facial blending shapes and symbols, and additional geometry including eyes, jaw, teeth, and tongue. This allows for the sharing of geometric and feature representations for any humanoid character (with appropriate deformations allowing for shape adaptation). This shared representation simplifies the mapping and animation processing of any animation to different forms or contexts.

[0113] Virtual characters in scene description

[0114] In the scene description, we extend the existing glTF node “MPEG_node_avatar” element by adding the above attributes.

[0115] Because the glTF standard allows the definition of 3D meshes, skeletal hierarchies, and skinning weights, no extensions are needed to implement the visual appearance of virtual characters. The proposed extension indicates which type of virtual character the node "MPEG_node_avatar" references (i.e., whether it uses "MPEG Reference Virtual Character" as a template or another template), adds additional information about the representation of the virtual character, and adds interaction constraints to the virtual character and corresponding elements.

[0116] Figure 7 This demonstrates the contribution of MPEG_node_avatar to MPEG-I SD (see the article by E. Thomas et al. entitled “[SD]Signalling for Avatars in Scene Description” (ISO / IEC JTC 1 / SC 29 / WG 3m62027, hereinafter referred to as “m62027”).

[0117] Specifically Figure 7 This represents a glTF file structure. The entry point is the "scene" node (730), which contains "node" nodes (735), which can be "camera" (710), "light" (765), or "mesh" (740). MPEG-ISD has defined a new Boolean attribute "MPEG_node_avatar" (735), which indicates whether the node corresponds to a virtual character node, as shown in Table 6. The purpose of this Boolean attribute is to define which node defines the user representation.

[0118] This approach has limitations because it does not provide any visual or animation information about virtual character assets. Such information must be explicitly defined at the following nodes. For appearance, it is defined as “material” (770), “technique” (775), “program” (780), and “shader” (785), where “material” has a “texture” (790) referenced by either a direct “source” (795) or “image” (798). For animation, it is defined in “accessor” (745), which has “animation” (720) and “skin” (725). Both are displayed using “bufferView” (750) and “buffer” node information (755). “MPEG_media” (760) can redefine some existing attributes, such as geometry, appearance, sound, haptics, and / or animation, by referencing external data.

[0119] Table 6

[0120] name type usage default value describe isAvatar Boolean value M real Indicates whether the target is an active virtual character.

[0121] We propose to refine the contributions made in m62027 with a new extension to glTF node elements to represent a wide range of possible virtual character representations using the node “MPEG_node_avatar”. This is to normalize virtual character representations in scene descriptions while retaining the newly introduced features previously presented.

[0122] The metadata contains all the elements previously described for defining / identifying the virtual character (identity, age, name, gender, etc.), as shown in Table 7.

[0123] Table 7

[0124]

[0125] Table 8 shows the format description of the “geometry” attribute proposed in Table 2.

[0126] Table 8

[0127]

[0128] Table 9 presents a semantic list of available names from the virtual character model. These names are used for the mapping attributes presented in Table 8.

[0129] Table 9

[0130]

[0131] glTF Scheme Example

[0132] The following glTF example demonstrates instantiating a new virtual character attribute in a client that supports "MPEG_node_avatar". Note that depending on the application, there may be numerous instantiations. This example is provided for illustrative purposes.

[0133] The following example demonstrates how to indicate the level of detail in models and textures, as well as the semantic definition of the head relative to the vertex index.

[0134]

[0135] Here we provide an illustrative example demonstrating the capability extension attribute and how to define it in the glTF format using MPEG extensions. This attribute indicates the virtual character's capabilities to the client. In this scene, the virtual character is capable of jumping, flying, and running by default. This signal provides semantics to the animation engine, which then generates and processes these capabilities. The engine's goal is to ensure this animation is usable for the end user. From a format perspective, this signal can be used to assess whether the conditions of the 3D environment can support and conform to the virtual character's default capabilities.

[0136]

[0137] Following the same principles mentioned above, the following examples are merely syntactic examples; explanations should be provided at the processing / engine model level. Here we provide an illustrative example of the barrier extension attribute.

[0138]

[0139] Here we provide an illustrative example of the purpose of the emotional extension attribute.

[0140]

[0141] Here we provide an illustrative example of the purpose of personalized extended attributes.

[0142]

[0143] Here we provide an illustrative example of the purpose of a complete metadata extended attribute.

[0144]

[0145] For each metadata example shown above, there are several ways for the user / application to interpret the metadata. A common approach and supported technology is the interaction between scene objects and virtual characters. The scene will consist of nodes, each representing a 3D mesh. Virtual character nodes trigger events when they interact with nodes in the scene, so some node interactions can utilize metadata information stored at the virtual character node level. A practical example is that parental control attributes will be compared with the interacting node; if the virtual character's parental control attribute is higher than the parental control attribute present in that node, the interaction is allowed; otherwise, it stops. This behavior is selected by the processing / engine model, and various behaviors can be instantiated from this standard format.

[0146] application

[0147] Virtual characters are represented using 3D geometry, which constitutes most of the vision and manipulation mechanisms to generate realistic content in 3D applications. Problems arise when this user virtual representation becomes social and gains access to interact with other virtual users or objects, and when additional information is needed to facilitate, refine, and restrict such interactions. Therefore, this document provides additional metadata for virtual character representations in virtual 3D environments.

[0148] Figure 8 This demonstrates a processing model for handling such metadata according to an embodiment. The engine will process all dynamic (sensor input data) and static (user metadata input and scene description) data and generate a unique user avatar reconstruction and animation. Specifically, the human feature encoder (810) and Figure 6 The encoders proposed and demonstrated in this paper are related to the engine (835). The engine (835) refers to the application that uses the decoded (830) content provided by the human feature encoder and scene description + virtual character metadata (820) to create personalized virtual characters, such as, but not limited to, Unreal, Unity or Blender.

[0149] More specifically, Figure 8 This demonstrates the processing of dynamic sensor data and static user metadata to generate a personalized user virtual character experience. Sensors represent real-world device information, which is fed into a human encoder model to extract relevant information (e.g., human posture, motion, appearance) to generate the personalized virtual character. Scene description files and virtual character metadata are combined with encoded features to inform the engine how to render (840), reconstruct (850), animate (860), and interact with other virtual objects (870), thus defining the personalized virtual character.

[0150] Various numerical values ​​are used in this application. Specific values ​​are used for illustrative purposes, and the aspects described are not limited to these specific values.

[0151] This document describes various methods, each including one or more steps or actions for implementing the described method. The order and / or use of steps and / or actions can be modified or combined unless the correct operation of the method requires a specific order. Furthermore, terms such as "first" and "second" can be used in various embodiments to modify elements, components, steps, operations, etc., for example, "first decoding" and "second decoding." The use of such terms does not imply an ordering of the modified operations unless specifically required. Therefore, in this example, the first decoding does not need to be performed before the second decoding and can occur before, during, or within overlapping time periods of the second decoding.

[0152] The implementations and aspects described herein can be implemented, for example, in a method or process, apparatus, software program, data stream, or signal. Even if discussed only in the context of a single implementation form (e.g., discussed only as a method), the implementation of the discussed features can also be implemented in other forms (e.g., apparatus or program). An apparatus can be implemented, for example, in appropriate hardware, software, and firmware. A method can be implemented, for example, in an apparatus, such as a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as computers, mobile phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0153] The references to "an embodiment" or "an embodiment" or "an implementation" or "implementation" and their variations mean that a particular feature, structure, characteristic, etc., described in connection with that embodiment is included in at least one embodiment. Therefore, the phrases "in an embodiment" or "in an embodiment" or "in an implementation" or "in an implementation" and any other variations appearing throughout this application do not necessarily all refer to the same embodiment.

[0154] Furthermore, this application may involve "determining" various types of information. Determining information may include, for example, one or more of the following: estimation information, calculation information, prediction information, or information retrieved from memory.

[0155] Furthermore, this application may involve "accessing" various types of information. Accessing information may include, for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information, or one or more of these.

[0156] Furthermore, this application may relate to "receiving" various types of information. Like "accessing," "receiving" is a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) or one or more of them. Moreover, "receiving" typically involves some form of operation, such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0157] It should be understood that the use of any of the following “ / ”, “and / or”, and “at least one”, such as in the cases of “A / B”, “A and / or B”, and “at least one A and B”, is intended to cover selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one A, B, and C”, such wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). As will be known to those skilled in the art and related fields, this can be extended to any number of options listed.

[0158] As will be apparent to those skilled in the art, implementations can generate signals of various formats to carry information, which may be stored or transmitted. This information may include, for example, instructions for performing a method or data generated by one of the described implementations. For instance, the signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using a radio frequency portion of the spectrum) or a baseband signal. Formatting may, for example, include encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal can be transmitted via a variety of known wired or wireless links. The signal may be stored on a processor-readable medium.

Claims

1. A method comprising: Obtain at least one parameter representing the virtual character from the description used to extend the real-world scene, wherein the at least one parameter includes one or more of the following: -identity; - Represents the boundary area of ​​the interaction zone between the virtual character and other objects in the scene; - The obstacles of the virtual character; -The abilities of the virtual character; -The personality of the virtual character; and - The emotions of the virtual character; and Obtain the 3D geometric data and textures associated with the virtual character.

2. A method comprising: Encoding at least one parameter for representing a virtual character in the description used for extended reality scenarios, wherein the at least one parameter includes one or more of the following: -identity; - Represents the boundary area of ​​the interaction zone between the virtual character and other objects in the scene; - The obstacles of the virtual character; -The abilities of the virtual character; -The personality of the virtual character; and - The emotions of the virtual character; and Encode the 3D geometric data and textures associated with the virtual character.

3. The method according to claim 1 or 2, wherein the boundary area corresponds to one of the following: (1) an attribute defined by the virtual character in a social behavior state; (2) an area that allows the virtual character to physically manipulate or interact with 3D objects in the scene; (3) a space that restricts user interaction with the virtual character; (4) a space that allows the virtual character to move; and (5) an area set by parental controls.

4. The method according to any one of claims 1 to 3, wherein the 3D geometric data includes information about the level of detail.

5. The method according to any one of claims 1 to 4, wherein the 3D geometric data includes at least one of an eye model and a hair model.

6. The method according to any one of claims 1 to 5, wherein the texture includes information about the level of detail.

7. The method according to any one of claims 1 to 6, wherein the representation of the virtual character is based on a reference template model.

8. The method according to any one of claims 1 to 7, wherein the at least one parameter further includes an attribute defining a texture map covering the character's skin.

9. The method according to any one of claims 1 to 8, further comprising: Acquire input data from the sensor; as well as The reconstruction and animation of the virtual character are generated based on the input from the sensors.

10. An apparatus comprising: One or more processors; as well as At least one memory coupled to the one or more processors, wherein the one or more processors are configured to: Obtain at least one parameter representing the virtual character from the description used to extend the real-world scene, wherein the at least one parameter includes one or more of the following: -identity; - Represents the boundary area of ​​the interaction zone between the virtual character and other objects in the scene; - The obstacles of the virtual character; -The abilities of the virtual character; -The personality of the virtual character; and - The emotions of the virtual character; and Obtain the 3D geometric data and textures associated with the virtual character.

11. An apparatus comprising: One or more processors; as well as At least one memory coupled to the one or more processors, wherein the one or more processors are configured to: Encoding at least one parameter for representing a virtual character in the description used for extended reality scenarios, wherein the at least one parameter includes one or more of the following: -identity; - Represents the boundary area of ​​the interaction zone between the virtual character and other objects in the scene; - The obstacles of the virtual character; -The abilities of the virtual character; -The personality of the virtual character; and - The emotions of the virtual character; and Encode the 3D geometric data and textures associated with the virtual character.

12. The apparatus of claim 10 or 11, wherein the boundary region corresponds to one of: (1) an attribute defined by the virtual character in a social behavior state; (2) an area that allows the virtual character to physically manipulate or interact with 3D objects in the scene; (3) a space that restricts user interaction with the virtual character; (4) a space that allows the virtual character to move; and (5) an area set by parental controls.

13. The apparatus according to any one of claims 10 to 12, wherein the 3D geometric data includes information about the level of detail.

14. The apparatus according to any one of claims 10 to 13, wherein the 3D geometric data includes at least one of an eye model and a hair model.

15. The apparatus according to any one of claims 10 to 14, wherein the texture includes information about levels of detail.

16. The apparatus according to any one of claims 10 to 15, wherein the representation of the virtual character is based on a reference template model.

17. The apparatus according to any one of claims 10 to 16, wherein the at least one parameter further includes an attribute defining a texture map covering the character's skin.

18. The apparatus according to any one of claims 10 to 17, wherein the one or more processors are further configured to: Acquire input data from sensors; and The reconstruction and animation of the virtual character are generated based on the input from the sensors.

19. A non-transitory computer-readable medium comprising instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 1 to 9.