Avatar angular gaze in scene descriptions
The avatar gaze descriptor in scene description formats addresses the lack of standardized gaze animation methods, enabling more realistic and customizable avatar animations by specifying gaze animation data within the scene description format.
Patent Information
- Application Number
- PCT/EP2024/084340
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-11
- Filing Date
- 2024-12-02
- Publication Date
- 2025-06-19
AI Technical Summary
Current scene description formats, such as glTF, lack standardized methods for animating avatar components, specifically gaze directions, which limits the ability of rendering applications to accurately render avatars with dynamic gaze animations.
The introduction of an avatar gaze descriptor within the scene description format allows for the specification of gaze animation data, including horizontal, vertical, and roll rotations, enabling rendering applications to animate avatar components according to predefined gaze directions.
This solution enables more realistic and customizable avatar animations in rendering applications, allowing for alignment of avatar gaze with user directions in virtual reality and other applications, enhancing user interaction and immersion.
Smart Images

Figure EP2024084340_19062025_PF_FP_ABST
Abstract
Description
AVATAR ANGULAR GAZE IN SCENE DESCRIPTIONSCROSS REFERENCE TO RELATED APPLICATIONS[1] This application claims the benefit of European Application No. 23307173.7, filed on December 11, 2023, which is incorporated herein by reference in its entirety.BACKGROUND[2] Scene description formats allow for efficient representations and delivery of scene elements to rendering applications. A common scene description format is the graphics library transmission format (glTF) that enables both designers of digital assets and application developers to generate and to use, respectively, the digital assets. Generally, glTF defines nodes representing (among other scene elements) objects, including respective meshes and animation descriptors, based on which an application can render the objects at the scene. The glTF allows for extensions, such as the MPEG node avatar extension introduced by the MPEG-I Scene Description (SD) format to define an avatar object. Using such an extension, a rendering application can identify a node, parsed from a glTF file, as one that describes an avatar object (that is, an avatar node) and render it accordingly. Incorporating directive data into that avatar node can be useful in further guiding the rendering application on how to handle the avatar object.SUMMARY[3] Aspects disclosed in the present disclosure describe methods for rendering an avatar. The methods comprise obtaining a scene description including node data representing an avatar, extracting from the node data gaze data that specify animation of a component of the avatar, and rendering the avatar using the scene description. The rendering includes animating the component of the avatar according to the gaze data. Aspects disclosed in the present disclosure also describe methods for generating a scene description. The methods comprise generating node data representing an avatar, including generating gaze data specifying animation of a component of the avatar.[4] Aspects disclosed in the present disclosure describe apparatuses for rendering an avatar. The apparatuses comprise at least one processor and memory storing instructions. The instructions, when executed by the at least one processor, cause the apparatuses to obtain a scene description including node data representing an avatar, to extract from the node data gazedata that specify animation of a component of the avatar, and to render the avatar using the scene description. The rendering includes animating the component of the avatar according to the gaze data. Aspects disclosed in the present disclosure also describe apparatuses for generating a scene description. The apparatuses comprise at least one processor and memory storing instructions. The instructions, when executed by the at least one processor, cause the apparatuses to generate node data representing an avatar, including generating gaze data specifying animation of a component of the avatar.[5] Aspects disclosed in the present disclosure describe a non-transitory computer-readable medium comprising instructions executable by at least one processor to perform methods for rendering an avatar. The methods comprise obtaining a scene description including node data representing an avatar, extracting from the node data gaze data that specify animation of a component of the avatar, and rendering the avatar using the scene description. The rendering includes animating the component of the avatar according to the gaze data. Aspects disclosed in the present disclosure also describe a non-transitory computer-readable medium comprising instructions executable by at least one processor to perform methods for generating a scene description. The methods comprise generating node data representing an avatar, including generating gaze data specifying animation of a component of the avatar.[6] This Summary is provided to introduce a selection of concepts in a simplified form that is further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to limitations that solve any or all disadvantages noted in any part of this disclosure.BRIEF DESCRIPTION OF THE DRAWINGS[7] FIG. 1 is a block diagram of an example system, according to aspects of the present disclosure.[8] FIG. 2 is a diagram illustrating a horizontal (yaw) rotation, according to aspects of the present disclosure.[9] FIG. 3 is a diagram illustrating a vertical (pitch) rotation, according to aspects of the present disclosure.
[0010] FIG. 4 is a diagram illustrating a roll rotation, according to aspects of the present disclosure.
[0011] FIG. 5 illustrates a forward looking avatar and a respective coordinate system, according to aspects of the present disclosure.
[0012] FIG. 6 illustrates horizontal and vertical animations of an avatar, according to aspects of the present disclosure.
[0013] FIG. 7 is a flow diagram illustrating the parsing of a node data structure, according to aspects of the present disclosure.
[0014] FIG. 8 is a flow diagram illustrating the parsing of a gaze data structure, according to aspects of the present disclosure.
[0015] FIG. 9 is a flow diagram of an example method for rendering an avatar, according to aspects of the present disclosure.
[0016] FIG. 10 is a diagram of an example method for generating a scene description, according to aspects of the present disclosure.DETAILED DESCRIPTION
[0017] Apparatuses and methods are presented herein for rendering avatars and for generating scene descriptions that can represent avatars using avatar gaze descriptors. An avatar gaze descriptor, as disclosed herein, facilitates animation of a component of an avatar. A system for processing and displaying content, with which various aspects and examples described herein may be implemented, is generally described in reference to FIG. 1, followed by description of the aspects of the present disclosure in reference to FIGS. 2-10.
[0018] FIG. 1 illustrates a block diagram of an example system 100. System 100 can be embodied as a device including the various components described below and can be configured to perform one or more of the aspects described in this application. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 100, singly or in combination, can be embodied in a single integrated circuit, multiple integrated circuits, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple integrated circuits and / or discrete components. In various embodiments, the system 100 is communicatively coupled to other systems, or to other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports. In various embodiments,the system 100 is configured to implement one or more of the aspects described in this application.
[0019] The system 100 includes at least one processor 110 that can be configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application. Processor 110 can include embedded memory, input and output interfaces, and various other circuitries as known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which can include non-volatile memory and / or volatile memory, including, for example, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. The storage device 140 can be an internal storage device, an attached storage device, and / or a network accessible storage device, for example.
[0020] System 100 includes an encoder / decoder module 130 configured to process data to provide encoded video data or decoded video data. The encoder / decoder module 130 can include its own processor and memory. The encoder / decoder module 130 represents module(s) that can be included in a device to perform encoding and / or decoding functions. Additionally, the encoder / decoder module 130 can be implemented as a separate element of system 100 or can be incorporated within processor 110 as a combination of hardware and software as known to those skilled in the art.
[0021] Program code that is to be loaded into processor 110 or into encoder / decoder 130 to perform the various aspects described in this application can be stored in a storage device 140 and subsequently loaded into memory 120 for execution by processor 110. In accordance with various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 can store one or more of various items during the performance of the processes described in this application. Such stored items can include, but are not limited to, the input video, the decoded video or portions of the decoded video, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0022] In several embodiments, memory inside of the processor 110 and / or the encoder / decoder module 130 is used to store instructions and to provide working memory for processing functions that are needed during encoding or decoding. In other embodiments, however, memory external to the processing device (where, for example, the processing device can be either the processor 110 or the encoder / decoder module 130) can be used for one ormore of these functions. The external memory can be the memory 120 and / or the storage device 140 that may comprise, for example, a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for video coding and decoding operations.
[0023] The input to the elements of system 100 can be provided through various input devices as indicated in block 105. Such input devices include, but are not limited to, (i) an RF portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Composite input terminal (COMP), (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0024] In various embodiments, the input devices of block 105 have associated respective input processing elements as known in the art. For example, the RF portion can be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down-converting the selected signal, (iii) band-limiting again to a narrower band of frequencies to select, for example, a signal frequency band which can be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements that perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion can include a tuner that performs some of these functions, including, for example, down-converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to a baseband. In one set-top box embodiment, the RF portion and its associated input processing element receive an RF signal transmitted over a wired (for example, cable) medium, and perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Added elements can include inserting elements in between existing elements, for example, inserting amplifiers and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna.
[0025] Additionally, the USB and / or HDMI terminals can include respective interface processors for connecting system 100 to other electronic devices across USB and / or HDMIconnections. It is to be understood that various aspects of input processing, for example, Reed- Solomon error correction, can be implemented, for example, within a separate input processing integrated circuit or within processor 110 as necessary. Similarly, aspects of USB or HDMI interface processing can be implemented within separate interface integrated circuits or within processor 110 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110, and encoder / decoder 130 operating in combination with the memory and storage elements to process the datastream as necessary for presentation on an output device.
[0026] Various elements of system 100 can be provided within an integrated housing. Within the integrated housing, the various elements can be interconnected and transmit data therebetween using a suitable connection arrangement 115, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards.
[0027] The system 100 includes a communication interface 150 that enables communication with other devices via communication channel 190. The communication interface 150 can include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190. The communication interface 150 can include, but is not limited to, a modem or network card. The communication channel 190 can be implemented, for example, within a wired and / or a wireless medium.
[0028] Data are streamed to the system 100, in various embodiments, using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signal of these embodiments is received over the communication channel 190 and the communication interface 150 which are adapted for WiFi communications. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 100 using a set-top box that delivers the data over the HDMI connection of the input block 105. Still other embodiments provide streamed data to the system 100 using the RF connection of the input block 105.
[0029] The system 100 can provide an output signal to various output devices, including a display device 165, an audio device (e.g., speaker(s)) 175, and other peripheral devices 185. The other peripheral devices 185 include, in various examples of embodiments, one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide a function based on the output of the system 100. In various embodiments, controlsignals are communicated between the system 100 and the display device 165, the audio device 175, or other peripheral devices 185 using signaling such as AV. link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices can be connected to system 100 using the communication channel 190 via the communication interface 150. The display device 165 and the audio device 175 can be integrated in a single unit with the other components of system 100 in an electronic device, for example, a television. In various embodiments, the display interface 160 includes a display driver, for example, a timing controller (T Con) chip.
[0030] The display device 165 and the audio device 175 can alternatively be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments in which the display device 165 and the audio device 175 are external components, the output signal can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0031] Aspects of this disclosure introduce an avatar gaze descriptor that may be incorporated into an avatar node of a scene description. The scene description may be encoded in a format such as the glTF (see, https: / / registry.khronos.Org / glTF / specs / 2.0 / glTF-2.0.pdf). Using gaze data (recorded in the avatar gaze descriptor), a rendering application can animate an avatar’s gaze in three rotational directions as prescribed therein. Although aspects of this disclosure are described in the context of animating the eyes of a “human” avatar, these aspects are not limited to a “human” avatar nor to the eyes of an avatar. Hence, according to aspects, an avatar can be a representation of a human (e.g., a user of the rendering application) or of any other graphical objects. Furthermore, any component of an avatar can be animated according to the avatar gaze data described herein; that is, a component of an avatar may be any part of the avatar (such as the head, a limb, a bottom section, or a top section) or any object that extends from the avatar (such as a wand, a baseball bat, or a hockey stick).
[0032] In the current MPEG-I Scene Description (SD) format, a node extension is defined - namely, MPEG node avatar extension - that can be used to define an avatar object for rendering. Thanks to this extension, applications can determine which node in the SD represents an avatar mesh and render that avatar mesh at the scene as needed. For instance, if an avatar mesh represents a human body, an application can place it at a scene and let the user interact with it (e.g., changing the avatar’s appearance or its pose at the scene). Typically, anavatar includes components (e.g., eyes, head, hands, or other attached objects) that can be rendered facing or pointing at a certain direction. However, there is no information in the extention that specifies how to animate these components.
[0033] Generally, an animation descriptor (or an animation property), that can be used to direct the animation of a respective object, can be defined in a glTF file. Such an animation descriptor can be associated with the eyes of an avatar by naming - that is, one can name an animation descriptor to associate it with avatar’s eyes (e.g., "eyesHorizontalGaze"). In this manner, an application can recognize the animation descriptor and can apply it to animate the avatar’s eyes. However, there are no standards or conventions for such naming, and no specific information about how to apply the animation descriptor. For instance, the glTF file does not specify to the application within what rotational range an eye is to be animated. Any such information, if required, should be provided to the application from other sources.
[0034] According to aspects, an avatar gaze descriptor is provided. The avatar gaze descriptor can be encoded according to glTF or according to any other suitable formats (e.g., XML or USD). When the glFT is used, the avatar gaze descriptor can be incorporated into the MPEG_node_avatar extension. Aspects are described with respect to axes in a three- dimensional space that are oriented according to the right-hand rule, as is the case in the glTF. In the glTF, the default axes are Z for the forward direction, Y for the upward direction, and - X for the right direction. Other axes can be defined and can be used by the rendering application or by other glTF extensions, like those proposed in European Application No. 23306725.5, filed on October 06, 2023, which is incorporated herein by reference in its entirety.
[0035] Information recorded in the avatar gaze descriptor specifies the manner in which the avatar’s gaze should be animated; for example, the manner in which the eyes of the avatar should be rotated around up to three rotation axes. The rotation axes - namely, forward axis, bottom-up axis, and right axis - are defined relative to the avatar pose axes. The animated rotations, applied counterclockwise (or clockwise in other formats), can extend within a rotation range of a minimum angle and a maximum angle. For instance, using this rotation range, a “human” avatar’s gaze can be limited to forward viewing. In an aspect, the minimum angle can be between -180 and 0 degrees, and the maximum angle can be between 0 and 180 degrees. A horizontal (yaw) gaze, a vertical (pitch) gaze, and a roll gaze are described below with respect to FIGS. 2-4.
[0036] FIG. 2 is a diagram illustrating a horizontal (yaw) rotation 200. The rotation axis in thiscase extends from bottom upward (i.e., the bottom-up axis 210). A rotation around the bottom- up axis 210 determines the horizontal gaze 200 (e.g., changing it from right 260 to left 240). A rotational angle of zero 250 corresponds to a gaze direction that is aligned with the forward axis 230. A negative rotational angle corresponds to a gaze direction to the right and a positive rotational angle corresponds to a gaze direction to the left.
[0037] FIG. 3 is a diagram illustrating a vertical (pitch) rotation 300. The rotation axis in this case extends from left to right (i.e., the right axis 320). A rotation around the right axis 320 determines the vertical gaze 300 (e.g., changing it from downward 360 to upward 340). A rotational angle of zero 350 corresponds to a gaze direction that is aligned with the forward axis 330. A negative rotational angle corresponds to a gaze direction aiming downward and a positive rotational angle corresponds to a gaze direction aiming upward.
[0038] FIG. 4 is a diagram illustrating a roll rotation 400. The rotation axis in this case extends from backward to forward (i.e., the forward axis 430). A rotation around the forward axis 430 determines the roll gaze 400. In the example of FIG. 4, rotational angle of zero 450 corresponds to a gaze direction that is aligned with the right axis.
[0039] FIG. 5 illustrates a forward looking avatar and a respective coordinate system 500. In the example of FIG. 5, the avatar gaze 550 is aligned with the forward axis (i.e., the -Y axis 530). The bottom-up axis (i.e., the Z axis 510) and the right axis (i.e., the X axis 520) are also shown in FIG. 5. If the avatar’s pose can be derived from a canonical pose descriptor (e.g., a "canonicalPose" property defined in the glTF file), then the rotations 200, 300, 400 are relative to the derived pose. If the canonical pose descriptor includes a transform descriptor, then the avatar can be first transformed according to the transform descriptor, next the avatar eyes can be animated to rotate the eyes according to the avatar gaze descriptor, and then the pose of the animated avatar can be inverse transformed. If a canonical pose descriptor is not available (e.g., a "canonicalPose" property is not defined in the glTF file), then the pose of the avatar should be provided to the application from another source. The animation of the avatar’s gaze in horizontal and vertical directions are further discussed with respect to FIG. 6.
[0040] FIG. 6 illustrates horizontal and vertical animations of an avatar 600. FIG. 6 demonstrates animating the eyes 610, 620 according to a horizontal (yaw) gaze 200; that is, rotating the gaze direction around the Z axis (i.e., the bottom-up axis 210), see FIG. 2 and FIG. 5. When the angle of this rotation is zero 250, the gaze direction is aligned with the Y axis (i.e., the forward axis 230). When the rotation angle is negative 260, the avatar looks to the right610, and when the rotation angle is positive 240, the avatar looks to the left 620. FIG. 6 also demonstrates animating the eyes 630, 640 according to a vertical (pitch) gaze 300; that is, rotating the gaze direction around the X axis (i. e. , the right axis 320), see also FIG. 3 and FIG. 5. When the angle of this rotation is zero 350, the gaze direction is aligned with the Y axis (i. e. , the forward axis 330). When the rotation angle is negative 360, the avatar looks downward 630, and when the rotation angle is positive 340, the avatar looks upward 640.
[0041] According to aspects, an avatar gaze descriptor is added to a node data structure (such as the MPEG node avatar). The avatar gaze descriptor contains information that specifies the manner in which the avatar’s eyes (or other components of the avatar) can be animated by a rendering application. Table 1 presents a node data structure where each line describes a descriptor (or a property). In the first line, an avatar node descriptor, named “isAvatar,” having a Boolean data structure, indicates whether this node represents an avatar. In the second line, an avatar representation descriptor, named “type,” having a string data structure, indicates the type of the avatar representation, e.g., provided as uniform resource name (URN). In the third line, an avatar mapping descriptor, named “mappings,” having an array data structure of AvatarMapping, indicates the mapping between N child nodes and semantics. These three descriptors (or properties) already appear in the MPEG node avatar extension and are mandatory therein. A new descriptor is introduced herein to the node data structure; that is, the avatar gaze descriptor shown in the fourth line. The avatar gaze descriptor, named “gaze,” having a Gaze data structure (defined in Table 2), represents the avatar gaze. Note that the avatar gaze descriptor is not mandatory.
[0042] Table 1: Node Data Structure
[0043] Table 2 describes the data structure of the avatar gaze descriptor. In the first line, a gaze iotype descriptor, named "type", having a string data structure, indicates the type of the gaze. In this case, the gaze type is "angular," however, other gaze types can be defined. The gaze data structure also contains at least one of a horizontal rotation descriptor, named "horizontal" (or can be named as a "yaw"), a vertical rotation descriptor, named “vertical” (or can be named as a "pitch"), and a roll rotation descriptor, named “roll.” These descriptors, having a GazeAnimation data structure (defined in Table 3), indicate respective gaze animation properties. Note that only the gaze type descriptor is mandatory.
[0044] Table 2: Gaze Data Structure
[0045] Table 3 describes the GazeAnimation data structure - that is, the data structure of each of the horizontal rotation, the vertical rotation, and the roll rotation descriptors. The gaze animation data structure contains an animation reference descriptor, named “animation,” having an integer data structure, that references (by an index) an animation out of animation assets included in the scene description (e.g., glTF file). The gaze animation data structure also contains a rotation range descriptor, named “range,” having an array data structure of two elements, that indicates a range of angles extending between a minimum angle (stored in number|0|) and a maximum angle (stored in numberfl]). Thus, for example, an avatar’s eyes can be horizontally animated according to the horizontal rotation descriptor (see second line of Table 2) using information in the respective gaze animation data structure. That is, the eyes can be animated using the animation referenced by the animation reference descriptor (see first line of Table 3), rotating the eyes within the rotation range (see second line of Table 3).
[0046] Table 3: A Gaze Animation Data Structure
[0047] For example, a rotation angle 9 between a minimum angle Gmin(the first value of the rotation range descriptor) and a maximum angle Qmax(the second value of the rotation range descriptor) can be determined according to a linear interpolation. One can map the rotation angle 9 to the corresponding animation time using the following formula:where tmaxis the maximum time value of the animation and t is the time value corresponding the rotation angle 9.
[0048] Hence, the horizontal rotation descriptor specifies the manner in which the avatar’s gaze should be animated around the bottom-up axis 219, as shown in FIG. 2. This animation can use any valid transforms (e.g., rotation, translation, morph target weights, or extensions like KHR_extension_pointer) as long as the result corresponds to a gaze moving from right 269, 619 to left 249, 629. The vertical rotation descriptor specifies the manner in which the avatar’s gaze should be animated around right axis 329, as shown in FIG. 3. This animation too can use any valid transforms (e.g., rotation, translation, morph target weights, or extensions like KHR_extension_pointer) as long as the result corresponds to a gaze moving from downward 369, 639 to upward 369, 649. When both animations (the horizontal and the vertical) are to be applied, the horizontal rotation is animated first and then the vertical rotation is animated. In an aspect, a roll rotation can be further animated according to the roll rotation descriptor. Consequently, the two (or three) animations (pointed to by the respective animation reference descriptor (of Table 3) for each of the horizontal, vertical, and roll descriptors (of Table 2) should be designed to support this processing order.
[0049] The information provided by the avatar gaze descriptor (as described herein) does not change the scene as may be described in a glTF file. Moreover, using this information to animate a respective avatar is optional and is in the control of the application. It is similar to animation assets: they may be provided in a glTF file, but the application determines whether to use them or not. The manner in which information provided by the avatar gaze descriptor is used can vary according to the requirements of each use case, as further described below.
[0050] When an avatar is rendered by a virtual reality (or a mixed-reality) rendering application, an avatar gaze descriptor associated with the avatar provides gaze data to the renderingapplication that can be applied to the animation of a component of the avatar. The gaze data can be determined so that the animation of a component of the avatar aligns the component with a direction of a respective component of a user associated with the avatar. For example, based on the gaze data, an application can animate the avatar eyes’ gaze to align with a viewing direction of a user associated with the avatar. In an aspect, the viewing direction of the user can be the direction of the user’s camera (e.g., a handheld camera or a camera attached to the user’s body). In another aspect, the viewing direction of the user can be the direction of a head mounted display. In yet another aspect, the viewing direction of the user can be detected (e.g., tracked by another camera or indicated by the user with a pointing device such as a mouse or a gamepad). Thus, based on the gaze data, the avatar’s eyes can be animated to follow a direction that corresponds to a user’s gaze. Alternatively, the gaze data can be used to animate the head of the avatar so that it follows a direction that corresponds to a user’s gaze and / or head directions. As a result, other users of the virtual reality rendering application can see the avatar’s eyes or the avatar’s head oriented to a direction associated with the avatar’s owner or with other events.
[0051] Avatars may also be rendered by a video conference application. A video conference application may be set to replace the image of a conference participant, captured at a respective client end, with an avatar that is rendered at another end (of another conference participant). Gaze data, provided by an avatar gaze descriptor associated with that avatar, can be used by the video conference application to animate the gaze of the avatar. Based on the gaze data, the avatar gaze can be oriented forward, and, possibly, small eye movements can be added to provide a realistic presentation. Furthermore, the gaze data may reflect the participant’s eye movements (captured at the client end) and can be used to accordingly animate the avatar eyes’ gaze when rendered at the other participant’s end.
[0052] Aspects described herein can be used by a 3D asset creation application. A 3D asset creation application may produce a scene description (e.g., encoded in glTF), including avatars and respective avatar gaze descriptors. An avatar gaze descriptor allows the 3D asset creation application to prescribe the manner in which a component (e.g., eyes) of a respective avatar should be animated by a rendering application to achieve a certain effect (e.g., aligning an avatar’s gaze with a direction associated with the avatar’s owner or with other events). According to aspects, an avatar gaze descriptor can be used to extend an MPEG node avatar encoded in a glTF file. A 3D asset creation tool supporting such an extension can detect the avatar gaze descriptor and, through a dedicated graphic user interface (GUI), for example, acontent creator can set the avatar gaze descriptor with gaze data that specifies the animation of a component of a respective avatar.
[0053] The encoding of gaze data, using the data structures presented in Tables 1 -3, into a glTF file is illustrated below. For clarity, only sections of the glTF file that are related to an avatar gaze are shown. It is assumed that the glTF file contains data representing a scene including an avatar, such as the one 500 shown in FIG. 5.
[0054] A glTF parser first loads all content except for the content of the extensions (i.e., the above content of the structure named “extensions”). Then, it loads the content within the first node structure, named "MPEG node avatar." Information recorded within this first node structure indicates: 1) the node represents an avatar ("isAvatar" is true); 2) the avatar representation ("type" is the string "um:mpeg:sd:2023:avatar"); 3) the avatar mappings include mappings associated with the head, the left eye, and the right eye of the avatar (see respective “path” and “node” within the "mappings" structure); 4) the avatar gaze is angular ("type": "angular"); 5) the horizontal rotation is to be animated by an animation indexed by index value 0, with a minimum angle of -45 degrees and with a maximum angle of 45 degrees ( "horizontal" : { "animation": 0, "range": [-45.0, 45.0]}); and 6) the vertical rotation is to be animated by an animation indexed by index value 1 , with a minimum angle of -30 degrees and with a maximum angle of 30 degrees ("vertical": {"animation": 1, "range": [-30.0, 30.0]}). The parsing process is further demonstrated with respect to FIG. 7 and FIG. 8, where the parsing of a node data structure (Table 1) and the parsing of a gaze data structure (Tables 2 and 3) are, respectively, demonstrated.
[0055] FIG. 7 is a flow diagram illustrating the parsing of a node data structure 700. Generally, the parsing of a node, in step 710, can be triggered each time the node is being loaded by a glTF loader employed by a rendering application. Accordingly, in step 720, a glTF loader first decodes standard properties and content recorded within extension structures, such as the MPEG node avatar extension node. If the MPEG node avatar extension node is not present 730, the parsing ends 790. Otherwise parsing proceeds to steps 740-760 to parse the node descriptor (see Table 1) - that is, to decode the avatar node property (named "isAvatar"), the avatar representation property (named "type"), and the avatar mapping property (named "mappings"). Since these properties 740-760 are mandatory, if these properties are not present in the glTF then an error is returned. Next, if the avatar gaze property (named "gaze") is not present 770, the parsing ends 790. Otherwise, in step 780, the avatar gaze property is parsed as illustrated in FIG. 8.
[0056] FIG. 8 is a flow diagram illustrating the parsing of a gaze data structure 800. The parsing, in step 810, is triggered each time the glTF loader finds an avatar gaze property in anMPEG_node_avatar extension. If the gaze type property (named "type") is not "angular" 820, an error is returned in step 830. If the gaze type property is "angular" 820, it is next tested whether a horizontal rotation property (named “horizontal”) 840, a vertical rotation property (named “vertical”) 850, or a roll rotation property (named “roll”) 860 are present; if not present, respective parsing ends 870. If the horizontal rotation property is present 840, then, in step 845, it is parsed. If the vertical rotation property is present 850, then, in step 855, it is parsed. And, if the roll rotation property is present 860, then, in step 865, it is parsed. As shown in the example of FIG. 8, the parsing in steps 845, 855, and 865 includes the parsing of the respective animation reference property (named “animation") and the respective animation range property (named "range").
[0057] The gaze data decoded from the MPEG node avatar extension (as illustrated by FIGS. 7-8) can be stored in a dedicated structure and can be applied to animate the gaze of the respective avatar as needed by the application. In an aspect, the application can modify the gaze data to align it with a direction associated with a user that owns the avatar or with other events.
[0058] FIG. 9 is a flow diagram of an example method for rendering an avatar 900. The method 900 may begin, in step 910, by obtaining a scene description (e.g., encoded in aglTF) including node data (e.g., stored in a node data structure as shown in Table 1) that represent an avatar. Next, in step 920, gaze data (e.g., stored in a gaze data structure as shown in Table 2) can be extracted from the node data. According to aspects, the extracted node data specify animation of a component of the avatar. Then, in step 930, the avatar can be rendered using the scene description, where the rendering can include animating the component of the avatar according to the gaze data. The gaze data may provide at least one of a horizontal rotation, a vertical rotation, and a roll rotation that can be used to animate the component of the avatar as described herein. For each rotation, the gaze data may provide respective data (e.g., stored in a gaze animation data structure as shown in Table 3) that indicate a reference to animation data provided by the scene description and a rotation range between a minimum angle and a maximum angle. Accordingly, the component of the avatar can be animated using the referenced animation data within the rotation range.
[0059] In an aspect, the component of the avatar can be an object associated with the avatar, such as a wand, a baseball bat, or a hockey stick. In another aspect, the component of the avatar can be a part of the avatar, such as eyes, the head, and / or limb(s). In a further aspect, the rendering of the avatar, in step 930, is according to a canonical pose provided by the scenedescription, where the component of the avatar is animated relative to the canonical pose.
[0060] FIG. 10 is a diagram of an example method for generating a scene description 1000. The method 1000 can be applied to generate 1010 node data (e.g., organized in a node data structure as shown in Table 1) representing an avatar. The method 1000 can be further applied to extend 1020 the node data by gaze data (e.g., organized in a gaze data structure as shown in Table 2). The generated gaze data specify the animation of a component of the avatar. At least one of a horizontal rotation, a vertical rotation, and a roll rotation may be provided in the gaze data. Accordingly, the animation specified by the gaze data can be applied to the component of the avatar using at least one of these rotations. For each rotation, the gaze data may provide respective data (e.g., organized in a gaze animation data structure as shown in Table 3) that indicate a reference to animation data provided by the scene description and a rotation range between a minimum angle and a maximum angle. Accordingly, the animation specified by the gaze data can be applied to the component of the avatar using the animation data within the rotation range.
[0061] According to aspects, the method 1000 can be applied to encoding the scene description according to glTF, including encoding a descriptor representing the gaze data. To that end, the encoding of the descriptor representing the gaze data can include encoding a descriptor indicating a gaze type and encoding a descriptor indicating a gaze animation for each of the horizontal, vertical, and roll rotations. Furthermore, the encoding of the descriptor indicating the gaze animation can include encoding a descriptor indicating the reference to the animation data and encoding a descriptor indicating the rotation range.
[0062] In an aspect, the component of the avatar can be animated (as specified by the gaze data) so that the component is aligned with a viewing direction of a user associated with the avatar. In another aspect, the generated gaze data 1020 can be generated by a GUI (e.g., of a 3D asset creation application) used by a user to determine the animation of the component of the avatar.
[0063] According to aspects, gaze can also be signaled in the avatar JSON interchange file (AJIF) format, as described in EP Application 24305094.5, filed on January 15, 2024, which is incorporated herein by reference in its entirety.
[0064] AJIF format allows the encoding of the same avatar at different levels of details (LOD). For each LOD, a different set of geometries, controllers, and other properties can be defined. It can also be the case for the gaze, with the introduction of a new “gaze” property in the LOD property of the AJIF, as shown in Table 4.
[0065] Table 3: AJIF LOD properties, including the “gaze” property.
[0066] The "gaze" property can be defined as shown in Table 5.
[0067] Table 4: The Gaze property
[0068] The "type" property defines the type of the gaze. Currently, the only possible value is "angular", but other extensions can propose new types. In some embodiments, the "horizontal" property can be replaced by a "yaw" property with the same specificities. In some embodiments, the "vertical" property can be replaced by a "pitch" property with the same specificities. The "horizontal" (or "yaw"), "vertical" (or "pitch") and "roll" properties can be defined as shown in Table 6.
[0069] Table 5: The GazeController property
[0070] Note that the GazeController property in the AJIF format is similar to GazeAnimation in the glTF format, except that “animation” is replaced by “controller”. This is because in the AJIF format, there is no animations but controllers instead. Both the glTF and the AJIF formats can be used to animate an avatar based on time and weight information, respectively.
[0071] The "controller" property references one of the controllers in the AJIF format. The lowest weight of the controller, denoted wmin. corresponds to the first angle of the "range" property, and the highest weight of the controller, denoted wmax. corresponds to the second angle of the "range" property. Angles in between the first and the second angles can be (linearly) interpolated. The rotation angle can be converted to the corresponding time using the following formula:where 9 is the rotation angle to be converted, 9minis the minimum rotation angle (first value of the "range" property), 9maxis the maximum rotation angle (second value of the "range" property), wmaxis the maximum weight of the controller, wminis the minimum weight of the controller, and w is the weight corresponding the angle 9.
[0072] The controller referenced by the "horizontal" property in the "gaze" property animates the avatar gaze along the horizontal axis for the current LOD. This controller can use any valid transforms (e.g., rotation, translation, or blendshape weights) as long as the result corresponds to a gaze moving from right to left. The controller referenced by the "vertical" property in the "gaze" property animates the avatar gaze along the vertical axis for the current LOD. This controller can use any valid transforms (e.g., rotation, translation, or blendshape weights) as long as the result corresponds to a gaze moving from bottom to top. The controller referenced by the "roll" property in the "gaze" property animates the avatar gaze along the forward axis for the current LOD. This controller can use any valid transforms (e.g., rotation, translation, or blendshape weights) as long as the result corresponds to a gaze moving around the forward axis. When both angles (horizontal and vertical) are not zero, the horizontal angle can be first applied, and then the vertical angle can be applied. As a result, the two corresponding controllers should be designed to support this processing order.
[0073] The illustrations of the aspects described herein are intended to provide a general understanding of the structure, function, and operation of the various aspects. The illustrations are not intended to serve as a complete description of all of the elements and features of apparatuses and systems that utilize the structures or methods described herein. Many other aspects may be apparent to those of skill in the art upon reviewing the disclosure. Other aspects may be utilized and derived from the disclosure, such that structural and logical substitutions and changes may be made without departing from the scope of the disclosure. Accordingly, the disclosure and the figures are to be regarded as illustrative rather than restrictive.
[0074] The description of the aspects is provided to enable the making or use of the aspects. Various modifications to these aspects will be readily apparent, and the generic principles defined herein may be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope possible consistent with the principles and novel features as defined by the following claims.
Claims
What is claimed is:
1. A method for rendering an avatar, comprising: obtaining a scene description including node data representing an avatar; extracting gaze data from the node data, wherein the gaze data specify animation of a component of the avatar; and rendering the avatar using the scene description, including animating the component of the avatar according to the gaze data.
2. The method according to claim 1, wherein at least one of a horizontal rotation, a vertical rotation, and a roll rotation is provided by the gaze data, and wherein the animating of the component of the avatar is according to the at least one of the horizontal rotation, the vertical rotation, and the roll rotation.
3. The method according to claim 2, wherein, for each of the at least one of the horizontal, vertical, and roll rotations, the gaze data indicate: a reference to animation data provided by the scene description, and a rotation range between a minimum angle and a maximum angle, wherein the animating is applied to the component of the avatar using the referenced animation data within the rotation range.
4. The method according to any one of claims 1 to 3, wherein the component of the avatar comprises an object associated with the avatar.
5. The method according to any one of claims 1 to 3, wherein the component of the avatar comprises a part of the avatar.
6. The method according to any one of claims 1 to 5, wherein the rendering of the avatar further comprising: rendering the avatar according to a canonical pose provided by the scene description, wherein the animating of the component of the avatar is relative to the canonical pose.
7. A method for generating a scene description, comprising: generating node data representing an avatar, including:generating gaze data specifying animation of a component of the avatar.
8. The method according to claim 7, wherein at least one of a horizontal rotation, a vertical rotation, and a roll rotation is provided in the gaze data, and wherein the animation specified by the gaze data is to be applied to the component of the avatar using at least one of the horizontal rotation, the vertical rotation, and the roll rotation.
9. The method according to claim 8, wherein, for each of the at least one of the horizontal, vertical, and roll rotations, the gaze data indicate: a reference to animation data provided by the scene description, and a rotation range between a minimum angle and a maximum angle, wherein the animation specified by the gaze data is to be applied to the component of the avatar using the referenced animation data within the rotation range.
10. The method according to claim 9, further comprising: encoding the scene description according to a graphics library transmission format (glTF), including encoding a descriptor representing the gaze data.
11. The method according to claim 10, wherein the encoding of the descriptor representing the gaze data comprises: encoding a descriptor indicating a gaze type, and encoding a descriptor indicating a gaze animation for each of the at least horizontal, vertical, and roll rotations.
12. The method according to claim 11, wherein the encoding of the descriptor indicating the gaze animation comprises: encoding a descriptor indicating the reference to the animation data, and encoding a descriptor indicating the rotation range.
13. The method according to any one of claims 7 to 12, wherein the specified animation of the component of the avatar aligns the component of the avatar with a viewing direction of a user associated with the avatar.
14. The method according to any one of claims 7 to 11, wherein the gaze data is generated by a graphic user interface used by a user to determine the animation of the component of the avatar.
15. An apparatus for rendering an avatar, comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the apparatus to: obtain a scene description including node data representing an avatar, extract gaze data from the node data, wherein the gaze data specify animation of a component of the avatar, and render the avatar using the scene description, including animating the component of the avatar according to the gaze data.
16. The apparatus according to claim 15, wherein at least one of a horizontal rotation, a vertical rotation, and a roll rotation is provided by the gaze data, and wherein the animating of the component of the avatar is according to the at least one of the horizontal rotation, the vertical rotation, and the roll rotation.
17. The apparatus according to claim 16, wherein, for each of the at least one of the horizontal, vertical and roll rotations, the gaze data indicate: a reference to animation data provided by the scene description, and a rotation range between a minimum angle and a maximum angle, wherein the animating is applied to the component of the avatar using the referenced animation data within the rotation range.
18. The apparatus according to any one of claims 15 to 17, wherein the component of the avatar comprises at least one of an object associated with the avatar and a part of the avatar.
19. An apparatus for generating a scene description, comprising: at least one processor; andmemory storing instructions that, when executed by the at least one processor, cause the apparatus to: generate node data representing an avatar, including: generating gaze data specifying animation of a component of the avatar.
20. The apparatus according to claim 19, wherein at least one of a horizontal rotation, a vertical rotation, and a roll rotation is provided in the gaze data, and wherein the animation specified by the gaze data is to be applied to the component of the avatar using at least one of the horizontal rotation, the vertical rotation, and the roll rotation.
21. The apparatus according to claim 20, wherein, for each of the at least one of the horizontal, vertical, and roll rotations, the gaze data indicate: a reference to animation data provided by the scene description, and a rotation range between a minimum angle and a maximum angle, wherein the animation specified by the gaze data is to be applied to the component of the avatar using the referenced animation data within the rotation range.
22. The apparatus according to claim 21, wherein the instructions further cause the apparatus to: encode the scene description according to a graphics library transmission format (glTF), including encoding a descriptor representing the gaze data.
23. The apparatus according to claim 22, wherein the encoding of the descriptor representing the gaze data comprises: encoding a descriptor indicating a gaze type, and encoding a descriptor indicating a gaze animation for each of the at least horizontal, vertical, and roll rotations.
24. The apparatus according to claim 23, wherein the encoding of the descriptor indicating the gaze animation comprises: encoding a descriptor indicating the reference to the animation data, and encoding a descriptor indicating the rotation range.
25. The apparatus according to any one of claims 19 to 24, wherein the specified animation of the component of the avatar aligns the component of the avatar with a viewing direction of a user associated with the avatar.
26. The apparatus according to any one of claims 19 to 25, wherein the gaze data is generated by a graphic user interface used by a user to determine the animation of the component of the avatar.
27. A non-transitory computer-readable medium comprising instructions executable by at least one processor to perform a method for rendering an avatar, the method comprising: obtaining a scene description including node data representing an avatar; extracting gaze data from the node data, wherein the gaze data specify animation of a component of the avatar; and rendering the avatar using the scene description, including animating the component of the avatar according to the gaze data.
28. A non-transitory computer-readable medium comprising instructions executable by at least one processor to perform a method for generating a scene description, the method comprising: generating node data representing an avatar, including: generating gaze data specifying animation of a component of the avatar.
Citation Information
Patent Citations
EP24305094A
EP23306725A