Avatar format validation in scene descriptions
A validation procedure for 3D avatar models in immersive environments checks against a set of definitions to ensure compatibility, addressing the inefficiencies of current formats by validating candidate avatars against reference formats, thus enhancing compatibility and flexibility.
Patent Information
- Application Number
- PCT/EP2025/050214
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2025-01-07
- Publication Date
- 2025-07-17
AI Technical Summary
Current scene description formats do not provide an efficient way to validate 3D avatar models against a collection of properties, requiring prior knowledge of all properties to instantiate an avatar model correctly and limiting compatibility with different avatar models.
A validation procedure that checks a candidate avatar format against a reference avatar format by comparing it to a set of definitions, ensuring the candidate avatar satisfies properties such as mesh topology, semantic topology, deformation, skeleton, material, and interactivity, allowing validation against a subset of definitions.
Ensures that candidate avatars meet specific properties, enabling applications to handle various avatar models by validating them against supported reference avatars, enhancing compatibility and flexibility in immersive environments.
Smart Images

Figure EP2025050214_17072025_PF_FP_ABST
Abstract
Description
AVATAR FORMAT VALIDATION IN SCENE DESCRIPTIONSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims benefit of European Patent Application No. EP24305073, entitled "AVATAR FORMAT VALIDATION IN SCENE DESCRIPTIONS” and filed January 11 , 2024, which is hereby incorporated by reference in its entirety.BACKGROUND
[0002] This application applies to 3D scene and user representation of interactions within immersive environments. Representations of users such as avatars are in common usage in 3D immersive environments. MPEG-I Scene Description (SD) (23090-14) provides an interactivity framework in support of use of user representations (e.g., avatars) in these environments, and in extended reality (XR) such as virtual reality (VR), augmented reality (AR), and / or mixed reality (MR).SUMMARY
[0003] A first example method in accordance with some embodiments may include: obtaining candidate avatar format data for an avatar, obtaining reference avatar format data for the avatar, wherein the reference avatar format data includes a first set of one or more definitions; performing a validation process for each of the one or more definitions of the first set, wherein the validation process includes: retrieving a current definition from the first set of one or more definitions; generating a comparison result by determining if the candidate avatar format data satisfies the current definition; and exiting the validation process if the comparison result is false; and indicating the candidate avatar format is valid if the comparison result is true for each pass through the validation process.
[0004] For some embodiments of the first example method, the first set includes at least one definition regarding a mesh topology of the avatar.
[0005] For some embodiments of the first example method, the first set includes at least one definition regarding a semantic topology of the avatar.
[0006] For some embodiments of the first example method, the first set includes at least one definition regarding a deformation of the avatar.
[0007] For some embodiments of the first example method, the first set includes at least one definition regarding a skeleton of the avatar.
[0008] For some embodiments of the first example method, the first set includes at least one definition regarding a material of the avatar.
[0009] For some embodiments of the first example method, the first set includes at least one definition regarding interactivity of the avatar.
[0010] For some embodiments of the first example method, the first set includes at least one definition regarding a first sub-format and at least one definition regarding a second sub-format.
[0011] Some embodiments of the first example method may further include: obtaining a second reference avatar format data for an avatar, wherein the second reference avatar format data includes a second set of one or more definitions; repeating the validation process using the second reference avatar format data, and indicating the candidate avatar format is valid only if the candidate avatar format is indicated as valid using at least one of the reference avatar format data or the second reference avatar format data.
[0012] For some embodiments of the first example method, at least one of the reference avatar format data, the second reference avatar format data, and the candidate avatar format data satisfies MPEG_node_avatar extension criteria.
[0013] A first example apparatus in accordance with some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform any one of the methods listed above.
[0014] An example method of validating a candidate avatar format against a reference avatar format in accordance with some embodiments may include: comparing one or more features of the candidate avatar format with one or more definitions of the reference avatar format; and validating the candidate avatar format if the one or more features of the candidate avatar format satisfy the one or more definitions.
[0015] A third example method in accordance with some embodiments may include: performing a validation process for each of one or more definitions of a first set within reference avatar format data, wherein the reference avatar format data includes the first set of one or more definitions, and wherein the validation process includes: retrieving a current definition from the first set of one or more definitions; determining a comparison result by comparing the current definition with a set of candidate avatar format data; and exiting the validation process if the comparison result is false; and indicating the candidate avatar format is valid if the comparison result is true for each pass through the validation process.
[0016] A third example apparatus in accordance with some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform any one of the methods listed above.
[0017] A fourth example method in accordance with some embodiments may include: obtaining reference avatar format data for an avatar, wherein the reference avatar format data includes a first set of one or more definitions; obtaining candidate avatar format data for the avatar; performing a validation process for each of the one or more definitions of the first set, wherein the validation process includes: retrieving a current definition from the first set of one or more definitions; determining a comparison result by comparing the current definition with the candidate avatar format data; and exiting the validation process if the comparison result is false; and indicating the candidate avatar format is valid if the comparison result is true for each pass through the validation process.
[0018] A fourth example apparatus in accordance with some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform any one of the methods listed above.
[0019] A fifth example method in accordance with some embodiments may include: obtaining reference avatar format data for an avatar, wherein the reference avatar format data includes a first set of one or more definitions; obtaining candidate avatar format data for the avatar; performing a validation process for each of the one or more definitions of the first set, wherein the validation process includes: retrieving a current definition from the first set of one or more definitions; determining a comparison result by comparing the current definition with the candidate avatar format data; indicating the candidate avatar format is invalid if the comparison result is false; exiting the validation process if the comparison result is false; returning to a beginning of the validation process if the comparison result is true and more definitions from the first set are left to be retrieved; and indicating the candidate avatar format is valid if the comparison result is true and all of the definitions from the first set have been retrieved.
[0020] A fifth example apparatus in accordance with some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform any one of the methods listed above.
[0021] A sixth example apparatus in accordance with some embodiments may include at least one processor configured to perform any one of the methods listed above.
[0022] A seventh example apparatus in accordance with some embodiments may include a computer- readable medium storing instructions for causing one or more processors to perform any one of the methods listed above.
[0023] An eighth example apparatus in accordance with some embodiments may include at least one processor and at least one non-transitory computer-readable medium storing instructions for causing the at least one processor to perform any one of the methods listed above.
[0024] An example bitstream in accordance with some embodiments may include a bitsream generated according to any one of the methods listed above.BRIEF DESCRIPTION OF THE DRAWINGS
[0025] FIG. 1 A is a schematic side view illustrating an example waveguide display that may be used with extended reality (XR) applications according to some embodiments.
[0026] FIG. 1 B is a schematic side view illustrating an example alternative display type that may be used with extended reality applications according to some embodiments.
[0027] FIG. 1 C is a schematic side view illustrating an example alternative display type that may be used with extended reality applications according to some embodiments.
[0028] FIG. 1 D is a system diagram illustrating an example set of interfaces for a system according to some embodiments.
[0029] FIG. 2 is a schematic illustration showing an example MORGAN skeleton model according to some embodiments.
[0030] FIG. 3 is a flowchart illustrating an example avatar format validation process according to some embodiments.
[0031] FIG. 4 is a flowchart illustrating an example avatar format validation process according to some embodiments.
[0032] The entities, connections, arrangements, and the like that are depicted in— and described in connection with— the various figures are presented by way of example and not by way of limitation. As such, any and all statements or other indications as to what a particular figure "depicts,” what a particular element or entity in a particular figure "is” or "has,” and any and all similar statements— that may in isolation and out of context be read as absolute and therefore limiting— may only properly be read as being constructively preceded by a clause such as "In at least one embodiment, ... " For brevity and clarity of presentation, this implied leading clause is not repeated ad nauseum in the detailed description.DETAILED DESCRIPTION
[0033] FIG. 1 A is a schematic side view illustrating an example waveguide display that may be used with extended reality (XR) applications according to some embodiments. An image is projected by an image generator 102. The image generator 102 may use one or more of various techniques for projecting an image. For example, the image generator 102 may be a laser beam scanning (LBS) projector, a liquid crystal display (LCD), a light-emitting diode (LED) display (including an organic LED (OLED) or micro LED (pi LED) display), a digital light processor (DLP), a liquid crystal on silicon (LCoS) display, or other type of image generator or light engine.
[0034] Light representing an image 112 generated by the image generator 102 is coupled into a waveguide 104 by a diffractive in-coupler 106. The in-coupler 106 diffracts the light representing the image 112 into one or more diffractive orders. For example, light ray 108, which is one of the light rays representing a portion of the bottom of the image, is diffracted by the in-coupler 106, and one of the diffracted orders 110 (e.g. the second order) is at an angle that is capable of being propagated through the waveguide 104 by total internal reflection. The image generator 102 displays images as directed by a control module 124, which operates to render image data, video data, point cloud data, or other displayable data.
[0035] At least a portion of the light 110 that has been coupled into the waveguide 104 by the diffractive in-coupler 106 is coupled out of the waveguide by a diffractive out-coupler 114. At least some of the light coupled out of the waveguide 104 replicates the incident angle of light coupled into the waveguide. For example, in the illustration, out-coupled light rays 116a, 116b, and 116c replicate the angle of the in-coupled light ray 108. Because light exiting the out-coupler replicates the directions of light that entered the in-coupler, the waveguide substantially replicates the original image 112. A user's eye 118 can focus on the replicated image.
[0036] In the example of FIG. 1A, the out-coupler 114 out-couples only a portion of the light with each reflection allowing a single input beam (such as beam 108) to generate multiple parallel output beams (such as beams 116a, 116b, and 116c). In this way, at least some of the light originating from each portion of the image is likely to reach the user's eye even if the eye is not perfectly aligned with the center of the out- coupler. For example, if the eye 118 were to move downward, beam 116c may enter the eye even if beams 116a and 116b do not, so the user can still perceive the bottom of the image 112 despite the shift in position. The out-coupler 114 thus operates in part as an exit pupil expander in the vertical direction. The waveguide may also include one or more additional exit pupil expanders (not shown in FIG. 1 A) to expand the exit pupil in the horizontal direction.
[0037] In some embodiments, the waveguide 104 is at least partly transparent with respect to light originating outside the waveguide display. For example, at least some of the light 120 from real-world objects (such as object 122) traverses the waveguide 104, allowing the user to see the real-world objects while using the waveguide display. As light 120 from real-world objects also goes through the diffraction grating 114, there will be multiple diffraction orders and hence multiple images. To minimize the visibility of multiple images, it is desirable for the diffraction order zero (no deviation by 114) to have a great diffraction efficiency for light 120 and order zero, while higher diffraction orders are lower in energy. Thus, in addition to expanding and out-coupling the virtual image, the out-coupler 114 is preferably configured to let through the zero order of the real image. In such embodiments, images displayed by the waveguide display may appear to be superimposed on the real world.
[0038] FIG. 1 B is a schematic side view illustrating an example alternative display type that may be used with extended reality applications according to some embodiments. In an XR head-mounted display device 130, a control module 132 controls a display 134, which may be an LCD, to display an image. The headmounted display includes a partly-reflective surface 136 that reflects (and in some embodiments, both reflects and focuses) the image displayed on the LCD to make the image visible to the user. The partly-reflective surface 136 also allows the passage of at least some exterior light, permitting the user to see their surroundings.
[0039] FIG. 1 C is a schematic side view illustrating an example alternative display type that may be used with extended reality applications according to some embodiments. In an XR head-mounted display device 140, a control module 142 controls a display 144, which may be an LCD, to display an image. The image is focused by one or more lenses of display optics 146 to make the image visible to the user. In the example of FIG. 1 C, exterior light does not reach the user's eyes directly. However, in some such embodiments, an exterior camera 148 may be used to capture images of the exterior environment and display such images on the display 144 together with any virtual content that may also be displayed.
[0040] The embodiments described herein are not limited to any particular type or structure of XR display device.
[0041] FIG. 1 D is a system diagram illustrating an example set of interfaces for a system according to some embodiments. An extended reality display device, together with its control electronics, may be implemented using a system such as the system of FIG. 1 D. System 150 can be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this document. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimediaset top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 150, singly or in combination, can be embodied in a single integrated circuit ( I C) , multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 150 are distributed across multiple ICs and / or discrete components. In various embodiments, the system 150 is communicatively coupled to one or more other systems, or other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports. In various embodiments, the system 1000 is configured to implement one or more of the aspects described in this document.
[0042] The system 150 includes at least one processor 152 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this document. Processor 152 may include embedded memory, input output interface, and various other circuitries as known in the art. The system 150 includes at least one memory 154 (e.g., a volatile memory device, and / or a non-volatile memory device). System 150 may include a storage device 158, which can include non-volatile memory and / or volatile memory, including, but not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), flash, magnetic disk drive, and / or optical disk drive. The storage device 158 can include an internal storage device, an attached storage device (including detachable and non-detachable storage devices), and / or a network accessible storage device, as non-limiting examples.
[0043] System 150 includes an encoder / decoder module 156 configured, for example, to process data to provide an encoded video or decoded video, and the encoder / decoder module 156 can include its own processor and memory. The encoder / decoder module 156 represents module(s) that can be included in a device to perform the encoding and / or decoding functions. As is known, a device can include one or both of the encoding and decoding modules. Additionally, encoder / decoder module 156 can be implemented as a separate element of system 150 or can be incorporated within processor 152 as a combination of hardware and software as known to those skilled in the art.
[0044] Program code to be loaded onto processor 152 or encoder / decoder 156 to perform the various aspects described in this document can be stored in storage device 158 and subsequently loaded onto memory 154 for execution by processor 152. In accordance with various embodiments, one or more of processor 152, memory 154, storage device 158, and encoder / decoder module 156 can store one or more of various items during the performance of the processes described in this document. Such stored items can include, but are not limited to, the input video, the decoded video or portions of the decoded video, thebitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0045] In some embodiments, memory inside of the processor 152 and / or the encoder / decoder module 156 is used to store instructions and to provide working memory for processing that is needed during encoding or decoding. In other embodiments, however, a memory external to the processing device (for example, the processing device can be either the processor 152 or the encoder / decoder module 152) is used for one or more of these functions. The external memory can be the memory 154 and / or the storage device 158, for example, a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of, for example, a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for video coding and decoding operations, such as for MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also referred to as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or WC (Versatile Video Coding, a new standard being developed by JVET, the Joint Video Experts Team).
[0046] The input to the elements of system 150 can be provided through various input devices as indicated in block 172. Such input devices include, but are not limited to, (i) a radio frequency (RF) portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Component (COMP) input terminal (or a set of COMP input terminals), (ill) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples, not shown in FIG. 1 C, include composite video.
[0047] In various embodiments, the input devices of block 172 have associated respective input processing elements as known in the art. For example, the RF portion can be associated with elements suitable for (I) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) downconverting the selected signal, (ill) band-limiting again to a narrower band of frequencies to select (for example) a signal frequency band which can be referred to as a channel in certain embodiments, (iv) demodulating the downconverted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion can include a tuner that performs various of these functions, including, for example, downconverting the received signal to a lower frequency (for example, an intermediatefrequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Adding elements can include inserting elements in between existing elements, such as, for example, inserting amplifiers and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna.
[0048] Additionally, the USB and / or HDMI terminals can include respective interface processors for connecting system 150 to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed-Solomon error correction, can be implemented, for example, within a separate input processing IC or within processor 152 as necessary. Similarly, aspects of USB or HDMI interface processing can be implemented within separate interface ICs or within processor 152 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 152, and encoder / decoder 156 operating in combination with the memory and storage elements to process the datastream as necessary for presentation on an output device.
[0049] Various elements of system 150 can be provided within an integrated housing, Within the integrated housing, the various elements can be interconnected and transmit data therebetween using suitable connection arrangement 174, for example, an internal bus as known in the art, including the Inter- IC (I2C) bus, wiring, and printed circuit boards.
[0050] The system 150 includes communication interface 160 that enables communication with other devices via communication channel 162. The communication interface 160 can include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 162. The communication interface 160 can include, but is not limited to, a modem or network card and the communication channel 162 can be implemented, for example, within a wired and / or a wireless medium.
[0051] Data is streamed, or otherwise provided, to the system 150, in various embodiments, using a wireless network such as a Wi-Fi network, for example IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal of these embodiments is received over the communications channel 162 and the communications interface 160 which are adapted for Wi-Fi communications. The communications channel 162 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 150 using a set-topbox that delivers the data over the HDMI connection of the input block 172. Still other embodiments provide streamed data to the system 150 using the RF connection of the input block 172. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, for example a cellular network or a Bluetooth network.
[0052] The system 150 can provide an output signal to various output devices, including a display 176, speakers 178, and other peripheral devices 180. The display 176 of various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 176 can be for a television, a tablet, a laptop, a cell phone (mobile phone), or other device. The display 176 can also be integrated with other components (for example, as in a smart phone), or separate (for example, an external monitor for a laptop). The other peripheral devices 180 include, in various examples of embodiments, one or more of a stand-alone digital video disc (or digital versatile disc) (DVR, for both terms), a disk player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 180 that provide a function based on the output of the system 150. For example, a disk player performs the function of playing the output of the system 150.
[0053] In various embodiments, control signals are communicated between the system 150 and the display 176, speakers 178, or other peripheral devices 180 using signaling such as AV. Link, Consumer Electronics Control (CEC), or other communications protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to system 1000 via dedicated connections through respective interfaces 164, 166, and 168. Alternatively, the output devices can be connected to system 150 using the communications channel 162 via the communications interface 160. The display 176 and speakers 178 can be integrated in a single unit with the other components of system 150 in an electronic device such as, for example, a television. In various embodiments, the display interface 164 includes a display driver, such as, for example, a timing controller (T Con) chip.
[0054] The display 176 and speaker 178 can alternatively be separate from one or more of the other components, for example, if the RF portion of input 172 is part of a separate set-top box. In various embodiments in which the display 176 and speakers 178 are external components, the output signal can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0055] The system 150 may include one or more sensor devices 168. Examples of sensor devices that may be used include one or more GPS sensors, gyroscopic sensors, accelerometers, light sensors, cameras, depth cameras, microphones, and / or magnetometers. Such sensors may be used to determine information such as user's position and orientation. Where the system 150 is used as the control module for an extended reality display (such as control modules 124, 132), the user's position and orientation may be used indetermining how to render image data such that the user perceives the correct portion of a virtual object or virtual scene from the correct point of view. In the case of head-mounted display devices, the position and orientation of the device itself may be used to determine the position and orientation of the user for the purpose of rendering virtual content. In the case of other display devices, such as a phone, a tablet, a computer monitor, or a television, other inputs may be used to determine the position and orientation of the user for the purpose of rendering content. For example, a user may select and / or adjust a desired viewpoint and / or viewing direction with the use of a touch screen, keypad or keyboard, trackball, joystick, or other input. Where the display device has sensors such as accelerometers and / or gyroscopes, the viewpoint and orientation used for the purpose of rendering content may be selected and / or adjusted based on motion of the display device.
[0056] The embodiments can be carried out by computer software implemented by the processor 152 or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. The memory 154 can be of any type appropriate to the technical environment and can be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. The processor 152 can be of any type appropriate to the technical environment, and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non-limiting examples.
[0057] Described herein is the encoding of avatars for 3D scene representations, along with a procedure for validating an avatar 3D model (mesh, materials, structure, and / or related items) with respect to an avatar reference model.
[0058] Current scene description formats (such as gITF and USD) may be used to encode 3D models that represent an avatar. These models may be manipulated by a user for moving or interacting in a virtual world. Some embodiments may indicate and / or define which part of the scene is an avatar, like in gITF with the MPEG_node_avatar extension.
[0059] As understood, the current standards do not define how to validate a model that may be considered as an avatar satisfying a collection of properties. There does not seem to be a way to encode all those properties in an efficient way. Consequently, an application must know all these properties a priori to correctly instantiate an avatar model.
[0060] Additional background information may be found in references P. Ekman and W.V. Friesen, Measuring Facial Movement, 1 :1 ENVIRONMENTAL PSYCHOLOGY AND NONVERBAL BEHAVIOR 56-75 (1976) andFACS Cheat Sheet, MELINDAOZEL<DOT>COM, available at melindaozel<dot>com / facs-cheat-sheet / (last visited Jan. 6, 2024).
[0061] A validation procedure may be used to ensure that a candidate avatar format satisfies the same properties as a reference avatar. If a candidate avatar is validated, an application may use the validated avatar as a reference avatar. Thanks to this approach, an application may handle many different avatar models, as long as they are validated with one of the reference avatars that the application supports. As understood, other systems have a priori knowledge and are only able to handle data exactly suited for them and may not be able to use avatar data from an outside system.
[0062] A validation procedure may be based on a collection of definitions an avatar format must satisfy. Each definition is identified by a unique id, like MT1 , MT2, or ST1. Each definition concerns one or more parameters or concepts in the model. For instance, MT1 is about the number of vertices.
[0063] A validation procedure may use a reference avatar format that lists the definitions to consider and may eventually include one or more specific values attached to a definition. For instance, an example MORGAN avatar format includes the MT 1 definition specifying that the avatar mesh must consist of 53,698 vertices, which may be written in the following syntax: "MTI(Morgan) = 53,698 vertices”. A validation procedure is not required to include all the definitions listed herein. In accordance with some embodiments the candidate avatar format needs to satisfy only a subset of reference avatar format definitions. In general, such a validation procedure processes only a subset of these definitions.
[0064] Validation of a candidate avatar format ensures that each definition of a reference avatar format is satisfied. For instance, a candidate avatar format validated regarding the MORGAN avatar format must have 53,698 vertices (e.g., MT 1 (Morgan) definition). A candidate avatar format must satisfy all the reference avatar definitions. For the MORGAN example, the candidate avatar format must validate MT1 (Morgan) as well as the other MORGAN definitions, such as MT2(Morgan) and MT4(Morgan).Mesh Topology Definition
[0065] An avatar format may be a mesh topology that, for example, may include mesh topology (MT) identifiers MT1 to MT5. Identifier MT1 may indicate the number of vertices. Identifier MT2 may indicate the number ef faces (polygons). Identifier MT3 may indicate the order of vertices. Identifier MT4 may indicate the order of faces. Identifier MT5 may indicate the values of vertex indices in the face buffer.
[0066] The term "vertices" refers to the "POSITION" attribute in gITF mesh primitives, and the term "faces" refers to the "indices" property of a gITF mesh. These properties may be found in a single gITF mesh primitive, or spread across several primitives for some embodiments, in which they are found from a componentreferenced by the MPEG_node_avatar. For example, an avatar mesh may be split into many gITF meshes, in which each one is referenced by a gITF node, and the MPEG_node_avatar has all these nodes as children. See the specification gITF 2.0, KHRONOS 3D FORMATS WORKING GROUP (2021 ) for more detail regarding gITF.Semantic Topology Definition
[0067] An avatar format may define how an avatar mesh is organized at a semantic level. This may include, for example, semantic topology (ST) identifiers ST1 to ST4.
[0068] Identifier ST 1 indicates the presence of parts, like "head" or "leg". This presence may be mandatory or optional. For instance, a format may ask for at least one arm and up to four arms. Parts may be limited to a single point on the mesh (a.k.a. landmarks). Different types of parts are possible for some embodiments. For instance, a format may have a mesh segmentation as well as landmarks.
[0069] Identifier ST2 indicates that each part may be defined by a unique identifier (such as an integer or a string). For instance, "head" may be the identifier for the head. Identifiers may handle hierarchical relations, like "full_body / upper_body / head" is the head inside the upper body inside the full body.
[0070] Identifier ST3 indicates that each part may be associated with low-level properties, like a range of face indices or vertex indices.
[0071] Identifier ST4 indicates semantic structure definitions, such as an indication that the head must be connected to the neck. The structure may be a partition or a hierarchy. For instance, the head part may be included in the "upper body" part.Deformations Definition
[0072] An avatar format may specify processes that may change the shape of an avatar. For example, the avatar format may include deformation (DF) identifiers DF1 to DF3.
[0073] Identifier DF1 may indicate the presence of mesh deformation objects, like animations or morph targets. For instance, a format may indicate morph targets based on the Facial Action Coding System (FACS) system model and specify which elements are being used. See references P. Ekman and W.V. Friesen, Measuring Facial Movement, 1 :1 ENVIRONMENTAL PSYCHOLOGY AND NONVERBAL BEHAVIOR 56-75 (1976) and FACS Cheat Sheet, MELINDA OZEL, available at melindaozel<dot>com / facs-cheat-sheet / (last visited Jan. 6, 2024) regarding FACS.
[0074] Identifier DF2 may indicate the localization of each deformation object in the gITF description, specified by means of an index into an array or a name. For example, the names of morph targets may bethe ones found in the "targetNames" of the "extras" property of a gITF mesh or provided by a gITF extension. For instance, the first morph target may be action unit #1 (AU1 ), the FACS element that raises inner brows.
[0075] Identifier DF3 may indicate the morph target weight basis, e.g. the vector space that spans the morph targets. For instance, the FLAME face model may be used as described in Tianye, Li, et al., Learning a Model of Facial Shape and Expression from 4D Scans, 38:6 ACM TRANSACTIONS ON GRAPHICS 194:1-17 (2017) (Tianye).
[0076] The morph targets may be attached to some or all the gITF meshes of the avatar mesh. The semantics of morph targets attached to a mesh may be the same or different in gITF meshes. For example, the morph targets in all of the gITF meshes for an avatar mesh may be related to the same semantic. For instance, the first morph target of a gITF mesh may be related to AU1 from FACS, but there may be morph targets only in the gITF meshes that move the inner brows.Skeleton Definition
[0077] An avatar format may define one or more skeletons using one or more of the skeleton (SK) identifiers SK1 to SK7. Identifier SK1 identifies the family or class of skeleton. Identifier SK2 identifies the number of joints. Identifier SK3 identifies the semantics of each joint. Identifier SK4 identifies the connection between joints. Identifier SK5 identifies the semantics of each connection between joints. Identifier SK6 identifies the degrees of freedom of each joint, alongside their value ranges. Identifier SK7 identifies the number of possible joints related to a body part. For instance, three or five joints may be indicated for the right arm.Materials Definition
[0078] An avatar format may define one or more material properties using one or more of the materials (M) identifiers M1 to M3.
[0079] Identifier M1 identifies a specific gITF material, in which some properties are static and others are dynamic. For example, a specific gITF material may be defined in which all properties are static except for the "baseColorFactor" field in the "pbrMetallicRoughness" component of the gITF material. An avatar format may define a range of correct values for a property. For instance, the "baseColorFactor" field may allow values from [0.5, 0.5, 0.5, 0] to [1 .0, 1 .0, 1 .0, 1 .0],
[0080] Identifier M2 identifies a set of allowed gITF materials. Some gITF materials may have dynamic properties for some embodiments, and likewise some gITF materials may be static for some embodiments, or a combination of dynamic and static properties.
[0081] Identifier M3 identifies parametric properties. For instance, texture images used by the avatar material may be the result of a linear combination of textures. For human faces, the FLAME model may be used (See Tianye).Other Identifiers
[0082] An avatar format may define components related to interactivity. For example, a body part may be dedicated to a particular action. For instance, the hands may be used to take and / or pick up objects in a virtual world.
[0083] For some embodiments, the previously-discussed identifiers may be applied to a single avatar format.
[0084] For some embodiments, an avatar format may include several sub-formats. Each sub-format may use the identifiers discussed in the previous sections. A valid avatar may be one that satisfies all of the specifications / defi nitions of one of the sub-formats.Avatar Format Data
[0085] The MORGAN avatar format (as defined in MPEG) is discussed below. The example avatar format has data sections for parts, vertices, faces, and skeletons. The MORGAN format is currently part of the MPEG specification series.
[0086] Example parts data is shown in Table 1 . This section may include the names of all body parts, the number of vertices, the number of faces, and an index range for the faces. The names of all body parts are listed in the first column of Table 1. These body parts are not necessarily partitions because some parts may contain other parts. The number of vertices in each part are listed in the second column of Table 1. The number of faces in each part are listed in the third column of Table 1 . The index range of faces are listed in the fourth column of Table 1 .Table 1.
[0087] For this example, the vertices data section indicates that there are 53,698 vectors (x, y, z). This value is indicated in the first line of Table 1 as the vertex count for the full body. Vertices are encoded here as vectors. In this example, the term "vectors” is used to explan that the axis is in an (x, y, z) system. For some embodiments, the term "vectors” may be replaced with the term "vertices”.
[0088] The faces data section for this example lists 106,952 3-tuples (index of vertex 1 , index of vertex 2, index of vertex 3). This value is indicated in the first line of Table 1 as the face count for the full body.
[0089] An example skeleton is shown in FIG. 2 and referenced in Table 2. Table 2 shows skeleton data that corresponds to FIG. 2. As indicated in Table 2, there are 52 hierarchized joints. Each joint has a unique name and an associated transform. Furthermore, most of the joints listed in Table 2 list a corresponding child body part.Table 2.
[0090] FIG. 2 is a schematic illustration showing an example MORGAN skeleton model according to some embodiments. FIG. 2 shows a pictorial version of the MORGAN skeleton model 200 for the skeletal body parts and joints listed in Table 2. Mesh Topology Definition
[0091] Using the data listed in Table 1 , several mesh topology identifiers are set. The MT1 (Morgan) identifier indicates that there are 53,698 vertices.
[0092] The MT2(Morgan) identifier indicates that there are 106,952 triangles.
[0093] The MT4(Morgan) identifier indicates the order of faces as defined in the MORGAN avatar representation format specification [SD] Update of Annex H on Potential Improvement CDAM223090- 14 document - Reference Avatar, MPEG Meeting #143, Input Document m63997 (July 2023) (“m63997”).
[0094] The MT5(Morgan) indicator field indicates the values of vertex indices in the face buffer as defined in the MORGAN avatar representation format specification (m63997).Semantic Topology Definition
[0095] Using the data listed in Table 1 , several semantic topology identifiers are set. The ST1 (Morgan) identifier indicates that there are 46 parts, which are shown in the first column of Table 1 . Not counting the “full_body” row, there are 46 rows in Table 1 , which correspond to the 46 parts.
[0096] The ST2(Morgan) identifier lists unique strings that are used to identify each part. See the first column of Table 1.
[0097] The ST3(Morgan) identifier lists the predefined number of vertices for each part, which is shown in the second column of Table 1 . The ST3(Morgan) identifier also lists the predefined face index range, which is shown in the fourth column of T able 1 .
[0098] The ST4(Morgan) identifier lists parts that form a hierarchy. The slash symbol " / " in part names encodes the hierarchy in parts. For instance, all IDs starting with "full_body / upper_body / head / ..." are parts inside "full_body / upper_body / head".Skeleton Definition
[0099] Using the example data listed in Table 2, several skeleton identifiers are set. The avatar format may define one or more skeletons, with one or more definitions, such as definitions for identifiers SK2, SK3, and SK4.
[0100] The SK2(Morgan) identifier indicates that there are 52 joints.
[0101] The SK3(Morgan) identifier indicates that each joint has a unique name and that each name corresponds to a body part.
[0102] The SK4(Morgan) identifier indicates the hierarchy of joints that define their connections.Validation
[0103] For this example, a candidate avatar format is valid with respect to the MORGAN reference avatar format if the candidate satisfies all of the following definitions:• Mesh Topology definitions MTI(Morgan), MT2(Morgan), MT4(Morgan), and MT5(Morgan)• Semantic Topology definitions ST1 (Morgan), ST2(Morgan), ST3(Morgan), and ST4(Morgan)• Skeleton definitions SK2(Morgan), SK3(Morgan), and SK4(Morgan)
[0104] By picking these definitions, an application may be able to make assumptions and provide specific features. For example, any avatar that validates this Morgan reference may assume that the hands are always at the same location in the mesh (e.g., in the same vertex index range), even if the geometry of the hands is different (e.g., the (x, y, z) values of the hand vertices may be different). New definitions (such as MTx, or even a new definition category, such as YYx) may be used with this validation process for some embodiments. Numerous other combinations of definitions (even ones not listed above) could be selected for a reference avatar format.
[0105] For some embodiments, there is no minimum number of definitions that a reference avatar format must have. For some embodiments, a reference avatar format may be limited to a single definition, such as only the number of vertices (MT 1 in a MORGAN format).
[0106] FIG. 3 is a flowchart illustrating an example avatar format validation process according to some embodiments. In processing blocks 301 and 302, the candidate avatar format (to be validated or not) and the reference avatar format (of which definitions are used for validation) are received and / or retrieved as inputs. In processing block 303, the next definition D in the reference avatar format is retrieved. Each definition in the reference avatar format (for example, "MT1 preference avatar format name>)”) is used to analyze the format.
[0107] In decision block 304, the candidate avatar data is analyzed and assessed relative to the current definition D of the reference avatar format. The assessment may be direct or indirect. For example, for definition MT1 , if the candidate format explicitly defines the number of vertices, then MT1 is validated. In another case, if a vertex buffer is provided with the format, the number of vertices may be computed. Such a computation may be done to satisfy / validate the MT1 definition. If the current definition is not satisfied / validated, the example process goes to block 305 to indicate that the candidate format is not valid. In decision block 306, an assessment is made to determine if more definitions exist in the reference avatar format. If no definitions exist, the process goes to process block 307 to indicate that the candidate format is valid. If more definitions exist, the process returns back to retrieve the next definition from the reference format.
[0108] For some embodiments, a validation process may be done offline prior to run-time. For some embodiments, the validation process may be done at run-time of potential use of a candidate avatar format.
[0109] For some embodiments, a reference avatar format data block and / or candidate avatar format data block may conform to the MPEG_node_avatar extension specification. For some embodiments, each avatar URN (in the "type” property of MPEG_node_avatar) may be "bound” to a reference avatar. If a standard URN is used, an avatar model may be used under an assumption that the avatar model meets all of the requirements of the reference model. For some embodiments, the user of the avatar model may run all the checks to be sure that the avatar model satisfies all of the requirements of the reference model.
[0110] For some embodiments, a validation process may be used with an MPEG_node_avatar to validate the content of the gITF regarding a URN, which is in the "type” property of MPEG_node_avatar. For instance, if the "type" field is "urn:mpeg:sd:2023:avatar" (e.g., MORGAN), then the avatar rig inside the gITF must satisfy all of the MORGAN's definitions (MT1 , ...). The validation process also may be used with other URNs or new standard avatars. For instance, a URN of "urn:mpeg:sd:2024:avatar" may be used for a MORGAN 2.0 format, which may be the current MORGAN format with updates and new / updated definitions, such as MT1 or other definition identifiers.
[0111] FIG. 4 is a flowchart illustrating an example avatar format validation process according to some embodiments. For some embodiments, an example process 400 may include obtaining 402 candidate avatar format data for an avatar. For some embodiments, the example process 400 may further include obtaining 404 reference avatar format data for the avatar, wherein the reference avatar format data includes a first set of one or more definitions. For some embodiments, the example process 400 may further include performing 406 a validation process for each of the one or more definitions of the first set. For some embodiments, the validation process of the example process 400 may include retrieving 408 a current definition from the first set of one or more definitions. For some embodiments, the validation process of the example process 400 may further include generating 410 a comparison result by determining if the candidate avatar format data satisfies the current definition. For some embodiments, the validation process of the example process 400 may further include exiting 412 the validation process if the comparison result is false. For some embodiments, the example process 400 may further include determining 414 if more definitions are left in the first set. If more definitions are left in the first set, the example process 400 may return to the top of a loop to retrieve 408 a current definition from the first set of one or definitions. Otherwise, the example process 400 continues. For some embodiments, the example process 400 may further include indicating 416 the candidate avatar format is valid if the comparison result is true for each pass through the validation process.
[0112] While the methods and systems in accordance with some embodiments are generally discussed in context of extended reality (XR), some embodiments may be applied to any XR contexts such as, e.g., virtual reality (VR) / mixed reality (MR) / augmented reality (AR) contexts. Also, although the term "headmounted display (HMD)” is used herein in accordance with some embodiments, some embodiments may be applied to a wearable device (which may or may not be attached to the head) capable of, e.g., XR, VR, AR, and / or MR for some embodiments.
[0113] A first example method in accordance with some embodiments may include: obtaining candidate avatar format data for an avatar, obtaining reference avatar format data for the avatar, wherein the reference avatar format data includes a first set of one or more definitions; performing a validation process for each of the one or more definitions of the first set, wherein the validation process includes: retrieving a current definition from the first set of one or more definitions; generating a comparison result by determining if the candidate avatar format data satisfies the current definition; and exiting the validation process if the comparison result is false; and indicating the candidate avatar format is valid if the comparison result is true for each pass through the validation process.
[0114] For some embodiments of the first example method, the first set includes at least one definition regarding a mesh topology of the avatar.
[0115] For some embodiments of the first example method, the first set includes at least one definition regarding a semantic topology of the avatar.
[0116] For some embodiments of the first example method, the first set includes at least one definition regarding a deformation of the avatar.
[0117] For some embodiments of the first example method, the first set includes at least one definition regarding a skeleton of the avatar.
[0118] For some embodiments of the first example method, the first set includes at least one definition regarding a material of the avatar.
[0119] For some embodiments of the first example method, the first set includes at least one definition regarding interactivity of the avatar.
[0120] For some embodiments of the first example method, the first set includes at least one definition regarding a first sub-format and at least one definition regarding a second sub-format.
[0121] Some embodiments of the first example method may further include: obtaining a second reference avatar format data for an avatar, wherein the second reference avatar format data includes a second set of one or more definitions; repeating the validation process using the second reference avatar format data, and indicating the candidate avatar format is valid only if the candidate avatar format is indicated as valid using at least one of the reference avatar format data or the second reference avatar format data.
[0122] For some embodiments of the first example method, at least one of the reference avatar format data, the second reference avatar format data, and the candidate avatar format data satisfies MPEG_node_avatar extension criteria.
[0123] A first example apparatus in accordance with some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform any one of the methods listed above.
[0124] An example method of validating a candidate avatar format against a reference avatar format in accordance with some embodiments may include: comparing one or more features of the candidate avatar format with one or more definitions of the reference avatar format; and validating the candidate avatar format if the one or more features of the candidate avatar format satisfy the one or more definitions.
[0125] A third example method in accordance with some embodiments may include: performing a validation process for each of one or more definitions of a first set within reference avatar format data, wherein the reference avatar format data includes the first set of one or more definitions, and wherein the validation process includes: retrieving a current definition from the first set of one or more definitions; determining a comparison result by comparing the current definition with a set of candidate avatar format data; and exiting the validation process if the comparison result is false; and indicating the candidate avatar format is valid if the comparison result is true for each pass through the validation process.
[0126] A third example apparatus in accordance with some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform any one of the methods listed above.
[0127] A fourth example method in accordance with some embodiments may include: obtaining reference avatar format data for an avatar, wherein the reference avatar format data includes a first set of one or more definitions; obtaining candidate avatar format data for the avatar; performing a validation process for each of the one or more definitions of the first set, wherein the validation process includes: retrieving a current definition from the first set of one or more definitions; determining a comparison result by comparing the current definition with the candidate avatar format data; and exiting the validation process if the comparison result is false; and indicating the candidate avatar format is valid if the comparison result is true for each pass through the validation process.
[0128] A fourth example apparatus in accordance with some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform any one of the methods listed above.
[0129] A fifth example method in accordance with some embodiments may include: obtaining reference avatar format data for an avatar, wherein the reference avatar format data includes a first set of one or more definitions; obtaining candidate avatar format data for the avatar; performing a validation process for each of the one or more definitions of the first set, wherein the validation process includes: retrieving a current definition from the first set of one or more definitions; determining a comparison result by comparing the current definition with the candidate avatar format data; indicating the candidate avatar format is invalid if the comparison result is false; exiting the validation process if the comparison result is false; returning to a beginning of the validation process if the comparison result is true and more definitions from the first set are left to be retrieved; and indicating the candidate avatar format is valid if the comparison result is true and all of the definitions from the first set have been retrieved.
[0130] A fifth example apparatus in accordance with some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform any one of the methods listed above.
[0131] A sixth example apparatus in accordance with some embodiments may include at least one processor configured to perform any one of the methods listed above.
[0132] A seventh example apparatus in accordance with some embodiments may include a computer- readable medium storing instructions for causing one or more processors to perform any one of the methods listed above.
[0133] An eighth example apparatus in accordance with some embodiments may include at least one processor and at least one non-transitory computer-readable medium storing instructions for causing the at least one processor to perform any one of the methods listed above.
[0134] An example bitstream in accordance with some embodiments may include a bitsream generated according to any one of the methods listed above.
[0135] This disclosure describes a variety of aspects, including tools, features, embodiments, models, approaches, etc. Many of these aspects are described with specificity and, at least to show the individual characteristics, are often described in a manner that may sound limiting. However, this is for purposes of clarity in description, and does not limit the disclosure or scope of those aspects. Indeed, all of the different aspects can be combined and interchanged to provide further aspects. Moreover, the aspects can be combined and interchanged with aspects described in earlier filings as well.
[0136] The aspects described and contemplated in this disclosure can be implemented in many different forms. While some embodiments are illustrated specifically, other embodiments are contemplated, and thediscussion of particular embodiments does not limit the breadth of the implementations. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a bitstream generated or encoded. These and other aspects can be implemented as a method, an apparatus, a computer readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the methods described, and / or a computer readable storage medium having stored thereon a bitstream generated according to any of the methods described.
[0137] In the present disclosure, the terms "reconstructed” and "decoded” may be used interchangeably, the terms "pixel” and "sample” may be used interchangeably, the terms "image,” "picture” and "frame” may be used interchangeably. Usually, but not necessarily, the term "reconstructed” is used at the encoder side while "decoded” is used at the decoder side.
[0138] The terms HDR (high dynamic range) and SDR (standard dynamic range) often convey specific values of dynamic range to those of ordinary skill in the art. However, additional embodiments are also intended in which a reference to HDR is understood to mean "higher dynamic range” and a reference to SDR is understood to mean "lower dynamic range.” Such additional embodiments are not constrained by any specific values of dynamic range that might often be associated with the terms "high dynamic range” and "standard dynamic range.”
[0139] Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as "first”, "second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., such as, for example, a "first decoding” and a "second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.
[0140] Various numeric values may be used in the present disclosure, for example. The specific values are for example purposes and the aspects described are not limited to these specific values.
[0141] Embodiments described herein may be carried out by computer software implemented by a processor or other hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. The processor can be of any type appropriate to the technical environment and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as nonlimiting examples.
[0142] Various implementations involve decoding. "Decoding”, as used in this disclosure, can encompass all or part of the processes performed, for example, on a received encoded sequence in order to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, inverse transformation, and differential decoding. In various embodiments, such processes also, or alternatively, include processes performed by a decoder of various implementations described in this disclosure, for example, extracting a picture from a tiled (packed) picture, determining an upsampling filter to use and then upsampling a picture, and flipping a picture back to its intended orientation.
[0143] As further examples, in one embodiment "decoding” refers only to entropy decoding, in another embodiment "decoding” refers only to differential decoding, and in another embodiment "decoding” refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process” is intended to refer specifically to a subset of operations or generally to the broader decoding process will be clear based on the context of the specific descriptions.
[0144] Various implementations involve encoding. In an analogous way to the above discussion about "decoding”, "encoding” as used in this disclosure can encompass all or part of the processes performed, for example, on an input video sequence in order to produce an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, for example, partitioning, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes also, or alternatively, include processes performed by an encoder of various implementations described in this disclosure.
[0145] As further examples, in one embodiment "encoding” refers only to entropy encoding, in another embodiment "encoding” refers only to differential encoding, and in another embodiment "encoding” refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process” is intended to refer specifically to a subset of operations or generally to the broader encoding process will be clear based on the context of the specific descriptions.
[0146] When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of a corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of a corresponding method / process.
[0147] Various embodiments refer to rate distortion optimization. In particular, during the encoding process, the balance or trade-off between the rate and distortion is usually considered, often given the constraints of computational complexity. The rate distortion optimization is usually formulated as minimizing a rate distortion function, which is a weighted sum of the rate and of the distortion. There are differentapproaches to solve the rate distortion optimization problem. For example, the approaches may be based on an extensive testing of all encoding options, including all considered modes or coding parameters values, with a complete evaluation of their coding cost and related distortion of the reconstructed signal after coding and decoding. Faster approaches may also be used, to save encoding complexity, in particular with computation of an approximated distortion based on the prediction or the prediction residual signal, not the reconstructed one. A mix of these two approaches can also be used, such as by using an approximated distortion for only some of the possible encoding options, and a complete distortion for other encoding options. Other approaches only evaluate a subset of the possible encoding options. More generally, many approaches employ any of a variety of techniques to perform the optimization, but the optimization is not necessarily a complete evaluation of both the coding cost and related distortion.
[0148] The implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed can also be implemented in other forms (for example, an apparatus or program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs”), and other devices that facilitate communication of information between end-users.
[0149] Reference to "one embodiment” or "an embodiment” or "one implementation” or "an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase "in one embodiment” or "in an embodiment” or "in one implementation” or "in an implementation”, as well any other variations, appearing in various places throughout this disclosure are not necessarily all referring to the same embodiment.
[0150] Additionally, this disclosure may refer to "determining” various pieces of information. Determining the information can include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
[0151] Further, this disclosure may refer to "accessing” various pieces of information. Accessing the information can include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0152] Additionally, this disclosure may refer to "receiving” various pieces of information. Receiving is, as with "accessing”, intended to be a broad term. Receiving the information can include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, "receiving” is typically involved, in one way or another, during operations such as, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0153] It is to be appreciated that the use of any of the following “ / ”, "and / or”, and "at least one of, for example, in the cases of “A / B”, "A and / or B” and "at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C” and "at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended for as many items as are listed.
[0154] Also, as used herein, the word "signal” refers to, among other things, indicating something to a corresponding decoder. For example, in certain embodiments the encoder signals a particular one of a plurality of parameters for region-based filter parameter selection for de-artifact filtering. In this way, in an embodiment the same parameter is used at both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways. For example, one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments. While the preceding relates to the verb form of the word "signal”, the word "signal” can also be used herein as a noun.
[0155] Implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal can be formatted to carry the bitstream of a described embodiment. Such a signal can be formatted, for example, as anelectromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
[0156] We describe a number of embodiments. Features of these embodiments can be provided alone or in any combination, across various claim categories and types. Further, embodiments can include one or more of the following features, devices, or aspects, alone or in any combination, across various claim categories and types:• Adapting residues at an encoder according to any of the embodiments discussed.• A bitstream or signal that includes one or more of the described syntax elements, or variations thereof.• A bitstream or signal that includes syntax conveying information generated according to any of the embodiments described.• Inserting in the signaling syntax elements that enable the decoder to adapt residues in a manner corresponding to that used by an encoder.• Creating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the described syntax elements, or variations thereof.• Creating and / or transmitting and / or receiving and / or decoding according to any of the embodiments described.• A method, process, apparatus, medium storing instructions, medium storing data, or signal according to any of the embodiments described.• A TV, set-top box, cell phone, tablet, or other electronic device that performs adaptation of filter parameters according to any of the embodiments described.• A TV, set-top box, cell phone, tablet, or other electronic device that performs adaptation of filter parameters according to any of the embodiments described, and that displays (e.g. using a monitor, screen, or other type of display) a resulting image.• A TV, set-top box, cell phone, tablet, or other electronic device that selects (e.g. using a tuner) a channel to receive a signal including an encoded image, and performs adaptation of filter parameters according to any of the embodiments described.• A TV, set-top box, cell phone, tablet, or other electronic device that receives (e.g. using an antenna) a signal over the air that includes an encoded image, and performs adaptation of filter parameters according to any of the embodiments described.
[0157] Note that various hardware elements of one or more of the described embodiments are referred to as "modules” that carry out (i.e., perform, execute, and the like) various functions that are described herein in connection with the respective modules. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memory devices) deemed suitable by those of skill in the relevant art for a given implementation. Each described module may also include instructions executable for carrying out the one or more functions described as being carried out by the respective module, and it is noted that those instructions could take the form of or include hardware (i.e., hardwired) instructions, firmware instructions, software instructions, and / or the like, and may be stored in any suitable non-transitory computer-readable medium or media, such as commonly referred to as RAM, ROM, etc.
[0158] Although features and elements are described above in particular combinations, one of ordinary skill in the art will appreciate that each feature or element can be used alone or in any combination with the other features and elements. In addition, the methods described herein may be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, a read only memory (ROM), a random access memory (RAM), a register, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks, and digital versatile disks (DVDs). A processor in association with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
CLAIMS1. A method comprising: obtaining candidate avatar format data for an avatar, obtaining reference avatar format data for the avatar, wherein the reference avatar format data comprises a first set of one or more definitions; performing a validation process for each of the one or more definitions of the first set, wherein the validation process comprises: retrieving a current definition from the first set of one or more definitions; generating a comparison result by determining if the candidate avatar format data satisfies the current definition; and exiting the validation process if the comparison result is false; and indicating the candidate avatar format is valid if the comparison result is true for each pass through the validation process.
2. The method of claim 1 , wherein the first set comprises at least one definition regarding a mesh topology of the avatar.
3. The method of any one of claims 1-2, wherein the first set comprises at least one definition regarding a semantic topology of the avatar.
4. The method of any one of claims 1-3, wherein the first set comprises at least one definition regarding a deformation of the avatar.
5. The method of any one of claims 1-4, wherein the first set comprises at least one definition regarding a skeleton of the avatar.
6. The method of any one of claims 1-5, wherein the first set comprises at least one definition regarding a material of the avatar.
7. The method of any one of claims 1-6, wherein the first set comprises at least one definition regarding interactivity of the avatar.
8. The method of any one of claims 1-7, wherein the first set comprises at least one definition regarding a first sub-format and at least one definition regarding a second sub-format.
9. The method of any one of claims 1-8, further comprising:obtaining a second reference avatar format data for an avatar, wherein the second reference avatar format data comprises a second set of one or more definitions; repeating the validation process using the second reference avatar format data, and indicating the candidate avatar format is valid only if the candidate avatar format is indicated as valid using at least one of the reference avatar format data or the second reference avatar format data.
10. The method of any one of claims 1-9, wherein at least one of the reference avatar format data, the second reference avatar format data, and the candidate avatar format data satisfies MPEG_node_avatar extension criteria.11 . An apparatus comprising: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform the method of any one of claims 1 through 10.
12. A method of validating a candidate avatar format against a reference avatar format, comprising: comparing one or more features of the candidate avatar format with one or more definitions of the reference avatar format; and validating the candidate avatar format if the one or more features of the candidate avatar format satisfy the one or more definitions.
13. A method comprising: performing a validation process for each of one or more definitions of a first set within reference avatar format data, wherein the reference avatar format data comprises the first set of one or more definitions, and wherein the validation process comprises: retrieving a current definition from the first set of one or more definitions; determining a comparison result by comparing the current definition with a set of candidate avatar format data; and exiting the validation process if the comparison result is false; and indicating the candidate avatar format is valid if the comparison result is true for each pass through the validation process.
14. An apparatus comprising:a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform the method of claim 13.
15. A method comprising: obtaining reference avatar format data for an avatar, wherein the reference avatar format data comprises a first set of one or more definitions; obtaining candidate avatar format data for the avatar; performing a validation process for each of the one or more definitions of the first set, wherein the validation process comprises: retrieving a current definition from the first set of one or more definitions; determining a comparison result by comparing the current definition with the candidate avatar format data; and exiting the validation process if the comparison result is false; and indicating the candidate avatar format is valid if the comparison result is true for each pass through the validation process.
16. An apparatus comprising: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform the method of claim 15.
17. A method comprising: obtaining reference avatar format data for an avatar, wherein the reference avatar format data comprises a first set of one or more definitions; obtaining candidate avatar format data for the avatar; performing a validation process for each of the one or more definitions of the first set, wherein the validation process comprises: retrieving a current definition from the first set of one or more definitions; determining a comparison result by comparing the current definition with the candidate avatar format data; indicating the candidate avatar format is invalid if the comparison result is false; exiting the validation process if the comparison result is false;returning to a beginning of the validation process if the comparison result is true and more definitions from the first set are left to be retrieved; and indicating the candidate avatar format is valid if the comparison result is true and all of the definitions from the first set have been retrieved.
18. An apparatus comprising: a processor; and a non-transitory computer-readable medium storing instructions operative, when executed by the processor, to cause the apparatus to perform the method of claim 17.
19. An apparatus comprising at least one processor configured to perform the method of any one of claims 1-10, 12, 13, 15, and 17.
20. An apparatus comprising a computer-readable medium storing instructions for causing one or more processors to perform the method of any one of claims 1-10, 12, 13, 15, and 17.
21. An apparatus comprising at least one processor and at least one non-transitory computer-readable medium storing instructions for causing the at least one processor to perform the method of any one of claims 1-10, 12, 13, 15, and 17.
22. A signal including a bitstream generated according to any one of claims 1-10, 12, 13, 15, and 17.
Citation Information
Patent Citations
Avatar creation system and method
GB2516241A
Avatar appearance transformation in a virtual universe
US8108774B2
Cited By
UDIM support for scene description
WO2026171852A1