Geometric avatar media codec for transmission

By using avatar geometry encoders and decoders, the problems of poor interoperability of avatar data formats and privacy protection in XR applications are solved. Standardized encoding and decoding of avatar data are achieved, improving data interoperability and privacy protection across different systems.

CN122270916APending Publication Date: 2026-06-23INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTERDIGITAL CE PATENT HOLDINGS SAS
Filing Date
2024-10-08
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing video coding systems, when processing virtual content in extended reality (XR), especially digital human representation, lack consideration for the interaction of social human behavior and environmental information, and fail to effectively address social and human privacy issues, resulting in poor interoperability of data formats across different systems.

Method used

Employing an avatar geometry encoder and decoder, it generates and parses avatar data structures by extracting and encoding/decoding avatar semantic information, geometric metadata, skeletal information, landmark information, or artificial intelligence information. It supports binary compression and decompression, ensuring data format standardization and interoperability.

Benefits of technology

It enables standardized encoding and decoding of data in XR applications, improves data interoperability across different systems, solves the problem of missing social human behavior and environmental information interaction, and ensures data privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122270916A_ABST
    Figure CN122270916A_ABST
Patent Text Reader

Abstract

A system, method, and tool for a geometria media codec for transmission are disclosed. An exemplary device (e.g., an encoder) can receive a descriptive avatar media file. The device can extract a first type of geometria information and a second type of geometria information from the descriptive avatar media file. The device can encode the first type of geometria information and the second type of geometria information into an avatar data structure. The device can send the avatar data structure, along with an indication of the file format contained within the avatar data structure, to the decoder.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims the benefit of European provisional patent application No. EP23306756.0, filed on October 10, 2023; European provisional application No. EP23306759.4, filed on October 10, 2023; and European provisional application No. EP23306760.2, filed on October 10, 2023, the contents of which are incorporated herein by reference. Background Technology

[0002] Video coding systems can be used to compress digital video signals, for example, to reduce the storage and / or transmission bandwidth required for such signals. Video coding systems can include, for example, block-based, wavelet-based, and / or object-based systems.

[0003] Extended Reality (XR) is a technology that enables interactive experiences where real-world environments and / or video content are enhanced with virtual content (e.g., which can be defined across multiple sensory modalities, including visual, auditory, tactile, etc.). During application runtime, virtual content (e.g., 3D content or audio / video files) can be rendered in real-time in a manner consistent with the user's context (e.g., environment, viewpoint, device, etc.). Scene graphs (e.g., Khronos / glTF and its extensions defined in the MPEG scene description format or Apple / USDZ) can be used to represent the content to be rendered. Scene graphs can combine a descriptive description of the scene structure connecting real-world objects and virtual objects with a binary representation of the virtual content. Summary of the Invention

[0004] A system, method, and tools for a geometrization media codec for transmission are disclosed. Techniques for implementing digital human representation can be achieved through synthetic 3D models. Synthetic representations are easier to manipulate and customize to represent specific human anatomy, visual effects, and / or social parameters. Synthetic representations can facilitate the generation of animation in immersive realities (e.g., given that all parameters are previously known). Synthetic models can facilitate the generalization and / or stylization of appearances (e.g., professionals can use them for geometry, animation, and / or streaming).

[0005] Virtual content in XR applications can include avatars. An avatar is a user's digital virtual representation and can take several forms, including meshes, volumes, voxels, point clouds, images, videos, and sounds. To create and deliver digital media content for XR applications, templated digital human formalization can be used to capture representations of different avatars and provide a cross-platform interoperable representation format. This template can be a general human model with skeletal structure, an object-specific model, a statistical shape model, or metadata representing an individual's social attributes. These methods can provide (e.g., accurately provide) a human image as an initial phase capable of statistically representing human representations (e.g., different human representations).

[0006] Three-dimensional (3D) synthetic models can be excellent representations. In some examples, 3D models may lack standard definitions and / or structures that can be used to and / or match other digital human representations. Such structures may include (e.g., in addition to) user-specific details such as virtual identity, social status, or input device control, semantic structure of geometric properties (e.g., geometric animation), shape and skeletal anatomy, and / or semantics of animation parameters, etc.

[0007] Some approaches to standardizing digital human content may disregard social human behavior and / or interaction with environmental information, and / or social and human privacy issues (e.g., these can be public characteristics that can be adopted within social technologies in the real world). In examples (e.g., real-time distribution or communication use cases), the encoding and / or data structure of the digital human content can be defined. In one example (e.g., in a proprietary system), the data may be known. In open systems, the data format can be specified and / or standardized to enable interoperability between different systems.

[0008] An exemplary device may have a processor configured to perform one or more actions. The device (e.g., an encoder, such as an avatar geometry encoder) may receive a descriptive avatar media file. The device may extract a first type of geometry-related avatar information and a second type of geometry-related avatar information from the descriptive avatar media file. The device may encode the first type of geometry-related avatar information and the second type of geometry-related avatar information into an avatar data structure. The device may send the avatar data structure, along with an indication of the file format contained within the avatar data structure, to a decoder.

[0009] At least one of the first type of geometrically related avatar information and the second type of geometrically related avatar information can be: avatar semantic information; avatar geometric metadata; avatar skeletal information; avatar landmark information; or artificial intelligence avatar information.

[0010] The avatar data structure can include binary files. The device can perform binary compression on both types of geometry-related information (first and second types) to generate compressed information. The device can then package this compressed information to generate a binary file.

[0011] The avatar data structure can include human-readable files. The device can encode first-type and second-type geometry-related avatar information into an avatar data structure by generating human-readable files based on first-type and second-type geometry-related information.

[0012] The first type of geometry-related avatar information can be associated with a first file format contained in the avatar data structure. The second type of geometry-related avatar information can be associated with a second file format contained in the avatar data structure. The first file format can be different from the second file format.

[0013] An exemplary device (e.g., a decoder, such as an avatar geometry decoder) can receive an avatar data structure and an indication of the file format contained within the avatar data structure from an encoder. The device can decode the avatar data structure. Based on the decoded avatar data structure, the device can determine a first type of geometry-related avatar information and a second type of geometry-related avatar information. The device can generate a descriptive avatar media file based on the first type of geometry-related avatar information, the second type of geometry-related information, and the indicated file format.

[0014] At least one of the first type of geometrically related avatar information or the second type of geometrically related avatar information may include: avatar semantic information; avatar geometric metadata; avatar skeletal information; avatar landmark information; or artificial intelligence avatar information.

[0015] The data structure can include binary files. This device can unpack binary files to generate compressed information. This device can also perform binary decompression on the compressed information.

[0016] The device can map a first type of geometry-related information to a first file format contained in the avatar data structure. The device can also map a second type of geometry-related information to a second file format contained in the avatar data structure, where the first file format differs from the second file format.

[0017] The avatar data structure may include a human-readable file containing encoded versions of a first type of geometry-related avatar information and a second type of geometry-related avatar information.

[0018] A device (e.g., an encoder, such as an avatar geometry encoder) can receive a descriptive avatar media file that includes a first type of geometry-related information and a second type of geometry-related information. The device can encode the descriptive avatar media file based on the first and second types of geometry-related information to generate an avatar data structure. The device can then send the avatar data structure, along with an indication of the file format contained within it, to a decoder.

[0019] The avatar data structure can be a binary file. Encoding a descriptive avatar media file based on a first type of geometrically related information and a second type of geometrically related information to generate an avatar data structure may involve: analyzing the format of the descriptive avatar media file to extract the first type of geometrically related information and the second type of geometrically related information from the descriptive avatar media file; performing binary compression on the first type of geometrically related information and the second type of geometrically related information to generate compressed information; and packaging the compressed information to generate a binary file.

[0020] The avatar data structure can be a human-readable file. Encoding a descriptive avatar media file based on a first type of geometrically relevant information and a second type of geometrically relevant information to generate the avatar data structure may involve: analyzing the format of the descriptive avatar media file to extract the first type of geometrically relevant information and the second type of geometrically relevant information from the descriptive avatar media file; and generating a human-readable file based on the first type of geometrically relevant information and the second type of geometrically relevant information.

[0021] An exemplary device (e.g., a decoder, such as an avatar geometry decoder) can receive an avatar data structure and an indication of the file format contained within the data structure from an encoder. Based on this indication, the device can determine that the avatar data structure includes a first type of geometry-related information and a second type of geometry-related information. The device can decode the avatar data structure based on the first type of geometry-related information and the second type of geometry-related information to generate a descriptive avatar media file. The device can then render the avatar based on the descriptive avatar media file.

[0022] The avatar data structure can be a binary file. Decoding the avatar data structure based on first-type and second-type geometry-related information to generate a descriptive avatar media file may involve: unpacking the binary file to generate compressed information; performing binary decompression on the compressed information to generate first-type and second-type geometry-related information; and identifying the formats of the first-type and second-type geometry-related information to generate the descriptive avatar media file.

[0023] The avatar data structure can be a human-readable file. Decoding the avatar data structure based on first-type and second-type geometric-related information to generate a descriptive avatar media file may involve: mapping a human-readable file to a file format of the descriptive avatar media file based on an instruction; and generating the descriptive avatar media file based on the file format of the descriptive avatar media file.

[0024] At least one of the first type of geometry-related information or the second type of geometry-related information can be: avatar semantic information; avatar geometric metadata; avatar skeletal information; avatar landmark information; or artificial intelligence avatar information. Attached Figure Description

[0025] Figure 1A This is a system diagram illustrating an exemplary communication system that can implement one or more of the disclosed embodiments.

[0026] Figure 1B This illustrates that, according to an embodiment, it is possible to Figure 1A A system diagram of an exemplary wireless transmit / receive unit (WTRU) used within the communication system shown.

[0027] Figure 1C This illustrates that, according to an embodiment, it is possible to Figure 1A The diagram shows an exemplary radio access network (RAN) and an exemplary core network (CN) used within the communication system.

[0028] Figure 1D This illustrates that, according to an embodiment, it is possible to Figure 1A The system diagram shown is another exemplary RAN and another exemplary CN used in the communication system.

[0029] Figure 2 An exemplary video encoder is shown.

[0030] Figure 3 An example video decoder is shown.

[0031] Figure 4 Examples of systems in which various aspects and examples can be implemented are shown.

[0032] Figure 5 An overview of an exemplary avatar geometric media framework is shown.

[0033] Figure 6 An exemplary avatar geometry encoder architecture is shown.

[0034] Figure 7 An exemplary avatar geometry decoder architecture is shown.

[0035] Figure 8 An overview of an exemplary avatar-style media framework is shown.

[0036] Figure 9 An exemplary incarnation-form encoder architecture is shown.

[0037] Figure 10 An exemplary incarnation-form decoder architecture is shown.

[0038] Figure 11 An overview of an exemplary avatar animation media framework is shown.

[0039] Figure 12 An exemplary avatar animation encoder architecture is shown.

[0040] Figure 13 An exemplary avatar animation decoder architecture is shown. Detailed Implementation

[0041] A more detailed understanding can be obtained from the following description, which is given by way of example in conjunction with the accompanying drawings.

[0042] Figure 1A This diagram illustrates an exemplary communication system 100 that may implement one or more of the disclosed embodiments. The communication system 100 may be a multiple access system providing content such as voice, data, video, messaging, and broadcasting to multiple wireless users. The communication system 100 enables multiple wireless users to access such content through shared system resources including wireless broadband. For example, the communication system 100 may employ one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), Single Carrier FDMA (SC-FDMA), Zero-Tail Unique Word DFT Extended OFDM (ZT UW DTS-s OFDM), Unique Word OFDM (UW-OFDM), Resource Block Filtered OFDM, Filter Bank Multicarrier (FBMC), etc.

[0043] like Figure 1AAs shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, public switched telephone network (PSTN) 108, Internet 110, and other networks 112. However, it will be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d can be any type of device configured to operate and / or communicate in a wireless environment. For example, WTRUs 102a, 102b, 102c, and 102d (any of which may be referred to as a “station” and / or “STA”) may be configured to transmit and / or receive wireless signals and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearable devices, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in the context of industrial and / or automated processing chains), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 102a, 102b, 102c, and 102d may be interchangeably referred to as a UE.

[0044] The communication system 100 may also include base station 114a and / or base station 114b. Each of base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks such as CN 106 / 115, Internet 110, and / or other networks 112. For example, base stations 114a and 114b may be base transceiver stations (BTS), Node-B, eNode B, home Node-B, home eNode B, gNB, NR NodeB, site controller, access point (AP), wireless router, etc. Although base stations 114a and 114b are each depicted as a single element, it will be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.

[0045] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage for a specific geographic area that may be relatively fixed or may change over time. A cell may also be divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Therefore, in one embodiment, base station 114a may include three transceivers, i.e., one transceiver per sector of the cell. In embodiments, base station 114a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.

[0046] Base stations 114a and 114b can communicate with one or more of WTRUs 102a, 102b, 102c, and 102d via air interface 116, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). Any suitable radio access technology (RAT) can be used to establish air interface 116.

[0047] More specifically, as described above, the communication system 100 can be a multiple access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base station 114a in RAN 104 / 113 and WTRUs 102a, 102b, 102c can implement radio technologies, such as using Wideband CDMA (WCDMA) to establish Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA) for air interfaces 115 / 116 / 117. WCDMA can include communication protocols such as High-Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High-Speed ​​UL Packet Access (HSUPA).

[0048] In the embodiment, base station 114a and WTRUs 102a, 102b, 102c may implement radio technologies, such as using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro) to establish Evolved UMTS Terrestrial Radio Access (E-UTRA) for air interface 116.

[0049] In the embodiments, base station 114a and WTRUs 102a, 102b, 102c may implement radio technologies, such as using New Radio (NR) to establish NR radio access for air interface 116.

[0050] In the embodiments, base station 114a and WTRUs 102a, 102b, and 102c can implement various radio access technologies. For example, base station 114a and WTRUs 102a, 102b, and 102c can, for instance, use the dual connectivity (DC) principle to jointly implement LTE radio access and NR radio access. Therefore, the air interface used by WTRUs 102a, 102b, and 102c can be characterized by various types of radio access technologies and / or by transmissions sent to / from various types of base stations (e.g., eNBs and gNBs).

[0051] In other embodiments, base station 114a and WTRUs 102a, 102b, 102c may implement radio technologies such as IEEE 802.11 (i.e., WiFi), IEEE 802.16 (i.e., WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate GSM Evolution (EDGE), GSMEDGE (GERAN), etc.

[0052] Figure 1ABase station 114b can be, for example, a wireless router, a home Node-B, a home eNode-B, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in localized areas such as commercial locations, homes, vehicles, campuses, industrial facilities, air corridors (e.g., for use by drones), roads, etc. In one embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.11 to establish a wireless local area network (WLAN). In another embodiment, base station 114b and WTRUs 102c, 102d can implement radio technologies such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, base station 114b and WTRUs 102c, 102d can utilize cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish picocells or femtocells. Figure 1A As shown, base station 114b can be directly connected to Internet 110. Therefore, base station 114b does not need to access Internet 110 via CN 106 / 115.

[0053] RAN 104 / 113 can communicate with CN 106 / 115, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRU 102a, 102b, 102c, and 102d. Data can have different Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. CN 106 / 115 can provide call control, billing services, location-based services, prepaid calling, internet connectivity, video distribution, etc., and / or perform advanced security functions such as user authentication. Although Figure 1A As not shown, but will be understood, RAN 104 / 113 and / or CN 106 / 115 can communicate directly or indirectly with other RANs using the same RAT as or a different RAT than RAN 104 / 113. For example, in addition to being connected to RAN 104 / 113 which may be utilizing NR radio technology, CN 106 / 115 can also communicate with another RAN (not shown) using GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.

[0054] CN 106 / 115 can also serve as a gateway for WTRU 102a, 102b, 102c, 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing Common Old-Style Telephone Service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) from the TCP / IP Internet Protocol suite. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs, which may use the same RAT as RAN 104 / 113 or a different RAT.

[0055] Some or all of the WTRUs 102a, 102b, 102c, and 102d in communication system 100 may include multi-mode capabilities (e.g., WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different wireless networks via different wireless links). For example, Figure 1A The WTRU 102c shown can be configured to communicate with a base station 114a that can use cellular-based radio technology and with a base station 114b that can use IEEE 802 radio technology.

[0056] Figure 1B This is a system diagram illustrating an exemplary WTRU 102. (See diagram below.) Figure 1B As shown, WTRU 102 may include a processor 118, a transceiver 120, a transmitting / receiving element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and / or other peripheral devices 138, etc. It will be understood that, while remaining consistent with the embodiments, WTRU 102 may include any sub-combination of the foregoing elements.

[0057] Processor 118 can be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. Processor 118 can perform signal encoding, data processing, power control, input / output processing, and / or any other functions that enable WTRU 102 to operate in a wireless environment. Processor 118 can be coupled to transceiver 120, which can be coupled to transmitting / receiving element 122. Although Figure 1B While the processor 118 and transceiver 120 are depicted as separate components, it will be understood that the processor 118 and transceiver 120 can be integrated together in an electronic package or chip.

[0058] Transmitting / receiving element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via air interface 116. For example, in one embodiment, transmitting / receiving element 122 can be an antenna configured to transmit and / or receive RF signals. In embodiments, for example, transmitting / receiving element 122 can be a transmitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, transmitting / receiving element 122 can be configured to transmit and / or receive both RF signals and optical signals. It will be understood that transmitting / receiving element 122 can be configured to transmit and / or receive any combination of wireless signals.

[0059] Although the transmitting / receiving element 122 is in Figure 1B While depicted as a single element, WTRU 102 may include any number of transmit / receive elements 122. More specifically, WTRU 102 may employ MIMO technology. Thus, in one embodiment, WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via air interface 116.

[0060] Transceiver 120 can be configured to modulate signals transmitted by transmitting / receiving element 122 and demodulate signals received by transmitting / receiving element 122. As described above, WTRU 102 can have multi-mode capability. Therefore, transceiver 120 can include multiple transceivers for enabling WTRU 102 to communicate via various RATs (e.g., such as NR and IEEE 802.11).

[0061] The processor 118 of WTRU 102 can be coupled to and receive user input data from: a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) unit or an organic light-emitting diode (OLED) display unit). The processor 118 can also output user data to the speaker / microphone 124, keypad 126, and / or display / touchpad 128. Additionally, the processor 118 can access information and store data from any suitable type of memory, such as non-removable memory 130 and / or removable memory 132. Non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 132 may include a subscriber identity module (SIM) card, memory stick, secure digital storage (SD) card, etc. In other embodiments, the processor 118 can access information and store data from memory not actually located on WTRU 102, such as on a server or home computer (not shown).

[0062] The processor 118 may receive power from the power supply 134 and may be configured to distribute power to other components in the WTRU 102 and / or control power to those other components. The power supply 134 may be any suitable device for powering the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.

[0063] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 102. In addition to or instead of information from the GPS chipset 136, the WTRU 102 may receive location information from base stations (e.g., base stations 114a, 114b) via air interface 116 and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be understood that, while remaining consistent with the embodiments, the WTRU 102 may acquire location information using any suitable location determination method.

[0064] The processor 118 can also be connected to other peripheral devices 138, which may include one or more software and / or hardware modules that provide additional features, functions, and / or wired or wireless connectivity. For example, peripheral devices 138 may include accelerometers, electronic compasses, satellite transceivers, digital cameras (for photos and / or videos), Universal Serial Bus (USB) ports, vibration devices, television transceivers, hands-free headsets, Bluetooth® modules, FM radio units, digital music players, media players, video game player modules, internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripheral devices 138 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors; geolocation sensors; altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors.

[0065] WTRU 102 may include a full-duplex radio, wherein the transmission and reception of some or all of the signals (e.g., associated with a specific subframe of both UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference via hardware (e.g., a choke) or signal processing via a processor (e.g., a separate processor (not shown) or via processor 118). In embodiments, WTRU 102 may include a half-duplex radio, wherein the transmission and reception of some or all of the signals (e.g., associated with a specific subframe of both UL (e.g., for transmission) or downlink (e.g., for reception)) may be concurrent and / or simultaneous.

[0066] Figure 1C This is a system diagram illustrating RAN 104 and CN 106 according to an embodiment. As described above, RAN 104 can employ E-UTRA radio technology to communicate with WTRUs 102a, 102b, and 102c via air interface 116. RAN 104 can also communicate with CN 106.

[0067] RAN 104 may include eNode-Bs 160a, 160b, and 160c, but it will be understood that RAN 104 may include any number of eNode-Bs while remaining consistent with the embodiments. eNode-Bs 160a, 160b, and 160c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, eNode-Bs 160a, 160b, and 160c may implement MIMO technology. Therefore, for example, eNode-B 160a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a.

[0068] Each of the eNode-B 160a, 160b, and 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, etc. Figure 1C As shown, eNode-B 160a, 160b, and 160c can communicate with each other via the X2 interface.

[0069] Figure 1C The CN 106 shown may include a Mobility Management Entity (MME) 162, a Serving Gateway (SGW) 164, and a Packet Data Network (PDN) Gateway (or PGW) 166. While each of the foregoing elements is described as part of CN 106, it will be understood that any of these elements may be owned and / or operated by an entity other than a CN operator.

[0070] The MME 162 can connect to each of the eNode-Bs 162a, 162b, and 162c in RAN 104 via the S1 interface and can act as a control node. For example, the MME 162 can be responsible for authenticating users of WTRUs 102a, 102b, and 102c, activating / deactivating bearers, selecting a specific serving gateway during the initial attachment of WTRUs 102a, 102b, and 102c, etc. The MME 162 can provide control plane functions for handover between RAN 104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.

[0071] The SGW 164 can connect to each of the eNode Bs 160a, 160b, and 160c in RAN 104 via the S1 interface. The SGW 164 can typically route and forward user data packets to or from WTRUs 102a, 102b, and 102c. The SGW 164 can perform other functions, such as anchoring the user plane during eNode-B handover, triggering paging when DL data is available to WTRUs 102a, 102b, and 102c, and managing and storing the context of WTRUs 102a, 102b, and 102c.

[0072] SGW 164 can be connected to PGW 166, which can provide WTRU 102a, 102b, 102c with access to packet-switched networks (such as Internet 110) to facilitate communication between WTRU 102a, 102b, 102c and IP-enabled devices.

[0073] CN 106 can facilitate communication with other networks. For example, CN 106 can provide WTRUs 102a, 102b, and 102c with access to circuit-switched networks (such as PSTN 108) to facilitate communication between WTRUs 102a, 102b, and 102c and traditional terrestrial line communication equipment. For example, CN 106 may include an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) or be able to communicate with such an IP gateway as an interface between CN 106 and PSTN 108. Additionally, CN 106 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0074] Despite WTRU in Figures 1A to 1D While described as a wireless terminal, it is envisioned that, in some representative embodiments, such a terminal may (e.g., temporarily or permanently) use a wired communication interface with a communication network.

[0075] In a representative embodiment, the other network 112 may be a WLAN.

[0076] A WLAN in Infrastructure Basic Services Set (BSS) mode may have an Access Point (AP) for the BSS and one or more Stations (STAs) associated with the AP. The AP may have an interface to a Distribution System (DS) or another type of wired / wireless network that loads traffic into and / or loads traffic out of the BSS. Traffic originating outside the BSS destined for a STA can be delivered to the AP. Traffic from a STA to a destination outside the BSS can be transmitted to the AP for delivery to the appropriate destination. Traffic between STAs within the BSS can be transmitted via the AP, for example, where a source STA can transmit traffic to the AP, and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS can be considered and / or referred to as point-to-point traffic. Point-to-point traffic can be transmitted between a source STA and a destination STA using a Direct Link Setup (DLS) (e.g., transmitted directly between them). In some representative embodiments, the DLS may use 802.11e DLS or 802.11z Tunneled DLS (TDLS). WLANs using the Independent BSS (IBSS) mode can function without access points (APs), and STAs within the IBSS or using the IBSS (e.g., all STAs) can communicate directly with each other. The IBSS communication mode may sometimes be referred to as a "self-organizing" communication mode in this document.

[0077] When operating in 802.11ac infrastructure mode or a similar mode, the AP can transmit beacons on a fixed channel, such as the primary channel. The primary channel can be of fixed width (e.g., a bandwidth of 20 MHz) or dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by the STA to establish a connection with the AP. In some representative embodiments, Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) can be implemented, for example, in an 802.11 system. For CSMA / CA, each STA, including the AP, can sense the primary channel. If a particular STA senses / detects that the primary signal is busy and / or determines that the primary signal is busy, that particular STA can back off. In a given BSS, at any given time, only one STA (e.g., only one station) can transmit.

[0078] High-throughput (HT) STAs can communicate using a 40 MHz wide channel, for example, by combining a primary 20 MHz channel with adjacent or non-adjacent 20 MHz channels.

[0079] Very High Throughput (VHT) STAs can support channels with widths of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz. 40 MHz and / or 80 MHz channels can be formed by combining consecutive 20 MHz channels. A 160 MHz channel can be formed by combining eight consecutive 20 MHz channels, or by combining two non-consecutive 80 MHz channels, which can be referred to as an 80+80 configuration. In the 80+80 configuration, after channel coding, data can be passed through a fragment parser, which divides the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time-domain processing can be performed on each stream separately. The streams can be mapped onto the two 80 MHz channels, and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the above operations of the 80+80 configuration can be reversed, and the combined data can be sent to the Media Access Control (MAC).

[0080] 802.11af and 802.11ah support operating modes below 1 GHz. The channel operating bandwidth and carrier are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV Blank (TVWS) spectrum, and 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to representative embodiments, 802.11ah can support instrument-type control / machine-type communication, such as MTC devices in macro coverage areas. MTC devices may have certain capabilities, such as limited capabilities, including support (e.g., only support) certain and / or limited bandwidths. MTC devices may include batteries with a battery life exceeding a threshold (e.g., to maintain a very long battery life).

[0081] WLAN systems that can support multiple channels and channel bandwidths (such as 802.11n, 802.11ac, 802.11af, and 802.11ah) include a channel that can be designated as the primary channel. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by the STAs operating in the BSS that support the minimum bandwidth operating mode. In the 802.11ah example, for STAs that support (e.g., only support) the 1 MHz mode (e.g., MTC type devices), the primary channel can be 1 MHz wide, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier Sense and / or Network Allocation Vector (NAV) settings can depend on the status of the primary channel. If the primary channel is busy, for example, due to STAs (which only support the 1 MHz operating mode) transmitting to the AP, the entire available band may be considered busy even if most of the band remains idle and potentially available.

[0082] In the United States, the available frequency band for 802.11ah is 902 MHz to 928 MHz. In South Korea, the available frequency band is 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is 916.5 MHz to 927.5 MHz. The total available bandwidth for 802.11ah is 6 MHz to 26 MHz, depending on the country code.

[0083] Figure 1D This is a system diagram illustrating RAN 113 and CN 115 according to one embodiment. As described above, RAN 113 may employ NR radio technology to communicate with WTRUs 102a, 102b, and 102c via air interface 116. RAN 113 may also communicate with CN 115.

[0084] RAN 113 may include gNBs 180a, 180b, and 180c, but it will be understood that RAN 113 may include any number of gNBs while remaining consistent with the embodiments. gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with WTRUs 102a, 102b, and 102c via air interface 116. In one embodiment, gNBs 180a, 180b, and 180c may implement MIMO technology. For example, gNBs 180a and 180b may utilize beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, and 180c. Thus, for example, gNB 180a may use multiple antennas to transmit radio signals to and / or receive radio signals from WTRU 102a. In one embodiment, gNBs 180a, 180b, and 180c may implement carrier aggregation technology. For example, gNB 180a can transmit multiple component carriers to WTRU 102a (not shown). A subset of these component carriers may be located on unlicensed spectrum, while the remaining component carriers may be located on licensed spectrum. In one embodiment, gNBs 180a, 180b, and 180c may implement Coordinated Multipoint (CoMP) technology. For example, WTRU 102a can receive coordinated transmissions from gNBs 180a and 180b (and / or gNB 180c).

[0085] WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using transmissions associated with a scalable digital architecture. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing can be varied for different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using subframes of various lengths or scalable lengths, or transmission time intervals (TTIs) (e.g., containing different numbers of OFDM symbols and / or absolute times of varying durations).

[0086] gNBs 180a, 180b, and 180c can be configured to communicate with WTRUs 102a, 102b, and 102c in standalone and / or non-standalone configurations. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c without accessing other RANs (e.g., eNode-Bs 160a, 160b, and 160c). In standalone configuration, WTRUs 102a, 102b, and 102c can use one or more of gNBs 180a, 180b, and 180c as mobile anchors. In standalone configuration, WTRUs 102a, 102b, and 102c can communicate with gNBs 180a, 180b, and 180c using signals in unlicensed frequency bands. In a non-standalone configuration, WTRUs 102a, 102b, and 102c can communicate / connect with gNBs 180a, 180b, and 180c while also communicating / connecting with another RAN (such as eNode-Bs 160a, 160b, and 160c). For example, WTRUs 102a, 102b, and 102c can implement DC principles to communicate substantially simultaneously with one or more gNBs 180a, 180b, and 180c and one or more eNode-Bs 160a, 160b, and 160c. In a non-standalone configuration, eNode-Bs 160a, 160b, and 160c can act as mobile anchors for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c can provide additional coverage and / or throughput to serve WTRUs 102a, 102b, and 102c.

[0087] Each of gNBs 180a, 180b, and 180c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support for network slicing, dual connectivity, interoperability between NR and E-UTRA, routing of user plane data to User Plane Functions (UPF) 184a and 184b, routing of control plane information to Access and Mobility Management Functions (AMF) 182a and 182b, etc. Figure 1D As shown, gNB 180a, 180b, and 180c can communicate with each other via the Xn interface.

[0088] Figure 1DThe CN 115 shown may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. While each of the foregoing elements is described as part of the CN 115, it will be understood that any of these elements may be owned and / or operated by an entity other than a CN operator.

[0089] AMF 182a and 182b can connect to one or more of the gNBs 180a, 180b, and 180c in RAN 113 via the N2 interface and can act as control nodes. For example, AMF 182a and 182b can be responsible for authenticating users of WTRU 102a, 102b, and 102c, supporting network slicing (e.g., handling different PDU sessions with different requirements), selecting specific SMFs 183a and 183b, managing registration areas, terminating NAS signaling, mobility management, etc. AMF 182a and 182b can use network slicing to customize CN support for WTRU 102a, 102b, and 102c based on the service types being utilized by WTRU 102a, 102b, and 102c. For example, different network slices can be established for different use cases, such as services dependent on Ultra Reliable Low Latency (URLLC) access, services dependent on Enhanced Massive Mobile Broadband (eMBB) access, services for Machine Type Communication (MTC) access, etc. AMF 162 can provide control plane functions for handover between RAN 113 and other RANs (not shown) that employ other radio technologies (such as LTE, LTE-A, LTE-A Pro) and / or non-3GPP access technologies (such as WiFi).

[0090] SMFs 183a and 183b can connect to AMFs 182a and 182b in CN 115 via the N11 interface. SMFs 183a and 183b can also connect to UPFs 184a and 184b in CN 115 via the N4 interface. SMFs 183a and 183b can select and control UPFs 184a and 184b, and configure traffic routing through UPFs 184a and 184b. SMFs 183a and 183b can perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, or Ethernet-based.

[0091] UPF 184a and 184b can connect via the N3 interface to one or more of the gNBs 180a, 180b, and 180c in RAN 113. These gNBs can provide WTRU 102a, 102b, and 102c with access to packet-switched networks (such as the Internet 110) to facilitate communication between WTRU 102a, 102b, and 102c and IP-enabled devices. UPF 184 and 184b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multihomed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.

[0092] CN 115 can facilitate communication with other networks. For example, CN 115 may include or be able to communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN 115 and PSTN 108. Additionally, CN 115 can provide WTRUs 102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRUs 102a, 102b, and 102c can be connected to DN 185a and 185b via UPF 184a and 184b through the N3 interface to UPF 184a and 184b and the N6 interface between UPF 184a and 184b and local data networks (DNs) 185a and 185b.

[0093] Given Figures 1A to 1D and Figures 1A to 1D The corresponding descriptions can be performed by one or more emulation devices (not shown) that perform one or more of the functions described herein with respect to: WTRU 102a to 102d, base stations 114a to 114b, eNode-B 160a to 160c, MME 162, SGW 164, PGW 166, gNB 180a to 180c, AMF 182a to 182b, UPF 184a to 184b, SMF 183a to 183b, DN 185a to 185b, and / or any other devices described herein. An emulation device can be one or more devices configured to emulate one or more of the functions described herein. For example, an emulation device can be used to test other devices and / or simulate network and / or WTRU functions.

[0094] Simulation devices can be designed to perform one or more tests on other devices in laboratory and / or carrier network environments. For example, one or more simulation devices may perform one or more functions when fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more simulation devices may perform one or more functions when temporarily implemented / deployed as part of a wired and / or wireless communication network. Simulation devices may be directly coupled to another device for testing purposes and / or use over-the-air wireless communication to perform tests.

[0095] One or more simulation devices may perform one or more functions when implemented / deployed without being part of a wired and / or wireless communication network. For example, simulation devices may be used to test scenarios in a laboratory and / or undeployed (e.g., tested) wired and / or wireless communication networks to perform tests on one or more components. One or more simulation devices may be test equipment. Simulation devices may transmit and / or receive data using direct RF connections and / or wireless communication via RF circuitry (e.g., which may include one or more antennas).

[0096] This application describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are described in a specific manner and, at least for the purpose of illustrating the individual features, are generally described in a way that may sound limiting. However, this is for the purpose of clarity of description and does not limit the application or scope of these aspects. In fact, all the different aspects can be combined and interchanged to provide other aspects. Furthermore, these aspects can also be combined and interchanged with aspects described in previous documents.

[0097] The aspects described and envisioned in this application can be implemented in many different forms. Figures 5 to 13 Some examples can be provided, but other examples are envisioned. Figures 5 to 13 The discussion does not limit the breadth of implementations. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatuses, computer-readable storage media having instructions thereon stored thereon for encoding or decoding video data according to any of the methods, and / or computer-readable storage media having bitstreams generated according to any of the methods stored thereon.

[0098] In this application, the terms “reconstruction” and “decoding” are used interchangeably, the terms “pixel” and “sample” are used interchangeably, and the terms “image”, “picture” and “frame” are used interchangeably.

[0099] This document describes various methods, each of which includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, in various examples, terms such as "first," "second," etc., may be used to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding." Unless specifically required, the use of such terms does not imply a sequence of operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding, but can occur, for example, before, during, or in a time period overlapping with the second decoding.

[0100] The various methods and other aspects described in this application can be used to modify, for example... Figure 2 and Figure 3 The illustrated video encoder 200 and decoder 300 modules include, for example, a decoding module. Furthermore, the subject matter disclosed herein can be applied to, for example, any type, format, or version of video encoding, whether described in standards or recommendations (whether pre-existing or future-developed) and any extensions to such standards and recommendations. Unless otherwise stated or technically excluded, these aspects described in this application may be used alone or in combination.

[0101] Various numerical values, such as 1, 2, 4, 7, 8, 16, 32, 64, etc., are used in the examples described in this application. These and other specific values ​​are for illustrative purposes only, and the aspects described are not limited to these specific values.

[0102] Figure 2 This is a diagram illustrating an exemplary video encoder 200. Variations of the exemplary encoder 200 are envisioned, but for clarity, the encoder 200 is described below without describing all anticipated variations.

[0103] Before being encoded, the video sequence may undergo pre-coding processing 201, such as applying a color transformation to the input color image (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input image components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata (e.g., which may include film grain parameters determined through pre-processing as described herein) may be associated with the pre-processing and appended to the bitstream.

[0104] In encoder 200, the image is encoded by encoder elements, as described below. The image to be encoded is partitioned (202) and processed in units, for example, coding units (CUs). Each unit is encoded using, for example, an intra-frame mode or an inter-frame mode. When a unit is encoded in intra-frame mode, intra-frame prediction (260) is performed. In inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which mode, intra-frame mode or inter-frame mode, to use to encode the unit and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (210) the prediction block from the original image block.

[0105] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, motion vectors, and other syntactic elements are entropy encoded (245) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., decode the residual directly without applying the transform or quantization process.

[0106] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (255) to reconstruct the image block. An in-loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (Sample Adaptive Shift) filtering, thereby reducing coding artifacts. The filtered image is stored in a reference image buffer 280.

[0107] Figure 3 This is a diagram illustrating an example of a video decoder. In the exemplary decoder 300, the bitstream is decoded by decoder elements, as described below. The video decoder 300 typically performs operations similar to... Figure 2 The encoding process shown is the reverse of the decoding process. Encoder 200 typically also performs video decoding as part of the encoded video data.

[0108] Specifically, the decoder's input includes a video bitstream, which can be generated by the video encoder 200. First, entropy decoding (330) is performed on the bitstream to obtain transform coefficients, motion vectors, and other encoded information. Image partitioning information indicates how the image should be partitioned. Therefore, the decoder can partition (335) the image based on the decoded image partitioning information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (355) to reconstruct the image blocks. Prediction blocks (370) can be obtained from intra-frame prediction (360) or motion-compensated prediction (i.e., inter-frame prediction) (375). An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference image buffer 380.

[0109] The decoded image can also undergo post-decoding processing 385, such as inverse color transformation (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping, which performs the inverse of the remapping process performed in pre-encoding processing 201. Post-decoding processing can use metadata derived in pre-encoding processing and signaled in the bitstream. In the example, the decoded image (e.g., after applying an in-loop filter (365) and / or, in the case of post-decoding processing, after post-decoding processing (385)) can be transmitted to a display device for presentation to the user.

[0110] Figure 4 This is a diagram illustrating examples of systems in which the various aspects and examples described herein may be implemented. System 400 may be embodied as a device including the various components described below and configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 400 may be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one example, the processing elements and encoder / decoder elements of system 400 are distributed across multiple ICs and / or discrete components. In various examples, system 400 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various examples, system 400 is configured to implement one or more aspects described in this document.

[0111] System 400 includes at least one processor 410 configured to execute instructions loaded thereon to implement various aspects, such as those described in this document. Processor 410 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 400 includes at least one memory 420 (e.g., a volatile memory device and / or a non-volatile memory device). System 400 includes a storage device 440 that may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 440 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0112] System 400 includes an encoder / decoder module 430 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 430 may include its own processor and memory. The encoder / decoder module 430 represents a module that can be included in a device to perform encoding and / or decoding functions. It is well known that a device may include one or both encoding and decoding modules. Alternatively, the encoder / decoder module 430 may be implemented as a separate element of system 400, or it may be incorporated within processor 410 as a combination of hardware and software known to those skilled in the art.

[0113] Program code to be loaded onto processor 410 or encoder / decoder 430 to execute the various aspects described in this document may be stored in storage device 440 and subsequently loaded onto memory 420 for execution by processor 410. Depending on various examples, one or more of processor 410, memory 420, storage device 440, and encoder / decoder module 430 may store one or more of various items during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from equations, formulas, operations, and operational logic processing.

[0114] In some examples, the memory within processor 410 and / or encoder / decoder module 430 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other examples, external memory (e.g., the processing device could be processor 410 or encoder / decoder module 430) is used for one or more of these functions. External memory could be memory 420 and / or storage device 440, such as volatile memory and / or non-volatile flash memory. In several examples, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one example, fast external volatile memory such as RAM is used as working memory for video encoding and decoding operations.

[0115] Input to the components of system 400 can be provided through various input devices, as indicated in box 445. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster, (ii) a component (COMP) input terminal (or a set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 4 Other examples not shown include composite video.

[0116] In various examples, the input device of block 445 has corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting the signal band to a band), (ii) down-converting the selected signal, (iii) further band-limiting to a narrower band to select (e.g.,) a signal band that may be referred to as a channel in some examples), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and / or (vi) demultiplexing to select the desired data packet stream. The RF section of various examples includes one or more elements performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners performing various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box example, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and filtering the signal again to the desired frequency band. Various examples rearrange the order of the aforementioned (and other) components, remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as inserting amplifiers and analog-to-digital converters. In various examples, the RF section includes an antenna.

[0117] USB and / or HDMI terminals may include corresponding interface processors for connecting system 400 to other electronic devices via USB and / or HDMI connections. It should be understood that aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within processor 410, as needed. Similarly, aspects of USB or HDMI interface processing may be implemented, as needed, within a separate interface IC or within processor 410. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 410 and an encoder / decoder 430 operating in conjunction with memory and storage elements, to process the data stream as needed for presentation on an output device.

[0118] Various components of system 400 can be housed within an integrated housing. Within the integrated housing, the various components can be interconnected and transmit data between them using suitable connection devices 425 (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards).

[0119] System 400 includes a communication interface 450 that enables communication with other devices via a communication channel 460. The communication interface 450 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 460. The communication interface 450 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 460 may be implemented, for example, within a wired and / or wireless medium.

[0120] In various examples, data is streamed to or otherwise provided to system 400 using a wireless network, such as a Wi-Fi network, for example, IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). In these examples, the Wi-Fi signal is received via a communication channel 460 and a communication interface 450 adapted for Wi-Fi communication. The communication channel 460 in these examples is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other examples use a set-top box to provide streaming data to system 400, delivering data via an HDMI connection to input block 445. Still other examples use an RF connection to input block 445 to provide streaming data to system 400. As indicated above, various examples provide data in a non-streaming manner. Additionally, various examples use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth® networks.

[0121] System 400 can provide output signals to various output devices, including display 475, speaker 485, and other peripheral devices 495. Various examples of display 475 include one or more of, for example, touchscreen displays, organic light-emitting diode (OLED) displays, curved displays, and / or foldable displays. Display 475 can be used in televisions, tablets, laptops, mobile phones, or other devices. Display 475 can also be integrated with other components (e.g., in a smartphone) or standalone (e.g., an external monitor for a laptop computer). In various examples, other peripheral devices 495 include one or more of standalone digital video discs (or digital multifunction discs) (DVDs, for both terms), disc players, stereo systems, and / or lighting systems. Various examples use one or more peripheral devices 495 that provide functionality based on the output of system 400. For example, a disc player performs the function of playing the output of system 400.

[0122] In various examples, signaling (such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that implement device-to-device control with or without user intervention) is used to transmit control signals between system 400 and display 475, speaker 485, or other peripheral devices 495. Output devices may be communicatively coupled to system 400 via dedicated connections through corresponding interfaces (470, 480, and 490). Alternatively, output devices may be connected to system 400 via communication interface 450 using communication channel 460. Display 475 and speaker 485 may be integrated into a single unit along with other components of system 400 in electronic devices such as televisions. In various examples, display interface 470 includes a display driver, such as a timing controller (TCon) chip.

[0123] For example, if the RF input section 445 is part of a separate set-top box, the display 475 and speaker 485 can alternatively be separate from one or more of the other components. In various examples where the display 475 and speaker 485 are external components, the output signal can be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0124] The example can be implemented by processor 410 or by computer software implemented by hardware or a combination of hardware and software. As a non-limiting example, the example can be implemented by one or more integrated circuits. As a non-limiting example, memory 420 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 410 can be of any type suitable for the technical environment and can encompass one or more microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0125] Various implementations involve decoding. As used herein, “decoding” can encompass all or part of a process performed, for example, on a received encoded sequence to produce a final output suitable for display. In various examples, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various examples, such a process also (or alternatively) includes processes performed by a decoder of various embodiments described herein, such as receiving an avatar data structure from an encoder and an indication of the file format contained in the avatar data structure; determining, based on the indication, that the avatar data structure includes a first type of geometry-related information and a second type of geometry-related information; decoding the avatar data structure based on the first type of geometry-related information and the second type of geometry-related information to generate a descriptive avatar media file; rendering the avatar based on the descriptive avatar media file; and so on.

[0126] As further examples, in one example, "decoding" refers only to entropy decoding; in another example, "decoding" refers only to differential decoding; and in yet another example, "decoding" refers to a combination of entropy decoding and differential decoding. Based on the specific context of the description, it will be clear whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to refer to a broader decoding process, and it is believed that those skilled in the art will readily understand this.

[0127] Various implementations involve encoding. Similar to the discussion above regarding “decoding,” “encoding” as used herein can include, for example, all or part of the processing performed on an input video sequence to produce an encoded bitstream. In various examples, such processes include one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various examples, such processes also (or alternatively) include processes performed by an encoder of various embodiments described herein, such as receiving a descriptive avatar media file containing first-type and second-type geometrically relevant information; encoding the descriptive avatar media file based on the first-type and second-type geometrically relevant information to generate an avatar data structure; sending the avatar data structure and an indication of the file format contained in the avatar data structure to a decoder; and so on.

[0128] As further examples, in one example, "encoding" refers only to entropy encoding; in another example, "encoding" refers only to differential encoding; and in yet another example, "encoding" refers to a combination of entropy encoding and differential encoding. Based on the specific context of the description, it will be clear whether the phrase "encoding process" is intended to specifically refer to a subset of operations or to refer to a broader encoding process, and it is believed that those skilled in the art will readily understand this.

[0129] It should be noted that the syntactic elements used in this paper (e.g., the encoding syntax for intensity ranges, granular parameters, block offsets, scaling factors, etc.) are descriptive terms. Therefore, they do not preclude the use of other syntactic element names.

[0130] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / device.

[0131] The embodiments and aspects described herein can be implemented, for example, in methods or processes, apparatuses, software programs, data streams, or signals. Even if discussed only in the context of a single embodiment (e.g., discussed only as a method), embodiments of the discussed features can also be implemented in other forms (e.g., apparatuses or programs). Apparatuses can be implemented, for example, with suitable hardware, software, and firmware. Methods can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0132] The references to “an example” or “example” or “an implementation” or “implementation”, and their variations, mean that a particular feature, structure, characteristic, etc., described in connection with the example is included in at least one example. Therefore, the phrases “in an example” or “in a sample” or “in an implementation” or “in an implementation”, and any other variations, appearing throughout this application, do not necessarily refer to the same example.

[0133] Additionally, this application may relate to "determining" various types of information. Determining information may include one or more of the following: for example, estimated information, calculated information, predicted information, or information retrieved from memory. Obtaining may include receiving, retrieving, constructing, generating, and / or determining.

[0134] Furthermore, this application may relate to "accessing" various types of information. Accessing information may include one or more of the following: for example, receiving information, retrieving information (e.g., retrieving from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0135] Additionally, this application may relate to "receiving" various types of information. Like "access," the intent to receive is a broad term. Receiving information may include one or more of the following: for example, accessing information or retrieving information (e.g., retrieving from memory). Furthermore, "receiving" is generally referred to in one or more ways during operation, such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0136] It should be understood that, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one” is intended to cover selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As yet another example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” this wording is intended to include selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or selecting all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to a large number of listed items.

[0137] Furthermore, as used herein, the term "signaling" specifically refers to instructing the corresponding decoder to do something. Encoder signals may include, for example, the number of intensity intervals, the number of model values, particle parameters, particle identification, scaling factors, etc. In this way, in the examples, the same parameters are used on both the encoder and decoder sides. Thus, for example, the encoder can send (explicitly signal) a specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, signaling can be used without sending (implicitly signaling) to simply allow the decoder to know and select the specific parameter. Bit savings are implemented in various examples by avoiding the transmission of any actual functionality. It should be understood that signaling can be implemented in various ways. For example, in various examples, one or more syntactic elements, flags, etc., are used to send information to the corresponding decoder. Although the verb form of the word "signal" was mentioned above, the word "signal" can also be used as a noun in this article.

[0138] As will be apparent to those skilled in the art, implementations can produce various signals formatted to carry, for example, signals that can be stored or transmitted. Information may include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, a signal may be formatted to carry a bit stream of the described example. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of a spectrum) or baseband signals. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. It is well known that signals can be transmitted via a variety of different wired or wireless links. Signals may be stored on, or accessed or received from, a processor-readable medium.

[0139] This document describes numerous examples. Features of the examples may be provided individually or in any combination across various claim classes and types. Furthermore, examples may include one or more of the features, devices, or aspects described herein, individually or in any combination across various claim classes and types. For example, the features described herein may be implemented in a bitstream or signal that includes information generated as described herein. This information may allow a decoder to decode the bitstream, the encoder, bitstream, and / or decoder being any of the embodiments described. For example, the features described herein may be implemented by creating and / or transmitting and / or receiving and / or decoding a bitstream or signal. For example, the features described herein may be implemented by a method, process, apparatus, medium storing instructions, medium storing data, or signal. For example, the features described herein may be implemented by a TV, set-top box, mobile phone, tablet computer, or other electronic device performing decoding. The TV, set-top box, mobile phone, tablet computer, or other electronic device may display (e.g., using a monitor, screen, or other type of display) a resulting image (e.g., an image reconstructed from the residual of a video bitstream). The TV, set-top box, mobile phone, tablet computer, or other electronic device may receive a signal including an encoded image and perform decoding.

[0140] The features described herein relate to the compression and representation of avatar and / or user information that allows for interoperability between systems and applications used for sending and delivering digital content. Avatar and / or user geometry formats can be compressed and sent between systems, devices, and / or applications, etc.

[0141] Digital humans can be represented in several forms (e.g., meshes, volumes, voxels, point clouds, images, videos, sounds, etc.). These representations can be used to create and formalize digital media content. Templates capturing human representations (e.g., all the different representations) can be formalized (e.g., to allow interoperable representations). These templates can be (e.g., with skeletal structures) general human models, object-specific models, statistical shape models, and / or metadata representing individual social attributes. Such templates can generate (e.g., accurately provide) human images as an initial stage. The human images in this initial stage can be statistically represented in different ways.

[0142] Synthetic 3D models can be used to create digital human representations. These representations can be manipulated and / or customized to represent specific human anatomy, visual effects, and / or social parameters. Synthetic representations can facilitate the generation of animations in immersive real-world environments (e.g., if all parameters are initially known). Synthetic models can also aid in the generalization and stylization of appearances (e.g., professionals can use them for animation and streaming).

[0143] 3D composite models can be very good representations. 3D composite models may lack standard definitions and / or structures that can be used and matched by another (e.g., any other) digital human representation (e.g., user-specific details such as virtual identity, social status or input device control, semantic structure of geometric properties, semantic structure of shape and skeletal anatomy, and / or semantics of animation parameters, etc.).

[0144] The standardization of digital human content may disregard human social behavior, interaction with environmental information, and / or social and human privacy issues (e.g., these can be public characteristics that can be adopted within social technologies in the real world). In real-time distribution or communication use cases, the encoding and / or data structure of digital human content may not be defined. In some systems (e.g., proprietary systems), the data may be known. In other systems (e.g., open systems), the data format may have been specified and standardized (e.g., so that the data can interoperate between different systems).

[0145] This document provides exemplary media content formats for avatars and / or user geometric representations.

[0146] This article provides features associated with avatar media representation.

[0147] Exemplary representations of avatars can be compatible with any technology (e.g., technical systems used for VR, AR, streaming, gaming, interaction, collaboration, or communication). This representation can consider the delivery of digital human content (e.g., 2D / 3D video, or images containing bodies, faces, or voices, 3D technologies for rendering human attribute assets, animate or manipulate human attribute assets, or any other system containing social and contextual information about a real user or avatar, such as those in the social and communication media industry).

[0148] This document provides an exemplary format for avatared geometric media. It also provides exemplary meanings, encoding schemes, and uses associated with the avatared geometric media format.

[0149] The format described in this article can be similar to JSON. This format can be compatible with current MPEG standards (e.g., by providing means for compressing and transmitting media content). The meaning and purpose can be general, and / or encoded using other formats (e.g., OpenXR, XML, USD).

[0150] This article provides an exemplary framework for embodying geometric media. This article provides an exemplary framework architecture for embodying geometric media.

[0151] Figure 5 An exemplary geometry avatar media framework overview is shown (e.g., top-level component). Figure 5 An overview of exemplary components present in the Geometry Avatar Media Codec representation is provided. Exemplary components include "Semantics", "Geometry Metadata", "Skeleton", "Landmarks", and "AI". Components (e.g., each component) may represent a higher-level specification of the Geometry Avatar Media Codec.

[0152] Geometric information related to digital humans (e.g., avatars) (e.g., basic geometric information) can be encoded. The encoded geometric information can be represented as a stream between systems. Table 1 shows exemplary information (e.g., higher-level information) that can be used for streaming and representing the geometry of avatar media.

[0153]

[0154] Table 1 : Geometric avatars media codec formats.

[0155] This article provides exemplary semantics.

[0156] The semantics presented in this article can represent the geometry-related semantics of the avatar (e.g., properties that define the geometry-related semantics of the avatar). Geometry-related semantics can take the form of connecting or labeling vertices, labeling skeletal structures, labeling 3D / 2D markers, and / or contextual information about the geometry of the avatar's medium. The function "processSemantics()" can be used to retrieve this information.

[0157] Signaling can be used to inform the decoder that geometry-related semantic information is available for reading. This signaling can take one or more forms. Exemplary forms may include the encoding standard of the geometry information (e.g., MPEG-I SD, MPEG-AI, MPEG V-DMC, L-PCC, MPEG 4, H-anim, and / or JVET, etc.).

[0158]

[0159] Table 2: Description of the semantics of "geometry" if the signal is 1.

[0160]

[0161] Table 3: If the signal is 1, semantics "Semantic description."

[0162] This article provides features associated with geometric metadata.

[0163] Geometric metadata for an avatar can represent the geometry-related metadata of an avatar asset (e.g., attributes defining geometry-related metadata). Geometric metadata can be presented as connected or disconnected vertices. The function "processGeometryMetadata()" can be used to retrieve this information.

[0164] Signaling can be used to inform the decoder that geometry-related metadata information is available for reading. This signaling can take one or more forms. Exemplary forms may include the encoding standard of the geometry information (e.g., MPEG-I SD, MPEG-AI MPEGV-DMC, L-PCC, MPEG 4, H-anim, and / or JVET, etc.).

[0165]

[0166] Table 4: Description of the semantics of "geometry" if the signal is 1.

[0167]

[0168] Table 5: Description of the semantics of "geometryMetadata" if the signal is 1.

[0169] This article provides features associated with the skeleton of the avatar.

[0170] The skeleton of an avatar can represent the avatar's skeletal structure (e.g., defining the properties of the avatar's skeletal structure). For example, the skeletal structure can represent joint positions, rotations, and / or hierarchical structures. The function "processSkeleton()" can be used to retrieve this information.

[0171] Signaling can be used to inform the decoder that skeleton information is available for reading. This signaling can take one or more forms. Exemplary forms may include coding standards for geometric and complementary information (e.g., MPEG-I SD, MPEG-AI, MPEG V-DMC, L-PCC, MPEG 4, H-anim, and / or JVET, etc.).

[0172]

[0173] Table 6: Description of the semantics of "geometry" if the signal is 1.

[0174]

[0175] Table 7: Description of the semantics of "skeleton" if the signal is 1.

[0176] This article provides features associated with landmarks.

[0177] Landmarks in an avatar can represent the location of a 2D or 3D landmark in the avatar (e.g., defining properties of the 2D or 3D landmark location). For example, landmarks can include facial keymarks. The function "processLandmarks()" can be used to retrieve this information.

[0178] Signaling can be used to inform the decoder that landmark information is available for reading. This signaling can take one or more forms. Exemplary forms may include coding standards for geometric and complementary information (e.g., MPEG-I SD, MPEG-AI, MPEG V-DMC, L-PCC, MPEG 4, H-anim, and / or JVET, etc.).

[0179]

[0180] Table 8: Description of the semantics of "geometry" if the signal is 1.

[0181]

[0182] Table 9: Description of the semantics of "landmarks" if the signal is 1.

[0183] This article provides features associated with artificial intelligence (AI).

[0184] AI that embodies geometric media can represent the neural representation of the embodiment (e.g., define the properties of the neural representation of the embodiment). For example, the neural representation can include a description of a neural network architecture or a pre-trained neural network model with respect to the properties of the embodiment media representation.

[0185] AI-generated information related to the geometry of avatar media can take different forms.

[0186] The geometry of the avatar can be represented by latent codes from an autoencoder neural network. An autoencoder neural network can include an encoder neural network and a decoder neural network. The encoder neural network can be fed with some representation of the avatar's geometry (e.g., a mesh or point cloud) and output latent codes. The latent codes can provide a more compact representation of the avatar's geometry as a vector of floating-point numbers. The decoder can reconstruct the input representation of the avatar's geometry from the latent codes.

[0187] The geometry of an avatar can be represented using a neural network. The 3D position in a coordinate system associated with the avatar and / or the viewing direction represented by two angles in that coordinate system can be used as input to the neural network. The network can output an opacity value, indicating whether the avatar occupies a position at the input 3D position and / or the color of a portion of the avatar at that position (e.g., when viewed from the input direction). The neural network can be used to reconstruct the complete 3D geometry of the avatar (e.g., by querying a discretized set of input positions within its enclosing volume for the avatar's occupancy and color at those positions). The function "processAI()" can be used to retrieve this information.

[0188] The purpose of this signal is to inform the decoder that AI information is available for reading. This signaling can take one or more forms. Exemplary forms may include encoding standards used for AI-generated information.

[0189]

[0190] Table 10: Description of the semantics of "geometry" if the signal is 1.

[0191]

[0192] Table 11: Description of the semantics of "AI" if the signal is 1.

[0193] The processGeometry() function can be used to process one or more (e.g., all) subcomponents. For example, processGeometry() can invoke subcomponent processing of any subcomponent described herein (e.g., via the syntax in Table 1).

[0194] An exemplary device may have a processor configured to perform one or more actions. The device (e.g., an encoder, such as an avatar geometry encoder) may receive a descriptive avatar media file. The device may extract a first type of geometry-related avatar information and a second type of geometry-related avatar information from the descriptive avatar media file. The device may encode the first type of geometry-related avatar information and the second type of geometry-related avatar information into an avatar data structure. The device may send the avatar data structure, along with an indication of the file format contained within the avatar data structure, to a decoder.

[0195] At least one of the first type of geometrically related avatar information and the second type of geometrically related avatar information can be: avatar semantic information; avatar geometric metadata; avatar skeletal information; avatar landmark information; or artificial intelligence avatar information.

[0196] The avatar data structure can include binary files. The device can perform binary compression on both types of geometry-related information (first and second types) to generate compressed information. The device can then package this compressed information to generate a binary file.

[0197] The avatar data structure can include human-readable files. The device can encode first-type and second-type geometry-related avatar information into an avatar data structure by generating human-readable files based on first-type and second-type geometry-related information.

[0198] The first type of geometry-related avatar information can be associated with a first file format contained in the avatar data structure. The second type of geometry-related avatar information can be associated with a second file format contained in the avatar data structure. The first file format can be different from the second file format.

[0199] An exemplary device (e.g., a decoder, such as an avatar geometry decoder) can receive an avatar data structure and an indication of the file format contained within the data structure from an encoder. The device can decode the avatar data structure. The device can determine a first type of geometry-related avatar information and a second type of geometry-related avatar information based on the decoded avatar data structure. The device can generate a descriptive avatar media file based on the first type of geometry-related avatar information, the second type of geometry-related information, and the indicated file format.

[0200] At least one of the first type of geometrically related avatar information and the second type of geometrically related avatar information may include: avatar semantic information; avatar geometric metadata; avatar skeletal information; avatar landmark information; or artificial intelligence avatar information.

[0201] The data structure is represented by binary files. This device can unpack binary files to generate compressed information. This device can also perform binary decompression on the compressed information.

[0202] The device can map a first type of geometry-related information to a first file format contained in the avatar data structure. The device can also map a second type of geometry-related information to a second file format contained in the avatar data structure, where the first file format differs from the second file format.

[0203] The avatar data structure may include a human-readable file containing encoded versions of a first type of geometry-related avatar information and a second type of geometry-related avatar information.

[0204] This article provides an exemplary processing model. It describes exemplary components of encoder and decoder architectures used to emulate media codecs.

[0205] Figure 6 An exemplary encoder architecture is illustrated. This encoder may be able to process one or more (e.g., multiple) types of input files. For example, the encoder may be able to process descriptive avatar media files (e.g., .x3d, .gltf, .json, .xml, .txt, etc.) containing the information described herein. Output data from the encoder may have one or more (e.g., multiple) formats. Figure 6 As shown, various formats can be used (e.g., two different formats). Exemplary formats may include human-readable formats (e.g., JSON-based) and binary formats. Human-readable formats can provide metadata information for elements (e.g., each element) of an avatar data structure (e.g., semantics, skeleton, landmarks, etc.). Human-readable formats can provide references to files (e.g., appropriate files). Binary formats can compress data (e.g., all data) into (e.g., a single) binary file. In this format, data can be packaged to allow independent access to elements (e.g., each element) of the avatar data structure. Other such encoding schemes and associated data formats can be used to implement the avatar representation described herein.

[0206] Figure 6 The format analysis and formatting modules shown can be parsing mechanisms. These parsing mechanisms can be used to extract relevant information in different file formats (e.g., so that the information can be equally packaged and compressed by a binary compression module, or converted into a human-readable format).

[0207] The input file can have a single format or a collection of files with different formats. If a collection of multiple file formats is used (e.g., ...), ... Figure 6As shown), module format analysis and formatting can be used to obtain relevant information. These modules can arrange information in such a way that the original file format can be recovered (e.g., in binary or human-readable format). These modules may contain header sections or data packets to tell the decoder how to decompress and recover the original different formats. Figure 6 The JSON file demonstrates that semantic attributes can be encoded in .x3d, metadata attributes in .xml, skeletal attributes in .gltf, and landmark attributes in .json. Header sections or data packets can be referenced in the same way to decode back to their original format in a lossless manner.

[0208] Figure 6 The formatting module in the library can generate all human-readable formats (e.g., not binary). For example, a human-readable format could be a JSON file format, or any other human-readable format.

[0209] Figure 6 The binary compression module can apply lossless compression (e.g., using the hierarchical tree (SPIHT) algorithm and / or canonical set partitioning in arithmetic coding (AC)). The binary compression module can also convert data into a bitstream.

[0210] Packaging and binary compression of .json files can follow standards (e.g., RFC 8949 Simplified Binary Object Representation Standard). Packaging and binary compression of .x3d files can follow standards (e.g., ISO / IEC 19776-3.2:2011 Extensible 3D Encoding Standard). Packaging and binary compression of .gltf files can follow standards (e.g., the Open glTF 2.0 specification from the Khronos Group for generating binary .glb files). Packaging and binary compression of ".xml" files can follow standards (e.g., the ISO / IEC 23001-1:2006 Binary MPEG Format for XML).

[0211] Figure 7 An exemplary decoder architecture is shown. The decoder can take a binary file as input. The decoder can output the original file format (e.g., these formats are interoperable).

[0212] The input binary format can be unpacked and decompressed for each of the different file formats (e.g., following the same specifications described in this paper regarding encoders). Data can be extracted from the file. The data can be mapped to data structures (e.g., the data structures described in this paper). The application can then have information related to the avatar media codec.

[0213] Human-readable format can be extracted. Human-readable format can be mapped back to the original file format (e.g., with the help of additional header information about attributes and the file format from which they were extracted).

[0214] The features described herein relate to the compression and representation of avatar and / or user information that enables interoperability between systems and applications used for sending and delivering digital content. Avatar and / or user-style formats can be compressed and sent between systems, devices, applications, etc.

[0215] Digital humans can be represented in several forms (e.g., meshes, volumes, voxels, point clouds, images, videos, sounds, etc.). These representations can be used to create and formalize digital media content. Templates capturing human representations (e.g., all the different representations) can be formalized (e.g., to allow interoperable representations). These templates can be (e.g., with skeletal structures) general human models, object-specific models, statistical shape models, and / or metadata representing individual social attributes. Such templates can generate (e.g., accurately provide) human images as an initial stage. The human images in this initial stage can be statistically represented in different ways.

[0216] Synthetic 3D models can be used to create digital human representations. These representations can be manipulated and / or customized to represent specific human anatomy, visual effects, and / or social parameters. Synthetic representations can facilitate the generation of animations in immersive real-world environments (e.g., if all parameters are initially known). Synthetic models can also aid in the generalization and stylization of appearances (e.g., professionals can use them for animation and streaming).

[0217] 3D composite models can be a good representation. However, 3D composite models may lack standard definitions and / or structures that can be used and matched by another (e.g., any other) digital human representation (e.g., user-specific details such as virtual identity, social status, or input device control, semantic structure of geometric properties, semantic structure of shape and skeletal anatomy, and / or semantics of animation parameters, etc.).

[0218] The standardization of digital human content may disregard human social behavior, interaction with environmental information, and / or social and human privacy issues (e.g., these can be public characteristics that can be adopted within social technologies in the real world). In real-time distribution or communication use cases, the encoding and / or data structure of digital human content may not be defined. In some systems (e.g., proprietary systems), the data may be known. In other systems (e.g., open systems), the data format may have been specified and standardized (e.g., so that the data can interoperate between different systems).

[0219] This document provides exemplary media content formats for avatars and / or user geometric representations.

[0220] This article provides features associated with avatar media representation.

[0221] Exemplary representations of avatars can be compatible with any technology (e.g., technical systems used for VR, AR, streaming, gaming, interaction, collaboration, or communication). This representation can consider the delivery of digital human content (e.g., 2D / 3D video, or images containing body, face, or voice, 3D technologies for rendering, animate, or manipulating human attribute assets, or any other system containing social and contextual information about a real user or avatar, such as in the social and communication media industry).

[0222] This document provides an exemplary format for avatar-style media. It also provides exemplary meanings, encoding schemes, and uses associated with the avatar-style media format.

[0223] The format described in this article can be similar to JSON. This format can be compatible with current MPEG standards (e.g., by providing means for compressing and transmitting media content). The meaning and purpose can be general, and / or encoded using other formats (e.g., OpenXR, XML, USD).

[0224] This article provides an exemplary avatar-style media framework. This article provides an exemplary framework architecture for avatar-style media.

[0225] Figure 8 An exemplary style avatar media framework overview is shown (e.g., top-level component). Figure 8 An overview of exemplary components present in the stylized media codec representation is provided. Exemplary components include “Garment,” “Wearables,” “Shape,” “Attributes,” and “AI.” Components (e.g., each component) may represent a higher-level specification of the stylized media codec.

[0226] Style-related information about digital humans (e.g., avatars) (e.g., basic style-related information) can be encoded. The encoded style-related information can be represented as a stream between systems. Table 1 shows exemplary information (e.g., higher-level information) that can be used for streaming and representing the style of avatar media.

[0227]

[0228] Table 12: Stylized Media Codec Formats.

[0229] This article provides features associated with clothing.

[0230] Avatar media style clothing can represent the avatar's clothes (e.g., defining the attributes of the avatar's clothes). Avatar media style can be presented as geometric, semantic, visual, or contextual information about the clothing attributes of the avatar media. The function "processGarment()" can be used to retrieve this information.

[0231] Signaling can be used to inform the decoder that clothing information is available for reading. This signaling can take one or more forms. Exemplary forms may include the encoding standard of geometric information (e.g., MPEG-I SD, MPEG-AI, MPEG V-DMC, L-PCC, MPEG 4, H-anim, and / or JVET, etc.).

[0232]

[0233] Table 13: Description of the semantics of "style" if the signal is 1.

[0234]

[0235] Table 14: Description of the semantics of "garment" if the signal is 1.

[0236] This article provides features associated with wearable devices.

[0237] Wearable devices styled as avatar media can represent the avatar's accessories (e.g., defining the attributes of accessories for avatar assets). Wearable devices can be presented in the form of geometric, semantic, visual, or contextual information about the wearable device attributes of the avatar media. The function "processWearables()" can be used to retrieve this information.

[0238] Signaling can be used to inform the decoder that wearable device information is available for reading. This signaling can take one or more forms. Exemplary forms may include the encoding standard of geometric information (e.g., MPEG-I SD, MPEG-AI, MPEG V-DMC, L-PCC, MPEG 4, H-anim, and / or JVET, etc.).

[0239]

[0240] Table 15: Description of the semantics of "style" if the signal is 1.

[0241]

[0242] Table 16: Description of the semantics of "wearables" if the signal is 1.

[0243] This article provides features associated with the shape of the avatar media style.

[0244] The shape of the avatar media style can represent the avatar style shape (e.g., defining the attributes of the avatar style shape). This style shape can take the form of geometric, semantic, visual, or contextual information about the avatar's shape (e.g., height, weight, BMI, etc.), face shape, and / or characteristics or hair attributes of the avatar media. The function "processShape()" can be used to retrieve this relevant information.

[0245] Signaling can be used to inform the decoder that shape information is available for reading. This signaling can take one or more forms. Exemplary forms may include coding standards for style, geometry, and / or complementary information (e.g., MPEG-I SD, MPEG-AI, MPEG V-DMC, L-PCC, MPEG 4, H-anim, and / or JVET, etc.).

[0246]

[0247] Table 17: Description of the semantics of "style" if the signal is 1.

[0248]

[0249] Table 18: Description of the semantics of "shape" if the signal is 1.

[0250] This article provides features associated with characteristics.

[0251] The characteristics of an avatar's media style can represent the avatar's personality, skills, and / or abilities (for example, attributes that can define the avatar's personality, skills, and abilities). The function "processAttributes()" can be used to retrieve this information.

[0252] Signaling can be used to inform the decoder that attribute (e.g., capability) information is available for reading. This signaling can take one or more forms. Exemplary forms may include encoding standards for style, geometry, and / or complementary (e.g., haptic) information (e.g., MPEG-I SD, MPEG-AI, MPEG V-DMC, L-PCC, MPEG 4, H-anim, and / or JVET, etc.).

[0253]

[0254] Table 19: Description of the semantics of "style" if the signal is 1.

[0255]

[0256] Table 20: Description of the semantics of "attributes" if the signal is 1.

[0257] This article provides features associated with artificial intelligence (AI).

[0258] AI for avatar-style media can represent the neural representation of the avatar (e.g., define the properties of the neural representation of the avatar). For example, the neural representation can include a description of the neural network architecture or a pre-trained neural network model of the properties of the avatar media representation.

[0259] AI-generated information related to the style of the avatar media can take different forms.

[0260] The style of an avatar can be represented by a set of feature vectors generated by a neural network that encodes the complex geometry and deformation of the clothing on the moving avatar's body.

[0261] The function "processAI()" can be used to obtain this information.

[0262] The purpose of this signal is to inform the decoder that AI information is available for reading. This signaling can take one or more forms. Exemplary forms may include encoding standards used for AI-generated information.

[0263]

[0264] Table 21: Description of the semantics of "style" if the signal is 1.

[0265]

[0266] Table 22: Description of the semantics of "AI" if the signal is 1.

[0267] The processStyle() function can be used to process one or more (e.g., all) subcomponents. For example, processStyle() can invoke the subcomponent processing of any subcomponent described herein (e.g., via the syntax in Table 1).

[0268] An exemplary device may have a processor configured to perform one or more actions. The device (e.g., an encoder, such as an avatar style encoder) may receive a descriptive avatar media file. The device may extract first-type and second-type style-related avatar information from the descriptive avatar media file. The device may encode the first-type and second-type style-related avatar information into an avatar data structure. The device may send the avatar data structure, along with an indication of the file format contained within the avatar data structure, to a decoder.

[0269] At least one of the first type of style-related avatar information and the second type of style-related avatar information can be: avatar clothing information; avatar wearable device metadata; avatar shape information; avatar attribute information; or artificial intelligence avatar information.

[0270] The avatar data structure can include binary files. The device can perform binary compression on both the first type and the second type of style-related avatar information to generate compressed information. The device can then package the compressed information to generate a binary file.

[0271] The avatar data structure can include human-readable files. The device can encode the first type of style-related avatar information and the second type of style-related avatar information into an avatar data structure by generating human-readable files based on the first type of style-related avatar information and the second type of style-related avatar information.

[0272] The first type of style-related avatar information can be associated with a first file format contained in the avatar data structure. The second type of style-related avatar information can be associated with a second file format contained in the avatar data structure. The first file format can be different from the second file format.

[0273] An exemplary device (e.g., a decoder, such as an avatar style decoder) can receive an avatar data structure and an indication of the file format contained in the data structure from an encoder. The device can decode the avatar data structure. The device can determine a first type of style-related avatar information and a second type of style-related avatar information based on the decoded avatar data structure. The device can generate a descriptive avatar media file based on the first type of style-related avatar information, the second type of style-related avatar information, and the indicated file format.

[0274] At least one of the first type of style-related avatar information and the second type of style-related avatar information can be: avatar clothing information; avatar wearable device metadata; avatar shape information; avatar attribute information; or artificial intelligence avatar information.

[0275] The data structure can include binary files. This device can unpack binary files to generate compressed information. This device can also perform binary decompression on the compressed information.

[0276] The device can map first-type style-related avatar information to a first file format contained in the avatar data structure. The device can also map second-type style-related avatar information to a second file format contained in the avatar data structure, wherein the first file format differs from the second file format.

[0277] The avatar data structure may include a human-readable file containing encoded versions of a first type of style-related avatar information and a second type of style-related avatar information.

[0278] This article provides an exemplary processing model. It describes exemplary components of the encoder and decoder architecture used to emulate media codecs.

[0279] Figure 9 An exemplary encoder architecture is illustrated. This encoder may be able to process one or more (e.g., multiple) types of input files. For example, the encoder may be able to process descriptive avatar media files (e.g., .x3d, .gltf, .json, .xml, .txt, etc.) containing the information described herein. The encoder's output data may have one or more (e.g., multiple) formats. Figure 9 As shown, various formats can be used (e.g., two different formats). Exemplary formats may include human-readable formats (e.g., JSON-based) and binary formats. Human-readable formats can provide metadata information for elements (e.g., each element) of an avatar data structure (e.g., clothing, wearable devices, etc.). Human-readable formats can provide references to files (e.g., appropriate files). Binary formats can compress data (e.g., all data) into (e.g., a single) binary file. In this format, data can be packaged to allow independent access to elements of the avatar data structure (e.g., each element). Other such encoding schemes and associated data formats can be used to implement the avatar representation described herein.

[0280] Figure 9 The format analysis and formatting modules shown can be parsing mechanisms. These parsing mechanisms can be used to extract relevant information in different file formats (e.g., so that the information can be equally packaged and compressed by a binary compression module, or converted into a human-readable format).

[0281] The input file can have a single format or a collection of files with different formats. If a collection of multiple file formats is used (e.g., ...), ... Figure 9 As shown), module format analysis and formatting can be used to obtain relevant information. These modules can arrange information in such a way that the original file format can be recovered (e.g., in binary or human-readable format). These modules may contain header sections or data packets to tell the decoder how to decompress and recover the original different formats. Figure 9 The JSON file shows that clothing attributes can be encoded in .x3d, wearable device attributes in .xml, shape attributes in .gltf, and feature attributes in .json. Header sections or data packets can be referenced in the same way to decode back to their original format in a lossless manner.

[0282] Figure 9 The formatting module in the library can generate all human-readable formats (e.g., not binary). For example, a human-readable format could be a JSON file format, or any other human-readable format.

[0283] Figure 9 The binary compression module can apply lossless compression (e.g., using the hierarchical tree (SPIHT) algorithm and / or canonical set partitioning in arithmetic coding (AC)). The binary compression module can also convert data into a bitstream.

[0284] Packaging and binary compression of .json files can follow standards (e.g., RFC 8949 Simplified Binary Object Representation Standard). Packaging and binary compression of .x3d files can follow standards (e.g., ISO / IEC 19776-3.2:2011 Extensible 3D Encoding Standard). Packaging and binary compression of .gltf files can follow standards (e.g., the Open glTF 2.0 specification from the Khronos Group for generating binary .glb files). Packaging and binary compression of ".xml" files can follow standards (e.g., the ISO / IEC 23001-1:2006 Binary MPEG Format for XML).

[0285] Figure 10 An exemplary decoder architecture is shown. The decoder can take a binary file as input. The decoder can output the original file format (e.g., these formats are interoperable).

[0286] The input binary format can be unpacked and decompressed for each of the different file formats (e.g., following the same specifications described in this paper regarding encoders). Data can be extracted from the file. The data can be mapped to data structures (e.g., the data structures described in this paper). The application can then have information related to the avatar media codec.

[0287] Human-readable format can be extracted. Human-readable format can be mapped back to the original file format (e.g., with the help of additional header information about attributes and the file format from which they were extracted).

[0288] The features described herein relate to the compression and representation of avatars and / or user information that are interoperable between systems and applications used for sending and delivering digital content. Avatar and / or user animation formats can be compressed and sent between systems, devices, applications, etc.

[0289] Digital humans can be represented in several forms (e.g., meshes, volumes, voxels, point clouds, images, videos, sounds, etc.). These representations can be used to create and formalize digital media content. Templates capturing human representations (e.g., all the different representations) can be formalized (e.g., to allow for interoperable representations). These templates can be (e.g., with skeletal structures) general human models, object-specific models, statistical shape models, metadata representing individual social attributes, etc. Such templates can (e.g., accurately provided) generate human images as an initial stage. The human images in this initial stage can be statistically represented in different ways.

[0290] Synthetic 3D models can be used to create digital human representations. These representations can be manipulated and / or customized to represent specific human anatomy, visual effects, and / or social parameters. Synthetic representations can facilitate the generation of animations in immersive real-world environments (e.g., if all parameters are initially known). Synthetic models can also aid in the generalization and stylization of appearances (e.g., professionals can use them for animation and streaming).

[0291] 3D composite models can be very good representations. 3D composite models may lack standard definitions and / or structures that can be used and matched by another (e.g., any other) digital human representation (e.g., user-specific details such as virtual identity, social status or input device control, semantic structure of geometric properties, semantic structure of shape and skeletal anatomy, and / or semantics of animation parameters, etc.).

[0292] The standardization of digital human content may disregard human social behavior, interaction with environmental information, and / or social and human privacy issues (e.g., these can be public characteristics that can be adopted within social technologies in the real world). In real-time distribution or communication use cases, the encoding and / or data structure of digital human content may not be defined. In some systems (e.g., proprietary systems), the data may be known. In other systems (e.g., open systems), the data format may have been specified and standardized (e.g., so that the data can interoperate between different systems).

[0293] This document provides exemplary media content formats for avatar and / or user animation representations.

[0294] This article provides features associated with avatar media representation.

[0295] Exemplary representations of avatars can be compatible with any technology (e.g., technical systems used for VR, AR, streaming, gaming, interaction, collaboration, or communication). This representation can consider the delivery of digital human content (e.g., 2D / 3D video, or images containing body, face, or voice, 3D technologies for rendering, animate, or manipulating human attribute assets, or any other system containing social and contextual information about a real user or avatar, such as in the social and communication media industry).

[0296] This document provides an exemplary format for avatar animation media. It also provides exemplary meanings, encoding schemes, and uses associated with the avatar animation media format.

[0297] The format described in this article can be similar to JSON. This format can be compatible with current MPEG standards (e.g., by providing means for compressing and transmitting media content). The meaning and purpose can be general, and / or encoded using other formats (e.g., OpenXR, XML, USD).

[0298] This article provides an exemplary framework for avatar animation media. This article provides an exemplary framework architecture for avatar animation media.

[0299] Figure 11 An exemplary animated avatar media framework overview is shown (e.g., top-level component). Figure 11 An overview of exemplary components present in the Animated Avatar Media Codec representation is provided. Exemplary components include "Motion Style", "Dynamic Mesh", "Skeletal", "2D Video", and "AI". Components (e.g., each component) may represent a higher-level specification of the Animated Avatar Media Codec.

[0300] Animation-related information about digital humans (e.g., avatars) (e.g., basic animation-related information) can be encoded. The encoded animation-related information can be represented as a stream between systems. Table 1 shows exemplary information (e.g., higher-level information) that can be used for streaming and representing animation in avatar media.

[0301]

[0302] Table 23: Animated Avatar Media Codec Formats.

[0303] This article provides features associated with athletic style.

[0304] The motion style of an avatar can represent the style animation of the avatar (e.g., defining the properties of the avatar's style animation). Style animation can take the form of a semantic description of motion related to the avatar's body shape, emotions, personality, mental state, or physical disability. Style animation can take the form of a descriptor that allows for motion style parameterization of geometric animations or rendered content. The function "processMotionStyle()" can be used to retrieve this information.

[0305] Signaling can be used to inform the decoder that motion style information is available for reading. This signaling can take one or more forms. Exemplary forms may include encoding standards for geometric animation information (e.g., MPEG-I SD, MPEG-AI, MPEG V-DMC, L-PCC, MPEG 4, H-anim, and / or JVET, etc.).

[0306]

[0307] Table 24: Description of the semantics of "animation" if the signal is 1.

[0308]

[0309] Table 25: Description of the semantics of "motionStyle" if the signal is 1.

[0310] This article provides features associated with dynamic meshes.

[0311] A dynamic mesh for an avatar can represent the avatar's geometric animation (e.g., defining the properties of the avatar asset's geometric animation). For example, a dynamic mesh can represent vertex shifts, facial blending shapes, non-rigid motion of vertices (e.g., soft tissue deformation, hair movement, facial expressions, etc.), and so on. The function "processDynamicMesh()" can be used to retrieve this information.

[0312] Signaling can be used to inform the decoder that dynamic mesh information is available for reading. This signaling can take one or more forms. Exemplary forms may include encoding standards for geometric animation information (e.g., MPEG-I SD, MPEG-AI, MPEG V-DMC, L-PCC, MPEG 4, H-anim, and / or JVET, etc.).

[0313]

[0314] Table 26: Description of the semantics of "animation" if the signal is 1.

[0315]

[0316] Table 27: Description of the semantics of "dynamicMesh" if the signal is 1.

[0317] This article provides features associated with the skeletal animation of the avatar.

[0318] The skeleton of an avatar media can represent avatar skeletal animation (e.g., defining properties of avatar skeletal animation). For example, skeletal animation can include joint positions, rotations, displacements, linearly blended skin, etc. The function "processSkeletal()" can be used to retrieve this information.

[0319] Signaling can be used to inform the decoder that skeletal animation information is available for reading. This signaling can take one or more forms. Exemplary forms may include encoding standards for geometric animation and complementary information (e.g., MPEG-I SD, MPEG-AI, MPEG V-DMC, L-PCC, MPEG 4, H-anim, and / or JVET, etc.).

[0320]

[0321] Table 28: Description of the semantics of "animation" if the signal is 1.

[0322]

[0323] Table 29: Description of the semantics of "skeletal" if the signal is 1.

[0324] This article provides features associated with 2D videos.

[0325] A 2D video representing an avatar can be a 2D video animation representing the avatar (e.g., defining properties of the avatar's 2D video animation). For example, a 2D video animation could include motion flow on facial keymarks. The function "process2DVideo()" can be used to retrieve this information.

[0326] Signaling can be used to inform the decoder that 2D video animation information is available for reading. This signaling can take one or more forms. Exemplary forms may include encoding standards for geometric animation and complementary (e.g., haptic) information (e.g., MPEG-ISD, MPEG-AI, MPEG V-DMC, L-PCC, MPEG 4, H-anim, and / or JVET, etc.).

[0327]

[0328] Table 30: Description of the semantics of "animation" if the signal is 1.

[0329]

[0330] Table 31: Description of the semantics of "2DVideo" if the signal is 1.

[0331] This article provides features associated with artificial intelligence (AI).

[0332] AI representing the avatar of an animated medium can represent the neural representation of the avatar (e.g., define the properties of the neural representation of the avatar). For example, the neural representation can include a description of a neural network architecture or a pre-trained neural network model of properties related to the avatar media representation.

[0333] AI-generated information related to animations in avatar media can take different forms.

[0334] An avatar's animation can be represented by latent codes from an autoencoder neural network. An autoencoder neural network can include an encoder neural network and a decoder neural network. The encoder neural network can be fed with a representation of the avatar's geometry (e.g., a geometric animation) animation. The encoder neural network can output latent codes. The latent codes can provide a representation of the avatar's animation as a vector of floating-point numbers (e.g., a more compact representation). The decoder can reconstruct the input representation of the avatar's geometry (e.g., a geometric animation) from the latent codes.

[0335] In some examples (such as the neural radiation field method), the geometry of the avatar (e.g., geometric animation) can be represented by a neural network. Given a 3D position in a coordinate system associated with the avatar as input (e.g., the viewing direction represented by two angles in the coordinate system), the network can output an opacity value indicating whether the avatar's body occupies that position at the input 3D position. The output can indicate the color of a portion of the body at a position (e.g., when viewed from the input direction). The neural network's knowledge can allow the system to reconstruct a complete 3D geometric model of the avatar (e.g., geometric animation) (e.g., by querying a discretized set of input positions within its enclosing volume for the avatar's occupancy and / or color at those positions). The function "processAI()" can be used to obtain this relevant information.

[0336] The purpose of this signal is to inform the decoder that AI information is available for reading. This signaling can take one or more forms. Exemplary forms may include the encoding standards for information generated by AI (e.g., geometry, geometric animation, and / or complementary information, such as MPEG-I SD, MPEG-AI, MPEG V-DMC, L-PCC, MPEG 4, H-anim, and / or JVET, etc.).

[0337]

[0338] Table 32: Description of the semantics of "animation" if the signal is 1.

[0339]

[0340] Table 33: Description of the semantics of "AI" if the signal is 1.

[0341] The processAnimation() function can be used to process one or more (e.g., all) subcomponents. For example, processAnimation() can invoke the subcomponent processing of any subcomponent described herein (e.g., via the syntax in Table 1).

[0342] An exemplary device may have a processor configured to perform one or more actions. The device (e.g., an encoder, such as an avatar animation encoder) may receive a descriptive avatar media file. The device may extract first-type and second-type animation-related avatar information from the descriptive avatar media file. The device may encode the first-type and second-type animation-related avatar information into an avatar data structure. The device may send the avatar data structure, along with an indication of the file format contained within the avatar data structure, to a decoder.

[0343] At least one of the first type of animation-related avatar information and the second type of animation-related avatar information can be: avatar motion style information; avatar dynamic mesh metadata; avatar skeletal information; avatar 2D video information; or artificial intelligence avatar information.

[0344] The avatar data structure can include binary files. The device can perform binary compression on both types of animation-related avatar information (first type and second type) to generate compressed information. The device can then package the compressed information to generate a binary file.

[0345] The avatar data structure can include human-readable files. The device can encode the first type of animation-related avatar information and the second type of animation-related avatar information into an avatar data structure by generating human-readable files based on the first type of animation-related avatar information and the second type of animation-related avatar information.

[0346] The first type of animation-related avatar information can be associated with a first file format contained in the avatar data structure. The second type of animation-related avatar information can be associated with a second file format contained in the avatar data structure. The first file format can be different from the second file format.

[0347] An exemplary device (e.g., a decoder, such as an avatar animation decoder) can receive an avatar data structure and an indication of the file format contained in the data structure from an encoder. The device can decode the avatar data structure. The device can determine a first type of animation-related avatar information and a second type of animation-related avatar information based on the decoded avatar data structure. The device can generate a descriptive avatar media file based on the first type of animation-related avatar information, the second type of animation-related avatar information, and the indicated file format.

[0348] At least one of the first type of animation-related avatar information and the second type of animation-related avatar information can be: avatar motion style information; avatar dynamic mesh metadata; avatar skeletal information; avatar 2D video information; or artificial intelligence avatar information.

[0349] The data structure can include binary files. This device can unpack binary files to generate compressed information. This device can also perform binary decompression on the compressed information.

[0350] The device can map a first type of animation-related avatar information to a first file format contained in the avatar data structure. The device can also map a second type of animation-related avatar information to a second file format contained in the avatar data structure, wherein the first file format differs from the second file format.

[0351] The avatar data structure may include a human-readable file containing encoded versions of a first type of animation-related avatar information and a second type of animation-related avatar information.

[0352] This article provides an exemplary processing model. It describes exemplary components of the encoder and decoder architecture used to emulate media codecs.

[0353] Figure 12 An exemplary encoder architecture is illustrated. This encoder may be able to process one or more (e.g., multiple) types of input files. For example, the encoder may be able to process descriptive avatar media files (e.g., .x3d, .gltf, .json, .xml, .txt, etc.) containing the information described herein. The encoder's output data may have one or more (e.g., multiple) formats. Figure 12 As shown, various formats can be used (e.g., two different formats). Exemplary formats may include human-readable formats (e.g., JSON-based) and binary formats. Human-readable formats can provide metadata information for elements (e.g., each element) of an avatar data structure (e.g., motion style, dynamic mesh, skeleton, etc.). Human-readable formats can provide references to files (e.g., appropriate files). Binary formats can compress data (e.g., all data) into (e.g., a single) binary file. In this format, data can be packaged to allow independent access to elements of the avatar data structure (e.g., each element). Other such encoding schemes and their associated data formats can be used to implement the avatar representation described herein.

[0354] Figure 12The format analysis and formatting modules shown can be parsing mechanisms. These parsing mechanisms can be used to extract relevant information in different file formats (e.g., so that the information can be equally packaged and compressed by a binary compression module, or converted into a human-readable format).

[0355] The input file can have a single format or a collection of files with different formats. If a collection of multiple file formats is used (e.g., ...), ... Figure 12 As shown), module format analysis and formatting can be used to obtain relevant information. These modules can arrange information in such a way that the original file format can be recovered (e.g., in binary or human-readable format). These modules may contain header sections or data packets to tell the decoder how to decompress and recover the original different formats. Figure 12 The JSON file shows that motion style attributes can be encoded in .x3d, dynamic mesh attributes in .xml, bone attributes in .gltf, and 2D video attributes in .json. Header sections or data packets can be referenced in the same way to decode back to their original format in a lossless manner.

[0356] Figure 12 The formatting module in the library can generate all human-readable formats (e.g., not binary). For example, a human-readable format could be a JSON file format, or any other human-readable format.

[0357] Figure 12 The binary compression module can apply lossless compression (e.g., using the hierarchical tree (SPIHT) algorithm and / or canonical set partitioning in arithmetic coding (AC)). The binary compression module can also convert data into a bitstream.

[0358] Packaging and binary compression of .json files can follow standards (e.g., RFC 8949 Simplified Binary Object Representation Standard). Packaging and binary compression of .x3d files can follow standards (e.g., ISO / IEC 19776-3.2:2011 Extensible 3D Encoding Standard). Packaging and binary compression of .gltf files can follow standards (e.g., the Open glTF 2.0 specification from the Khronos Group for generating binary .glb files). Packaging and binary compression of ".xml" files can follow standards (e.g., the ISO / IEC 23001-1:2006 Binary MPEG Format for XML).

[0359] Figure 13 An exemplary decoder architecture is shown. The decoder can take a binary file as input. The decoder can output the original file format (e.g., these formats are interoperable).

[0360] The input binary format can be unpacked and decompressed for each of the different file formats (e.g., following the same specifications described in this paper regarding encoders). Data can be extracted from the file. The data can be mapped to data structures (e.g., the data structures described in this paper). The application can then have information related to the avatar media codec.

[0361] Human-readable format can be extracted. Human-readable format can be mapped back to the original file format (e.g., with the help of additional header information about attributes and the file format from which they were extracted).

[0362] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in a computer program, software, or firmware incorporated into a computer-readable medium for execution by a computer or processor. Examples of computer-readable media include electronic signals (transmitted via a wired or wireless connection) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM discs and digital multifunction disks (DVDs). The processor associated with the software can be used to implement a radio frequency transceiver for a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. A geometric encoder, comprising: Processor, the processor being configured to: Receive descriptive avatar media files; Extract the first type of geometry-related avatar information and the second type of geometry-related avatar information from the descriptive avatar media file; The first type of geometrically related avatar information and the second type of geometrically related avatar information are encoded into an avatar data structure; as well as The avatar data structure and the file format indication contained in the avatar data structure are sent to the decoder.

2. The geometry encoder of claim 1, wherein at least one of the first type of geometry-related avatar information and the second type of geometry-related avatar information comprises: Transformation into semantic information; Transform into geometric metadata; Transformation into skeletal information; Transformed into landmark information; or Artificial intelligence is transformed into information.

3. The geometry encoder of claim 1, wherein the avatar data structure comprises a binary file, and wherein the processor is configured to encode the first type of geometry-related avatar information and the second type of geometry-related avatar information into the avatar data structure comprising: The processor is configured to: The first type of geometry-related information and the second type of geometry-related information are binary compressed to generate compressed information; as well as The compressed information is packaged to generate the binary file.

4. The geometry encoder of claim 1, wherein the avatar data structure comprises a human-readable file, and the processor is configured to encode the first type of geometry-related avatar information and the second type of geometry-related avatar information into the avatar data structure comprising: The processor is configured to generate the human-readable file based on the first type of geometry-related information and the second type of geometry-related information.

5. The geometry encoder of claim 1, wherein the first type of geometry-related avatar information is associated with a first file format contained in the avatar data structure, the second type of geometry-related avatar information is associated with a second file format contained in the avatar data structure, and the first file format is different from the second file format.

6. A geometry decoder, comprising: Processor, the processor being configured to: Receive avatar data structure and an indication of the file format contained in the avatar data structure from the encoder; Decode the avatar data structure; The first type of geometrically related avatar information and the second type of geometrically related avatar information are determined based on the decoded avatar data structure; as well as A descriptive avatar media file is generated based on the first type of geometry-related avatar information, the second type of geometry-related information, and the indicated file format.

7. The geometry decoder of claim 6, wherein at least one of the first type of geometry-related avatar information and the second type of geometry-related avatar information comprises: Transformation into semantic information; Transform into geometric metadata; Transformation into skeletal information; Transformed into landmark information; or Artificial intelligence is transformed into information.

8. The geometry decoder of claim 6, wherein the avatar data structure comprises a binary file, and wherein the processor is configured to decode the avatar data structure comprising: The processor is configured to: The binary file is unpacked to generate compressed information; as well as The compressed information is then decompressed in binary form.

9. The geometry decoder of claim 6, wherein the processor is configured to generate the descriptive avatar media file based on the first type of geometry-related avatar information, the second type of geometry-related information, and an indicated file format, comprising: The processor is configured to: Map the geometry-related information of the first type to the first file format contained in the avatar data structure; as well as The second type of geometry-related information is mapped to a second file format contained in the avatar data structure, wherein the first file format is different from the second file format.

10. The geometry decoder of claim 6, wherein the avatar data structure comprises a human-readable file, the human-readable file comprising encoded versions of the first type of geometry-related avatar information and the second type of geometry-related avatar information.

11. A method executed by a geometric encoder, the method comprising: Receive descriptive avatar media files; Extract the first type of geometry-related avatar information and the second type of geometry-related avatar information from the descriptive avatar media file; The first type of geometrically related avatar information and the second type of geometrically related avatar information are encoded into an avatar data structure; as well as The avatar data structure and the file format indication contained in the avatar data structure are sent to the decoder.

12. The method of claim 10, wherein at least one of the first type of geometrically related avatar information and the second type of geometrically related avatar information comprises: Transformation into semantic information; Transform into geometric metadata; Transformation into skeletal information; Transformed into landmark information; or Artificial intelligence is transformed into information.

13. The method of claim 10, wherein the avatar data structure comprises a binary file, and wherein encoding the first type of geometry-related avatar information and the second type of geometry-related avatar information into the avatar data structure comprises: The first type of geometry-related information and the second type of geometry-related information are binary compressed to generate compressed information; as well as The compressed information is packaged to generate the binary file.

14. The method of claim 10, wherein the avatar data structure comprises a human-readable file, and encoding the first type of geometry-related avatar information and the second type of geometry-related avatar information into the avatar data structure comprises: The human-readable file is generated based on the first type of geometry-related information and the second type of geometry-related information.

15. The method of claim 10, wherein the first type of geometry-related avatar information is associated with a first file format contained in the avatar data structure, the second type of geometry-related avatar information is associated with a second file format contained in the avatar data structure, and the first file format is different from the second file format.

16. A method executed by a geometry decoder, the method comprising: Receive avatar data structure and an indication of the file format contained in the avatar data structure from the encoder; Decode the avatar data structure; The first type of geometrically related avatar information and the second type of geometrically related avatar information are determined based on the decoded avatar data structure; as well as A descriptive avatar media file is generated based on the first type of geometry-related avatar information, the second type of geometry-related information, and the indicated file format.

17. The method of claim 16, wherein at least one of the first type of geometrically related avatar information and the second type of geometrically related avatar information comprises: Transformation into semantic information; Transform into geometric metadata; Transformation into skeletal information; Transformed into landmark information; or Artificial intelligence is transformed into information.

18. The method of claim 16, wherein the avatar data structure comprises a binary file, and wherein decoding the avatar data structure comprises: The binary file is unpacked to generate compressed information; as well as The compressed information is then decompressed in binary form.

19. The method of claim 16, wherein generating the descriptive avatar media file based on the first type of geometry-related avatar information, the second type of geometry-related information, and the indicated file format comprises: Map the geometry-related information of the first type to the first file format contained in the avatar data structure; as well as The second type of geometry-related information is mapped to a second file format contained in the avatar data structure, wherein the first file format is different from the second file format.

20. The method of claim 16, wherein the avatar data structure comprises a human-readable file, the human-readable file comprising encoded versions of the first type of geometry-related avatar information and the second type of geometry-related avatar information.