Methods and apparatus for base grid entropy coding in video-based dynamic mesh coding (V-DMC), methods and apparatus for base grid entropy coding in V-DMC

CN122847868APending Publication Date: 2026-09-29SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202580015052.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2025-04-01
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0005]然而,基础网格熵编解码可以是复杂的,并且期望附加的简单性

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122847868A_ABST
    Figure CN122847868A_ABST
Patent Text Reader

Abstract

An improved apparatus is provided for fundamental grid entropy coding in a fundamental grid frame for inter-frame coding. The apparatus decodes the fundamental grid frame. The apparatus performs arithmetic decoding on one or more codewords corresponding to one or more prediction errors associated with the fundamental grid frame, wherein the one or more prediction errors are associated with a fine-grained class or a coarse-grained class. In an embodiment, the apparatus may further allocate one or more contexts for decoding the one or more codewords corresponding to the one or more prediction errors. In an embodiment, the apparatus also shares one or more contexts that will be used for the one or more prediction errors associated with the fine-grained class or the coarse-grained class.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to improvements in video-based compression of dynamic grids, and more specifically, to improvements, for example but not limited to, basic grid entropy encoding and decoding. Background Technology

[0002] Currently, ISO / IEC SC29 / WG07 is working on developing a standard for video-based compression for dynamic meshes. For example, the committee is working on a standard for video-based dynamic mesh coding and decoding (V-DMC), which specifies the syntax, semantics and decoding, basic mesh coding and decoding, Moving Picture Experts Group (MPEG) edgebreaker static mesh coding and decoding, and arithmetic coded shifts for V-DMC. As an example, the eighth test model (version 8.0) for the V-DMC mesh test model (TMM) was established at the 14th meeting of ISO / IEC SC29 / WG07 in June 2024. Draft specifications for video-based compression for dynamic meshes are also available.

[0003] In the example, a mesh is a fundamental element in a three-dimensional (3D) computer graphics model. In an embodiment, the mesh consists of several polygons that describe the boundary surfaces of volumetric objects. In such an embodiment, each polygon is defined by its vertices in three-dimensional (3D) space, and information about how the vertices are connected is called connectivity information. Furthermore, vertex attributes can be associated with mesh vertices. For example, vertex attributes can include color, normals, etc. In an embodiment, attributes are also associated with the mesh's surface using mapping information that describes the mesh's parameterization on a two-dimensional (2D) region of a plane. In an embodiment, this mapping is described by a set of parametric coordinates, referred to as (U,V) coordinates or texture coordinates. In an embodiment, if the connectivity or attribute information changes, the mesh is called a dynamic mesh. In an embodiment, a dynamic mesh contains a large amount of data and is therefore standardized according to MPEG.

[0004] In some examples, the base mesh has fewer vertices compared to the original mesh. For example, the base mesh is created and compressed in a lossy or lossless manner. In embodiments, the reconstructed base mesh undergoes subdivision, and then the displacement field between the original mesh and the subdivided reconstructed base mesh is calculated. In embodiments, the base mesh is encoded during inter-frame encoding / decoding of mesh frames by sending vertex motions instead of directly compressing the base mesh.

[0005] However, basic grid entropy encoding and decoding can be complex, and additional simplicity is expected.

[0006] The descriptions set forth in the Background section should not be assumed to be prior art simply because they are set forth therein. The Background section may describe aspects or embodiments of this disclosure. Summary of the Invention

[0007] Technical solution In embodiments, this disclosure may relate to improvements in the encoding and decoding of the underlying mesh entropy. Specifically, this disclosure may relate to improvements related to prediction error information, such as geometric prediction errors or texture coordinate prediction errors.

[0008] In this embodiment, the Moving Picture Experts Group (MPEG) edgebreaker static mesh codec, introduced in the test model V-DMC TMM 8.0, can be used. This MPEG edgebreaker static mesh codec allows for arithmetic decoding of prediction errors and assigns them one or more contexts for decoding, as described herein.

[0009] One aspect of this disclosure provides a computer-implemented method for decoding a base grid frame. The method includes: performing arithmetic decoding on one or more codewords corresponding to one or more prediction errors associated with a base grid frame, wherein the one or more prediction errors are associated with a fine-grained class or a coarse-grained class; allocating one or more contexts for decoding the one or more codewords corresponding to the one or more prediction errors; and sharing the one or more contexts to be used for the one or more prediction errors associated with the fine-grained class or the coarse-grained class.

[0010] One aspect of this disclosure provides an apparatus for decoding a grid frame, including a processor. In an embodiment, the processor is configured to: receive a bitstream including a prediction error of an arithmetic code for current coordinates of the grid frame; determine one or more contexts of the prediction error of the arithmetic code for the current coordinates; perform arithmetic decoding on the prediction error of the arithmetic code based on the one or more contexts to determine a prediction error for the current coordinates; determine a prediction value for the current coordinates; and determine a coordinate value for the current coordinates based on the prediction error for the current coordinates and the prediction value for the current coordinates, wherein one or more contexts of the prediction error for at least one coordinate associated with a fine category are shared for the prediction error for at least one coordinate associated with a coarse category.

[0011] One aspect of this disclosure provides an apparatus for encoding a grid frame, including a processor. The processor is configured to: determine a predicted value for current coordinates of the grid frame; determine a prediction error for the current coordinates based on the value of the current coordinates and the predicted value for the current coordinates; determine one or more contexts for the prediction error for the current coordinates; perform arithmetic encoding on the prediction error for the current coordinates based on the one or more contexts to generate an arithmetic-encoded prediction error for the current coordinates; and transmit a bitstream including the arithmetic-encoded prediction error, wherein one or more contexts for the prediction error of at least one coordinate associated with a fine category are shared for the prediction error of at least one coordinate associated with a coarse category. Attached Figure Description

[0012] Figure 1 An example communication system 100 according to an embodiment of the present disclosure is shown.

[0013] Figure 2 and Figure 3 An example electronic device is shown according to an embodiment of the present disclosure.

[0014] Figure 4 A block diagram of an encoder for encoding intra-frames according to an embodiment is shown.

[0015] Figure 5 A block diagram for a decoder according to an embodiment is shown.

[0016] Figure 6 and Figure 7 A block diagram illustrating parallelogram grid prediction according to an embodiment is shown.

[0017] Figure 8A , Figure 8B , Figure 9A and Figure 9B An example prediction error context is shown according to an embodiment.

[0018] Figure 10A , Figure 10B , Figure 11A , Figure 11B , Figure 12A , Figure 12B , Figure 12C , Figure 12D , Figure 13A , Figure 13B , Figure 13C , Figure 13D , Figure 14A , Figure 14B , Figure 15A , Figure 15B , Figure 16A , Figure 16B , Figure 16C , Figure 16D , Figure 17A , Figure 17B , Figure 17C , Figure 17D , Figure 18A , Figure 18B , Figure 18C and Figure 18D An example illustrating a simplified prediction error context is shown according to an embodiment.

[0019] Figure 19 A flowchart illustrating the operation of a basic mesh encoder according to an embodiment is shown.

[0020] Figure 20 A flowchart illustrating the operation of a basic mesh decoder according to an embodiment is shown.

[0021] Figure 21 A flowchart illustrating the operation of a basic mesh encoder according to an embodiment is shown.

[0022] In one or more embodiments, not all components depicted in each figure may be required, and one or more embodiments may include additional components not shown in the figures. Variations in the arrangement and type of components may be made without departing from the scope of this subject matter disclosure. Within the scope of this subject matter disclosure, additional components, different components, or fewer components may be utilized. Detailed Implementation

[0023] One aspect of this disclosure provides a computer-implemented method for decoding a base grid frame. The method includes: performing arithmetic decoding on one or more codewords corresponding to one or more prediction errors associated with a base grid frame, wherein the one or more prediction errors are associated with a fine-grained class or a coarse-grained class; allocating one or more contexts for decoding the one or more codewords corresponding to the one or more prediction errors; and sharing the one or more contexts to be used for the one or more prediction errors associated with the fine-grained class or the coarse-grained class.

[0024] In an embodiment, the one or more prediction errors include at least one of fine geometry prediction error, coarse geometry prediction error, fine texture prediction error, or coarse texture prediction error. The method also includes sharing the one or more contexts used among at least one of the fine geometry prediction error, coarse geometry prediction error, fine texture prediction error, or coarse texture prediction error.

[0025] In an embodiment, the one or more codewords include one or more portions, wherein a portion of the one or more portions is associated with a truncated unary binarization, and wherein multiple bit positions within the truncated unary binarization share the same context in the one or more contexts.

[0026] In an embodiment, there are three contexts among the one or more contexts associated with the portion of the codeword, wherein the three contexts are used for coarse categories, and wherein a subset of the three contexts is used for fine categories.

[0027] In an embodiment, there is a first number of binary bits for truncated unary binarization associated with fine texture prediction error, and a second number of binary bits for truncated unary binarization associated with coarse texture prediction error, wherein the first number is greater than the second number.

[0028] In an embodiment, the one or more codewords include one or more portions, wherein a portion of the one or more portions is associated with an exponential Golomb prefix binarization, and wherein multiple bit positions within the exponential Golomb prefix binarization share the same context in the one or more contexts.

[0029] In an embodiment, there are five contexts among the one or more contexts associated with the portion of the codeword, wherein the five contexts are used for a first prediction error type of the one or more prediction errors, and wherein a subset of the five contexts are used for a second prediction error type of the one or more prediction errors.

[0030] In an embodiment, the one or more codewords include one or more portions, wherein a portion of the one or more portions is associated with an exponential Golomb postfix binarization, and wherein multiple bit positions within the exponential Golomb postfix binarization share the same context in the one or more contexts.

[0031] In an embodiment, there are five contexts among the one or more contexts associated with the portion of the codeword, wherein the five contexts are used for a first prediction error type of the one or more prediction errors, and wherein a subset of the five contexts are used for a second prediction error type of the one or more prediction errors.

[0032] In an embodiment, the one or more prediction errors include at least one of fine normal prediction error, coarse normal prediction error, fine attribute prediction error, or coarse attribute prediction error.

[0033] One aspect of this disclosure provides an apparatus for decoding a grid frame, including a processor. In an embodiment, the processor is configured to: receive a bitstream including a prediction error of an arithmetic code for current coordinates of the grid frame; determine one or more contexts of the prediction error of the arithmetic code for the current coordinates; perform arithmetic decoding on the prediction error of the arithmetic code based on the one or more contexts to determine a prediction error for the current coordinates; determine a prediction value for the current coordinates; and determine a coordinate value for the current coordinates based on the prediction error for the current coordinates and the prediction value for the current coordinates, wherein one or more contexts of the prediction error for at least one coordinate associated with a fine category are shared for the prediction error for at least one coordinate associated with a coarse category.

[0034] In an embodiment, one or more contexts of the prediction error for at least one geometric coordinate are shared for the prediction error of at least one texture coordinate.

[0035] In an embodiment, one or more contexts of a truncated unary portion of the prediction error for at least one coordinate associated with a fine category are shared for a truncated unary portion of the prediction error for at least one coordinate associated with a coarse category.

[0036] In at least one example, one or more contexts of the prefix portion of the prediction error for at least one coordinate associated with the fine category are shared for the prefix portion of the prediction error for at least one coordinate associated with the coarse category.

[0037] In an embodiment, one or more contexts of the suffix portion of the prediction error for at least one coordinate associated with a fine category are shared for the suffix portion of the prediction error for at least one coordinate associated with a coarse category.

[0038] One aspect of this disclosure provides an apparatus for encoding a grid frame, including a processor. The processor is configured to: determine a predicted value for current coordinates of the grid frame; determine a prediction error for the current coordinates based on the value of the current coordinates and the predicted value for the current coordinates; determine one or more contexts for the prediction error for the current coordinates; perform arithmetic encoding on the prediction error for the current coordinates based on the one or more contexts to generate an arithmetic-encoded prediction error for the current coordinates; and transmit a bitstream including the arithmetic-encoded prediction error, wherein one or more contexts for the prediction error of at least one coordinate associated with a fine category are shared for the prediction error of at least one coordinate associated with a coarse category.

[0039] In an embodiment, one or more contexts of the prediction error for at least one geometric coordinate are shared for the prediction error of at least one texture coordinate.

[0040] In an embodiment, one or more contexts of a truncated unary portion of the prediction error for at least one coordinate associated with a fine category are shared for a truncated unary portion of the prediction error for at least one coordinate associated with a coarse category.

[0041] In an embodiment, one or more contexts of the prefix portion of the prediction error for at least one coordinate associated with a fine category are shared for the prefix portion of the prediction error for at least one coordinate associated with a coarse category.

[0042] In an embodiment, one or more contexts of the suffix portion of the prediction error for at least one coordinate associated with a fine category are shared for the suffix portion of the prediction error for at least one coordinate associated with a coarse category.

[0043] Invention patterns The detailed description set forth below with reference to the accompanying drawings is intended to describe various embodiments and not to represent the only possible implementation of the subject matter. Rather, the detailed description includes specific details for the purpose of providing a thorough understanding of the subject matter of the invention. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the scope of this disclosure. Therefore, the drawings and description are to be considered illustrative rather than restrictive in nature. The same reference numerals designate the same elements.

[0044] In this embodiment, 360° video and 3D volumetric video are emerging as new ways to experience immersive content due to the readily available availability of powerful handheld devices such as smartphones. In this embodiment, 360° video provides consumers with an immersive, “real-life,” “being there” experience by capturing a 360° outside-in view of the world, while 3D volumetric video can provide a full “six degrees of freedom” (6DoF) experience of presence and movement within the content. In this embodiment, users can interactively change their viewpoint and dynamically view any part of the captured scene or object they desire. Display and navigation sensors can track the user’s head movements in real time to determine the area of ​​360° video or volumetric content that the user wants to view or interact with. Multimedia data that is inherently three-dimensional (3D), such as point clouds or 3D polygon meshes, can be used in immersive environments.

[0045] In this embodiment, a point cloud is a set of 3D points and attributes representing the surface or volume of an object, such as color, normals, reflectivity, point size, etc. Point clouds are common in various applications, including games, 3D maps, visualization, medical applications, augmented reality, virtual reality, autonomous driving, multi-view replay, 6DoF immersive media, and a few others. Uncompressed point clouds typically require significant bandwidth for transmission. Therefore, due to high bit rate requirements, point clouds are usually compressed before transmission. In at least one example, compressing 3D objects such as point clouds typically requires dedicated hardware. To avoid requiring dedicated hardware to compress 3D point clouds, they can be transformed into traditional two-dimensional (2D) frames, compressed, and later reconstructed for the user's viewing.

[0046] In embodiments, polygonal 3D meshes, particularly triangular meshes, are another popular format for representing 3D objects. A mesh typically comprises a set of vertices, edges, and faces representing the surface of a 3D object. A triangular mesh is a simple polygonal mesh where the faces are simple triangles covering the surface of a 3D object. In embodiments, one or more attributes may exist associated with the mesh. In one scene, one or more attributes may be associated with each vertex in the mesh. For example, texture attributes (RGB) may be associated with each vertex. In another scene, each vertex may be associated with a pair of coordinates (u, v). The (u, v) coordinates may point to a position in a texture map associated with the mesh. For example, the (u, v) coordinates may refer to the row and column indices in the texture map, respectively. A mesh can be thought of as a point cloud with additional connectivity information.

[0047] Point clouds or meshes can be dynamic, meaning they can change over time. In these cases, a point cloud or mesh at a specific moment can be referred to as a point cloud frame or a mesh frame, respectively. Because point clouds and meshes contain a large amount of data, they need to be compressed for efficient storage and transmission. This is especially true for dynamic point clouds and meshes, which may contain 60 frames or more per second.

[0048] The figures discussed below and the various embodiments used to describe the principles of this disclosure in this patent document are for illustrative purposes only and should not be construed as limiting the scope of this disclosure in any way. Those skilled in the art will understand that the principles of this disclosure can be implemented in any suitably arranged system or apparatus.

[0049] Figure 1 An example communication system 100 according to an embodiment of the present disclosure is shown. Figure 1 The embodiment of the communication system 100 shown is for illustrative purposes only. Other embodiments of the communication system 100 may be used without departing from the scope of this disclosure.

[0050] In one embodiment, the communication system 100 includes a network 102 that facilitates communication between various components within the communication system 100. For example, the network 102 may transmit IP packets, Frame Relay frames, Asynchronous Transfer Mode (ATM) cells, or other information between network addresses. The network 102 may include one or more local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), all or part of a global network such as the Internet, or any other one or more communication systems located in one or more locations.

[0051] In this example, network 102 facilitates communication between server 104 and various client devices 106-116. Client devices 106-116 may be, for example, smartphones, tablets, laptops, personal computers, TVs, interactive displays, wearable devices, head-mounted displays (HMDs), etc. Server 104 may represent one or more servers. Each server 104 includes any suitable computing or processing means that can provide computing services to one or more client devices, such as client devices 106-116. Each server 104 may, for example, include one or more processing means, one or more memories storing instructions and data, and one or more network interfaces facilitating communication through network 102. As described in more detail below, server 104 may send compressed bitstreams representing point clouds or meshes to one or more display devices, such as client devices 106-116. In some embodiments, each server 104 may include an encoder.

[0052] Each client device 106-116 represents any suitable computing or processing device that interacts with at least one server (such as server 104) or other (one or more) computing devices via network 102. Client devices 106-116 include, but are not limited to, desktop computer 106, mobile phone or mobile device 108 (such as a smartphone), personal digital assistant (PDA) 110, laptop computer 112, tablet computer 114 (e.g., with a touchscreen or stylus), and HMD 116. However, any other or additional client devices may be used in the communication system 100. A smartphone represents a class of mobile devices 108 that are handheld devices with a mobile operating system and integrated mobile broadband cellular network connectivity for voice, short message service (SMS), and internet data communications. In embodiments, HMD 116 may display a 360° scene including one or more dynamic or static 3D point clouds. In some embodiments, any of client devices 106-116 may include an encoder, a decoder, or both. For example, mobile device 108 can record 3D volumetric video and then encode the video so that it can be sent to one of client devices 106-116. In another example, laptop computer 112 can be used to generate 3D point clouds or meshes, which are then encoded and sent to one of client devices 106-116.

[0053] In this example, some client devices 108-116 communicate indirectly with network 102. For example, mobile device 108 and PDA 110 communicate via one or more base stations (e.g., BS) 118 (such as cellular base stations or eNodeBs (eNBs), or fifth-generation (5G) base stations implementing new radio (NR) technology, or gNodeBs (gNb)). Additionally, laptop computer 112, tablet computer 114, and HMD 116 communicate via one or more wireless access points 120 (such as IEEE 802.11 wireless access points). Note that these are for illustrative purposes only, and each client device 106-116 may communicate directly with network 102 or indirectly with network 102 via any suitable intermediary device(s) or network(s). In some embodiments, server 104 or any client device 106-116 may be used to compress point clouds or meshes, generate bitstreams representing point clouds or meshes, and send the bitstreams to another client device (such as any client device 106-116).

[0054] In some embodiments, any of client devices 106-114 securely and efficiently transmits information to another device (such as, for example, server 104). Furthermore, any of client devices 106-116 can trigger information transmission between itself and server 104. Any of client devices 106-114, when attached to a head-mounted device via a bracket, can function as a virtual reality (VR) display and operate similarly to HMD 116. For example, mobile device 108 can function similarly to HMD 116 when attached to a bracket system and worn on a user's eyes. Mobile device 108 (or any other client device 106-116) can trigger information transmission between itself and server 104.

[0055] In some embodiments, any of client devices 106-116 or server 104 may create a 3D point cloud or mesh, compress a 3D point cloud or mesh, send a 3D point cloud or mesh, receive a 3D point cloud or mesh, decode a 3D point cloud or mesh, render a 3D point cloud or mesh, or a combination thereof. For example, server 104 may then compress the 3D point cloud or mesh to generate a bitstream, and then send the bitstream to one or more of client devices 106-116. As another example, one of client devices 106-116 may compress a 3D point cloud or mesh to generate a bitstream, and then send the bitstream to another of client devices 106-116 or to server 104.

[0056] although Figure 1 An example of a communication system 100 is shown, but it is possible to compare it with other systems. Figure 1 Various changes can be made. For example, communication system 100 can include any number of each component in any suitable arrangement. Typically, computing and communication systems have a wide variety of configurations, and Figure 1 This disclosure is not intended to limit the scope to any particular configuration. Although Figure 1 An operating environment is shown that can use the various features disclosed in this patent document, but these features can be used in any other suitable system.

[0057] Figure 2 and Figure 3 An example electronic device according to an embodiment of the present disclosure is shown. In particular, Figure 2 Example server 200 is shown, and server 200 can be represented as referenced. Figure 1 The described server 104. In embodiments, server 200 may represent one or more encoders, decoders, local servers, remote servers, cluster computers, and components acting as a single seamless resource pool, cloud-based servers, etc. Server 200 may be comprised of... Figure 1One or more of the client devices 106-116 or another server can access the server.

[0058] Server 200 may represent one or more local servers, one or more compression servers, or one or more encoding servers (such as encoders). In some embodiments, the encoder may perform decoding. Figure 2 As shown, server 200 includes bus system 205, which supports communication between at least one processing device (such as processor 210), at least one storage device 215, at least one communication interface 220 and at least one input / output (I / O) unit 225.

[0059] Processor 210 executes instructions that can be stored in memory 230. Processor 210 may include any suitable number and type of processors or other devices in any suitable arrangement. Example types of processor 210 include microprocessors, microcontrollers, digital signal processors, field-programmable gate arrays, application-specific integrated circuits (ASICs), and discrete circuits.

[0060] In some embodiments, processor 210 may encode a 3D point cloud or mesh stored in storage device 215. In some embodiments, encoding the 3D point cloud also includes decoding the 3D point cloud or mesh to ensure that when the point cloud or mesh is reconstructed, the reconstructed 3D point cloud or mesh matches the 3D point cloud or mesh before encoding.

[0061] Memory 230 and permanent storage 235 are examples of storage device 215, which represents any one or more structures capable of storing and facilitating the retrieval of information, such as data, program code, or other suitable information on a temporary or permanent basis. Memory 230 may represent random access memory or any other suitable one or more volatile or non-volatile storage devices. For example, instructions stored in memory 230 may include instructions for decomposing a point cloud into patches, instructions for packing patches onto 2D frames, instructions for compressing 2D frames, and instructions for encoding 2D frames in a certain order to generate a bitstream. Instructions stored in memory 230 may also include instructions for storing information, such as through a VR headset (e.g., ...). Figure 1 The HMD 116 provides instructions for rendering point clouds on an omnidirectional 360° scene. Persistent storage 235 may contain one or more components or devices that support long-term storage of data, such as read-only memory, hard disk drive, flash memory, or optical disc.

[0062] Communication interface 220 supports communication with other systems or devices. For example, communication interface 220 may include features that facilitate communication via... Figure 1The network interface card or wireless transceiver of network 102. Communication interface 220 can support communication via any suitable physical or wireless communication link(s). For example, communication interface 220 can send a bitstream containing a 3D point cloud to another device (such as one of client devices 106-116).

[0063] I / O unit 225 allows for data input and output. For example, I / O unit 225 can provide connectivity for user input via a keyboard, mouse, keypad, touchscreen, or other suitable input device. I / O unit 225 can also send output to a display, printer, or other suitable output device. However, note that I / O unit 225 can be omitted, such as when I / O interaction with server 200 occurs via a network connection.

[0064] Note that, although Figure 2 Described as representing Figure 1 The server 104 can be a single unit, but the same or similar structure can be used in one or more of various client devices 106-116. For example, desktop computer 106 or laptop computer 112 can have the same structure as... Figure 2 The structures shown are the same or similar.

[0065] Figure 3 An example electronic device 300 is shown, and the electronic device 300 can represent Figure 1 One or more of the client devices 106-116. Electronic device 300 may be a mobile communication device, such as, for example, a mobile station, a subscriber station, a wireless terminal, or a desktop computer (similar to...). Figure 1 Desktop computers 106), portable electronic devices (similar to) Figure 1 Mobile devices 108, PDA 110, laptop computer 112, tablet computer 114, or HMD 116, etc. In some embodiments, Figure 1 One or more of the client devices 106-116 may include configurations identical or similar to those of electronic device 300. In some embodiments, electronic device 300 is an encoder, decoder, or both. For example, electronic device 300 may be used for data transmission, image or video compression, image or video decompression, encoding, decoding, and media rendering applications.

[0066] like Figure 3As shown, electronic device 300 includes an antenna 305, a radio frequency (RF) transceiver 310, a transmit (TX) processing circuitry 315, a microphone 320, and a receive (RX) processing circuitry 325. The RF transceiver 310 may include, for example, an RF transceiver, a Bluetooth transceiver, a Wi-Fi transceiver, a Zigbee transceiver, an infrared transceiver, and various other wireless communication signals. Electronic device 300 also includes a speaker 330, a processor 340, an input / output (I / O) interface (IF) 345, an input 350, a display 355, a memory 360, and one or more sensors 365. The memory 360 includes an operating system (OS) 361 and one or more applications 362.

[0067] In this embodiment, RF transceiver 310 receives incoming RF signals from antenna 305 from an access point (such as a base station, Wi-Fi router, or Bluetooth device) or other devices on network 102 (such as Wi-Fi, Bluetooth, cellular, 5G, LTE, LTE-A, WiMAX, or any other type of wireless network). RF transceiver 310 down-converts the incoming RF signals to generate an intermediate frequency (IF) or baseband signal. The IF or baseband signal is sent to RX processing circuitry 325, which generates a processed baseband signal by filtering, decoding, and / or digitizing the baseband or IF signal. RX processing circuitry 325 sends the processed baseband signal to speaker 330 (e.g., for voice data) or to processor 340 for further processing (e.g., for web browsing data).

[0068] The TX processing circuit 315 receives analog or digital voice data from the microphone 320 or other outgoing baseband data from the processor 340. The outgoing baseband data may include network data, email, or interactive video game data. The TX processing circuit 315 encodes, multiplexes, and / or digitizes the outgoing baseband data to generate a processed baseband or intermediate frequency (IF) signal. The RF transceiver 310 receives the processed outgoing baseband or IF signal from the TX processing circuit 315 and up-converts the baseband or IF signal into an RF signal transmitted via the antenna 305.

[0069] Processor 340 may include one or more processors or other processing devices. Processor 340 may execute instructions (such as OS 361) stored in memory 360 to control the overall operation of electronic device 300. For example, processor 340 may control the reception of forward channel signals and the transmission of reverse channel signals via RF transceiver 310, RX processing circuitry 325, and TX processing circuitry 315 according to well-known principles. Processor 340 may include any suitable number and type of processors or other devices in any suitable arrangement. For example, in some embodiments, processor 340 includes at least one microprocessor or microcontroller. Example types of processor 340 include microprocessors, microcontrollers, digital signal processors, field-programmable gate arrays, application-specific integrated circuits (ASICs), and discrete circuits.

[0070] Processor 340 is also capable of executing other processes and programs residing in memory 360, such as operations for receiving and storing data. Processor 340 can move data into or out of memory 360 as needed during execution. In some embodiments, processor 340 is configured to execute one or more applications 362 based on OS 361 or in response to signals received from one or more external sources or operators. For example, applications 362 may include encoders, decoders, VR or augmented reality (AR) applications (e.g., devices from extended reality fields), camera applications (for still images and video), video call applications, email clients, social media clients, SMS messaging clients, virtual assistants, etc. In some embodiments, processor 340 is configured to receive and transmit media content.

[0071] The processor 340 is also coupled to an I / O interface 345, which provides the electronic device 300 with the ability to connect to other devices, such as client devices 106-114. The I / O interface 345 is the communication path between these accessories and the processor 340.

[0072] Processor 340 is also coupled to input 350 and display 355. An operator of electronic device 300 can use input 350 to input data or information into electronic device 300. Input 350 may be a keyboard, touchscreen, mouse, trackball, voice input, or other device capable of acting as a user interface to allow the user to interact with electronic device 300. For example, input 350 may include voice recognition processing, allowing the user to input voice commands. In another example, input 350 may include a touch panel, (digital) pen sensor, key, or ultrasonic input device. Touch panel can recognize touch input in at least one of the following methods: capacitive, pressure-sensitive, infrared, or ultrasonic. By providing additional input to processor 340, input 350 may be associated with one or more sensors 365 and / or cameras. In some embodiments, sensor 365 includes one or more inertial measurement units (IMUs) (such as accelerometers, gyroscopes, and magnetometers), motion sensors, optical sensors, cameras, pressure sensors, heart rate sensors, altimeters, etc. Input 350 may also include control circuitry. In a capacitive scheme, input 350 can detect touch or proximity.

[0073] Display 355 may be a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic LED (OLED), an active-matrix OLED (AMOLED), or other display capable of rendering text and / or graphics such as those from websites, videos, games, images, etc. The size of display 355 may be suitable for inclusion within an HMD. Display 355 may be a single display screen or multiple display screens capable of creating stereoscopic displays. In some embodiments, display 355 is a head-up display (HUD). Display 355 may display 3D objects, such as 3D point clouds or meshes.

[0074] Memory 360 is coupled to processor 340. A portion of memory 360 may include random access memory (RAM), and another portion of memory 360 may include flash memory or other read-only memory (ROM). Memory 360 may include permanent storage (not shown), which refers to any one or more structures capable of storing and facilitating the retrieval of information such as data, program code, and / or other suitable information. Memory 360 may contain one or more components or devices supporting long-term storage of data, such as read-only memory, hard disk drive, flash memory, or optical disk. Memory 360 may also contain media content. Media content may include various types of media, such as images, videos, 3D content, VR content, AR content, 3D point clouds, meshes, etc.

[0075] The electronic device 300 also includes one or more sensors 365, which can measure physical quantities or detect the activation state of the electronic device 300 and convert the measured or detected information into electrical signals. For example, the sensor 365 may include one or more buttons for touch input, a camera, a gesture sensor, an IMU sensor (such as a gyroscope or gyroscope sensor and accelerometer), an eye-tracking sensor, a barometric pressure sensor, a magnetic sensor or magnetometer, a grip sensor, a proximity sensor, a color sensor, a biophysical sensor, a temperature / humidity sensor, an illuminance sensor, an ultraviolet (UV) sensor, an electromyography (EMG) sensor, an electroencephalography (EEG) sensor, an electrocardiography (ECG) sensor, an IR sensor, an ultrasound sensor, an iris sensor, a fingerprint sensor, a color sensor (such as a red-green-blue (RGB) sensor), etc. The sensor 365 may also include control circuitry for controlling any of the included sensors.

[0076] As discussed in more detail below, one or more of these sensors(s)365 can be used to control the user interface (UI), detect UI input, determine the user's orientation and facing direction for 3D content display recognition, etc. Any of these sensors(s)365 may be located within the electronic device 300, within an auxiliary device operatively connected to the electronic device 300, within a head-mounted device configured to house the electronic device 300, or within a single device in which the electronic device 300 includes a head-mounted device.

[0077] Electronic device 300 can create media content, such as generating virtual objects or capturing (or recording) content via a camera. Electronic device 300 can encode the media content to generate a bitstream, allowing the bitstream to be sent directly to another electronic device or, for example, via... Figure 1 Network 102 is transmitted indirectly. Electronic device 300 can receive bit streams directly from another electronic device, or through, for example, via... Figure 1 The network 102 is indirectly received.

[0078] although Figure 2 and Figure 3 Examples of electronic devices are shown, but more can be found on... Figure 2 and Figure 3 Make various changes. For example, Figure 2 and Figure 3 The various components can be combined, further subdivided, or omitted, and additional components can be added as needed. As a specific example, processor 340 can be divided into multiple processors, such as one or more central processing units (CPUs) and one or more graphics processing units (GPUs). Furthermore, as with computing and communication, electronic devices and servers can have a wide variety of configurations, and Figure 2 and Figure 3 This disclosure is not limited to any particular electronic device or server.

[0079] Additionally, ISO / IEC SC29 / WG07 is currently working on developing a standard for video-based compression for dynamic meshes. In this embodiment, the eighth test model—V-DMC Mesh Test Model (TMM) 8.0—represents the current state of the standard, and was established at the 14th meeting of ISO / IEC SC29 / WG07 in June 2024. In this embodiment, a software implementation of V-DMC TTM 8.0 can be available as software from a git repository. In this embodiment, a committee draft (CD) specification for video-based compression for dynamic meshes is also available.

[0080] The following documents are incorporated herein by reference as if fully set forth herein: i) V-DMCTMM 8.0, ISO / IEC SC29 WG07 N00874, June 2024; ii) CD of V-DMC (ISO / IEC SC29 WG07 N00885, June 2024); iii) CD of V-DMC (ISO / IEC SC29 WG07 N01027, December 2024); and iv) V-DMC8.0, ISO / IEC SC29 WG07 N01099, February 2025.

[0081] Figure 4 and Figure 5 Block diagrams for the V-DMC encoder and decoder are shown respectively.

[0082] like Figure 4As shown, system 400 may include a preprocessing unit 410 that communicates with one or more encoders (e.g., with an atlas encoder 435, a base mesh encoder 440, a displacement encoder 445, and a video encoder 450). In one embodiment, system 400 illustrates the encoding of a dynamic mesh sequence 405, which is multiplexed and transmitted as a visual volumetric video encoded (V3C) bitstream 497. In one embodiment, for each mesh frame, system 400 may create a base mesh 420, which may include fewer vertices than the original mesh. In one embodiment, the base mesh is compressed in a lossy or lossless manner to create a base mesh sub-bitstream 460. In one embodiment, the base mesh 420 is intra-frame encoded, e.g., encoded without predictions from neighboring base mesh frames. In other embodiments, the base mesh 420 is inter-frame encoded, e.g., encoded using predictions from neighboring base mesh frames. In one embodiment, the reconstructed base mesh undergoes subdivision, and then the displacement field between the original mesh and the subdivided reconstructed base mesh is calculated, compressed, and transmitted.

[0083] For example, preprocessing unit 410 may receive dynamic mesh sequence 405. In an embodiment, preprocessing unit 410 may convert dynamic mesh sequence 405 into components: graph 415, base mesh 420, displacement 425, and attributes 430. That is, dynamic mesh sequence 405 may include information about connectivity, geometry, mapping, vertex attributes, and attribute graph. In an embodiment, connectivity information refers to the connections between vertices of dynamic mesh sequence 405. In an embodiment, geometric information refers to the position of each vertex in 3D space, represented as coordinates. In an embodiment, attribute 430 information includes information about vertex or mesh face color, material information, normal direction, texture coordinates, etc. In an embodiment, dynamic mesh sequence 405 may be referred to as dynamic if one or more of connectivity, geometry, mapping, vertex attributes, and / or attribute graph changes.

[0084] In an embodiment, the preprocessing unit 410 may receive a dynamic mesh sequence 405 and send portions of the dynamic mesh sequence to multiple encoders. For example, the dynamic mesh sequence 405 may include a portion of a map 415 that has been preprocessed and sent to the map encoder 435. In an embodiment, the map 415 refers to a set of two-dimensional (2D) bounding boxes placed on a rectangular frame and corresponding to volumes in three-dimensional (3D) space of the rendered volume data, along with their related information, and a list of metadata corresponding to a portion of the mesh surface in 3D space. In an embodiment, the map 415 may include information about geometry (e.g., depth) or texture (e.g., texture map). In an embodiment, the system 400 may utilize the metadata of the map 415 to generate a bitstream 497. For example, the map 415 component provides information about how to perform inverse reconstruction; for example, the map 415 may describe how to perform subdivision of the base mesh 420, how to apply displacement vectors 425 to the subdivided mesh, or how to apply attributes 430 to the reconstructed mesh.

[0085] In an embodiment, the base mesh 420 may be referred to as a simplified low-resolution approximation of the original mesh, and is encoded using any mesh codec.

[0086] In this embodiment, the displacement information provides a displacement vector that can be encoded as a VC3 geometric video component using any video codec.

[0087] In this embodiment, attribute 430 provides additional features and can be encoded by any video codec.

[0088] In one embodiment, preprocessing unit 410 can create a base mesh 420 from a dynamic mesh sequence 405. In another embodiment, preprocessing unit 410 can transform the original mesh into the base mesh based on a series of displacements 425 according to an attribute graph 430. For example, the original dynamic mesh sequence 405 can be downsampled to reduce the number of vertices, for example, to create an extracted mesh. In another embodiment, the extracted mesh undergoes reparameterization to generate the base mesh 420 by applying graph information 415 and a graph encoder 435. In yet another embodiment, subdivision is then applied to the base mesh 420, in part based on the displacement information 425.

[0089] In this embodiment, the graph encoder 435 generates a graph sub-bitstream 455, the base grid encoder 440 generates a base grid sub-bitstream 460, and the video encoder 450 generates an attribute sub-bitstream 470. In this embodiment, the sub-bitstreams are multiplexed at multiplexer 495 to generate and transmit a bitstream 497.

[0090] Figure 5 A block diagram for a decoder according to an embodiment is shown.

[0091] like Figure 5 As shown, system 500 may include a demultiplexer 510 that communicates with one or more decoders (e.g., with a graph decoder 520, a base grid decoder 525, a displacement decoder 530, and a video decoder 535). In one embodiment, system 500 shows decoding a visual volumetric video encoded (V3C) bitstream 505 into a reconstructed dynamic grid sequence 570. In one embodiment, system 500 decodes a base grid sub-bitstream 514 to form a reconstructed base grid 542. In one embodiment, the reconstructed base grid 542 undergoes subdivision in the decoder. In one embodiment, a received displacement field is decompressed and added to the reconstructed base grid to generate the final reconstructed grid in the decoder.

[0092] For example, demultiplexer 510 can receive bitstream 505 and determine spectrogram sub-bitstream 512, base grid sub-bitstream 514, displacement sub-bitstream 516, and attribute sub-bitstream 518. In an embodiment, spectrogram decoder 520 processes spectrogram sub-bitstream 512 information and sends encoded information to base grid processing 550. In an embodiment, base grid decoder 525 decodes base grid sub-bitstream 514 information to generate reconstructed base grid 542. In an embodiment, displacement decoder 530 can decompress displacement sub-bitstream 516 information and send decoded bits 544 to displacement processing unit 555. In an embodiment, system 500 reconstructs grid 560 to generate reconstructed grid 565 by processing base grid 542, decompressed spectrogram sub-bitstream 512 information, and using the processed output and displacement information generated by displacement processing 555. In an embodiment, video decoder 535 can decompress attribute sub-bitstream 518 information and send this information to reconstruction unit 565. In an embodiment, the reconstruction unit 565 can generate a reconstructed dynamic mesh sequence 570 based on the reconstruction mesh 560 and attribute information 546.

[0093] In an embodiment, Figure 6 and Figure 7 An example parallelogram grid prediction is shown. In the embodiment, Figure 6 and Figure 7 This illustrates an intra-frame encoded base grid, for example, that utilizes predictions from neighboring vertices within the same base grid frame. For example, as... Figure 6As shown, the vertex position is predicted based on the positions of available neighboring vertices. In this embodiment, vertex “V”625 is predicted. In such an example, a prediction factor “P”620 for “V”625 is calculated from the available neighboring vertices (vertex 605 “A”, vertex 610 “B”, and vertex 615 “C”). In this embodiment, available neighboring vertices may refer to vertices that have already been sent. For example, the triangle formed by vertices 605 “A”, vertices 610 “B”, and vertices 615 “C” (e.g., the shaded area) may have already been sent when the prediction factor “P”620 was predicted.

[0094] In this embodiment, a parallelogram prediction algorithm is used. Different prediction factors can be used in this embodiment, for example, the average of the vertices, previous vertices, left vertices, right vertices, etc. In the embodiment using parallelogram prediction, the prediction factor "P"620 is determined by the following equation (Equation 1): P = B + C - A In this embodiment, the geometric prediction error “D” is determined by taking the difference between vertex “V”625 and prediction factor “P”620, as shown in the following equation (Equation 2): D = V - P In this embodiment, the prediction error is calculated and transmitted. In this embodiment, each vertex is represented by three-dimensional coordinates (e.g., in X, Y, Z geometric coordinates).

[0095] In the embodiments, multiple parallelograms can be predicted, such as Figure 7 As shown. For example, using a parallelogram to predict from three neighboring triangles (e.g., as shown) Figure 7 The predicted factors “P1” 715, “P2” 720, and “P3” 725 are calculated from the vertices of the triangle (shown in the shaded area). In this embodiment, the final predicted factor “P” 710 is calculated as the average of predicted factors “P1” 715, “P2” 720, and “P3” 725. In this embodiment, the predicted factor is determined by Equation 2 shown above. Figure 7 The parallelogram prediction shown is associated with geometric errors. In an embodiment, the prediction error is determined to occur as referenced. Figure 4 The basic mesh encoder described at position 440 or as referenced Figure 5 The basic mesh decoder is described at position 525.

[0096] Figure 8A and Figure 8B This illustrates the context of a binary arithmetic coding scheme for geometric prediction errors in video-based dynamic mesh coding and decoding (V-DMC).

[0097] For reference Figure 6 and Figure 7 The prediction error can be calculated and transmitted based on the generated base grid. In an embodiment, the prediction error is encoded (e.g., at an entropy encoder or arithmetic encoder) and then transmitted. In an embodiment, the prediction error value is converted to a positive number determined before encoding; for example, non-positive integers (e.g., x ≤ 0) are mapped to odd integers -2x + 1, while positive integers x > 0 are mapped to even integers 2x, as referenced in Table I.1 of V-DMC TMM 8.0. In an embodiment, the prediction error is encoded using an arithmetic coding scheme to generate the transmitted prediction error codeword. In an embodiment, the prediction error codeword may consist of a combination of truncated unary (TU) codes and exponential Golomb (EG) codes. That is, the prediction error codeword may include a first part associated with the TU code, a second part associated with a prefix of the EG code, and a third part associated with a suffix of the EG code. In an embodiment, binary arithmetic encoding and decoding use different contexts to encode and decode different bits of the TU+EG codeword. In this embodiment, the binarization of information allows context modeling to be applied to each binary bit (e.g., each bit position). In this embodiment, the context model is a probabilistic model for one or more binary bits of a TU+EG codeword; for example, the context model stores the probability that each binary bit is "1" or "0".

[0098] In the embodiments, various different types of prediction errors may exist. For example, there may be a reference error. Figure 6 and Figure 7 The geometric prediction error is described. In one example, in V-DMC TMM 7.0, the geometric prediction error is classified into two categories, "fine" and "coarse". In an embodiment, the "fine" category refers to a vertex that is part of at least one parallelogram whose three remaining vertices are available (e.g., have been sent). In an embodiment, the "coarse" category refers to the remaining vertices (e.g., have one or two available vertex neighbors, or are on a boundary, etc.). In an embodiment, no explicit symbol is associated with either category (e.g., with "fine" or "coarse"). In an embodiment, the category can be inferred from neighbor information (e.g., whether the remaining vertices are available).

[0099] In the embodiment, regarding geometric prediction error, Figure 8A Show the context used for the "refined" category. Figure 8B The context used for the "rough" category is shown.

[0100] like Figure 8AAs shown, for the “fine” category of geometric prediction error, there can be up to seven (7) bits of two contexts (e.g., A0 or A1) using TU context 805. In an embodiment, the EG prefix context 810 portion of the codeword uses up to twelve (12) bits, which use twelve (12) contexts (e.g., B0-B11). In an embodiment, the EG suffix context 815 portion of the codeword uses up to twelve (12) bits, which use twelve (12) contexts (e.g., C0-C11). That is, the prediction error codeword can have a variable length, and the actual codeword can use a subset of TU context 805, EG prefix context 810, and EG suffix context 815. In an embodiment, the maximum number of bits used for TU context 805, EG prefix context 810, and EG suffix context 815 is different from seven (7) or twelve (12), respectively. In other words, the maximum number of binary bits can be any number greater than zero. For example, the maximum number of binary bits can be 1, 2, 3, 4, 5, 6, 7, etc.

[0101] like Figure 8B As shown, for the “coarse” category of geometric prediction error, a maximum of seven (7) bits can exist using three contexts (e.g., D0, D1, or D2) of TU context 820. In an embodiment, the EG prefix context 825 portion of the codeword uses twelve (12) bits, which use twelve (12) contexts (e.g., E0-E11). In an embodiment, the EG suffix context 835 portion of the codeword uses twelve (12) bits, which use twelve (12) contexts (e.g., F0-F11). In an embodiment, the maximum number of bits used for TU context 820, EG prefix context 825, and EG suffix context 835 is different from seven (7) or twelve (12), respectively. That is, the maximum number of bits can be any number greater than zero, for example, the maximum number of bits can be 1, 2, 3, 4, 5, 6, 7, etc.

[0102] Figure 9A and Figure 9B This illustrates the context of a binary arithmetic coding scheme for texture coordinate prediction errors in video-based dynamic mesh coding and decoding (V-DMC).

[0103] For reference Figure 8A and Figure 8B As described, there can be many different types of prediction errors. As an example, in V-DMC, for each vertex (e.g., as referenced...), Figure 6 and Figure 7Each vertex described is sent with material properties (e.g., texture coordinates). In an embodiment, texture coordinates map the vertex to a two-dimensional (2D) location in a texture image, which is then used for texture mapping when rendering a three-dimensional (3D) object. In an embodiment, the two-dimensional location in the texture image is typically represented by (U,V) coordinates. In an embodiment, texture coordinates are predicted from the texture coordinates and geometric coordinates of available neighboring vertices. In one example, a prediction error (e.g., actual texture coordinates (T) minus predicted texture coordinates (M)) is determined and sent. Texture prediction error Difference = TM In this embodiment, texture prediction errors are categorized into "fine" and "coarse" categories. For example, a "fine" category refers to a vertex that is part of at least one parallelogram whose three remaining vertices are available (e.g., have been sent), and a "coarse" category refers to the remaining vertices (e.g., have one or two available vertex neighbors, or are on a boundary, etc.). In this embodiment, no explicit symbol is associated with either category (e.g., "fine" or "coarse"). In this embodiment, the category can be inferred from neighbor information (e.g., whether the remaining vertices are available).

[0104] In this embodiment, the prediction error value of the texture coordinate prediction error is converted into a positive number determined before encoding. For example, non-positive integers (e.g., x ≤ 0) are mapped to odd integers -2x + 1, while positive integers x > 0 are mapped to even integers 2x, as referenced in Table I.1 of V-DMC TMM 8.0. In this embodiment, the texture coordinate prediction error is encoded using a binary arithmetic encoding / decoding scheme, that is, the prediction error for the texture coordinate prediction error has a value similar to that of a referenced... Figure 8A and Figure 8B The described geometric prediction error follows a similar format. For example, texture coordinate prediction error can utilize a combination of truncated unary (TU) codes and exponential Golomb (EG) codes. That is, the texture coordinate prediction error codeword can include a first part associated with the TU code, a second part associated with the prefix of the EG code, and a third part associated with the suffix of the EG code. In embodiments, binary arithmetic encoding and decoding use different contexts to encode and decode different bits of the TU+EG codeword. In embodiments, binarization of information allows context modeling to be applied to each bit (e.g., each bit position). In embodiments, the context model is a probabilistic model for one or more bits of the TU+EG codeword; for example, the context model stores the probability that each bit is "1" or "0". In embodiments, based on reference... Figure 7 The values ​​of the neighboring triangles shown are used to select the context (e.g., the context model).

[0105] In the embodiment, regarding texture coordinate prediction error, Figure 9A Show the context used for the "refined" category. Figure 9B The context used for the "rough" category is shown.

[0106] like Figure 9A As shown, for the “fine” category of texture coordinate prediction error, there may be up to seven (7) bits of two contexts (e.g., G0 or G1) using TU context 905. In an embodiment, the EG prefix context 910 portion of the codeword uses up to twelve (12) bits, which use twelve (12) contexts (e.g., H0-H11). In an embodiment, the EG suffix context 915 portion of the codeword uses up to twelve (12) bits, which use twelve (12) contexts (e.g., I0-I11). In an embodiment, the maximum number of bits used for TU context 905, EG prefix context 910, and EG suffix context 915 may differ from seven (7) and twelve (12), respectively. For example, the maximum number of bits may be any number greater than zero, such as 1, 2, 3, 4, 5, 6, etc.

[0107] like Figure 9B As shown, for the “coarse” category of texture coordinate prediction error, a maximum of seven (7) bits may exist for three contexts (e.g., J0, J1, or J2) using TU context 920. In an embodiment, the EG prefix context 925 portion of the codeword uses a maximum of twelve (12) bits, which use twelve (12) contexts (e.g., K0-K11). In an embodiment, the EG suffix context 930 portion of the codeword uses a maximum of twelve (12) bits, which use twelve (12) contexts (e.g., L0-L11). In an embodiment, the maximum number of bits used for TU context 920, EG prefix context 925, and EG suffix context 930 may differ from seven (7) and twelve (12), respectively. For example, the maximum number of bits can be any number greater than zero, such as 1, 2, 3, 4, 5, 6, etc.

[0108] Figure 10A , Figure 10B , Figure 11A and Figure 11B A simplified contextual scheme for binary arithmetic coding and decoding schemes for geometric prediction errors in video-based dynamic mesh coding and decoding (V-DMC) according to embodiments described herein is shown. In the embodiments, for both geometric prediction errors and texture coordinate prediction errors, Figure 10A and Figure 11A A simplified context is provided for the "refined" category. Figure 10B and Figure 11B Simplified contexts for the "rough" category are shown separately. In embodiments, Figures 10 and 11 illustrate examples of contexts for sharing the corresponding TU context, EG prefix context, or EG suffix context.

[0109] For example, in V-DMC TMM 7.0, such as Figure 8A , Figure 8B , Figure 9A and Figure 9B As shown, a total of 106 contexts were used for geometry prediction error and texture coordinate prediction error. However, as described herein (e.g., refer to...), Figure 10A , Figure 10B , Figure 11A and Figure 11B (This refers to using a context that reduces the number of elements. For example, refer to...) Figure 10A , Figure 10B , Figure 11A and Figure 11B A total of 58 contexts are shown. In this embodiment, the additional context model causes the problem of context dilution. In this embodiment, context dilution occurs when there are a large number of contexts but not enough data to suggest an accurate model for all contexts. Therefore, using... Figure 10A , Figure 10B , Figure 11A and Figure 11B The simplified context model reduces the likelihood of context dilution, lowers context memory requirements (e.g., less context is stored in memory), and can reduce the overall complexity of the system. In this embodiment, bits can also be saved because the context used for higher-order bits is better trained.

[0110] For example, such as Figure 10A As shown, for the “fine” category of geometric prediction error, there can be up to seven (7) bits of two contexts (e.g., A0 or A1) using TU context 805. In an embodiment, the EG prefix context 810 portion of the codeword uses twelve (12) bits, which use six (6) contexts (e.g., B0-B5). In such an embodiment, bits 0-5 have their own contexts (e.g., B0-B5), and the context of bit 5 (e.g., B5) is reused from bit 6 onwards (e.g., bits 6-11). That is, without using a reference Figure 8A and Figure 8BThe contexts B6-B11 are described. In an embodiment, the EG suffix context 815 portion of the codeword uses twelve (12) bits, which use six (6) contexts (e.g., C0-C5). In such an embodiment, bits 0-5 have their own contexts (e.g., C0-C5), and the context of bit 5 (e.g., C5) is reused from bit 6 (e.g., bits 6-11). That is, contexts C5-C11 are not used. In an embodiment, the maximum number of bits used for TU context 805, EG prefix context 810, and EG suffix context 815 can be different from seven (7) and twelve (12), respectively. For example, the maximum number of bits can be any number greater than zero, such as 1, 2, 3, 4, 5, 6, etc.

[0111] like Figure 10B As shown, for the “coarse” category of geometric prediction error, there may be up to seven (7) bits using three contexts (e.g., D0, D1, or D2) of TU context 820. In an embodiment, the EG prefix context 825 portion of the codeword uses twelve (12) bits, which use six (6) contexts (e.g., E0-E5). In such an embodiment, bits 0-5 have their own contexts (e.g., E0-E5), and the context of bit 5 (e.g., E5) is repeated starting from bit 6 (e.g., bits 6-11). In an embodiment, the EG suffix context 835 portion of the codeword uses twelve (12) bits, which use six (6) contexts (e.g., F0-F5). In such an embodiment, bits 0-5 have their own contexts (e.g., F0-F5), and the context of bit 5 (e.g., F5) is repeated starting from bit 6 (e.g., bits 6-11). In an embodiment, the maximum number of bits used for the TU context 820, the EG prefix context 825, and the EG suffix context 835 may be different from seven (7) and twelve (12), respectively. For example, the maximum number of bits may be any number greater than zero, such as 1, 2, 3, 4, 5, 6, etc.

[0112] For example, such as Figure 11AAs shown, for the “fine” category of texture coordinate prediction error, there may be up to seven (7) bits of two contexts (e.g., G0 or G1) using TU context 905. In an embodiment, the EG prefix context 910 portion of the codeword uses twelve (12) bits, which use three (3) contexts (e.g., H0-H2). In such an embodiment, bits 0-2 have their own contexts (e.g., H0, H1, H2), and the context of bit 2 (e.g., H2) is repeated from bit 3 (e.g., bits 3-11). In an embodiment, the EG suffix context 915 portion of the codeword uses twelve (12) bits, which use three (3) contexts (e.g., I0-I2). In such an embodiment, bits 0-2 have their own contexts (e.g., I0, I1, I2), and the context of bit 2 (e.g., I2) is reused starting from bit 3 (e.g., bits 3-11). In an embodiment, the maximum number of bits used for TU context 905, EG prefix context 910, and EG suffix context 915 can be different from seven (7) and twelve (12), respectively. For example, the maximum number of bits can be any number greater than zero, such as 1, 2, 3, 4, 5, 6, etc.

[0113] like Figure 11B As shown, for the “coarse” category of texture coordinate prediction error, there may be up to seven (7) bits using three contexts (e.g., J0, J1, or J2) of TU context 920. In an embodiment, the EG prefix context 925 portion of the codeword uses twelve (12) bits, which use three (3) contexts (e.g., K0-K2). In such an embodiment, bits 0-2 have their own contexts (e.g., K0, K1, K2), and the context of bit 2 (e.g., K2) is repeated from bit 3 (e.g., bits 3-11). In an embodiment, the EG suffix context 930 portion of the codeword uses twelve (12) bits, which use three (3) contexts (e.g., L0-L2). In such an embodiment, bits 0-2 have their own contexts (e.g., L0, L1, L2), and the context of bit 2 (e.g., L2) is reused starting from bit 3 (e.g., bits 3-11). In an embodiment, the maximum number of bits used for TU context 920, EG prefix context 925, and EG suffix context 930 can be different from seven (7) and twelve (12), respectively. For example, the maximum number of bits can be any number greater than zero, such as 1, 2, 3, 4, 5, 6, etc.

[0114] In this embodiment, texture coordinate prediction error can use a reduced amount of context compared to geometry prediction error because texture coordinate prediction error is biased towards lower values ​​due to better prediction. In this embodiment, geometry prediction uses geometric information from neighboring vertices, while texture coordinate prediction error uses both geometric and texture coordinate information from neighboring vertices.

[0115] In an embodiment, the prefix portion (e.g., EG prefix context) may use "N" contexts, where bits 0 through N-1 have their own contexts, and the context of bits N-1 is used starting from bit N, as described herein. In an embodiment, the value of N may be a predetermined constant, or it may be transmitted in the bitstream (e.g., in a sequence, frame, stripe, subgrid, etc.). In an embodiment, the value of N may vary based on whether it is a "fine" or "coarse" category, or based on geometric prediction error, texture prediction error, or other material property prediction error. In an embodiment, various values ​​of N may be transmitted in the bitstream (e.g., in a sequence, frame, stripe, subgrid, etc.). For example, as... Figure 11A and Figure 11B As shown, for the EG prefix context 910, the value of N can be four (4), such that the first three (e.g., 4-1 Each binary bit has its own context, and starting from binary bit 4, the context of the third binary bit is used (e.g., H2).

[0116] Figure 12A , Figure 12B , Figure 12C and Figure 12D A simplified contextual scheme for binary arithmetic coding and decoding schemes for geometric prediction errors in video-based dynamic mesh coding and decoding (V-DMC) according to embodiments described herein is shown. In the embodiments, Figure 12A and Figure 12C A simplified context is provided for the "fine" categories used for geometric prediction error and texture coordinate prediction error, respectively. In the embodiment, Figure 12B and Figure 12D Simplified contexts are shown for the “coarse” category used for geometric prediction error and texture coordinate prediction error, respectively.

[0117] In embodiments, contexts can be shared across “fine” and “coarse” categories, as well as across geometry prediction errors, texture prediction errors, or any other material property prediction errors. In at least one embodiment, “fine” and “coarse” categories can be combined, and contexts with a common set and a common number can be used for them. For example, for both the “fine” and “coarse” categories of geometry prediction errors and texture coordinate prediction errors, three (3) contexts (A0-A2) are used for the TU context portion (e.g., TU context 1205, TU context 1220, TU context 1235, and TU context 1250). In embodiments, subsets of these contexts (e.g., A0 and A1) are used for the TU context portion associated with the “fine” category of geometry prediction errors and texture coordinate prediction errors (e.g., TU context 1205 and TU context 1235).

[0118] In one embodiment, for both the “fine” and “coarse” categories of geometry prediction error and texture coordinate prediction error, six (6) contexts (B0-B5) are used for the EG prefix context portion (e.g., EG prefix context 1210, EG prefix context 1225, EG prefix context 1240, and EG prefix context 1255). In another embodiment, a subset of these contexts (e.g., B0-B2) is used for the EG prefix context portion associated with the “coarse” category of texture coordinate prediction error (e.g., EG prefix context 1240 and EG prefix context 1255).

[0119] In one embodiment, for both the “fine” and “coarse” categories of geometric prediction error and texture coordinate prediction error, six (6) contexts (C0-C5) are used for the EG suffix context portion (e.g., EG suffix context 1215, EG suffix context 1230, EG suffix context 1245, and EG suffix context 1260). In another embodiment, a subset of these contexts (e.g., C0-C2) is used for the EG suffix context portion associated with the “coarse” category of texture coordinate prediction error (e.g., EG suffix context 1245 and EG suffix context 1260).

[0120] As mentioned above, in V-DMC TMM 7.0, a total of 106 contexts are used for geometry prediction error and texture coordinate prediction error. However, as described herein (e.g., refer to...), Figures 12A-12D This uses a reduced number of contexts, for example, fifteen (15). Therefore, using a reduced number of contexts reduces context memory requirements (e.g., fewer contexts are stored in memory) and can reduce the overall complexity of the system. In the embodiment, bits can also be saved because the context used for higher-order bits is better trained.

[0121] Figure 13A , Figure 13B , Figure 13C and Figure 13D A simplified contextual scheme is shown for a binary arithmetic codec scheme for geometric prediction errors and texture coordinate prediction errors in video-based dynamic mesh coding and decoding (V-DMC) according to embodiments described herein. In the embodiments, Figure 13A and Figure 13C A simplified context is provided for the "fine" categories used for geometric prediction error and texture coordinate prediction error, respectively. In the embodiment, Figure 13B and Figure 13D Simplified contexts are shown for the “coarse” category used for geometric prediction error and texture coordinate prediction error, respectively.

[0122] In embodiments, the number of TU bits used can vary based on different types of prediction errors. For example, the number of TU bits used for different types of prediction errors (e.g., geometric prediction error, texture coordinate prediction error, etc.) is adaptable based on the prediction error type. (See also...) Figures 13A-13D An example is shown. As shown, TU contexts 1305, 1310, and 1320 use seven (7) TU bits, for example, seven TU bits are used for geometric prediction errors in the “fine” and “coarse” categories and for texture coordinate prediction errors in the “coarse” category. In this example, TU context 1315 utilizes ten (10) bits, for example, ten TU bits are used for texture coordinate prediction errors in the “fine” category. It should be noted that seven and ten bits are used only as examples. The system can implement any number of bits, for example, the system can use 1, 2, 3, 4, 5, etc., numbers of bits for TU contexts.

[0123] In an embodiment, the order "k" of the exponential Golomb code used is adapted based on the type of prediction error; for example, different types of prediction errors can utilize different orders of "k". For instance, geometric prediction errors in the "fine" and "coarse" categories, and texture prediction coordinates in the "coarse" category, can be... k = 2 Used for EG codes. In such an embodiment, the texture prediction coordinates in the "fine" category can be... k = 1 Used for EG codes. In an embodiment, the system can adaptively select the order "k", such as... Figures 13A-13D The adaptive TU length selection shown can be utilized Figures 12A-12D Optimizations (e.g., reduced context in the context of the EG prefix and EG suffix). In embodiments, this combination can save bits.

[0124] Figure 14A and Figure 14BA simplified contextual scheme for binary arithmetic coding and decoding schemes for geometric prediction errors in video-based dynamic mesh coding and decoding (V-DMC) according to embodiments described herein is shown. In the embodiments, Figure 14A A simplified context is provided for the "fine" category used to describe geometric prediction errors. In the embodiment, Figure 14B A simplified context is provided for the “coarse” category used for geometric prediction errors.

[0125] For example, such as Figure 14A As shown, for the “fine” category of geometric prediction error, there may be up to seven (7) bits of two contexts (e.g., A0 or A1) using TU context 1405. In an embodiment, the EG prefix context 1410 portion of the codeword uses twelve (12) bits, which use five (5) contexts (e.g., B0-B4). In such an embodiment, bits 0-4 have their own contexts (e.g., B0-B4), and the context of bit 4 (e.g., B4) is repeated from bit 5 (e.g., bits 5-11). In an embodiment, the EG suffix context 1415 portion of the codeword uses twelve (12) bits, which use five (5) contexts (e.g., C0-C4). In such an embodiment, bits 0-4 have their own contexts (e.g., C0-C4), and the context of bit 4 (e.g., C4) is repeated from bit 5 (e.g., bits 5-11). In an embodiment, the maximum number of bits used for TU context 1405, EG prefix context 1410 and EG suffix context 1415 may be different from seven (7) and twelve (12), respectively. For example, the maximum number of bits may be any number greater than zero, such as 1, 2, 3, 4, 5, 6, etc.

[0126] like Figure 14BAs shown, for the “coarse” category of geometric prediction error, there may be up to seven (7) bits using three contexts (e.g., D0, D1, or D2) of TU context 1420. In an embodiment, the EG prefix context 1425 portion of the codeword uses twelve (12) bits, which use five (5) contexts (e.g., E0-E4). In such an embodiment, bits 0-4 have their own contexts (e.g., E0-E4), and the context of bit 4 (e.g., E4) is repeated from bit 5 (e.g., bits 5-11). In an embodiment, the EG suffix context 1430 portion of the codeword uses twelve (12) bits, which use five (5) contexts (e.g., F0-F4). In such an embodiment, bits 0-4 have their own contexts (e.g., F0-F4), and the context of bit 4 (e.g., F4) is reused starting from bit 5 (e.g., bits 5-11). In an embodiment, the maximum number of bits used for TU context 1420, EG prefix context 1425, and EG suffix context 1430 can be different from seven (7) and twelve (12), respectively. For example, the maximum number of bits can be any number greater than zero, such as 1, 2, 3, 4, 5, 6, etc.

[0127] Figure 15A and Figure 15B A simplified contextual scheme for binary arithmetic coding and decoding schemes for geometric prediction errors in video-based dynamic mesh coding and decoding (V-DMC) according to embodiments described herein is shown. In the embodiments, Figure 15A A simplified context is provided for the "fine" category used to describe texture coordinate prediction errors. In the embodiment, Figure 15B This provides a simplified context for the "rough" category used to describe texture coordinate prediction errors.

[0128] For example, such as Figure 15AAs shown, for the “fine” category of texture coordinate prediction error, there may be up to seven (7) bits of two contexts (e.g., G0 or G1) using TU context 1505. In an embodiment, the EG prefix context 1510 portion of the codeword uses twelve (12) bits, which use four (4) contexts (e.g., H0-H3). In such an embodiment, bits 0-3 have their own contexts (e.g., H0, H1, H2, and H3), and the context of bit 3 (e.g., H3) is repeated from bit 4 (e.g., bits 4-11). In an embodiment, the EG suffix context 1515 portion of the codeword uses twelve (12) bits, which use four (4) contexts (e.g., I0-I3). In such an embodiment, bits 0-3 have their own contexts (e.g., I0, I1, I2, and I3), and the context of bit 3 (e.g., I3) is reused starting from bit 4 (e.g., bits 4-11). In an embodiment, the maximum number of bits used for TU context 1505, EG prefix context 1510, and EG suffix context 1515 may differ from seven (7) and twelve (12), respectively. For example, the maximum number of bits may be any number greater than zero, such as 1, 2, 3, 4, 5, 6, etc.

[0129] like Figure 15BAs shown, for the “coarse” category of texture coordinate prediction error, there may be up to seven (7) bits of three contexts (e.g., J0, J1, or J2) using TU context 1520. In an embodiment, the EG prefix context 1525 portion of the codeword uses twelve (12) bits, which use four (4) contexts (e.g., K0-K3). In such an embodiment, bits 0-3 have their own contexts (e.g., K0, K1, K2, and K3), and the context of bit 3 (e.g., K3) is repeated from bit 4 (e.g., bits 4-11). In an embodiment, the EG suffix context 1530 portion of the codeword uses twelve (12) bits, which use four (3) contexts (e.g., L0-L3). In such an embodiment, bits 0-3 have their own contexts (e.g., L0, L1, L2, and L3), and the context of bit 3 (e.g., L3) is reused starting from bit 4 (e.g., bits 4-11). In an embodiment, the maximum number of bits used for TU context 1520, EG prefix context 1525, and EG suffix context 1530 can be different from seven (7) and twelve (12), respectively. For example, the maximum number of bits can be any number greater than zero, such as 1, 2, 3, 4, 5, 6, etc.

[0130] In this embodiment, texture coordinate prediction error can use a reduced amount of context compared to geometry prediction error because texture coordinate prediction error is biased towards lower values ​​due to better prediction. In this embodiment, geometry prediction uses geometric information from neighboring vertices, while texture coordinate prediction error uses both geometric and texture coordinate information from neighboring vertices.

[0131] Figure 16A , Figure 16B , Figure 16C and Figure 16D A simplified contextual scheme for binary arithmetic coding and decoding schemes for geometric prediction errors in video-based dynamic mesh coding and decoding (V-DMC) according to embodiments described herein is shown. In the embodiments, Figure 16A and Figure 16C A simplified context is provided for the "fine" categories used for geometric prediction error and texture coordinate prediction error, respectively. In the embodiment, Figure 16B and Figure 16DSimplified contexts for the “coarse” category of geometric prediction error and texture coordinate prediction error are shown separately. It should be noted that while a maximum of seven (7) bits are shown for TU contexts 1205, 1220, 1235, and 1250 (e.g., for the “fine” and “coarse” categories of geometric or texture coordinate prediction errors), any maximum number of bits can be used. For example, the maximum number of bits for the TU portion of a codeword can be 1, 2, 3, 4, 5, 6, etc. Additionally, while a maximum of twelve (12) bits are shown for the EG prefix and EG suffix portions of a codeword (e.g., EG prefix contexts 1210, 1225, 1240, 1255, 1215, 1230, 1245, and 1260), any maximum number of bits can be used. For example, the maximum number of binary bits used for the EG prefix and EG suffix portions of a codeword can be 1, 2, 3, 4, 5, 6, etc.

[0132] In an embodiment, a number of N contexts (e.g., P0, P1, ..., PN-1) are reserved for encoding “fine” and “coarse” categories for geometric prediction errors, texture coordinate prediction errors, and other attribute prediction errors (e.g., normal prediction errors, etc.). In an embodiment, (e.g., as...) Figures 16A-16D As shown), the EG prefix portion (e.g., EG prefix context 1210, EG prefix context 1225, EG prefix context 1240, EG prefix context 1255) can use a subset M of these contexts, for example, where M ≤ N For example, binary bits 0 through M-1 use contexts P0, P1, ..., PM-1, respectively. In such an embodiment, the context of binary bit M-1 is used starting from binary bit M. In an embodiment, the value of M can vary based on different types of prediction errors (e.g., based on "fine" category, "coarse" category, geometry prediction error, texture coordinate prediction error, attribute prediction error, normal prediction error, etc.). In an embodiment, the EG suffix portion (e.g., EG suffix context 1215, EG suffix context 1230, EG suffix context 1245, EG suffix context 1260) can use a subset Q of context N, for example, where Q ≤ NFor example, bits 0 through Q-1 use contexts P0, P1, ..., PQ-1, respectively. In such an embodiment, the context of bit Q-1 is used starting from bit Q. In an embodiment, the value of Q can vary based on different types of prediction errors (e.g., based on "fine" category, "coarse" category, geometric prediction error, texture coordinate prediction error, attribute prediction error, normal prediction error, etc.). In an embodiment, the values ​​of M and Q for different prediction types can be predetermined constants, or they can be transmitted in a bitstream (e.g., in a sequence, frame, stripe, subgrid, etc.).

[0133] In embodiments, contexts can be shared across “fine” and “coarse” categories, as well as across geometry prediction errors, texture prediction errors, or any other material property prediction errors. In at least one embodiment, “fine” and “coarse” categories can be combined, and contexts with a common set and a common number can be used for them. For example, for both the “fine” and “coarse” categories of geometry prediction errors and texture coordinate prediction errors, three (3) contexts (A0-A2) are used for the TU context portion (e.g., TU context 1205, TU context 1220, TU context 1235, and TU context 1250). In embodiments, subsets of these contexts (e.g., A0 and A1) are used for the TU context portion associated with the “fine” category of geometry prediction errors and texture coordinate prediction errors (e.g., TU context 1205 and TU context 1235).

[0134] In one embodiment, for both the “fine” and “coarse” categories of geometric prediction error and texture coordinate prediction error, five (5) contexts (B0-B4) are used for the EG prefix context portion (e.g., EG prefix context 1210, EG prefix context 1225, EG prefix context 1240, and EG prefix context 1255). In another embodiment, a subset of these contexts (e.g., values ​​M, B0-B3) is used for the EG prefix context portion associated with the “coarse” category of texture coordinate prediction error (e.g., EG prefix context 1240 and EG prefix context 1255).

[0135] In an embodiment, for both the “fine” and “coarse” categories of geometric prediction error and texture coordinate prediction error, five (5) contexts (C0-C4) are used for the EG suffix context portion (e.g., EG suffix context 1215, EG suffix context 1230, EG suffix context 1245, and EG suffix context 1260). In an embodiment, a subset of these contexts (e.g., C0-C3) is used for the EG suffix context portion associated with the “coarse” category of texture coordinate prediction error (e.g., EG suffix context 1245 and EG suffix context 1260). In an embodiment, Figures 16A-16D The values ​​of M and Q differ from the value of N, for example, different amounts are used for subsets of texture coordinate prediction error and geometric prediction error, where the values ​​of M and Q are based on the type of prediction error.

[0136] Figure 17A , Figure 17B , Figure 17C and Figure 17D A simplified contextual scheme for binary arithmetic coding and decoding schemes for geometric prediction errors in video-based dynamic mesh coding and decoding (V-DMC) according to embodiments described herein is shown. In the embodiments, Figure 17A and Figure 17C A simplified context is provided for the "fine" categories used for geometric prediction error and texture coordinate prediction error, respectively. In the embodiment, Figure 17B and 17D Simplified contexts for the “coarse” category of geometric prediction error and texture coordinate prediction error are shown separately. It should be noted that while a maximum of seven (7) bits are shown for TU contexts 1705, 1720, 1735, and 1750 (e.g., for the “fine” and “coarse” categories of geometric or texture coordinate prediction errors), any maximum number of bits can be used. For example, the maximum number of bits for the TU portion of a codeword can be 1, 2, 3, 4, 5, 6, etc. Additionally, while a maximum of twelve (12) bits are shown for the EG prefix and EG suffix portions of a codeword (e.g., EG prefix contexts 1710, 1725, 1740, 1755, 1715, 1730, 1745, and 1760), any maximum number of bits can be used. For example, the maximum number of binary bits used for the EG prefix and EG suffix portions of a codeword can be 1, 2, 3, 4, 5, 6, etc.

[0137] In an embodiment, a number of N contexts (e.g., P0, P1, ..., PN-1) are reserved for encoding “fine” and “coarse” categories for geometric prediction errors, texture coordinate prediction errors, and other attribute prediction errors (e.g., normal prediction errors, etc.). In an embodiment, (e.g., as...) Figures 17A-17D As shown), the EG prefix portion (e.g., EG prefix context 1710, EG prefix context 1725, EG prefix context 1740, EG prefix context 1755) can use a subset M of these contexts, for example, where M ≤ NFor example, bits 0 through M-1 use contexts P0, P1, ..., PM-1, respectively. In such an embodiment, the context of bit M-1 is used starting from bit M, or optionally, it can be bypassed and encoded, such as... Figures 17A to 17D As shown in the example. In this embodiment, the value of M can vary based on different types of prediction errors (e.g., based on a "fine" category, a "coarse" category, geometric prediction error, texture coordinate prediction error, attribute prediction error, normal prediction error, etc.). In this embodiment, the EG suffix portion (e.g., EG suffix context 1715, EG suffix context 1730, EG suffix context 1745, EG suffix context 1760) can use a subset Q of context N, for example, where Q ≤ N For example, bits 0 through Q-1 use contexts P0, P1, ..., PQ-1, respectively. In such an embodiment, the context of bit Q-1 is used starting from bit Q, or optionally it can be bypassed and encoded, such as... Figures 17A to 17D As shown in the example. In this embodiment, the value of Q can vary based on different types of prediction errors (e.g., based on a "fine" category, a "coarse" category, geometric prediction error, texture coordinate prediction error, attribute prediction error, normal prediction error, etc.). In this embodiment, the values ​​of M and Q for different prediction types can be predetermined constants, or they can be transmitted in a bitstream (e.g., in a sequence, frame, stripe, subgrid, etc.).

[0138] In embodiments, context can be shared across “fine” and “coarse” categories, as well as across geometric prediction errors, texture prediction errors, or any other material property prediction errors. In embodiments, sharing context can lead to significant context and bit savings because the context used for higher-order bits is better trained, and the context is better initialized across different types of prediction errors.

[0139] For example, for both the “fine” and “coarse” categories of geometric prediction error and texture coordinate prediction error, three (3) contexts (A0-A2) are used in the TU context portion (e.g., TU context 1705, TU context 1720, TU context 1735, and TU context 1750). In at least one embodiment, a subset of these contexts (e.g., A0 and A1) is used in the TU context portion associated with the “fine” category of geometric prediction error and texture coordinate prediction error (e.g., TU context 1705 and TU context 1735).

[0140] In one embodiment, for both the “fine” and “coarse” categories of geometry prediction error and texture coordinate prediction error, five (5) contexts (B0-B4) are used for the EG prefix context portion (e.g., EG prefix context 1710, EG prefix context 1725, EG prefix context 1740, and EG prefix context 1755). In at least one embodiment, a subset of these contexts (e.g., B0-B3) is used for the EG prefix context portion associated with the “coarse” category of texture coordinate prediction error (e.g., EG prefix context 1740 and EG prefix context 1755). In at least one embodiment, portions of the EG prefix bits are bypassed, indicated by “B”. For example, for EG prefix 1710 and EG prefix 1175 (e.g., for the “fine” and “coarse” categories of geometry prediction error), bits 0-4 have their own contexts (e.g., B0-B4), and bits 5 onwards (e.g., bits 5-11) are bypassed. In another example, for EG prefix 1740 and EG prefix 1755 (e.g., for “fine” and “coarse” categories of texture coordinate prediction error), bits 0-3 have their own context (e.g., B0-B3), bit 4 reuses the context of bit 3 (e.g., B3), and bits 5 onwards (e.g., bits 5-11) are bypassed and encoded.

[0141] In one embodiment, for both the “fine” and “coarse” categories of geometry prediction error and texture coordinate prediction error, five (5) contexts (C0-C4) are used for the EG suffix context portion (e.g., EG suffix context 1715, EG suffix context 1730, EG suffix context 1745, and EG suffix context 1760). In at least one embodiment, a subset of these contexts (e.g., C0-C3) is used for the EG suffix context portion associated with the “coarse” category of texture coordinate prediction error (e.g., EG suffix context 1745 and EG suffix context 1760). For example, for EG suffix 1715 and EG suffix 1730 (e.g., for the “fine” and “coarse” categories of geometry prediction error), bits 0-4 have their own contexts (e.g., C0-C4), and the context of bit 4 (e.g., C4) is reused from bit 5 onwards (e.g., bits 5-11). In another example, for EG suffix 1745 and EG suffix 1760 (e.g., for the "fine" and "coarse" categories of texture coordinate prediction error), bits 0-3 have their own contexts (e.g., C0-C3), bit 4 reuses the context of bit 3 (e.g., C3), and bit 3 (e.g., C3) is reused from bit 5 onwards (e.g., bits 5-11). In this embodiment, in V-DMC TMM 8.0, a total of 78 contexts are used for geometry and texture coordinate prediction errors. In this embodiment, for example, as... Figures 17A-17D As shown, using a total of 13 contexts significantly reduces context memory storage and complexity.

[0142] In the embodiments, Table 1 below can show the requirements for having Figures 17A-17D Syntax elements of the binary arithmetic encoding / decoding scheme for the context shown: [Table 1]

[0143]

[0144] In at least one embodiment, syntax elements mesh_position_fine_residual This refers to the geometric prediction error of the "fine category", a syntax element. mesh_position_coarse_residual This refers to the geometric prediction error of the "rough category," a syntax element. mesh_attribute_fine_residual This refers to the attribute prediction error of the "fine category" (e.g., including texture coordinate prediction error). TEXCORD Normal prediction error ( NORMAL ) or material prediction error ( MATERIAL ID ), and syntax elements mesh_attribute_coarse_residual This refers to the attribute prediction error for the "roughness category" (e.g., including texture coordinate prediction error). TEXCORD Normal prediction error ( NORMAL In at least one embodiment, nbPfxCtx It can refer to the number of prefix contexts, and nbSfxCtx It can refer to the number of contexts associated with the suffix.

[0145] In at least one embodiment, CtxTbl Elements can refer to a context table, and CtxIdx Elements can refer to context identifiers. In at least one embodiment, when geometric prediction errors and texture coordinate prediction errors across "fine" and "coarse" categories share a context, CtxTbl The value can be one (1). In the embodiment, CtxIdx Part of the first column of identifier codewords, CtxIdx The second column identifies the position of the codeword (e.g., the binary bit number), and CtxIdx The third column identifies the context count. For example, Offset (Offset) can refer to the TU part of the codeword. Prefix (Prefix) can refer to the prefix part of a codeword, and Suffix (Suffix) can refer to the suffix portion of a codeword. In an embodiment, the position of the codeword is determined based on provided conditions. For example, a position column can indicate... Offset Partially spanning at least one binary bit to a maximum number determined by the number of TU binary bits used, then indicating Prefix The portion extends from the TU portion (e.g., at bit 3) to a maximum number determined by the number of suffix bits used (e.g., from bit 3 to the number of suffix bits used). In embodiments, Table 1 may also indicate when a context or bypass is reused. For example, BinIdxPfx<= 4 This can indicate the value used to assign the binary bits 0 to 4 to the EG prefix, while BinIdxPfx>4 This can indicate the value used to assign a prefix binary bit greater than 5. In an embodiment, the count can refer to the maximum number of context bits used for a given portion of the codeword.

[0146] In at least one embodiment, Table 1 is included as Table K-8 in the CD of V-DMC (ISO / IEC SC29 WG07N00885, June 2024), for example, a table of values ​​of CtxTbl and CtxIdx for syntax elements used in MPEG Edge Breaker binarized ae(v) encoding.

[0147] Figure 18A , Figure 18B , Figure 18C and Figure 18DA simplified contextual scheme for binary arithmetic coding and decoding schemes for geometric prediction errors in video-based dynamic mesh coding and decoding (V-DMC) according to embodiments described herein is shown. In the embodiments, Figure 18A and Figure 18C A simplified context is provided for the "fine" categories used for geometric prediction error and texture coordinate prediction error, respectively. In the embodiment, Figure 18B and Figure 18D Simplified contexts for the “coarse” category of geometric prediction error and texture coordinate prediction error are shown separately. It should be noted that while a maximum number of seven (7) bits is shown for TU contexts 1805, 1820, 1835, and 1850 (e.g., for the “fine” and “coarse” categories of geometric or texture coordinate prediction errors), any maximum number of bits can be used. For example, the maximum number of bits for the TU portion of a codeword can be 1, 2, 3, 4, 5, 6, etc. Additionally, while a maximum number of twelve (12) bits is shown for the EG prefix and EG suffix portions of a codeword (e.g., EG prefix contexts 1810, 1825, 1840, 1855, 1815, 1830, 1845, and 1860), any maximum number of bits can be used. For example, the maximum number of binary bits used for the EG prefix and EG suffix portions of a codeword can be 1, 2, 3, 4, 5, 6, etc.

[0148] For example, for both the “fine” and “coarse” categories of geometric prediction error and texture coordinate prediction error, three (3) contexts (A0-A2) are used in the TU context portion (e.g., TU context 1805, TU context 1820, TU context 1835, and TU context 1850). In at least one embodiment, a subset of these contexts (e.g., A0 and A1) is used in the TU context portion (e.g., TU context 1805 and TU context 1835) associated with the “fine” category of geometric prediction error and texture coordinate prediction error.

[0149] In one embodiment, for both the “fine” and “coarse” categories of geometry prediction error and texture coordinate prediction error, five (5) contexts (B0-B4) are used for the EG prefix context portion (e.g., EG prefix context 1810, EG prefix context 1825, EG prefix context 1840, and EG prefix context 1855). In at least one embodiment, a subset of these contexts (e.g., B0-B3) is used for the EG prefix context portion associated with the “coarse” category of texture coordinate prediction error (e.g., EG prefix context 1840 and EG prefix context 1855). In at least one embodiment, portions of the EG prefix bits are bypassed and indicated by “B”. For example, for EG prefix 1810 and EG prefix 1825 (e.g., for the “fine” and “coarse” categories of geometry prediction error), bits 0-4 have their own contexts (e.g., B0-B4), bit 5 reuses the context of bit 4 (e.g., B4), and bits 6 onwards (e.g., bits 6-11) are bypassed. In another example, for EG prefix 1840 and EG prefix 1855 (e.g., for “fine” and “coarse” categories of texture coordinate prediction error), bits 0-3 have their own context (e.g., B0-B3), bits 4 and 5 reuse the context of bit 3 (e.g., B3), and bits 6 onwards (e.g., bits 6-11) are bypassed and encoded.

[0150] In one embodiment, for both the “fine” and “coarse” categories of geometry prediction error and texture coordinate prediction error, five (5) contexts (C0-C4) are used for the EG suffix context portion (e.g., EG suffix context 1815, EG suffix context 1830, EG suffix context 1845, and EG suffix context 1860). In at least one embodiment, a subset of these contexts (e.g., C0-C3) is used for the EG suffix context portion associated with the “coarse” category of texture coordinate prediction error (e.g., EG suffix context 1845 and EG suffix context 1860). In at least one embodiment, a portion of the EG suffix binary bits is bypassed and indicated by “B”. For example, for EG suffix 1815 and EG suffix 1830 (e.g., for the “fine” and “coarse” categories of geometric prediction error), bits 0-4 have their own context (e.g., C0-C4), bit 5 reuses the context of bit 4 (e.g., C4), and bits 6 onwards (e.g., bits 6-11) are bypassed. In another example, for EG suffix 1845 and EG suffix 1860 (e.g., for the “fine” and “coarse” categories of texture coordinate prediction error), bits 0-3 have their own context (e.g., C0-C3), bits 4 and 5 reuse the context of bit 3 (e.g., C3), and bits 6 onwards (e.g., bits 6-11) are bypassed.

[0151] In an embodiment, Table 2 below can show the following table with Figures 18A-18D Syntax elements of the binary arithmetic encoding / decoding scheme for the context shown: [Table 2]

[0152]

[0153] In at least one embodiment, syntax elements mesh_position_fine_residual This refers to the geometric prediction error of the "fine category", a syntax element. mesh_position_coarse_residual This refers to the geometric prediction error of the "rough category," a syntax element. mesh_attribute_fine_residual This refers to the attribute prediction error of the "fine category" (e.g., including texture coordinate prediction error). TEXCORD ), normal prediction error ( NORMAL ) or material prediction error ( MATERIAL ID ), and syntax elements mesh_attribute_coarse_residual This refers to the attribute prediction error for the "roughness category" (e.g., including texture coordinate prediction error). TEXCORD ), normal prediction error ( NORMAL In at least one embodiment, nbPfxCtxIt can refer to the number of prefix contexts, and nbSfxCtx It can refer to the number of contexts associated with the suffix.

[0154] In at least one embodiment, CtxTbl Elements can refer to a context table, and CtxIdx Elements can refer to context identifiers. In at least one embodiment, when geometric prediction errors and texture coordinate prediction errors across "fine" and "coarse" categories share a context, CtxTbl The value can be one (1). In the embodiment, CtxIdx Part of the first column of identifier codewords, CtxIdx The second column identifies the position of the codeword (e.g., the binary bit number), and CtxIdx The third column indicates the context count. In an embodiment, when the context is reused across geometry and texture attribute prediction errors (e.g., and across “fine” and “coarse” categories), the count can be zero (0) for geometry prediction of the “coarse” category and for texture coordinate prediction (e.g., both “fine” and “coarse”). That is, because the context is reused, there is no additional context for counting.

[0155] In an embodiment, Offset (Offset) can refer to the TU part of the codeword. Prefix (Prefix) can refer to the prefix part of a codeword, and Suffix (Suffix) can refer to the suffix portion of a codeword. In an embodiment, the position of the codeword is determined based on provided conditions. For example, a position column can indicate... Offset Partially spanning at least one binary bit to a maximum number determined by the number of TU binary bits used, then indicating Prefix The portion extends from the TU portion (e.g., at bit 3) to a maximum number determined by the number of suffix bits used (e.g., the number of suffix bits from bit 3 to the number of suffix bits used). In embodiments, Table 2 may also indicate when a context or bypass is reused. For example, Suffix (BinIdxSfx≤5) and Suffix(BinIdxSfx>5) It can indicate when to use context when the binary number is 5 or less, and when to bypass when the binary number is 6 or more.

[0156] In at least one embodiment, a portion of Table 2 is included as Table K-12 in the DIS of V-DMC (ISO / IEC SC29 WG07N01099), for example, a table of values ​​for CtxTbl and CtxIdx of syntax elements used for MPEG Edge Breaker binarized ae(v) encoding.

[0157] Figure 19This is a flowchart illustrating the operation of a basic mesh decoder according to an embodiment. In at least one embodiment, it can be provided by, as referenced... Figure 5 The basic mesh decoder 525 described is used to perform reference. Figure 19 The described operation.

[0158] In operation 1905, the base grid decoder (e.g., a processor of the base grid decoder) performs arithmetic decoding on one or more codewords corresponding to one or more prediction errors associated with a base grid frame, wherein the one or more prediction errors are associated with a fine-grained or coarse-grained category. In at least one embodiment, the base grid frame is decoded using the Moving Picture Experts Group (MPEG) EdgeBreaker (MEB) static grid codec. In at least one embodiment, the one or more codewords comprise one or more portions. For example, the codewords may include portions associated with truncated unary binarization, portions associated with exponential Golomb prefix binarization, and portions associated with exponential Golomb suffix binarization, as referenced. Figures 18A-18D Description. In at least one embodiment, the one or more prediction errors include one of fine geometry prediction error, coarse geometry prediction error, fine texture prediction error, or coarse texture prediction error. In embodiments, the prediction error may include at least one of fine normal prediction error, coarse normal prediction error, fine attribute prediction error, or coarse attribute prediction error.

[0159] In operation 1910, the basic trellis decoder can assign one or more contexts for decoding one or more codewords corresponding to one or more prediction errors. In at least one embodiment, the one or more contexts can be context models, which are probabilistic models of one or more bits for one or more codewords (e.g., TU+EG codewords). That is, the context stores the probability that each bit is '1' or '0'.

[0160] In Operation 1915, the base grid decoder can share one or more contexts that will be used for one or more prediction errors associated with the fine or coarse category. That is, as referenced Figures 18A-18D The system can share context across “fine” and “coarse” categories, across geometric and texture prediction errors, and across multiple bit positions within a portion of a codeword.

[0161] For example, the base mesh decoder can share one or more contexts that will be used between at least one of fine geometry prediction error, coarse geometry prediction error, fine texture prediction error, or coarse texture prediction error.

[0162] In embodiments, multiple bit positions within a truncated unary binarization share the same context in one or more contexts. In some cases, there are three contexts associated with a portion of a codeword that is associated with the truncated unary binarization. In this case, three contexts are used for a coarse category, and a subset of the three contexts is used for a fine category; for example, the fine category uses two contexts. In embodiments, there is a first number of bits for the truncated unary binarization associated with the fine category, and a second number of bits for the truncated unary binarization associated with the coarse category.

[0163] In at least one embodiment, multiple bit positions within an exponential Golomb prefix binarization share the same context in one or more contexts. In one embodiment, there are five contexts associated with a portion of a codeword associated with the exponential Golomb prefix binarization. In such an embodiment, the five contexts are used for a first prediction error type of one or more prediction errors, and a subset of the five contexts is used for a second prediction error type of one or more prediction errors. In one example, the first prediction error type is a geometric prediction error, and the second prediction error type is a texture coordinate prediction error. In at least one embodiment, the subset of the five contexts consists of four contexts used for the second prediction error type.

[0164] In at least one embodiment, multiple bit positions within an exponential Golomb postfix binarization share the same context in one or more contexts. In one embodiment, there are five contexts associated with a portion of a codeword associated with the exponential Golomb postfix binarization. In such an embodiment, the five contexts are used for a first prediction error type of one or more prediction errors, and a subset of the five contexts is used for a second prediction error type of one or more prediction errors. In one example, the first prediction error type is a geometric prediction error, and the second prediction error type is a texture coordinate prediction error. In at least one embodiment, the subset of the five contexts consists of four contexts used for the second prediction error type.

[0165] Figure 20 This is a flowchart illustrating the operation of the V-DMC decoder according to an embodiment.

[0166] In Operation 2005, the V-DMC decoder receives a bitstream that includes the prediction error, which is arithmetically encoded for the current coordinates of the grid frame.

[0167] In an embodiment, the current coordinates may be one of the geometric coordinates associated with the fine category, the geometric coordinates associated with the coarse category, the texture coordinates associated with the fine category, or the texture coordinates associated with the coarse category. In another embodiment, the current coordinates may be one of the material property coordinates associated with the fine category, the material property coordinates associated with the coarse category, the normal coordinates associated with the fine category, or the normal coordinates associated with the coarse category.

[0168] In an embodiment, the prediction error of the arithmetic coding may be one of the following: the geometric coordinate prediction error mesh_position_fine_residual associated with the fine category, the geometric coordinate prediction error mesh_position_coarse_residual associated with the coarse category, the texture coordinate prediction error mesh_attribute_fine_residual associated with the fine category, or the texture coordinate prediction error mesh_attribute_coarse_residual associated with the coarse category.

[0169] In Operation 2010, the V-DMC decoder determines one or more contexts of the prediction error for the arithmetic encoding used for the current coordinates.

[0170] In an embodiment, the V-DMC decoder determines one or more contexts for the prediction error of the current coordinates, as described above, for example, in Figures 16A to 18D And in Tables 1 and 2. For example, when the V-DMC decoder performs arithmetic encoding on a corresponding bit of the prediction error, the V-DMC decoder may determine the context identified by the context index ctxIdx in the context table ctxTbl for the prediction error, or it may determine the bypass as the context for the prediction error, as shown in Table 1 or Table 2.

[0171] In an embodiment, one or more contexts of the prediction error for at least one geometric coordinate may be shared for the prediction error for at least one texture coordinate.

[0172] In an embodiment, one or more contexts of a truncated unary portion of the prediction error for at least one coordinate associated with a fine category may be shared for a truncated unary portion of the prediction error for at least one coordinate associated with a coarse category.

[0173] In an embodiment, one or more contexts of the prefix portion of the prediction error for at least one coordinate associated with a fine category may be shared for the prefix portion of the prediction error for at least one coordinate associated with a coarse category.

[0174] In an embodiment, one or more contexts of the suffix portion of the prediction error for at least one coordinate associated with a fine category may be shared for the suffix portion of the prediction error for at least one coordinate associated with a coarse category.

[0175] In Operation 2015, the V-DMC decoder performs arithmetic decoding on the prediction error of the arithmetic code based on one or more contexts to determine the prediction error for the current coordinates.

[0176] In Operation 2020, the V-DMC decoder determines the predicted values ​​for the current coordinates.

[0177] In Operation 2025, the V-DMC decoder determines the coordinate value of the current coordinate based on the prediction error and the predicted value for the current coordinate.

[0178] In an embodiment, the V-DMC decoder can determine the coordinate value of the current coordinate by summing the prediction error and the predicted value.

[0179] Figure 21 This is a flowchart illustrating the operation of the V-DMC encoder according to an embodiment.

[0180] In this embodiment, it can be performed by a V-DMC encoder. Figure 21 The operation.

[0181] In operation 2105, the V-DMC encoder determines the predicted values ​​for the current coordinates of the grid frame.

[0182] In an embodiment, the current coordinates may be one of the geometric coordinates associated with the fine category, the geometric coordinates associated with the coarse category, the texture coordinates associated with the fine category, or the texture coordinates associated with the coarse category. In another embodiment, the current coordinates may be one of the material property coordinates associated with the fine category, the material property coordinates associated with the coarse category, the normal coordinates associated with the fine category, or the normal coordinates associated with the coarse category.

[0183] In operation 2110, the V-DMC encoder determines the prediction error for the current coordinate based on the coordinate value of the current coordinate and the predicted value used for the current coordinate.

[0184] In one embodiment, the V-DMC encoder subtracts the predicted value from the coordinate value of the current coordinate to determine the prediction error for the current coordinate.

[0185] In operation 2115, the V-DMC encoder determines one or more contexts for the prediction error of the current coordinates.

[0186] In an embodiment, the V-DMC encoder determines one or more contexts for the prediction error of the current coordinates, as described above, for example, in Figures 16A to 18D And in Tables 1 and 2. For example, when the V-DMC encoder performs arithmetic encoding on a corresponding bit of the prediction error, the V-DMC encoder can determine the context identified by the context index ctxIdx in the context table ctxTbl for the prediction error, or it can determine the bypass as the context for the prediction error, as shown in Table 1 or Table 2.

[0187] In an embodiment, one or more contexts of the prediction error for at least one geometric coordinate may be shared for the prediction error for at least one texture coordinate.

[0188] In an embodiment, one or more contexts of a truncated unary portion of the prediction error for at least one coordinate associated with a fine category may be shared for a truncated unary portion of the prediction error for at least one coordinate associated with a coarse category.

[0189] In an embodiment, one or more contexts of the prefix portion of the prediction error for at least one coordinate associated with a fine category may be shared for the prefix portion of the prediction error for at least one coordinate associated with a coarse category.

[0190] In an embodiment, one or more contexts of the suffix portion of the prediction error for at least one coordinate associated with a fine category may be shared for the suffix portion of the prediction error for at least one coordinate associated with a coarse category.

[0191] In operation 2120, the V-DMC encoder performs arithmetic encoding on the prediction error for the current coordinates based on one or more contexts to generate an arithmetic-encoded prediction error for the current coordinates. In an embodiment, the arithmetic-encoded prediction error may be one of the following: a geometric coordinate prediction error mesh_position_fine_residual associated with a fine category, a geometric coordinate prediction error mesh_position_coarse_residual associated with a coarse category, a texture coordinate prediction error mesh_attribute_fine_residual associated with a fine category, or a texture coordinate prediction error mesh_attribute_coarse_residual associated with a coarse category.

[0192] In operation 2125, the V-DMC encoder sends a bit stream that includes the prediction error coded arithmetically.

[0193] Various descriptive blocks, units, modules, components, methods, operations, instructions, items, and algorithms can be implemented or executed using processing circuitry.

[0194] Unless otherwise specified, references to elements in the singular form are not intended to refer to one and only one, but rather to one or more. For example, a “one” module can refer to one or more modules. Without further constraints, elements beginning with “a,” “an,” “the,” or “the” do not preclude the existence of other identical elements.

[0195] Titles and subtitles (if any) are used for convenience only and do not limit the technical scope of the subject matter. The term "exemplary" is used to mean as an example or illustration. In the sense of using the terms "comprising," "having," "carrying," "including," etc., such terms are intended to be inclusive in a manner similar to the term "comprising," as interpreted when "comprising" is used as a transition word in the claims. Relational terms such as "first" and "second" may be used to distinguish one entity or action from another without necessarily requiring or implying any actual such relationship or order between these entities or actions.

[0196] Phrases such as aspect, that aspect, on the other hand, some aspects, one or more aspects, implementation, that implementation, another implementation, some implementations, one or more implementations, embodiment, that embodiment, another embodiment, some embodiments, one or more embodiments, configuration, that configuration, another configuration, some configurations, one or more configurations, subject matter, disclosure, this disclosure, other variations thereof, etc., are for convenience and do not imply that disclosures associated with such phrases(s) are essential to the subject matter, or that such disclosures apply to all configurations of the subject matter. Disclosures associated with such phrases(s)(s)(s) may apply to all configurations or one or more configurations. Disclosures associated with such phrases(s)(s)(s)(s)(s)(s)(s)) may provide one or more examples. Phrases such as aspect or some aspects may refer to one or more aspects, and vice versa, and this similarly applies to other foregoing phrases.

[0197] The phrase "at least one" preceding a list of items, separated by the terms "and" or "or," modifies the list as a whole, not each member of the list. The phrase "at least one of..." does not require selection of at least one item; instead, it allows for the inclusion of at least one of any one of the items, and / or at least one of any combination of items, and / or at least one of each of the items. For example, "at least one of A, B, and C" or "at least one of A, B, or C" means only A, only B, or only C; any combination of A, B, and C; and / or at least one of each of A, B, and C.

[0198] It will be understood that the specific order or hierarchy of the disclosed steps, operations, or processes is an illustration of exemplary methods. Unless otherwise expressly stated, it will be understood that the specific order or hierarchy of steps, operations, or processes may be performed in a different order. Some steps, operations, or processes may be performed simultaneously or as part of one or more other steps, operations, or processes. The appended method claims (if any) present the elements of various steps, operations, or processes in a sample order and are not intended to be limited to the specific order or hierarchy presented. These may be performed serially, linearly, in parallel, or in a different order. It should be understood that the described instructions, operations, and systems can generally be integrated together in a single software / hardware product or packaged into multiple software / hardware products.

[0199] This disclosure is provided to enable any person skilled in the art to practice the various aspects described herein. In some cases, well-known structures and components are shown in block diagram form to avoid obscuring the concept of the subject matter. This disclosure provides various examples of the subject matter, and the subject matter is not limited to these examples. Various modifications to these aspects will be apparent to those skilled in the art, and the principles described herein can be applied to other aspects.

[0200] All structural and functional equivalents of elements throughout the various aspects described in this disclosure are expressly incorporated herein by reference and are intended to be covered by the claims, and such structural and functional equivalents are known or would be known to one of ordinary skill in the art. Furthermore, nothing disclosed herein is intended to be offered to the public, whether or not such disclosure is explicitly stated in the claims.

[0201] The title, background information, description of the drawings, abstract, and figures are incorporated herein by reference and are provided as illustrative examples rather than as limiting descriptions. It is understood at the time of filing that they are not intended to limit the scope or meaning of the claims. Furthermore, in the detailed description, illustrative examples are provided, and various features may be grouped together in various embodiments for the purpose of simplifying the disclosure. The approach of this disclosure should not be construed as reflecting an intention to require more features than expressly recited in each claim. Rather, as reflected in the following claims, the inventive subject matter lies in all features of fewer than those in a single disclosed configuration or operation. The following claims are incorporated herein by reference, wherein each claim is independently claimed as a separate subject matter.

[0202] The embodiments are provided merely as examples for understanding the disclosed technology. They are not intended and should not be construed as limiting the scope of the disclosed technology in any way. Although certain embodiments and examples have been provided, it will be apparent to those skilled in the art based on the disclosure herein that changes may be made to the illustrated embodiments and examples without departing from the scope of the disclosed technology.

[0203] The claims are not intended to be limited to the aspects described herein, but will conform to the full scope consistent with the language claims and include all legal equivalents. Nevertheless, no claim is intended to include subject matter that does not meet the requirements of applicable patent law, nor should they be interpreted in this way.

Claims

1. A computer-implemented method (1900) for decoding a basic grid frame, comprising: Arithmetic decoding (1905) is performed on one or more codewords corresponding to one or more prediction errors, which are associated with a base grid frame, wherein the one or more prediction errors are associated with a fine class or a coarse class; Assignment (1910) to one or more contexts for decoding the one or more codewords corresponding to the one or more prediction errors; and Share (1915) will be used in the context of the one or more prediction errors associated with the fine category or the coarse category.

2. The computer-implemented method (1900) according to claim 1, wherein, The one or more prediction errors include at least one of fine geometry prediction error, coarse geometry prediction error, fine texture prediction error, or coarse texture prediction error, and wherein the method (1900) further includes: The shared context will be used among at least one of fine geometry prediction error, coarse geometry prediction error, fine texture prediction error, or coarse texture prediction error.

3. The computer-implemented method (1900) according to any one of claims 1 to 2, wherein, The one or more codewords include one or more portions, wherein a portion of the one or more portions is associated with a truncated unary binarization, and wherein multiple bit positions within the truncated unary binarization share the same context in the one or more contexts.

4. The computer-implemented method (1900) according to any one of claims 1 to 3, wherein, There are three contexts among the one or more contexts associated with the portion of the codeword, wherein the three contexts are used for coarse categories, and wherein a subset of the three contexts is used for fine categories.

5. The computer-implemented method (1900) according to any one of claims 1 to 4, wherein, There exists a first number of binary bits for truncated unary binarization associated with fine texture prediction error, and a second number of binary bits for truncated unary binarization associated with coarse texture prediction error, wherein the first number is greater than the second number.

6. The computer-implemented method (1900) according to any one of claims 1 to 5, wherein, The one or more codewords include one or more portions, wherein a portion of the one or more portions is associated with an exponential Golomb prefix binarization, and wherein multiple bit positions within the exponential Golomb prefix binarization share the same context in the one or more contexts.

7. The computer-implemented method (1900) according to any one of claims 1 to 6, wherein, There are five contexts among the one or more contexts associated with the portion of the codeword, wherein the five contexts are used for a first prediction error type of the one or more prediction errors, and wherein a subset of the five contexts are used for a second prediction error type of the one or more prediction errors.

8. The computer-implemented method (1900) according to any one of claims 1 to 7, wherein, The one or more codewords include one or more portions, wherein a portion of the one or more portions is associated with an exponential Golomb postfix binarization, and wherein multiple bit positions within the exponential Golomb postfix binarization share the same context in the one or more contexts.

9. The computer-implemented method (1900) according to any one of claims 1 to 8, wherein, There are five contexts among the one or more contexts associated with the portion of the codeword, wherein the five contexts are used for a first prediction error type of the one or more prediction errors, and wherein a subset of the five contexts are used for a second prediction error type of the one or more prediction errors.

10. An apparatus for decoding grid frames, comprising a processor configured such that: The receiver (2005) includes a bitstream of prediction error for arithmetic coding of the current coordinates of the grid frame; Determine one or more contexts (2010) for the prediction error of the arithmetic code used for the current coordinates; Arithmetic decoding (2015) is performed on the prediction error of the arithmetic code based on the one or more contexts to determine the prediction error for the current coordinates; Determine the predicted value (2020) for the current coordinates; as well as The coordinate value of the current coordinate (2025) is determined based on the prediction error used for the current coordinate and the predicted value used for the current coordinate. In this context, one or more contexts of the prediction error of at least one coordinate associated with the fine category are shared for the prediction error of at least one coordinate associated with the coarse category.

11. An apparatus for encoding grid frames, comprising a processor configured such that: Determine (2105) the predicted values ​​for the current coordinates of the grid frame; The prediction error for the current coordinates is determined based on the current coordinate value and the predicted value used for the current coordinates; Determine (2115) one or more contexts for the prediction error of the current coordinates; Arithmetic coding (2120) is performed on the prediction error for the current coordinates based on the one or more contexts to generate an arithmetic-coded prediction error for the current coordinates; as well as Send (2125) a bitstream including the arithmetic-coded prediction error. In this context, one or more contexts of the prediction error of at least one coordinate associated with the fine category are shared for the prediction error of at least one coordinate associated with the coarse category.

12. The device according to claim 11, wherein, One or more contexts of the prediction error for at least one geometric coordinate are shared for the prediction error for at least one texture coordinate.

13. The device according to any one of claims 11 to 12, wherein, One or more contexts of the truncated unary portion of the prediction error for at least one coordinate associated with a fine category are shared for the truncated unary portion of the prediction error for at least one coordinate associated with a coarse category.

14. The device according to any one of claims 11 to 13, wherein, One or more contexts of the prefix portion of the prediction error for at least one coordinate associated with the fine category are shared for the prefix portion of the prediction error for at least one coordinate associated with the coarse category.

15. The device according to any one of claims 11 to 14, wherein, One or more contexts of the suffix portion of the prediction error for at least one coordinate associated with the fine category are shared for the suffix portion of the prediction error for at least one coordinate associated with the coarse category.