Base grid coding in grid compression

Through base grid coding technology, the problem of low grid compression efficiency in the existing technology is solved, and efficient data compression and high-quality rendering effects are achieved.

CN120019416APending Publication Date: 2025-05-16TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004349.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-06-18
Filing Date
2024-06-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process complex 3D grid data in grid compression, especially when finding a balance between maintaining high quality and reducing the amount of data.

Method used

The base grid encoding technology is used to receive the base grid information in the code stream, determine the position prediction and attribute prediction of the current vertex, quantify the prediction residual, and use entropy encoding for compression.

Benefits of technology

Improves the compression efficiency of grid data, reduces the amount of data required to store and transfer, while maintaining high-quality 3D grid rendering effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019416A_ABST
    Figure CN120019416A_ABST
Patent Text Reader

Abstract

In one method, a location prediction of a current vertex of a base grid associated with a grid in a current grid frame is determined. And determining attribute prediction of the current vertex. The base grid includes a subset of a plurality of vertices of the grid. A position prediction residual for the position prediction of the current vertex is quantized to generate a quantized position prediction residual. An attribute prediction residual for the attribute prediction of the current vertex is quantized to generate a quantized attribute prediction residual. And entropy coding is carried out on the quantized position prediction residual error of the position prediction of the current vertex in the code stream. And entropy coding is carried out on the quantized attribute prediction residual error of the attribute prediction of the current vertex in the code stream.
Need to check novelty before this filing date? Find Prior Art

Description

Incorporation by reference

[0001] This application claims the benefit of priority to U.S. Patent Application No. 18 / 747,222, filed on June 18, 2024, entitled “Base Mesh Coding in Mesh Compression,” which claims the benefit of priority to U.S. Provisional Application No. 63 / 522,048, filed on June 20, 2023, entitled “Base Mesh Coding in Mesh Compression.” The entire disclosure of the prior application is incorporated herein by reference in its entirety. Technical Field

[0002] The present application discloses various aspects generally related to trellis coding. Background Art

[0003] The background description provided herein is for the purpose of presenting the context of the disclosure of the present application in general. The work of the currently named inventors described in this background technology section and other aspects of this specification that may not meet the standards of the prior art at the time of filing this application should not be admitted, either explicitly or implicitly, as the prior art disclosed in this application.

[0004] Image / video compression helps to transmit image / video data between different devices, storages, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra-frame prediction, which can compress an image based on spatial redundancy. For example, intra-frame prediction can use reference data from the current picture being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-frame prediction, which can compress an image based on temporal redundancy. For example, inter-frame prediction can predict samples in the current picture based on previously reconstructed pictures through motion compensation. Motion compensation can usually be represented by a motion vector (MV).

[0005] Advances in three-dimensional (3D) capture, modeling, and rendering have facilitated the ubiquity of 3D content across a variety of platforms and devices. For example, capturing a baby’s first steps on one continent, grandparents can see (and in some cases interact with) the child on another continent and enjoy a fully immersive experience with the child. To achieve this realism, models have become more complex, and large amounts of data are associated with the creation and consumption of these models. 3D meshes are widely used to represent this immersive content. Summary of the invention

[0006] Various aspects disclosed herein include code streams, methods, and apparatus for mesh processing. In some examples, an apparatus for mesh processing includes a processing circuit.

[0007] According to one aspect disclosed in the present application, a device for mesh decoding is provided. The device includes a processing circuit. The processing circuit is configured to receive a code stream, which includes base mesh information of a base mesh. The base mesh includes a subset of multiple vertices of a mesh in a current mesh frame. The processing circuit is configured to: determine (i) a position prediction of a current vertex of the base mesh and determine (ii) an attribute prediction of the current vertex. The processing circuit is configured to: determine (i) a position prediction residual of the position prediction of the current vertex and determine (ii) an attribute prediction residual of the attribute prediction of the current vertex of the base mesh. The processing circuit is configured to: reconstruct (i) the position of the current vertex of the base mesh based on the position prediction and the position prediction residual, and reconstruct (ii) the attribute of the current vertex of the base mesh based on the attribute prediction and the attribute prediction residual.

[0008] In an example, the attribute of the current vertex is a two-dimensional (2D) texture coordinate of the current vertex, the attribute prediction is a 2D texture coordinate prediction, and the attribute prediction residual is a 2D texture coordinate prediction residual.

[0009] In an example, the processing circuit is configured to determine the position prediction of the current vertex based on one or more encoded vertices according to a multi-parallelogram prediction. The processing circuit is configured to determine the 2D texture coordinate prediction of the current vertex based on the one or more encoded vertices according to a stretch prediction algorithm.

[0010] In an example, the processing circuit is configured to determine a quantized position prediction residual of the position prediction of the current vertex based on a first entropy encoding. The processing circuit is configured to dequantize the quantized position prediction residual of the position prediction of the current vertex to determine the position prediction residual based on a first quantization step value. The processing circuit is configured to determine the quantized 2D texture coordinate prediction residual of the 2D texture coordinate prediction of the current vertex based on a second entropy encoding. The processing circuit is configured to dequantize the quantized 2D texture coordinate prediction residual of the 2D texture coordinate prediction of the current vertex to determine the 2D texture coordinate prediction residual based on a second quantization step value.

[0011] In an example, the second quantization step value is determined based on the first quantization step value. The range of log2 of the denominator of the first quantization step value is between 0 and 7. The second quantization step value is one of 1, 2, 4, 8, 16, 32, 64 and 128. The first quantization step value is one of 1, 2, 4, 8, 16, 32, 64 and 128.

[0012] In an example, the second quantization step value is equal to a multiple of the first quantization step value.

[0013] In an example, the second quantization step value is equal to one of a linear function and a monotonically non-decreasing function of the first quantization step value.

[0014] In an example, the second quantization step value is equal to a monotonically non-decreasing function of a quantization error associated with the position prediction of the current vertex. The quantization error is the difference between the quantized position prediction residual and an unquantized position prediction residual of the current vertex.

[0015] In an example, the first quantization step value is determined based on the second quantization step value.

[0016] In an example, the first quantization step value is equal to a multiple of the second quantization step value.

[0017] In an example, the first quantization step value is equal to one of a linear function and a monotonically non-decreasing function of the second quantization step value.

[0018] In an example, the first quantization step value is equal to a monotonically non-decreasing function of a quantization error associated with the 2D texture coordinate prediction of the current vertex. The quantization error is the difference between the quantized 2D texture coordinate prediction residual and the unquantized 2D texture coordinate prediction residual of the current vertex.

[0019] In one aspect disclosed in the present application, a method for mesh encoding is provided. In the method, a position prediction of a current vertex of a base mesh associated with a mesh in a current mesh frame is determined. An attribute prediction of the current vertex is determined. The base mesh includes a subset of multiple vertices of the mesh. A position prediction residual of the position prediction of the current vertex is quantized to generate a quantized position prediction residual. An attribute prediction residual of an attribute prediction of the current vertex is quantized to generate a quantized attribute prediction residual. The quantized position prediction residual of the position prediction of the current vertex is entropy encoded in a bitstream. The quantized attribute prediction residual of the attribute prediction of the current vertex is entropy encoded in a bitstream.

[0020] In an example, the attribute of the current vertex is a two-dimensional (2D) texture coordinate of the current vertex. The attribute prediction is a 2D texture coordinate prediction. The attribute prediction residual is a 2D texture coordinate prediction residual.

[0021] In an example, the quantized position prediction residual of the position prediction of the current vertex is dequantized to obtain the position prediction residual of the position prediction of the current vertex. The quantized 2D texture coordinate prediction residual of the 2D texture coordinate prediction of the current vertex is dequantized to obtain the 2D texture coordinate prediction residual of the current vertex. The position of the current vertex is reconstructed based on the position prediction and the position prediction residual. The 2D texture coordinates of the current vertex are reconstructed based on the 2D texture coordinate prediction and the 2D texture coordinate prediction residual.

[0022] In an example, the position prediction of the current vertex is determined based on one or more encoded vertices according to a multi-parallelogram prediction. The 2D texture coordinate prediction of the current vertex is determined based on the one or more encoded vertices according to a stretch prediction algorithm.

[0023] In an example, the position prediction residual of the position prediction of the current vertex is quantized based on a first quantization step value. The 2D texture coordinate prediction residual of the 2D texture coordinate prediction of the current vertex is quantized based on a second quantization step value.

[0024] In one aspect disclosed in the present application, a method for processing mesh data is provided. In the method, a code stream of the mesh data is processed according to a format rule. In an example, the code stream includes base mesh information of a base mesh, wherein the base mesh includes a subset of multiple vertices of a mesh in a current mesh frame. The format rule specifies: (i) determining a position prediction of a current vertex of the base mesh, and (ii) determining an attribute prediction of the current vertex. The format rule specifies: (i) determining a position prediction residual of the position prediction of the current vertex, and (ii) determining an attribute prediction residual of the attribute prediction of the current vertex of the base mesh. The format rule specifies: (i) processing the position of the current vertex of the base mesh based on the position prediction and the position prediction residual, and (ii) processing the attribute of the current vertex of the base mesh based on the attribute prediction and the attribute prediction residual.

[0025] The various aspects disclosed in the present application also provide an apparatus for trellis coding. The apparatus for trellis coding includes a processing circuit configured to implement any of the methods for trellis coding described.

[0026] The various aspects disclosed in the present application also provide a method for trellis decoding, which includes any method implemented by a device for trellis decoding.

[0027] Aspects disclosed in the present application also provide a non-transitory computer-readable medium storing instructions, which, when executed by a computer, cause the computer to perform any of the described methods for grid decoding, methods for grid encoding, and methods for grid data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:

[0029] Figure 1 is a schematic diagram of an example of a block diagram of a communication system (100).

[0030] Figure 2 is a diagram showing an example of a block diagram of a decoder.

[0031] Figure 3 is a diagram of an example of a block diagram of an encoder.

[0032] Figure 4 is a schematic diagram of an example of an encoding process (400) for mesh processing according to one aspect of the present disclosure.

[0033] Figure 5 is a schematic diagram of an example of a pre-processing step (500) according to one aspect of the present disclosure.

[0034] Figure 6 is a schematic diagram of a decoding process (600) for grid processing according to one aspect disclosed in the present application.

[0035] Figure 7 is a schematic diagram of an example of base grid encoding according to one aspect disclosed in the present application.

[0036] Figure 8 is a schematic diagram of an example of base grid encoding with fine-grained quantization according to one aspect disclosed in the present application.

[0037] Fig. 9 A flow chart outlining a trellis decoding process according to some aspects disclosed herein is shown.

[0038] Fig.10 A flow chart outlining a trellis encoding process according to some aspects disclosed herein is shown.

[0039] Fig.11 is a schematic diagram of a computer system according to one aspect disclosed herein. DETAILED DESCRIPTION

[0040] Figure 1 A block diagram of a video processing system (100) in some examples is shown. The video processing system (100) is an example of an application of the disclosed subject matter in a video encoder and a video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications including, for example, video conferencing, digital television, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0041] The video processing system (100) includes an acquisition subsystem (113), which may include a video source (101). The video source (101) may include one or more images captured by a camera and / or generated by a computer. For example, a digital camera, for example, creates an uncompressed video picture stream (102). In an example, the video picture stream (102) includes samples taken by the digital camera. Compared to the encoded video data (104) (or the encoded video bitstream), the video picture stream (102) is depicted as a thick line to emphasize the high data volume of the video picture stream, and the video picture stream (102) can be processed by an electronic device (120), which includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. Compared to the video picture stream (102), the encoded video data (104) (or the encoded video bitstream) is depicted as a thin line to emphasize the lower amount of data of the encoded video data, which can be stored on the streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1 The client subsystems (106) and (108) in the video streaming server (105) can access the encoded video data (104) copies (107) and (109). The client subsystem (106) can include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy of the encoded video data (107) and creates an output video picture stream (111) that can be presented on a display (112) (e.g., a display screen) or other display device (not depicted). In some streaming systems, the encoded video data (104), (107) and (109) (e.g., video bitstreams) can be encoded according to certain video encoding / compression standards. Examples of such standards include ITU-T Recommendation H.265. In the example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter can be used in the context of the VVC standard.

[0042] It should be noted that the electronic devices (120) and (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may also include a video encoder (not shown).

[0043] Figure 2 An example of a block diagram of a video decoder (210) is shown. The video decoder (210) may be provided in an electronic device (230). The electronic device (230) may include a receiver (231) (eg, a receiving circuit). The video decoder (210) may be used instead of Figure 1 A video decoder (110) is shown in an example.

[0044] A receiver (231) may receive one or more encoded video sequences, such as included in a bitstream, to be decoded by a video decoder (210). In one aspect, one encoded video sequence is received at a time, wherein decoding of each encoded video sequence is independent of decoding of other encoded video sequences. The encoded video sequence may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not depicted). The receiver (231) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and an entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other cases, the buffer memory (215) may be located external to the video decoder (210) (not depicted). In other cases, a buffer memory (not depicted) may be provided external to the video decoder (210), for example, to prevent network jitter, and another buffer memory (215) may be provided internal to the video decoder (210), for example, to handle playout timing. When the receiver (231) receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (215) may not be required, or the buffer memory (215) may be made smaller. For use on a traffic packet network such as the Internet, a buffer memory (215) may be required, the buffer memory (215) may be relatively large and may advantageously have an adaptive size, and the buffer memory (215) may be implemented at least in part in an operating system or similar element (not depicted) external to the video decoder (210).

[0045] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. The categories of symbols include information for managing the operation of the video decoder (210) and potential information for controlling a display device such as a display device (212) (e.g., a display screen) that is not part of the electronic device (230) but can be coupled to the electronic device (230), such as Figure 2 As shown. The control information for one or more display devices may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not depicted). The parser (220) may parse / entropy decode the received coded video sequence. The encoding of the coded video sequence may be performed according to a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context-inspiredness, etc. The parser (220) may extract a subgroup parameter set for at least one subgroup of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (220) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.

[0046] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215), thereby creating symbols (221).

[0047] Depending on the type of coded video picture or part of coded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of symbol (221) may involve multiple different units. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the coded video sequence by parser (220). For the sake of brevity, such subgroup control information flow between parser (220) and multiple units below is not described.

[0048] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into a number of functional units as described below. In actual implementations operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purposes of describing the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.

[0049] The first unit is a sealer / inverse transform unit (251). The sealer / inverse transform unit (251) receives quantized transform coefficients as symbols (221) from the parser (220) and control information, including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit (251) can output a block including sample values, which can be input into an aggregator (255).

[0050] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates surrounding blocks of the same size and shape as the block being reconstructed using reconstructed information extracted from a current picture buffer (258). For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a per-sample basis.

[0051] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to an inter-frame coded and potentially motion compensated block. In this case, the motion compensated prediction unit (253) may access the reference picture memory (257) to extract samples for prediction. After the extracted samples are motion compensated according to the symbol (221) belonging to the block, these samples may be added to the output of the scaler / inverse transform unit (251) (in this case referred to as residual samples or residual signal) by the aggregator (255) to generate output sample information. The acquisition of the predicted samples by the motion compensated prediction unit (253) from the address in the reference picture memory (257) may be controlled by a motion vector, and the motion vector is provided to the motion compensated prediction unit (253) in the form of the symbol (221), which includes, for example, X, Y and reference picture components. Motion compensation may also include interpolation of sample values ​​extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.

[0052] The output samples of the aggregator (255) may be employed by various loop filtering techniques in a loop filter unit (256). The video compression techniques may include in-loop filter techniques controlled by parameters included in an encoded video sequence (also referred to as an encoded video bitstream) and available to the loop filter unit (256) as symbols (221) from the parser (220). The video compression may also be responsive to meta-information obtained during decoding of an encoded picture or a previous (in decoding order) portion of an encoded video sequence, and to previously reconstructed and loop filtered sample values.

[0053] The output of the loop filter unit (256) may be a sample stream that may be output to a display device (212) and stored in a reference picture memory (257) for subsequent inter-picture prediction.

[0054] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed, and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) may become part of the reference picture memory (257), and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.

[0055] The video decoder (210) may perform decoding operations according to, for example, the ITU-T Rec. H.265 standard or a predetermined video compression technology. The encoded video sequence may conform to the syntax specified by the video compression technology or standard used in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as the only tools available for use under the profile. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata of the HRD buffer management signaled in the encoded video sequence.

[0056] In one aspect, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be part of one or more encoded video sequences. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0057] Figure 3 An example of a block diagram of a video encoder (303) is shown. The video encoder (303) is provided in an electronic device (320). The electronic device (320) includes a transmitter (340) (eg, a transmission circuit). The video encoder (303) may be used instead of Figure 1 A video encoder (103) in an example.

[0058] The video encoder (303) can be used to obtain the video source (301) (not Figure 3 In another embodiment, the video source (301) is a part of the electronic device (320) in the example, and receives the video samples, wherein the video source can collect one or more video images to be encoded by the video encoder (503). In another embodiment, the video source (301) is a part of the electronic device (320).

[0059] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), wherein the digital video sample stream may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The pictures themselves may be constructed as a spatial pixel array, wherein each pixel may include one or more samples depending on the sampling structure, color space, etc. used. The following description focuses on the samples.

[0060] According to one aspect, the video encoder (303) can encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other time constraints required. Implementing the appropriate encoding speed is a function of the controller (350). In some aspects, the controller (350) controls other functional units as described below and is functionally coupled to these units. For the sake of simplicity, the coupling is not shown in the figure. The parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology, etc.), picture size, group of pictures (group of pictures, GOP) layout, maximum motion vector search range, etc. The controller (350) can be used to have other suitable functions that are related to the video encoder (303) optimized for a certain system design.

[0061] In some aspects, the video encoder (303) operates in an encoding loop. As a simple description, in an example, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and one or more reference pictures) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-accurately corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values ​​that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, such as due to channel errors) is also used in some related technologies.

[0062] The operation of the "local" decoder (333) may be similar to that already described above in conjunction with Figure 2 The operation of a "remote" decoder such as the video decoder (210) described in detail is the same. However, additional brief reference is made to Figure 2 , when symbols are available and the entropy encoder (345) and parser (220) are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and the parser (220), may not be fully implemented in the local decoder (333).

[0063] In one aspect, the decoder techniques present in the decoder, except parsing / entropy decoding, are present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on the decoder operation. The description of the encoder techniques can be simplified because the encoder techniques are mutually inverse to the decoder techniques described comprehensively. In some areas, a more detailed description is provided below.

[0064] During operation, in some examples, the source encoder (330) may perform motion compensated predictive coding. The motion compensated predictive coding predictively encodes an input picture with reference to one or more previously encoded pictures from a video sequence designated as "reference pictures." In this manner, the encoding engine (332) encodes the differences between pixel blocks of the input picture and pixel blocks of one or more reference pictures that may be selected as one or more prediction references for the input picture.

[0065] The local video decoder (333) may decode the encoded video data of the picture that may be designated as the reference picture based on the symbol created by the source encoder (330). The operation of the encoding engine (332) may advantageously be a lossy process. When the encoded video data is available at the video decoder ( Figure 3 When the video sequence is decoded at a remote location (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture memory (334). In this way, the video encoder (333) may locally store a copy of the reconstructed reference picture that has common content (absent transmission errors) with the reconstructed reference frame to be obtained by the remote video decoder.

[0066] The predictor (335) may perform a prediction search for the encoding engine (332). That is, for a new picture (or grid) to be encoded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may serve as appropriate prediction references for the new picture. The predictor (335) may operate on a pixel-by-pixel-block basis based on sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (335), it may be determined that the input picture may have prediction references taken from a plurality of reference pictures stored in the reference picture memory (334).

[0067] In an example, a mesh position quantization step size is defined based on a first parameter (e.g., mesh_position_quantization_step_size_log2_denominator) and a second parameter (e.g., mesh_position_quantization_step_size_numerator_minus1). In an example, mesh_position_quantization_step_size_log2_denominator is the log2 of the denominator of the position quantization step size, and its value is between 0 and 7, including the end value. In an example, mesh_position_quantization_step_size_numerator_minus1 plus 1 is the value of the numerator of the position quantization step size, and its value is between 1 and 256, including the end value. After position encoding in a mesh-based edgebreak (MEB), the reconstructed position value used for reference is scaled to within the inter-frame dynamic range.

[0068] An example of the grid attribute encoding parameter syntax is as follows:

[0069] An example of mesh attribute encoding parameter semantics is defined by a parameter such as mesh_attribute_quantization_step_size_log2_denominator[i]. In the example, mesh_atribute_quantization_step_size_log2_denominator[i] is the log2 of the denominator of the i-th attribute quantization step size, and its value is between 0 and 7, inclusive.

[0070] An example of a generic base grid sequence parameter set RBSP syntax is shown below:

[0071] An example of motion field quantization is defined by a first parameter (e.g., bmsps_inter_quantization_step_size_log2_denominator) and a second parameter (e.g., bmsps_inter_quantization_step_size_numerator_minus1). In the example, bmsps_inter_quantization_step_size_log2_denominator is the log2 of the denominator of the motion field quantization step size, and its value is between 0 and 7, including the end value. In the example, bmsps_inter_quantization_step_size_numerator_minus plus 1 is the value of the numerator of the motion field quantization step size, and its value is between 1 and 256, including the end value. After motion field encoding, the position value is reconstructed and will be used as a reference for future vertices. In the example, mesh attribute quantization is defined by a step size parameter (e.g., mesh_attribute_quantization_step_size_numerator[i]). In the example, mesh_attribute_quantization_step_size_numerator[i] is the value of the numerator of the quantization step size of the i-th attribute, and its value is between 0 and 255, including the end value. At the same time, the quantized prediction residual is dequantized. The reconstructed value is calculated by adding the dequantized prediction residual to the predicted value. The reconstructed value will be used as a reference for future vertices.

[0072] The controller (350) may manage encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.

[0073] The outputs of all the above functional units may be entropy encoded in an entropy encoder (345). The entropy encoder (345) performs lossless compression on the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into a coded video sequence.

[0074] The transmitter (340) may buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission over a communication channel (360), which may be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (340) may combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or ancillary data streams (source not shown).

[0075] The controller (350) may manage the operation of the video encoder (303). During encoding, the controller (350) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following picture types:

[0076] An intra picture (I picture) may be a picture that is encoded and decoded without using any other picture in the sequence as a prediction source.Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures.

[0077] A predictive picture (P picture) may be a picture that is encoded and decoded using intra prediction or inter prediction that predicts sample values ​​of each block using a motion vector and a reference index.

[0078] Bidirectional predictive pictures (B pictures) can be pictures that are encoded and decoded using intra prediction or inter prediction, which uses two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.

[0079] The source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined according to the coding allocation applied to the block's corresponding picture. For example, a block of an I picture may be non-predictively coded, or the block may be predictively coded (spatial prediction or intra prediction) with reference to already coded blocks of the same picture. A block of a P picture may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. A block of a B picture may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.

[0080] The video encoder (303) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T Rec. H.265. In operation, the video encoder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0081] In one aspect, the transmitter (340) may transmit additional data when transmitting the encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0082] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In an example, a particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.

[0083] In some aspects, bidirectional prediction techniques may be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be in the past and future, respectively, in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.

[0084] In addition, merge mode technology can be used in inter-picture prediction to improve coding efficiency.

[0085] According to some aspects disclosed in the present application, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks, such as polygonal or triangular blocks. For example, according to the HEVC standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU includes three coding tree blocks (CTBs), which are a luminance CTB and two chrominance CTBs. Each CTU can be recursively split into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be split into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In the example, each CU is analyzed to determine the prediction type for the CU, such as an inter-prediction type or an intra-prediction type. Depending on temporal and / or spatial predictability, the CU is split into one or more prediction units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In one aspect, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Taking the luma prediction block as an example, the prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, and the like.

[0086] It should be noted that the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using any suitable technology. In one aspect, the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using one or more integrated circuits. In another aspect, the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using one or more processors executing software instructions.

[0087] Various aspects disclosed herein include techniques for base grid encoding in grid compression.

[0088] A mesh may include several polygons that describe the surface of a volumetric object. Each polygon of a mesh may be defined by the vertices of the corresponding polygons in a three-dimensional (3D) space and information about how these vertices are connected (which may be referred to as connection information). In some aspects, vertex attributes (e.g., color, normal, etc.) may be associated with vertices (or mesh vertices). Attributes (or vertex attributes) may also be associated with the surface of a mesh by utilizing mapping information that parameterizes the mesh using a two-dimensional (2D) attribute map. Such a mapping may be described by a set of parameter coordinates (referred to as UV coordinates or texture coordinates) associated with mesh vertices. A 2D attribute map may be used to store high-resolution attribute information, such as textures, normals, displacements, etc. High-resolution attribute information may be used for various purposes, such as texture mapping and shading.

[0089] Dynamic mesh sequences may require a large amount of data because they may include a large amount of information that varies over time. Therefore, efficient compression techniques can be used to store and transmit such content. Mesh compression standards (such as Information and Communication (IC) mesh compression, MESHGRID and frame-based animated mesh compression (FAMC)) were previously developed by the Moving Picture Experts Group (MPEG) to address dynamic meshes with constant connections, time-varying geometry and vertex attributes. However, these standards may not take into account time-varying attribute mapping and connection information. Digital Content Creation (DCC) tools can generate such dynamic meshes. However, for volume acquisition technology, generating a dynamic mesh with constant connections may be challenging, especially under real-time constraints. Existing standards may not support this type of content (e.g., a dynamic mesh with constant connections). The present application discloses various aspects including a new mesh compression standard that can directly handle dynamic meshes with time-varying connection information and optional time-varying attribute mapping. Mesh compression can target lossy and lossless compression for various applications such as real-time communication, storage, free viewpoint video, augmented reality (AR) and virtual reality (VR). Features such as random access and scalable / progressive coding can also be considered.

[0090] The mesh geometry information may include vertex connection information, 3D coordinates, 2D texture coordinates, etc. The 3D vertex coordinates and 2D texture coordinates account for an important part of the mesh geometry information. Therefore, it is necessary to compress the 3D vertex coordinates (also called vertex positions) and 2D texture coordinates to reduce the amount of data required to store and / or transmit the mesh geometry information.

[0091] Figure 4 An example of an encoding process (400) for mesh processing based on a related video codec (eg, MPEG V-MeshTM V1.0) according to one aspect of the present disclosure is shown. Figure 4 As shown, the encoding process (400) may include a preprocessing step (400A) and an encoding step (400B). The preprocessing step (400A) may be configured to generate a base grid m(i) of a current frame and a displacement field d(i) of the current frame, the displacement field including a displacement vector, the displacement vector depending on the input grid M(i) of the current frame. The encoding step (400B) may be configured to encode the base grid m(i), the displacement field d(i), and the texture information of the base grid m(i). The displacement field d(i) of the current frame may include the displacement vector. The index i may refer to the current frame. In one aspect, a mode decision method may be performed in the encoding process (400) to determine whether to apply inter-frame coding (also known as inter-frame prediction or inter-frame mode), intra-frame coding (also known as intra-frame prediction or intra-frame mode), etc. to the current frame. For example, the mode decision method may compare the cost of the intra-frame mode with the cost of the inter-frame mode, and determine the encoding mode of the base grid m(i) of the current frame based on the mode with the smaller cost. In some examples, the skip mode is used to encode (e.g., encode or decode) the base grid m(i). In an example, the skip mode is a special mode of the inter-frame mode. For example, the base grid m(i) can be intra-coded, or inter-coded, or encoded using the SKIP mode.

[0092] Still reference Figure 4, the preprocessing step (400A) may include a mesh extraction process (402), a parameterization process (e.g., an atlas parameterization process (404)), and a subdivision surface fitting process (406). The mesh extraction process (402) is configured to downsample the vertices of the input mesh M(i) to generate an extracted mesh dm(i), and the extracted mesh may include multiple extracted (or downsampled) vertices. The number of the multiple extracted vertices is less than the number of vertices of the input mesh M(i). The parameterization process (e.g., an atlas parameterization process (404)) is configured to map the extracted mesh dm(i) to a planar domain, such as a UV atlas (or UV map), to generate a re-parameterized mesh pm(i). In an example, the atlas parameterization may be performed based on a video processing tool (e.g., a UV atlas tool). The subdivision surface fitting process (406) is configured to take the re-parameterized mesh pm(i) and the input mesh M(i) as input and produce a base mesh m(i) and a displacement field d(i), which includes a displacement vector or a set of displacements. In an example of the subdivision surface fitting process (406), pm(i) is subdivided using a subdivision scheme (e.g., iterative interpolation) to obtain a subdivided mesh. Iterative interpolation includes inserting a new point in the middle of each edge of the re-parameterized mesh pm(i) at each iteration. Any suitable subdivision scheme can be applied to subdivide pm(i). For each vertex of the subdivided mesh, the displacement field d(i) is calculated by determining the nearest point on the surface of the input mesh M(i).

[0093] Advantages of the subdivided mesh may include: the subdivided mesh has a subdivision structure that can provide a faithful approximation of the input mesh while being efficiently compressed. Compression efficiency can be improved due to the following properties. The decimated mesh dm(i) may have a small number of vertices and may be encoded and transmitted using a smaller number of bits than the input mesh M(i) or the subdivided mesh. Figure 4 , a base mesh m(i) may be generated from the extracted mesh dm(i). In the example, the base mesh m(i) is the extracted mesh dm(i). Since the subdivided mesh may be generated based on the subdivision method, the subdivided mesh may be automatically generated by the decoder when decoding the base mesh or the extracted mesh (e.g., without using any information other than the subdivision scheme and the subdivision iteration count). On the decoder side, the displacement field d(i) may be generated by decoding the displacement vectors associated with the vertices of the subdivided mesh. In addition to allowing spatial / quality scalability, the subdivision structure enables efficient transforms (e.g., wavelet decomposition), which can provide high compression performance.

[0094] For simplicity, the preprocessing step (400A) that can be applied to an input mesh (e.g., a 3D mesh) can be described using the preprocessing step (500) applied to a two-dimensional (2D) curve. The preprocessing step (400A) is similar to the preprocessing step (500), except that the 3D mesh can be replaced by a 2D curve.

[0095] Figure 5 An example of a preprocessing step (500) according to one aspect of the present disclosure is shown. Figure 5 As shown, an input 2D curve (represented by a 2D polyline) (502) may be downsampled to generate a base curve, such as a polyline, referred to as a "decimated" curve (504). A subdivision scheme may then be applied to the decimated polyline (504) to generate a "decision" curve (506). In an example, the subdivision scheme may be an iterative interpolation scheme. The iterative interpolation scheme may include inserting a new point in the middle of each edge of the polyline (or decimated curve) (504) at each iteration. For example, point (510) may be inserted onto edge (508) of the decimated curve (504). In an example, edge (508) is between point (512) and point (514). Additionally, point (522) may be added between point (512) and point (510), and point (516) may be added between point (510) and point (514). The decimalized polyline (506) may then be deformed to generate a displacement curve (518). The displacement curve (518) may be a better approximation of the input curve (502) than the subdivided curve (506). For example, a displacement vector (e.g., (520)) is calculated for each vertex (e.g., (510)) of the subdivided curve (506) so that the shape of the displacement curve (518) is as close as possible to the shape of the input curve (502). An advantage of the subdivided curve (506) is that the subdivided curve (506) has a subdivision structure that can provide a faithful approximation of the input curve (502) while being more efficiently compressed.

[0096] The decimation curve (504) may have a small number of points and may be encoded and transmitted using a limited number of bits. Since the decimation curve may be generated based on the decimation scheme, the decimation curve may be automatically generated by the decoder when decoding the base curve or the decimation curve (e.g., without using any information other than the decimation method and the decimation iteration count). The displacement curve is generated by decoding the displacement vectors associated with the vertices of the decimation curve. In addition to allowing spatial / quality scalability, the decimation structure enables efficient transforms (e.g., wavelet decomposition), which can provide high compression performance.

[0097] Still reference Figure 5In an example, the input mesh M(i) may include an input 2D curve (502). The base mesh m(i) may include a decimated curve (504) formed by down-sampling vertices of the input 2D curve (502). The displacement field dm(i) may include a plurality of displacement vectors, such as Figure 5 The displacement vector (520) shown in .

[0098] The encoding step (400B) may include base mesh encoding (408), displacement encoding (410), texture encoding (412), etc. The base mesh encoding (408) is configured to encode geometric information of a base mesh m(i) associated with a current frame. In intra-frame encoding, the base mesh m(i) may be first quantized (e.g., using uniform quantization) and then encoded, for example, using an encoding mode determined by a mode decision method. The encoding mode may be an inter-frame mode, an intra-frame mode, a skip mode, etc. An encoder for intra-frame encoding of the base mesh m(i) may be referred to as a static mesh encoder. In inter-frame encoding, a reference base mesh associated with a reference frame indicated by index j (e.g., a reconstructed quantized reference base mesh m'(j)) may be used to predict a base mesh m(i) associated with a current frame indicated by index i. The displacement encoding (410) is configured to encode a displacement field d(i) generated in the preprocessing step (400A). The displacement field d(i) may include a set of displacement vectors (or displacements) associated with subdivided mesh vertices. The texture encoding (412) is configured to encode attribute information of the base mesh m(i). The attribute information may include texture, normal, color, etc. The attribute information may be encoded based on a suitable codec, such as High-Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC).

[0099] In one aspect, reference Figure 4 , a mesh encoding process (e.g., encoding process (400)) begins with preprocessing (e.g., preprocessing step (400A)). The preprocessing may convert an input mesh (e.g., an input dynamic mesh) M(i) into a base mesh m(i) and a displacement field d(i) including a set of displacements (or a set of displacement vectors). The encoding step (400B) may compress the output from the preprocessing (e.g., m(i), d(i), etc.) and generate a compressed code stream b(i). The compressed code stream b(i) may include a compressed base mesh code stream, a compressed displacement field code stream, a compressed attribute code stream, etc.

[0100] Figure 6An example of a decoding process (600) for grid processing according to one aspect disclosed herein is shown. The decoding process (600) may include a decoding step (605) and a post-processing step (610). A compressed code stream b(i) may be fed to the decoding step (605). In an example, for example for lossless transmission, the compressed code stream b(i) is the output b(i) from the encoding process (400). The decoding step (605) may extract various sub-code streams, such as a compressed base grid sub-stream, a compressed displacement field sub-stream, a compressed attribute sub-stream, etc. The decoding step (605) may decompress the sub-code streams to generate the following components: patch metadata (metadata) indicated by metadata(i), a decoded base grid m"(i), a decoded displacement field (including displacement) d"(i), a decoded attribute map A"(i), etc.

[0101] In one aspect, a base grid substream may be fed to a grid decoder to generate a reconstructed quantized base grid m'(i). A decoded base grid (or reconstructed base grid) m"(i) may be obtained by applying inverse quantization to m'(i). A displacement field substream may be decoded by a video and / or image decoder, the displacement field substream comprising encoded packed and quantized wavelet coefficients. Image unpacking and inverse quantization may be applied to the reconstructed packed and quantized wavelet coefficients to obtain unpacked and unquantized transform coefficients (e.g., wavelet coefficients). An inverse wavelet transform may be applied to the unpacked and unquantized wavelet coefficients to generate a decoded displacement field (or reconstructed displacement) d"(i).

[0102] The decoded components (e.g., including: metadata(i), m"(i), d"(i), A"(i), etc.) may be fed to a post-processing step (610). A mesh (also referred to as a decoded / reconstructed mesh) M"(i) may be generated by the post-processing step (610) based on m"(t) and d"(i). In an example, a mesh M"(i) (also referred to as a reconstructed deformed mesh DM(i)) may be obtained by subdividing m"(i) using a subdivision scheme and applying reconstructed displacements d"(i) to vertices of the subdivided mesh. In an example, DM(i) may include a displacement curve (518). In an example, when the encoding process (400), the decoding process (600), and the transmission are all lossless, the mesh M"(i) may be the same as the input mesh M(i). When one of the encoding process (400), the decoding process (600), and the transmission is lossy, M"(i) is different from M(i). In various examples, the difference between M"(i) and M(i), if any, may be relatively small. In the example, an attribute map (map) A" (i) is also generated by the post-processing step (610).

[0103] In one aspect, the base grid may be intra-coded, inter-coded, or encoded using a skip mode, etc. In an example, the skip mode may be a special mode of the inter-frame mode, in which the base grid m(i) of the current frame indicated by the index i is the same as the base grid m(j) of the reference frame indicated by the index (also referred to as the frame index) j. When the inter-frame mode is applied to encode the base grid in the current frame, the encoder may generate a predicted base grid for the current frame based on the reconstructed base grid of the reference frame. In an example, such as in MPEG V-DMC WD 2.0, the reference frame is a frame immediately preceding the current frame in display order. The frame index i of the current frame indicates the display order. When the frame index of the current frame is i, the frame index of the reference frame is (i-1). In an example, the current frame and the reference frame are in the same group of frames (GoF).

[0104] In one aspect, for example, in MPEG V-DMC WD 2.0, the mesh encoding process begins with preprocessing. The preprocessing may convert the input dynamic mesh (denoted as M(i)) into a base mesh m(i) and a set of displacements d(i). The encoder may compress the base mesh m(i) and the displacements d(i) to generate a compressed code stream b(i).

[0105] like Figure 4 As shown, preprocessing may include mesh extraction, atlas parameterization, and subdivision surface fitting. Mesh extraction may use simplification techniques to extract the input mesh M(i) and generate an extracted mesh dm(i). The extracted mesh dm(i) may be reparameterized. The resulting mesh may be denoted as pm(i). Subdivision surface fitting may take the reparameterized mesh pm(i) and the input mesh M(i) as input to generate a base mesh m(i) and a set of displacements d(i).

[0106] In the present application disclosure, a method and system for base grid coding in grid compression is provided. In one aspect, base grid coding includes position coding and 2D texture coordinate coding. In an example, such as in MPEG V-DMC WD 2.0, position coding and 2D texture coordinate coding utilize a workflow including bit depth quantization, prediction and entropy coding. Figure 7 An example of a workflow (700) is shown in FIG.

[0107] like Figure 7 As shown, the workflow (700) includes bit depth quantization (702), prediction (704), and entropy encoding (706). In bit depth quantization (702), the position (or position coordinate) of the vertex in the base mesh and the 2D texture coordinate are quantized into bit depth values. The bit depth value may be specified (or defined) by the encoder. The base mesh may include a subset of multiple vertices of the mesh. In one aspect, if the bit depth value of the position is n and the bit depth value of the 2D texture coordinate is m, then the position is quantized into 0 to 2n -l, and the 2D texture coordinates are quantized to 0 to 2 m -1, where n and m are positive integers. For example, if the bit depth value of the position is 12 and the bit depth value of the 2D texture coordinate is 13, the position is quantized to an integer between 0 and 4095, and the 2D texture coordinate is quantized to an integer between 0 and 8191. The quantized value may be predicted at prediction (704) by a prediction mode or a prediction algorithm, such as by a multi-parallelogram algorithm for the position of the vertex and a stretch prediction algorithm for the 2D texture coordinate. The prediction residual of the position of the vertex and the prediction residual of the 2D texture coordinate of the vertex may be further compressed by entropy coding (706).

[0108] In one aspect, for example in MPEG V-DMC WD 2.0, the bit depth values ​​used for quantization are limited or constrained. In the present application disclosure, base grid coding with fine-grained quantization is provided. In one aspect, base grid coding with fine-grained quantization is provided with joint quantization. In joint quantization, the quantization of the position of the vertex and the quantization of the 2D texture coordinates of the vertex can be interdependent.

[0109] In the present application disclosure, base grid coding with fine-grained quantization is provided. The base grid coding may include prediction, quantization, inverse quantization and entropy coding. Figure 8 An example of base grid coding (800) is shown in FIG. Figure 8As shown, base mesh encoding (800) includes prediction (802), quantization (804), inverse quantization (808) and entropy encoding (806). In an example, base mesh encoding (800) may also include bit depth quantization (not shown) before prediction (802). Bit depth quantization may be configured to quantize the positions of the vertices and 2D texture coordinates of the base mesh into bit depth values. At prediction (802), when base mesh encoding (800) includes bit depth quantization, the bit depth values ​​generated at the bit depth quantization are predicted based on an algorithm or prediction mode, such as a multi-parallelogram algorithm for the positions of the vertices and a stretch prediction algorithm for the 2D texture coordinates. When base mesh encoding (800) does not include bit depth quantization, the positions of the vertices and 2D texture coordinates in the base mesh may be predicted based on a suitable prediction mode or algorithm. Therefore, position predictions of the vertices, 2D texture coordinate predictions, position prediction residuals of the position predictions, and 2D texture coordinate prediction residuals of the 2D texture coordinate predictions may be generated at prediction (802). The position prediction residuals for position prediction of the vertices and the 2D texture coordinate prediction residuals for 2D texture coordinate prediction of the vertices may be further quantized at quantization (804). Quantization (804) is configured to reduce the precision of the prediction residuals according to a quantization parameter (QP), such as a quantization step value. Thus, less important information is discarded and significant compression is achieved by quantization (804). The quantized prediction residuals are further compressed by entropy coding (806).

[0110] Still reference Figure 8 , the quantized prediction residual obtained at quantization (804) may be provided to inverse quantization (808). Inverse quantization (808) may inverse quantize the quantized prediction residual to obtain the prediction residual generated at prediction (802). The prediction residual may also be provided to prediction (802). The position and 2D texture coordinates of the vertex may be reconstructed based on the prediction and the prediction residual. The reconstructed position and the reconstructed 2D texture coordinates of the vertex may be used as reference information for subsequent vertices.

[0111] In one aspect, the position of the current vertex in the mesh (or base mesh) is predicted by the position of the encoded vertex according to a suitable prediction mode or prediction algorithm (e.g., multi-parallelogram prediction or position algorithm). The 2D texture coordinates of the current vertex are predicted by the 2D texture coordinates of the encoded vertex according to a suitable prediction mode or prediction algorithm (e.g., stretch prediction algorithm or 2D texture coordinate prediction algorithm). In an example, the position of the current vertex and the 2D texture coordinates of the current vertex are predicted by a prediction process (e.g., Figure 8 The prediction is made using the prediction (802) in .

[0112] In one aspect, the prediction residual is quantized by a quantization parameter (e.g., a quantization step value), and the prediction residual is the difference between the true value (or initial value) and the predicted value. The quantization step value can be a positive integer, a positive rational number, or a positive real number. In an example, the prediction residual of the position of the current vertex is the difference between the true value (or initial value) of the position and the predicted value of the position. In an example, the prediction residual of the 2D texture coordinate of the current vertex is the difference between the true value (or initial value) of the 2D texture coordinate of the current vertex and the predicted value of the 2D texture coordinate of the current vertex. In an example, the prediction residual of the position of the current vertex and the prediction residual of the 2D texture coordinate of the current vertex are quantized by a quantization process (e.g., Figure 8 quantization (804) in .

[0113] In one aspect, the quantized prediction residuals are compressed by entropy coding. Entropy coding may include fixed length coding, variable length coding, Huffman coding, arithmetic coding, etc. In an example, the quantized prediction residuals of the position of the current vertex and the quantized prediction residuals of the 2D texture coordinates of the current vertex are compressed by an entropy coding process (e.g., Figure 8 The entropy coding (806) in is used for entropy coding.

[0114] In one aspect, the quantized prediction residual of the current vertex may be dequantized. A reconstructed value is calculated by adding the dequantized prediction residual to the predicted value. The reconstructed value may be used as a reference value for a future vertex, such as a vertex that follows the current vertex according to the coding order. In an example, the quantized prediction residual of the position of the current vertex is dequantized to obtain a position prediction residual. The reconstructed position of the current vertex is calculated by adding the position prediction residual to the predicted value of the position of the current vertex. In an example, the quantized prediction residual of the position of the current vertex and the quantized prediction residual of the 2D texture coordinates of the current vertex are dequantized via a dequantization process (e.g., Figure 8 Dequantization is performed using the dequantization (808) in the example.

[0115] Joint quantization is provided in the present disclosure. In joint quantization, the quantization of the position of the vertex and the quantization of the 2D texture coordinates of the vertex may be dependent on each other. For example, the quantization parameter of the quantization of the position of the vertex and the quantization parameter of the quantization of the 2D texture coordinates of the vertex may be dependent on each other. In an example, the quantization of the position of the vertex depends on the quantization of the 2D texture coordinates. In an example, the quantization of the 2D texture coordinates depends on the quantization of the position of the vertex.

[0116] In one aspect, a quantization parameter (eg, a quantization step value) of a vertex's 2D texture coordinate is determined by a quantization process of the vertex's position.

[0117] In one aspect, the quantization step size of the 2D texture coordinates of the vertex depends on the quantization step size of the position of the vertex. For example, the quantization step size of the 2D texture coordinates of the vertex is equal to the quantization step size of the position of the vertex.

[0118] In one aspect, the quantization step value of the 2D texture coordinates of the vertex depends on a multiple of the quantization step value of the position of the vertex. For example, the quantization step value of the 2D texture coordinates of the vertex is equal to the multiple of the quantization step value of the position of the vertex. The multiple can be a positive integer, a positive rational number, or a positive real number.

[0119] In one aspect, the quantization step value of the 2D texture coordinate of the vertex depends on a function (e.g., a linear function or a nonlinear function) of the quantization step value of the position of the vertex. In an example, the quantization step value of the position is an input to the function. The quantization step value of the 2D texture coordinate of the vertex is equal to the output of the function of the quantization step value of the position of the vertex.

[0120] An example of a linear function is f(x)=ax+b, where a is a non-negative real number, b is a real number, x is the quantization step value of the vertex position, and f(x) is the output of the linear function and is equal to the quantization step value of the 2D texture coordinates of the vertex.

[0121] In one aspect, the quantization step value of the 2D texture coordinate of the vertex depends on a monotonically non-decreasing function of the quantization step value of the position of the vertex. In an example, the quantization step value of the position is an input to the monotonically non-decreasing function. The quantization step value of the 2D texture coordinate of the vertex is equal to the output of the monotonically non-decreasing function of the quantization step value of the position of the vertex.

[0122] Examples of monotone non-decreasing functions include f(x)=x, f(x)=e x and f(x) = x n , where n is a positive integer.

[0123] In one aspect, the quantization step value of the 2D texture coordinate of the vertex depends on a monotonically non-decreasing function of the quantization error of the position of the vertex. The quantization error can be the difference between an unquantized value (e.g., an unquantized prediction residual of the position of the vertex) and a quantized value (e.g., a quantized prediction residual of the position of the vertex).

[0124] In an example, the quantization error of the position of the vertex is an input to the monotonically non-decreasing function. The quantization step value of the 2D texture coordinate of the vertex is equal to the output of the monotonically non-decreasing function of the quantization error of the position of the vertex.

[0125] In the present disclosure, the quantization step value of the position of a vertex is determined by a quantization process of the 2D texture coordinates of the vertex.

[0126] In one aspect, the quantization step value of the position of the vertex depends on the quantization step value of the 2D texture coordinates of the vertex. In an example, the quantization step value of the position of the vertex is equal to the quantization step value of the 2D texture coordinates of the vertex.

[0127] In one aspect, the quantization step value of the position of the vertex depends on a multiple of the quantization step value of the 2D texture coordinates of the vertex, where the multiple can be a positive integer, a positive rational number, or a positive real number. For example, the quantization step value of the position of the vertex is equal to the multiple of the quantization step value of the 2D texture coordinates of the vertex.

[0128] In one aspect, the quantization step value of the position of the vertex depends on a function (e.g., a linear function or a nonlinear function) of the quantization step value of the 2D texture coordinates of the vertex. In an example, the quantization step value of the 2D texture coordinates is an input to the function. The quantization step value of the position of the vertex is equal to the output of the function of the quantization step value of the 2D texture coordinates of the vertex.

[0129] An example of a linear function is f(x)=ax+b, where a is a non-negative real number, b is a real number, x is the quantization step value of the 2D texture coordinates of the vertex, and f(x) is the quantization step value of the position of the vertex.

[0130] In one aspect, the quantization step value of the position of the vertex depends on a monotonically non-decreasing function of the quantization step value of the 2D texture coordinates of the vertex. In an example, the quantization step value of the 2D texture coordinates of the vertex is an input to the monotonically non-decreasing function. The quantization step value of the position of the vertex is equal to the output of the monotonically non-decreasing function of the quantization step value of the 2D texture coordinates of the vertex.

[0131] In one aspect, the quantization step value of the position of the vertex depends on a monotonically non-decreasing function of the quantization error of the 2D texture coordinates of the vertex. The quantization error can be the difference between an unquantized value (e.g., an unquantized prediction residual of the 2D texture coordinates of the vertex) and a quantized value (e.g., a quantized prediction residual of the 2D texture coordinates of the vertex). In an example, the quantization error of the 2D texture coordinates of the vertex is an input to the monotonically non-decreasing function. The quantization step value of the position of the vertex is equal to the output of the monotonically non-decreasing function of the quantization error of the 2D texture coordinates of the vertex.

[0132] Fig. 9 A flow chart outlining a process (900) according to one aspect disclosed herein is shown. The process (900) may be used in a video decoder. In various aspects, the process (900) is performed by a processing circuit, such as a processing circuit that performs the functions of a video decoder (110), a processing circuit that performs the functions of a video decoder (210), etc. In some aspects, the process (900) is implemented by software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (900). The process starts at (S901) and proceeds to (S910).

[0133] At (S910), a code stream is received, the code stream including base mesh information of a base mesh. The base mesh includes a subset of a plurality of vertices of a mesh in a current mesh frame.

[0134] At (S920), a position prediction of a current vertex of a base mesh and an attribute prediction of the current vertex are determined.

[0135] At (S930), a position prediction residual of a position prediction of the current vertex and an attribute prediction residual of an attribute prediction of the current vertex of the base mesh are determined.

[0136] At (S940), the position of the current vertex of the base mesh is reconstructed based on the position prediction and the position prediction residual. The attribute of the current vertex of the base mesh is reconstructed based on the attribute prediction and the attribute prediction residual.

[0137] In an example, the attribute of the current vertex is a two-dimensional (2D) texture coordinate of the current vertex. The attribute prediction is a 2D texture coordinate prediction. The attribute prediction residual is a 2D texture coordinate prediction residual.

[0138] In an example, according to the multi-parallelogram prediction, a position prediction of a current vertex is determined based on one or more encoded vertices. According to the stretch prediction algorithm, a 2D texture coordinate prediction of the current vertex is determined based on one or more encoded vertices.

[0139] In an example, a quantized position prediction residual of a position prediction of a current vertex is determined based on a first entropy encoding. The quantized position prediction residual of the position prediction of the current vertex is dequantized to determine the position prediction residual based on a first quantization step value. A quantized 2D texture coordinate prediction residual of a 2D texture coordinate prediction of the current vertex is determined based on a second entropy encoding. The quantized 2D texture coordinate prediction residual of the 2D texture coordinate prediction of the current vertex is dequantized to determine the 2D texture coordinate prediction residual based on a second quantization step value.

[0140] In an example, the second quantization step value is determined based on the first quantization step value. The log2 of the denominator of the first quantization step value is in the range of 0 to 7. The second quantization step value is one of 1, 2, 4, 8, 16, 32, 64 and 128. The first quantization step value is one of 1, 2, 4, 8, 16, 32, 64 and 128.

[0141] In an example, the second quantization step value is equal to a multiple of the first quantization step value.

[0142] In an example, the second quantization step value is equal to one of a linear function and a monotonically non-decreasing function of the first quantization step value.

[0143] In an example, the second quantization step value is equal to a monotonically non-decreasing function of a quantization error associated with the position prediction of the current vertex. The quantization error is the difference between a quantized position prediction residual and an unquantized position prediction residual for the current vertex.

[0144] In an example, the first quantization step value is determined based on the second quantization step value.

[0145] In an example, the first quantization step value is equal to a multiple of the second quantization step value.

[0146] In an example, the first quantization step value is equal to one of a linear function and a monotonically non-decreasing function of the second quantization step value.

[0147] In an example, the first quantization step value is equal to a monotonically non-decreasing function of a quantization error associated with the 2D texture coordinate prediction of the current vertex. The quantization error is the difference between the quantized 2D texture coordinate prediction residual and the unquantized 2D texture coordinate prediction residual for the current vertex.

[0148] Then, the process proceeds to (S999) and terminates.

[0149] The process (900) may be adjusted appropriately. One or more steps in the process (900) may be modified and / or omitted. One or more additional steps may be added. Any suitable order of implementation may be used.

[0150] Fig.10 A flow chart of an overview process (1000) according to one aspect disclosed herein is shown. The process (1000) may be used in a video encoder. In various aspects, the process (1000) is performed by a processing circuit, such as a processing circuit that performs the functions of a video encoder (103), a processing circuit that performs the functions of a video encoder (303), etc. In some aspects, the process (1000) is implemented by software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (1000). The process starts at (S1001) and proceeds to (S1010).

[0151] At (S1010), a position prediction of a current vertex of a base mesh associated with a mesh in a current mesh frame is determined. An attribute prediction of the current vertex is determined. The base mesh includes a subset of a plurality of vertices of the mesh.

[0152] At (S1020), the position prediction residual of the position prediction of the current vertex is quantized to generate a quantized position prediction residual. The attribute prediction residual of the attribute prediction of the current vertex is quantized to generate a quantized attribute prediction residual.

[0153] At (S1030), entropy encoding is performed on the quantized position prediction residual of the position prediction of the current vertex in the bitstream. Entropy encoding is performed on the quantized attribute prediction residual of the attribute prediction of the current vertex in the bitstream.

[0154] In an example, the attribute of the current vertex is a two-dimensional (2D) texture coordinate of the current vertex. The attribute prediction is a 2D texture coordinate prediction. The attribute prediction residual is a 2D texture coordinate prediction residual.

[0155] In the example, the quantized position prediction residual of the position prediction of the current vertex is dequantized to obtain the position prediction residual of the position prediction of the current vertex. The quantized 2D texture coordinate prediction residual of the 2D texture coordinate prediction of the current vertex is dequantized to obtain the 2D texture coordinate prediction residual of the current vertex. The position of the current vertex is reconstructed based on the position prediction and the position prediction residual. The 2D texture coordinates of the current vertex are reconstructed based on the 2D texture coordinate prediction and the 2D texture coordinate prediction residual.

[0156] The position of a vertex in a grid (or 3D grid) may refer to the position of the vertex in a virtual world or 3D space. The position of a vertex may be defined using a coordinate system, such as a Cartesian coordinate system having three axes X, Y, and Z. The position prediction of a vertex indicates a predicted value of the position of the vertex. In an example, the position of a vertex may be predicted by the positions of one or more neighboring vertices of the vertex.

[0157] The 2D texture coordinates (also called UV coordinates) of a vertex in a 3D model (e.g., a 3D mesh) specify the location on a 2D texture image that can be mapped to the vertex. The 2D texture coordinates can be used to indicate to the renderer how to "wrap" the texture (e.g., color and detail information) of the 3D model onto the vertex. The 2D texture coordinate prediction of a vertex indicates the predicted value of the 2D texture coordinate of the vertex.

[0158] Vertices of a mesh may have various attributes that provide additional information beyond the simple position of the vertex in 3D space. For example, these attributes define how the vertex interacts with light, textures, and other elements that contribute to the final appearance of the mesh. Examples of vertex attributes include: (1) the position of the vertex, (2) a normal vector that defines the surface orientation at the vertex, (3) 2D coordinates (u, v) that can be used to map a texture to the vertex on the surface of the mesh, (4) a color associated with the vertex, (5) tangent vectors and bitangent vectors used for advanced lighting calculations, such as tangent vectors and bitangent vectors used for bump mapping and normal mapping techniques, and (6) bone weights and indices for the vertex. In an example, an attribute prediction for a vertex indicates a predicted value for an attribute of the vertex. For example, an attribute of a vertex may be predicted based on the attributes of its neighboring vertices in a mesh.

[0159] In an example, according to the multi-parallelogram prediction, a position prediction of a current vertex is determined based on one or more encoded vertices. According to the stretch prediction algorithm, a 2D texture coordinate prediction of the current vertex is determined based on one or more encoded vertices.

[0160] In an example, the position prediction residual of the position prediction of the current vertex is quantized based on the first quantization step value, and the 2D texture coordinate prediction residual of the 2D texture coordinate prediction of the current vertex is quantized based on the second quantization step value.

[0161] Then, the process proceeds to (S1099) and terminates.

[0162] The process (1000) may be adjusted appropriately. One or more steps in the process (1000) may be modified and / or omitted. One or more additional steps may be added. Any suitable implementation order may be used.

[0163] In one aspect, a method of processing mesh data includes processing a codestream of the mesh data according to a format rule. For example, the codestream may be a codestream decoded by any of the decoding methods described herein and / or encoded by any of the encoding methods described herein. The format rule may specify one or more constraints of the codestream and / or one or more processes to be performed by a decoder and / or an encoder.

[0164] In an example, a code stream of mesh data is processed according to a format rule. The code stream includes base mesh information of a base mesh, wherein the base mesh includes a subset of multiple vertices of a mesh in a current mesh frame. The format rule specifies: (i) determining a position prediction of a current vertex of the base mesh, and (ii) determining an attribute prediction of the current vertex. The format rule specifies: (i) determining a position prediction residual of the position prediction of the current vertex, and (ii) determining an attribute prediction residual of the attribute prediction of the current vertex of the base mesh. The format rule specifies: (i) processing the position of the current vertex of the base mesh based on the position prediction and the position prediction residual, and (ii) processing the attribute of the current vertex of the base mesh based on the attribute prediction and the attribute prediction residual.

[0165] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig.11 A computer system (1100) suitable for implementing certain aspects of the disclosed subject matter is shown.

[0166] Computer software may be encoded using any suitable machine code or computer language, which may be assembled, compiled, linked or similarly constructed to create code comprising instructions that may be directly executed by one or more computer central processing units (CPU), graphics processing units (GPU), etc., or executed through interpretation, microcode execution, etc.

[0167] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0168] Fig.11 The components of the computer system (1100) shown are examples and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing aspects of the present disclosure. The configuration of components should also not be interpreted as having any dependency or requirement related to any one or combination of components shown in the example aspects of the computer system (1100).

[0169] The computer system (1100) may include certain human-machine interface input devices. Such human-machine interface input devices may be responsive to input from one or more human users, for example, through: tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, high fives), visual input (e.g., gestures), olfactory input (not depicted). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to a person's conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video), etc.

[0170] The human-machine interface input device may include one or more of the following devices (only one of each is depicted): keyboard (1101), mouse (1102), touchpad (1103), touch screen (1110), data gloves (not shown), joystick (1105), microphone (1106), scanner (1107), camera (1108).

[0171] The computer system (1100) may also include certain human interface output devices. Such human interface output devices may stimulate one or more senses of a human user, for example, through tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback through a touch screen (1110), a data glove (not shown), or a joystick (1105), but may also be a tactile feedback device that is not used as an input device), an audio output device (e.g., a speaker (1109), a headset (not depicted)), a visual output device (e.g., a screen (1110) including a cathode ray tube (CRT) screen, a liquid crystal display (LCD) screen, a plasma screen, an organic light-emitting diode (OLED) screen, each screen having or not having a touch screen input function, each screen having or not having a tactile feedback function, some of which may be capable of outputting two-dimensional visual output or output in more than three dimensions through devices such as stereoscopic picture output, virtual reality glasses (not depicted), a holographic display and a smoke box (not depicted), and a printer (not depicted)).

[0172] The computer system (1100) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1120) with CD / DVD or similar media (1121), thumb drives (1122), removable hard drives or solid-state drives (1123), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security dongles (not depicted), and the like.

[0173] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0174] The computer system (1100) may also include an interface (1154) to one or more communication networks (1155). The network may be, for example, a wireless network, a wired network, an optical network. The network may also be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, and the like. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including Global System for Mobile communications (GSM), 3G, 4G, 5G, Long-Term Evolution (LTE), etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CAN buses, and the like. Certain networks typically require an external network interface adapter (e.g., a USB port of the computer system (1100)) attached to certain general data ports or peripheral buses (1149). Other network interfaces are typically integrated into the kernel of the computer system (1100) by attaching to a system bus as described below (e.g., connected to an Ethernet interface in a PC computer system or connected to a cellular network interface in a smartphone computer system). The computer system (1100) can communicate with other entities using any of these networks. Such communications can be one-way receive-only (e.g., broadcast television), one-way send-only (e.g., CANbus to certain CANbus devices), or two-way, for example, connecting to other computer systems using a local area digital network or a wide area digital network. As described above, certain protocols and protocol stacks can be used on each of those networks and network interfaces.

[0175] The above-mentioned human-machine interface device, human-accessible storage device, and network interface may be attached to the kernel ( 1140 ) of the computer system ( 1100 ).

[0176] The core (1140) may include one or more central processing units (CPUs) (1141), graphics processing units (GPUs) (1142), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (1143), hardware accelerators (1144) for certain tasks, graphics adapters (1150), and the like. These devices, together with read-only memory (ROM) (1145), random access memory (1146), and internal mass storage (1147) such as internal non-user accessible hard drives, SSDs, etc., may be connected through a system bus (1148). In some computer systems, the system bus (1148) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, and the like. Peripheral devices may be attached directly to the core's system bus (1148) or to the core's system bus (1448) through a peripheral bus (1149). In an example, a screen (1110) may be connected to a graphics adapter (1150). The architecture of the peripheral bus includes PCI, USB, and the like.

[0177] The CPU (1141), GPU (1142), FPGA (1143) and accelerator (1144) can execute certain instructions, which can be combined to form the above-mentioned computer code. The computer code can be stored in ROM (1145) or RAM (1146). Transitional data can also be stored in RAM (1146), while permanent data can be stored, for example, in internal mass storage (1147). Fast storage and retrieval to any storage device can be achieved by using a cache, which can be closely associated with the following components: one or more CPUs (1141), GPUs (1142), mass storage (1147), ROM (1145), RAM (1146), etc.

[0178] The computer readable medium may have computer code thereon, which is used to perform various computer-implemented operations. The medium and computer code may be specially designed and constructed for the purposes disclosed in this application, or the medium and computer code may be of a type well known and available to those skilled in the art of computer software.

[0179] As an example and not limitation, a computer system having an architecture (1100), in particular a kernel (1140), is able to provide functionality due to the execution of software contained in one or more tangible computer-readable media by (one or more) processors (including CPUs, GPUs, FPGAs, accelerators, etc.). Such computer-readable media may be media associated with user-accessible mass storage as described above, as well as certain non-temporary memories of the kernel (1140), such as kernel internal mass storage (1147) or ROM (1145). Software that implements various aspects disclosed in the present application may be stored in such devices and executed by the kernel (1140). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software may enable the kernel (1140), in particular the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform specific processes or specific parts of specific processes described herein, including defining data structures stored in RAM (1146) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system is enabled to provide functionality due to logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1144)) that can run in place of software or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to portions of software may include logic and vice versa. Where appropriate, references to portions of computer-readable media may include circuitry (e.g., integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both. The present application discloses any suitable combination of hardware and software.

[0180] The use of "at least one" or "one" in the present disclosure is intended to include any one or combination of the listed elements. For example, reference to "at least one of A, B, or C", "at least one of A, B, and C", "at least one of A, B, and / or C", and "at least one of A to C" is intended to include only A, only B, only C, or any combination thereof. Reference to "one of A or B" and "one of A and B" is intended to include A or B, or (A and B). Where applicable, the use of "one of" does not exclude any combination of the listed elements, such as when the elements are not mutually exclusive.

[0181] Although the present application discloses a plurality of non-limiting embodiments, there are modifications, substitutions and various replacement equivalents that fall within the scope of the present application discloses. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods as described below, which, although not explicitly shown or described in the present application discloses, embody the principles disclosed in the present application and therefore fall within the spirit and scope of the present application discloses.

Claims

1. A device for grid decoding, comprising: A processing circuit, the processing circuit being configured to: receiving a code stream, the code stream comprising base grid information of a base grid, the base grid comprising a subset of a plurality of vertices of a grid in a current grid frame; Determining (i) a position prediction of a current vertex of the base mesh, and determining (ii) an attribute prediction of the current vertex; determining (i) a position prediction residual of the position prediction of the current vertex of the base mesh, and determining (ii) an attribute prediction residual of the attribute prediction of the current vertex of the base mesh; as well as (i) The position of the current vertex of the base mesh is reconstructed based on the position prediction and the position prediction residual, and (ii) The attribute of the current vertex of the base mesh is reconstructed based on the attribute prediction and the attribute prediction residual.

2. The device according to claim 1, wherein: The attribute prediction is a two-dimensional (2D) texture coordinate prediction; and The processing circuit is configured to: Determining the position prediction of the current vertex based on one or more encoded vertices according to a multi-parallelogram prediction; as well as The 2D texture coordinate prediction of the current vertex is determined based on the one or more encoded vertices according to a stretch prediction algorithm.

3. The device according to claim 1, wherein: The attribute prediction is a two-dimensional (2D) texture coordinate prediction, and the attribute prediction residual is a 2D texture coordinate prediction residual; and The processing circuit is configured to: determining a quantized position prediction residual of the position prediction of the current vertex based on a first entropy encoding; Dequantizing the quantized position prediction residual of the position prediction of the current vertex to determine the position prediction residual based on a first quantization step value; determining a quantized 2D texture coordinate prediction residual of the 2D texture coordinate prediction of the current vertex based on a second entropy encoding; as well as The quantized 2D texture coordinate prediction residual of the 2D texture coordinate prediction of the current vertex is dequantized to determine the 2D texture coordinate prediction residual based on a second quantization step value.

4. The device according to claim 3, wherein: The second quantization step value is determined based on the first quantization step value; The range of log2 of the denominator of the first quantization step value is between 0 and 7; The second quantization step value is one of 1, 2, 4, 8, 16, 32, 64 and 128; and The first quantization step value is one of 1, 2, 4, 8, 16, 32, 64 and 128.

5. The device according to claim 3, wherein: The second quantization step value is equal to a multiple of the first quantization step value.

6. The device according to claim 3, wherein: The second quantization step value is equal to one of a linear function and a monotonically non-decreasing function of the first quantization step value.

7. The device according to claim 3, wherein: The second quantization step value is equal to a monotonically non-decreasing function of a quantization error associated with the position prediction of the current vertex, the quantization error being a difference between the quantized position prediction residual and an unquantized position prediction residual of the current vertex.

8. The device according to claim 3, wherein: The first quantization step value is determined based on the second quantization step value.

9. The device according to claim 3, wherein: The first quantization step value is equal to a multiple of the second quantization step value.

10. The device according to claim 3, wherein: The first quantization step value is equal to one of a linear function and a monotonically non-decreasing function of the second quantization step value.

11. The device according to claim 3, wherein: The first quantization step value is equal to a monotonically non-decreasing function of a quantization error associated with the 2D texture coordinate prediction of the current vertex, the quantization error being a difference between the quantized 2D texture coordinate prediction residual and an unquantized 2D texture coordinate prediction residual for the current vertex.

12. A method for video encoding, comprising: determining (i) a position prediction of a current vertex of a base mesh associated with a mesh in a current mesh frame, and determining (ii) an attribute prediction of the current vertex; the base mesh comprising a subset of a plurality of vertices of the mesh; quantizing (i) a position prediction residual of the position prediction of the current vertex to generate a quantized position prediction residual, and quantizing (ii) an attribute prediction residual of the attribute prediction of the current vertex to generate a quantized attribute prediction residual; as well as In a bitstream, (i) the quantized position prediction residual of the position prediction of the current vertex is entropy encoded, and (ii) the quantized attribute prediction residual of the attribute prediction of the current vertex is entropy encoded.

13. The method according to claim 12, wherein: The attribute prediction residual is a two-dimensional (2D) texture coordinate prediction residual, and the attribute prediction is a 2D texture coordinate prediction; as well as The method further comprises: Dequantizing (i) the quantized position prediction residual of the position prediction of the current vertex to obtain the position prediction residual of the position prediction of the current vertex, and dequantizing (ii) the quantized 2D texture coordinate prediction residual of the 2D texture coordinate prediction of the current vertex to obtain the 2D texture coordinate prediction residual of the current vertex; as well as (i) The position of the current vertex is reconstructed based on the position prediction and the position prediction residual, and (ii) The 2D texture coordinates of the current vertex are reconstructed based on the 2D texture coordinate prediction and the 2D texture coordinate prediction residual.

14. The method of claim 12, wherein: The attribute prediction residual is a two-dimensional (2D) texture coordinate prediction residual, and the attribute prediction is a 2D texture coordinate prediction; The determining the position prediction residual comprises: determining the position prediction of the current vertex based on one or more encoded vertices according to multi-parallelogram prediction; and The determining the 2D texture coordinate prediction residual comprises: determining the 2D texture coordinate prediction of the current vertex based on the one or more encoded vertices according to a stretch prediction algorithm.

15. A method of processing visual media data, the method comprising: Processing the code stream of the visual media data according to the format rules, wherein: The code stream includes base mesh information of a base mesh, the base mesh including a subset of a plurality of vertices of a mesh in a current mesh frame; and The format rules specify: (i) determining a position prediction of a current vertex of the base mesh, and (ii) determining an attribute prediction of the current vertex; (i) determining a position prediction residual of the position prediction of the current vertex, and (ii) determining an attribute prediction residual of the attribute prediction of the current vertex of the base mesh; and (i) processing the position of the current vertex of the base mesh based on the position prediction and the position prediction residual, and (ii) processing the attribute of the current vertex of the base mesh based on the attribute prediction and the attribute prediction residual.