Quantization in mesh compression
By proposing a grid decoding method in grid compression, using quantization step value and entropy coding technology to effectively quantify and encode the grid vertex position and sports field, the problem of low data efficiency in the existing technology is solved and more efficient data transmission and storage is achieved.
Patent Information
- Application Number
- CN202480004331.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-07-02
- Filing Date
- 2024-07-03
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively quantify and encode the vertex positions, sports fields and texture coordinates of the grid in grid compression, resulting in low data transmission and storage efficiency.
A grid decoding method is proposed. By receiving the basic grid information in the code stream, the position prediction and sports field prediction of the current vertex are determined, and the quantization and solution quantization are performed based on the quantization step value, the position prediction residual and sports field prediction residual are generated, and further compressed by entropy coding.
It realizes efficient quantification and encoding of grid vertex positions and sports fields, reduces data volume and improves transmission and storage efficiency.
Smart Images

Figure CN120019653A_ABST
Abstract
Description
[0001] Join by reference
[0002] This application claims the benefit of priority to U.S. Patent Application No. 18 / 762,428, filed on July 2, 2024, entitled “Quantization in Mesh Compression,” which claims the benefit of priority to U.S. Provisional Application No. 63 / 542,886, filed on October 6, 2023, entitled “Quantization of Texture Coordinates in Mesh Compression,” and U.S. Provisional Application No. 63 / 525,134, filed on July 5, 2023, entitled “Quantization of Position and Motion Field in Mesh Compression.” The entire disclosures of these prior applications are incorporated herein by reference in their entirety. Technical Field
[0003] This disclosure describes aspects generally related to trellis coding. Background Art
[0004] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent that the work of the presently named inventors described in this background section and in various aspects of this specification was performed, it does not indicate that it qualifies as prior art at the time of filing, and it is never explicitly or implicitly admitted that it is prior art to the present disclosure.
[0005] Image / video compression can help transmit image / video data between different devices, storages, and networks with minimal quality degradation. In some examples, video codec techniques can compress video based on spatial redundancy and temporal redundancy. In one example, a video codec can use a technique called intra-frame prediction, which can compress an image based on spatial redundancy. For example, intra-frame prediction can use reference data from a current picture being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-frame prediction, which can compress an image based on temporal redundancy. For example, inter-frame prediction can predict samples in a current picture based on a picture previously reconstructed using motion compensation. Motion compensation can be indicated by a motion vector (MV).
[0006] Advances in three-dimensional (3D) capture, modeling, and rendering have facilitated 3D content across a variety of platforms and devices. For example, a baby's first steps can be filmed on one continent, and grandparents can see (and in some cases, interact with) and enjoy a fully immersive experience with the child on another continent. To achieve this sense of realism, models have become more complex, and a large amount of data is associated with the creation and use of these models. 3D meshes are widely used to represent such immersive content. Summary of the invention
[0007] Aspects of the present disclosure include code streams, methods, and apparatus for mesh processing.In some examples, the apparatus for mesh processing includes processing circuitry.
[0008] According to one aspect of the present disclosure, a mesh decoding method is provided. In the method, a code stream is received, the code stream including base mesh information of a base mesh. The base mesh includes a subset of multiple vertices of the mesh in a current mesh frame. A position prediction of a current vertex of the base mesh is determined. A motion field prediction of the current vertex of the base mesh is determined. A position prediction residual of the position prediction of the current vertex is determined based on a first quantization step value. A motion field prediction residual of the motion field prediction of the current vertex is determined based on a second quantization step value, wherein the second quantization step value depends on the first quantization step value. The position of the current vertex of the base mesh is reconstructed based on the position prediction and the position prediction residual. The motion field of the current vertex of the base mesh is reconstructed based on the motion field prediction and the motion field prediction residual.
[0009] In one aspect, to determine a position prediction residual, a quantized position prediction residual of a position prediction of a current vertex is determined based on a first entropy coding. The quantized position prediction residual of the position prediction of the current vertex is further dequantized based on a first quantization step value to determine a position prediction residual. In one aspect, to determine a motion field prediction residual, a quantized motion field prediction residual of a motion field prediction of a motion field of the current vertex is determined based on a second entropy coding. The quantized motion field prediction residual of the motion field prediction of the current vertex is dequantized based on a second quantization step value to determine a motion field prediction residual.
[0010] In one aspect, the second quantization step value is equal to a multiple of the first quantization step value.
[0011] In one aspect, the second quantization step value is equal to one of a linear function and a monotonically non-decreasing function of the first quantization step value.
[0012] In one aspect, the second quantization step size value is equal to a monotonically non-decreasing function of a quantization error associated with the position prediction of the current vertex. The quantization error is the difference between a quantized position prediction residual and an unquantized position prediction residual for the current vertex.
[0013] In one aspect, the first quantization step value is a positive binary rational number Definition, where m is a positive integer between 0 and 7, and n is a non-negative integer between 1 and 256.
[0014] In one aspect, when the first quantization step value and the second quantization step value are received at the sequence level, the first quantization step value and the second quantization step value are equal when a syntax element in the codestream indicates that the first quantization step value is equal to the second quantization step value.
[0015] On the one hand, when a first quantization step value and a second quantization step value are received at a frame level, when a syntax element in a code stream is a first value, the second quantization step value is equal to a quantization step value of a reference frame of a current grid frame, and when a syntax element in a code stream is a second value, the second quantization step value is equal to the first quantization step value.
[0016] In the method, a two-dimensional (2D) texture coordinate prediction of a current vertex of a base mesh is determined, a 2D texture coordinate prediction residual of the 2D texture coordinate prediction of the current vertex is determined, and the 2D texture coordinate of the current vertex of the base mesh is reconstructed based on the 2D texture coordinate prediction and the 2D texture coordinate prediction residual.
[0017] In one aspect, when the 2D texture coordinates of the current vertex are directly encoded, in order to determine the 2D texture coordinate prediction residual, a quantized 2D texture coordinate of the current vertex is determined, wherein the quantized 2D texture coordinate is encoded by fixed length coding. The quantized 2D texture coordinate of the current vertex is dequantized to obtain the 2D texture coordinate. The quantized 2D texture coordinate is obtained by rounding quantization, wherein the rounding quantization is configured to convert the 2D texture coordinate into an integer.
[0018] In one aspect, when predicting the 2D texture coordinates of the current vertex by the stretch prediction algorithm, in order to determine the 2D texture coordinate prediction residual, a quantized 2D texture coordinate prediction residual of the current vertex is determined, wherein the quantized 2D texture coordinate prediction residual is encoded by variable length coding. The quantized 2D texture coordinate prediction residual of the current vertex is dequantized to determine the 2D texture coordinate prediction residual. The quantized 2D texture coordinate prediction residual is obtained by rounding quantization, wherein the rounding quantization is configured to convert the 2D texture coordinate prediction residual into an integer.
[0019] On the one hand, rounding quantization is done by binary rational numbers Defined as follows, where m is a positive integer between 0 and 7, and n is a non-negative integer between 0 and 255. The quantized 2D texture coordinate prediction residual is obtained by rounding off a multiple of the 2D texture coordinate prediction residual and the reciprocal of a binary rational number.
[0020] According to another aspect of the present disclosure, a mesh encoding method is provided. In the method, a position prediction of a current vertex of a base mesh associated with a mesh in a current mesh frame is determined. A motion field prediction of a current vertex of the base mesh is determined. A position prediction residual of the position prediction of the current vertex is quantized based on a first quantization step value to generate a quantized position prediction residual. A motion field prediction residual of the motion field prediction of the current vertex is quantized based on a second quantization step value to generate a quantized motion field prediction residual. The second quantization step value depends on the first quantization step value. Entropy encoding is performed on the quantized position prediction residual of the position prediction of the current vertex. Entropy encoding is performed on the quantized motion field prediction residual of the motion field prediction of the current vertex.
[0021] In one aspect, the second quantization step value is equal to a multiple of the first quantization step value.
[0022] In one aspect, the second quantization step value is equal to one of a linear function and a monotonically non-decreasing function of the first quantization step value.
[0023] In the method, a 2D texture coordinate prediction of a 2D texture coordinate of a current vertex of a base mesh is determined. A 2D texture coordinate prediction residual of the 2D texture coordinate prediction of the current vertex is quantized to generate a quantized 2D texture coordinate prediction residual of the current vertex. The quantized 2D texture coordinate prediction residual of the current vertex is further entropy encoded.
[0024] In one aspect, when directly encoding the 2D texture coordinates of the current vertex, in order to quantize the 2D texture coordinate prediction residual, the 2D texture coordinates of the current vertex are quantized based on rounded quantization, wherein the rounded quantization is configured to convert the 2D texture coordinates into integers. The quantized 2D texture coordinates are further encoded based on fixed length coding.
[0025] On the one hand, when predicting the 2D texture coordinates of the current vertex by the stretch prediction algorithm, in order to determine the 2D texture coordinate prediction residual, the multiple of the 2D texture coordinate prediction residual and the binary rational number The 2D texture coordinate prediction residual of the current vertex is quantized by rounding the reciprocal of , where m is a positive integer between 0 and 7, and n is a non-negative integer between 0 and 255. The quantized 2D texture coordinate prediction residual is further encoded based on variable length coding.
[0026] In one aspect, in order to quantize the 2D texture coordinate prediction residual, an offset is added to the 2D texture coordinate prediction residual to generate an updated 2D texture coordinate prediction residual. Further, by quantizing the multiple of the updated 2D texture coordinate prediction residual and the binary rational number The updated 2D texture coordinate prediction residual of the current vertex is quantized by rounding the reciprocal of .
[0027] According to another aspect of the present disclosure, a method for processing dynamic mesh data is provided. In the method, a code stream of dynamic mesh data is processed according to format rules. The code stream includes base mesh information of a base mesh, and the base mesh includes a subset of multiple vertices of the mesh in the current mesh frame. The format rule specifies: determining (i) a position prediction of a current vertex of the base mesh and (ii) a motion field prediction of the current vertex of the base mesh. The format rule specifies: (i) determining a position prediction residual of the position prediction of the current vertex based on a first quantization step value, and (ii) determining a motion field prediction residual of the motion field prediction of the current vertex based on a second quantization step value, wherein the second quantization step value depends on the first quantization step value. The format rule specifies: (i) processing the position of the current vertex of the base mesh based on the position prediction and the position prediction residual, and (ii) processing the motion field of the current vertex of the base mesh based on the motion field prediction and the motion field prediction residual.
[0028] Aspects of the present disclosure also provide an apparatus for trellis coding. The apparatus for trellis coding comprises a processing circuit configured to implement any of the described methods for trellis coding.
[0029] Aspects of the present disclosure also provide an apparatus for trellis decoding. The apparatus for trellis decoding comprises a processing circuit configured to implement any of the described methods for trellis decoding.
[0030] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions which, when executed by a computer, cause the computer to perform any of the described methods for grid decoding, grid encoding, and grid data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Further features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0032] Figure 1 is a schematic illustration of an example of a block diagram of a communication system (100).
[0033] Figure 2 is a schematic illustration of an example of a block diagram of a decoder.
[0034] Figure 3 is a schematic illustration of an example of a block diagram of an encoder.
[0035] Figure 4 is a schematic illustration of an example of an encoding process (400) for mesh processing according to an aspect of the present disclosure.
[0036] Figure 5 is a schematic illustration of an example of a pre-processing step (500) according to an aspect of the present disclosure.
[0037] Figure 6 is a schematic illustration of a decoding process (600) for trellis processing according to an aspect of the present disclosure.
[0038] Figure 7 is a schematic illustration of an example of base grid encoding according to an aspect of the present disclosure.
[0039] Figure 8 is a schematic illustration of an example of base grid encoding with fine quantization granularity according to an aspect of the present disclosure.
[0040] Fig. 9 A flow chart outlining a trellis decoding process according to some aspects of the present disclosure is shown.
[0041] Fig.10 A flow chart outlining a trellis encoding process according to some aspects of the present disclosure is shown.
[0042] Fig.11 is a schematic illustration of a computer system according to an aspect of the present disclosure. DETAILED DESCRIPTION
[0043] Figure 1 A block diagram of a video processing system (100) in some examples is shown. The video processing system (100) is an example of an application of the disclosed subject matter, namely a video encoder and a video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., and the like.
[0044] The video processing system (100) includes an acquisition subsystem (113), which may include a video source (101). The video source (101) may include one or more images acquired by a camera and / or generated by a computer. For example, a digital camera creates, for example, an uncompressed video picture stream (102). In one example, the video picture stream (102) includes samples taken by the digital camera. The video picture stream (102), which is depicted as a thick line to emphasize the high amount of data compared to the encoded video data (104) (or the encoded video bitstream), can be processed by an electronic device (120), which includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. The encoded video data (104) (or encoded video bitstream), depicted as thin lines to emphasize the lower amount of data compared to the video picture stream (102), can be stored on the streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1 A client subsystem (106) and a client subsystem (108) in a streaming server (105) may access a copy (107) and a copy (109) of the encoded video data (104). The client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and creates an output video picture stream (111) that can be presented on a display (112) (e.g., a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), the video data (107), and the video data (109) (e.g., a video bitstream) may be encoded according to certain video encoding / compression standards. Examples of these standards include ITU-T H.265. In one example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of the VVC standard.
[0045] It should be noted that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may also include a video encoder (not shown).
[0046] Figure 2An example of a block diagram of a video decoder (210) is shown. The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used to replace Figure 1 A video decoder (110) in an example of FIG.
[0047] The receiver (231) may receive one or more encoded video sequences, for example, included in a bitstream to be decoded by the video decoder (210). In one aspect, the encoded video sequences are received one at a time, wherein the decoding of each encoded video sequence is independent of the decoding of the other encoded video sequences. The encoded video sequence may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data and other data, for example, encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not depicted). The receiver (231) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory (215) may be located external to the video decoder (210) (not depicted). In still other applications, a buffer memory (not depicted) may be provided external to the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) may be provided internal to the video decoder (210) to, for example, handle playout timing. When the receiver (231) receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (215) may not be required, or the buffer memory (215) may be made smaller. For use on a traffic packet network such as the Internet, the buffer memory (215) may be required, and the buffer memory (215) may be relatively large, advantageously may have an adaptive size, and may be implemented at least partially in an operating system or similar element (not depicted) external to the video decoder (210).
[0048] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. The categories of symbols include information for managing the operation of the video decoder (210) and potential information for controlling a display device such as a display device (212) (e.g., a display screen) that is not part of the electronic device (230) but can be coupled to the electronic device (230), such as Figure 2As shown. The control information for the display device may be a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not depicted). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may be performed according to a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract a subgroup parameter set for at least one subgroup of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (220) may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the coded video sequence.
[0049] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).
[0050] Depending on the type of coded video picture or part of coded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of symbol (221) may involve multiple different units. Which units are involved and how they are involved may be controlled by subgroup control information parsed by parser (220) from the coded video sequence. For clarity, such subgroup control information flow between parser (220) and the multiple units below is not depicted.
[0051] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into a number of functional units as described below. In actual implementations operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purposes of describing the disclosed subject matter, the conceptual subdivision into the following multiple functional units is appropriate.
[0052] The first unit is a sealer / inverse transform unit (251). The sealer / inverse transform unit (251) receives quantized transform coefficients as symbols (221) from the parser (220) and control information including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit (251) may output a block including sample values, which may be input into an aggregator (255).
[0053] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses surrounding reconstructed information extracted from a current picture buffer (258) to generate a block of the same size and shape as the block being reconstructed. For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a per-sample basis.
[0054] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to a block that is inter-coded and potentially motion compensated. In this case, the motion compensated prediction unit (253) may access the reference picture memory (257) to extract samples for prediction. After the extracted samples are motion compensated according to the symbols (221) belonging to the block, these samples may be added to the output of the scaler / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) by the aggregator (255), thereby generating output sample information. The extraction of prediction samples by the motion compensated prediction unit (253) from the address in the reference picture memory (257) may be controlled by a motion vector, which may be provided to the motion compensated prediction unit (253) in the form of symbols (221), which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.
[0055] The output samples of the aggregator (255) may be employed by various loop filtering techniques in a loop filter unit (256). The video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video sequence (also referred to as a coded video bitstream) and available to the loop filter unit (256) as symbols (221) from the parser (220). The video compression may also be responsive to meta-information obtained during decoding of a coded picture or a previous (in decoding order) portion of the coded video sequence, and to previously reconstructed and loop filtered sample values.
[0056] The output of the loop filter unit (256) may be a sample stream that may be output to a display device (212) and stored in a reference picture memory (257) for future inter-picture prediction.
[0057] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed, and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) may become part of the reference picture memory (257), and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.
[0058] The video decoder (210) may perform decoding operations according to a predetermined video compression technology or standard such as ITU-T H.265. The encoded video sequence may conform to the syntax specified by the video compression technology or standard used in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as the only tools available for use under the profile. For compliance, it is also required that the complexity of the encoded video sequence is within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata of the HRD buffer management signaled in the encoded video sequence.
[0059] In one aspect, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0060] Figure 3 An example of a block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used to replace Figure 1 A video encoder (103) in an example of FIG.
[0061] The video encoder (303) can be used to obtain the video source (301) (not Figure 3 In an example, the video source (301) is a part of the electronic device (320) to receive video samples, and the video source (301) can capture video images to be encoded by the video encoder (303). In another example, the video source (301) is a part of the electronic device (320).
[0062] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), the digital video sample stream may have any suitable bit depth (e.g. 8-bit, 10-bit, 12-bit, ...), any color space (e.g. BT.601 Y CrCB, RGB, ...) and any suitable sampling structure (e.g. Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The pictures themselves may be organized into a spatial pixel array, where each pixel may include one or more samples, depending on the sampling structure, color space, etc. used. The following description focuses on the samples.
[0063] According to one aspect, the video encoder (303) can encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other time constraints required. Implementing the appropriate encoding speed is a function of the controller (350). In some aspects, the controller (350) controls other functional units as described below and is functionally coupled to these other functional units. For clarity, the coupling is not depicted in the figure. The parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization techniques, etc.), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions related to the video encoder (303) optimized for a certain system design.
[0064] In some aspects, the video encoder (303) is configured to operate in a coding loop. As a simple description, in one example, the coding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Since the decoding of the symbol stream produces a bit-accurate result that is independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-accurately corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, such as due to channel errors) is also used in some related technologies.
[0065] The operation of the "local" decoder (333) may be similar to that described above in conjunction with Figure 2 The "remote" decoder of the video decoder (210) described in detail is identical. However, additional brief reference is made to Figure 2 , since the symbols are available and the entropy encoder (345) and the parser (220) can losslessly encode / decode the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (210) including the buffer memory (215) and the parser (220) may not be fully implemented in the local decoder (333).
[0066] On the one hand, except for the parsing / entropy decoding present in the decoder, the decoder technology is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on the decoder operation. The description of the encoder technology can be simplified because the encoder technology is mutually inverse to the decoder technology described comprehensively. In some areas, a more detailed description is provided below.
[0067] During operation, in some examples, the source encoder (330) may perform motion compensated predictive coding that predictively encodes an input picture by referencing one or more previously encoded pictures from a video sequence designated as "reference pictures." In this manner, the encoding engine (332) encodes the differences between pixel blocks of the input picture and pixel blocks of a reference picture that may be selected as a prediction reference for the input picture.
[0068] The local video decoder (333) may decode the encoded video data of the picture that may be designated as the reference picture based on the symbol created by the source encoder (330). The operation of the encoding engine (332) may advantageously be a lossy process. When the encoded video data may be decoded at the video decoder ( Figure 3 When the video sequence is decoded at a remote video decoder (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture memory (334). In this way, the video encoder (303) may store a copy of the reconstructed reference picture locally that has common content (absent transmission errors) with the reconstructed reference picture to be obtained by the remote video decoder.
[0069] The predictor (335) may perform a prediction search for the encoding engine (332). That is, for a new picture (or grid) to be encoded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may be used as appropriate prediction references for the new picture. The predictor (335) may operate on a pixel block by pixel block basis to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (335), it may be determined that the input picture may have prediction references taken from a plurality of reference pictures stored in the reference picture memory (334).
[0070] In one example, the mesh position quantization step size is defined based on a first parameter (e.g., mesh_position_quantization_step_size_log2_denominator) and a second parameter (e.g., mesh_position_quantization_step_size_numerator_minus1). In one example, mesh_position_quantization_step_size_log2_denominator is the logarithm (log2) of the denominator of the position quantization step size to the base 2, and its value is between 0 and 7 (including 0 and 7). In one example, mesh_position_quantization_step_size_numerator_minus1 plus 1 is the value of the numerator of the position quantization step size, and the value is between 1 and 256 (including 1 and 256). After the position encoding in the MEB, the reconstructed position value used for reference is scaled to the dynamic range of the inter-frame.
[0071] An example of the grid attribute encoding parameter syntax is as follows:
[0072]
[0073] An example of the semantics of mesh attribute encoding parameters is defined by a parameter such as mesh_attribute_quantization_step_size_log2_denominator[i]. In one example, mesh_attribute_quantization_step_size_log2_denominator[i] is the log2 of the denominator of the i-th attribute quantization step size, and its value is between 0 and 7 (inclusive).
[0074] An example of a generic base grid sequence parameter set RBSP syntax is shown below:
[0075]
[0076]
[0077]
[0078] An example of motion field quantization is defined by a first parameter (e.g., bmsps_inter_quantization_step_size_log2_denominator) and a second parameter (e.g., bmsps_inter_quantization_step_size_numerator_minus1). In one example, bmsps_inter_quantization_step_size_log2_denominator is the log2 of the denominator of the motion field quantization step size, and its value is between 0 and 7 (including 0 and 7). In one example, bmsps_inter_quantization_step_size_numerator_minus1 plus 1 is the value of the numerator of the motion field quantization step size, and the value is between 1 and 256 (including 1 and 256). After motion field encoding, the position value is reconstructed and can be used as a reference for future vertices. In one example, mesh attribute quantization is defined by a step size parameter (e.g., mesh_attribute_quantization_step_size_numerator[i]). In one example, mesh_attribute_quantization_step_size_numerator[i] is the value of the numerator of the i-th attribute quantization step size, which is between 0 and 255 (including 0 and 255). At the same time, the quantized prediction residual is dequantized. The reconstructed value is calculated by adding the dequantized prediction residual and the predicted value. The reconstructed value can be used as a reference for future vertices.
[0079] The controller (350) may manage encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.
[0080] The outputs of all the above functional units may be entropy encoded in an entropy encoder (345). The entropy encoder (345) applies lossless compression to the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into a coded video sequence.
[0081] The transmitter (340) may buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission over a communication channel (360), which may be a hardware / software link to a storage device that may store the encoded video data. The transmitter (340) may combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or ancillary data streams (source not shown).
[0082] The controller (350) may manage the operation of the video encoder (303). During encoding, the controller (350) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following picture types:
[0083] Intra pictures (I pictures) that can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures.
[0084] Predictive pictures (P pictures), which can be encoded and decoded using intra prediction or inter prediction, which uses a motion vector and reference index to predict the sample values of each block.
[0085] Bidirectional predictive pictures (B pictures), which can be encoded and decoded using intra prediction or inter prediction, which uses two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.
[0086] The source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined according to the coding allocation applied to the block's corresponding picture. For example, a block of an I picture may be non-predictively coded, or the block may be predictively coded (spatial prediction or intra prediction) with reference to an already coded block of the same picture. A block of a P picture may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. A block of a B picture may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.
[0087] The video encoder (303) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.
[0088] In one aspect, the transmitter (340) may transmit additional data when transmitting the encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0089] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In one example, a particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.
[0090] In some aspects, a bidirectional prediction technique may be used for inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture that precede the current picture in the video in decoding order (but may be in the past and future, respectively, in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.
[0091] In addition, merge mode technology can be used for inter-picture prediction to improve coding efficiency.
[0092] According to some aspects of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks (e.g., polygonal blocks or triangle blocks). For example, according to the High-Efficiency Video Coding (HEVC) standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU includes three coding tree blocks (CTBs), which are a luminance CTB and two chrominance CTBs. Each CTU can be recursively divided into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be divided into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type for the CU, such as an inter prediction type or an intra prediction type. The CU is divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Typically, each PU includes a luma prediction block (PB) and two chroma PBs. On the one hand, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Taking the luma prediction block as an example, the prediction block includes a matrix of pixel values (e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0093] It should be noted that the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using any suitable technology. In one aspect, the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using one or more integrated circuits. In another aspect, the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using one or more processors executing software instructions.
[0094] Aspects of the present disclosure include techniques for base mesh encoding in mesh compression. For example, base mesh encoding includes quantization of positions, motion fields, and texture coordinates of vertices of the base mesh.
[0095] A mesh may include multiple polygons that describe the surface of a volumetric object. Each polygon of a mesh may be defined by the vertices of the corresponding polygon in a three-dimensional (3D) space and information about how the vertices are connected (which may be referred to as connectivity information). In some aspects, vertex attributes (e.g., color, normal, etc.) may be associated with vertices (or mesh vertices). Attributes (or vertex attributes) may also be associated with the surface of a mesh by utilizing mapping information that uses a two-dimensional (2D) attribute map to parameterize the mesh. This mapping may be described by a set of parameter coordinates (referred to as UV coordinates or texture coordinates) associated with mesh vertices. A 2D attribute map may be used to store high-resolution attribute information, such as textures, normals, displacements, etc. High-resolution attribute information may be used for various purposes, such as texture mapping and shading.
[0096] Dynamic mesh sequences may require a large amount of data because dynamic meshes may include a large amount of information that changes over time. Therefore, efficient compression techniques can be used to store and transmit such content. Mesh compression standards (e.g., Information and Communication (IC) mesh compression, MESHGRID, frame-based animated mesh compression (FAMC)) were previously developed by the Moving Picture Experts Group (MPEG) to handle dynamic meshes with constant connectivity, time-varying geometry and vertex attributes. However, these standards do not consider time-varying attribute graphs and connectivity information. DCC (Digital Content Creation) tools can generate such dynamic meshes. However, for volume acquisition technology, it may be challenging to generate dynamic meshes with constant connectivity, especially to generate dynamic meshes with constant connectivity under real-time constraints. Existing standards may not support this type of content (e.g., dynamic meshes with constant connectivity). The present disclosure includes aspects of a new mesh compression standard that can directly handle dynamic meshes with time-varying connectivity information and optional time-varying attribute graphs. Mesh compression can target lossy and lossless compression for various applications, such as real-time communication, storage, free viewpoint video, augmented reality (AR) and virtual reality (VR). Features such as random access and scalable / progressive coding can also be considered.
[0097] The mesh geometry information may include vertex connectivity information, 3D coordinates, 2D texture coordinates, etc. The 3D vertex coordinates and 2D texture coordinates occupy a large portion of the mesh geometry information. Therefore, it is necessary to compress the 3D vertex coordinates (also called vertex positions) and 2D texture coordinates to reduce the amount of data required to store and / or transmit the mesh geometry information.
[0098] Figure 4 An example of an encoding process (400) for mesh processing based on a related video codec (e.g., MPEG V-Mesh TMv1.0) according to an aspect of the present disclosure is shown. Figure 4 As shown, the encoding process (400) may include a preprocessing step (400A) and an encoding step (400B). The preprocessing step (400A) may be configured to generate a base grid m(i) of the current frame and a displacement field d(i) of the current frame based on an input grid M(i) of the current frame, wherein the displacement field d(i) of the current frame includes a displacement vector. The encoding step (400B) may be configured to encode the base grid m(i), the displacement field d(i), and the texture information of the base grid m(i). The displacement field d(i) of the current frame may include a displacement vector. The index i may refer to the current frame. In one aspect, a mode decision method may be performed in the encoding process (400) to determine whether to apply inter-frame coding (also referred to as inter-frame prediction or inter-frame mode), intra-frame coding (also referred to as intra-frame prediction or intra-frame mode), etc. to the current frame. For example, the mode decision method may compare the cost of the intra-frame mode with the cost of the inter-frame mode, and determine the encoding mode of the base grid m(i) of the current frame based on which cost is smaller. In some examples, the base grid m(i) is encoded (e.g., encoded or decoded) using a skip mode. In one example, the skip mode is a special mode of the inter-frame mode. For example, the base grid m(i) can be intra-coded, or inter-coded, or encoded using a skip (SKIP) mode.
[0099] Still reference Figure 4, the preprocessing step (400A) may include a mesh extraction process (402), a parameterization process such as an atlas parameterization process (404), and a subdivision surface fitting process (406). The mesh extraction process (402) is configured to downsample the vertices of the input mesh M(i) to generate an extracted mesh dm(i) that may include a plurality of extracted (or downsampled) vertices. The number of the plurality of extracted vertices is less than the number of vertices of the input mesh M(i). The parameterization process such as the atlas parameterization process (404) is configured to map the extracted mesh dm(i) onto a planar domain, such as onto a UV atlas (or UV map), to generate a re-parameterized mesh pm(i). In one example, atlas parameterization can be performed based on a video processing tool (e.g., a UVAtlas tool). The subdivision surface fitting process (406) is configured to take the re-parameterized mesh pm(i) and the input mesh M(i) as input, and generate a base mesh m(i) and a displacement field d(i) including a displacement vector or a displacement set. In an example of the subdivision surface fitting process (406), pm(i) is subdivided using a subdivision scheme such as iterative interpolation to obtain a subdivided mesh. The iterative interpolation includes inserting a new point in the middle of each edge of the re-parameterized mesh pm(i) at each iteration. Any suitable subdivision scheme can be applied to subdivide pm(i). The displacement field d(i) is calculated by determining the nearest point on the surface of the input mesh M(i) for each vertex of the subdivided mesh.
[0100] Advantages of the subdivided mesh may include: the subdivided mesh has a subdivision structure that allows efficient compression while providing a reliable approximation to the input mesh. The improved compression efficiency may be obtained due to the following properties. The decimated mesh dm(i) may have a small number of vertices and may be encoded and transmitted using fewer bits than the input mesh M(i) or the number of subdivided meshes. Figure 4 , a base mesh m(i) can be generated based on the extracted mesh dm(i). In one example, the base mesh m(i) is the extracted mesh dm(i). Since the subdivided mesh can be generated based on the subdivision method, the subdivided mesh can be automatically generated by the decoder when the base mesh or the extracted mesh is decoded (for example, without using any information other than the subdivision scheme and the subdivision iteration count). On the decoder side, the displacement field d(i) can be generated by decoding the displacement vectors associated with the vertices of the subdivided mesh. In addition to enabling spatial / quality scalability, the subdivision structure enables efficient transformations such as wavelet decomposition, which can provide high compression performance.
[0101] For simplicity, the preprocessing step (400A) that can be applied to an input mesh (e.g., a 3D mesh) can be described using the preprocessing step (500) applied to a two-dimensional (2D) curve. The preprocessing step (400A) is similar to the preprocessing step (500), except that the 3D mesh can be replaced by a 2D curve.
[0102] Figure 5 An example of a preprocessing step (500) according to an aspect of the present disclosure is shown. Figure 5 As shown, an input 2D curve (represented by a 2D polyline) (502) can be downsampled to generate a base curve such as a polyline, which is referred to as a "decimated" curve (504). Then, a subdivision scheme can be applied to the decimated polyline (504) to generate a "decision" curve (506). In one example, the subdivision scheme can be an iterative interpolation scheme. The iterative interpolation scheme can include inserting a new point in the middle of each edge of the polyline (or decimated curve) (504) at each iteration. For example, a point (510) can be inserted in an edge (508) of the decimated curve (504). In one example, the edge (508) is located between a point (512) and a point (514). In addition, a point (522) can be added between the point (512) and the point (510), and a point (516) can be added between the point (510) and the point (514). Then, the decimal polyline (506) is deformed to generate a displacement curve (518). The displacement curve (518) may be a better approximation of the input curve (502) than the subdivided curve (506). For example, a displacement vector (e.g., (520)) is calculated for each vertex (e.g., (510)) of the subdivided curve (506) so that the shape of the displacement curve (518) is as close as possible to the shape of the input curve (502). An advantage of the subdivided curve (506) is that the subdivided curve (506) has a subdivision structure that allows for more efficient compression while providing a reliable approximation of the input curve (502).
[0103] The decimation curve (504) may have a smaller number of points and may be encoded and transmitted using a limited number of bits. Since the subdivision curve may be generated based on the subdivision scheme, the subdivision curve may be automatically generated by the decoder when decoding the base curve or the decimation curve (e.g., without using any information other than the subdivision method and the subdivision iteration count). The displacement curve may be generated by decoding the displacement vectors associated with the vertices of the subdivision curve. In addition to enabling spatial / quality scalability, the subdivision structure enables efficient transforms such as wavelet decomposition, which may provide high compression performance.
[0104] Still reference Figure 5In one example, the input mesh M(i) may include an input 2D curve (502). The base mesh m(i) may include a decimated curve (504), which is formed by downsampling the vertices of the input 2D curve (502). The displacement field dm(i) may include a plurality of displacement vectors, such as Figure 5 The displacement vector (520) is shown.
[0105] The encoding step (400B) may include base mesh encoding (408), displacement encoding (410), texture encoding (412), etc. The base mesh encoding (408) is configured to encode geometric information of a base mesh m(i) associated with a current frame. In intra-frame coding, the base mesh m(i) may first be quantized (e.g., quantized using uniform quantization) and then encoded, for example, by using a coding mode determined by a mode decision method. The coding mode may be an inter-frame mode, an intra-frame mode, a skip mode, etc. An encoder for intra-coding the base mesh m(i) may be referred to as a static mesh encoder. In inter-frame coding, a reference base mesh associated with a reference frame indicated by index j (e.g., a reconstructed quantized reference base mesh m'(j)) may be used to predict a base mesh m(i) associated with a current frame indicated by index i. The displacement encoding (410) is configured to encode a displacement field d(i) generated in the preprocessing step (400A). The displacement field d(i) may include a set of displacement vectors (or displacements) associated with vertices of a subdivided mesh. The texture encoding (412) is configured to encode attribute information of the base mesh m(i). The attribute information may include texture, normal, color, etc. The attribute information may be encoded based on a suitable codec (e.g., High Efficiency Video Coding (HEVC) or Next Generation Video Coding (VVC)).
[0106] On the one hand, reference Figure 4 , a mesh encoding process such as the encoding process (400) begins with preprocessing (e.g., preprocessing step (400A)). The preprocessing may convert an input mesh (e.g., an input dynamic mesh) M(i) into a base mesh m(i) and a displacement field d(i) including a displacement set (or a displacement vector set). The encoding step (400B) may compress the output from the preprocessing (e.g., m(i), d(i), etc.) and generate a compressed code stream b(i). The compressed code stream b(i) may include a compressed base mesh code stream, a compressed displacement field code stream, a compressed attribute code stream, etc.
[0107] Figure 6An example of a decoding process (600) for grid processing according to an aspect of the present disclosure is shown. The decoding process (600) may include a decoding step (605) and a post-processing step (610). A compressed code stream b(i) may be provided to the decoding step (605). In one example, for example for lossless transmission, the compressed code stream b(i) is the output b(i) from the encoding process (400). The decoding step (605) may extract various sub-code streams, such as a compressed base grid sub-stream, a compressed displacement field sub-stream, a compressed attribute sub-stream, etc. The decoding step (605) may decompress the sub-code streams to generate the following components: block (patch) metadata indicated by metadata (i) (metadata(i)), decoded base grid m" (i), decoded displacement field (including displacement) d" (i), decoded attribute map A" (i), etc.
[0108] On the one hand, the base grid substream may be provided to a grid decoder to generate a reconstructed quantized base grid m'(i). The decoded base grid (or reconstructed base grid) m"(i) may be obtained by applying inverse quantization to m'(i). The displacement field substream including encoded, packed and quantized wavelet coefficients may be decoded by a video and / or image decoder. Image unpacking and inverse quantization may be applied to the reconstructed, packed quantized wavelet coefficients to obtain unpacked and dequantized transform coefficients (e.g., wavelet coefficients). An inverse wavelet transform may be applied to the unpacked and dequantized wavelet coefficients to generate a decoded displacement field (or reconstructed displacement) d"(i).
[0109] The decoded components (e.g., including metadata(i), m"(i), d"(i), A"(i), etc.) can be provided to a post-processing step (610). A mesh (also referred to as a decoded / reconstructed mesh) M"(i) can be generated based on m"(i) and d"(i) by the post-processing step (610). In one example, the mesh M"(i) (also referred to as a reconstructed deformed mesh DM(i)) can be obtained by subdividing m"(i) using a subdivision scheme and applying reconstructed displacements d"(i) to the vertices of the subdivided mesh. In one example, DM(i) can include a displacement curve (518). In one example, when the encoding process (400), the decoding process (600), and the transmission are lossless, the mesh M"(i) can be the same as the input mesh M(i). When one of the encoding process (400), the decoding process (600), and the transmission is lossy, M"(i) is different from M(i). In various examples, the difference between M"(i) and M(i), if any, can be relatively small. In one example, a property graph A”(i) is also generated by a post-processing step (610).
[0110] In one aspect, the base grid may be intra-coded, inter-coded, or encoded using a SKIP mode, etc. In one example, the SKIP mode may be a special mode of the inter-mode, in which the base grid m(i) of the current frame indicated by index i is the same as the base grid m(j) of the reference frame indicated by index (also referred to as frame index) j. When the inter-mode is applied to encode the base grid in the current frame, the encoder may generate a predicted base grid for the current frame based on the reconstructed base grid of the reference frame. In one example, such as in MPEG V-DMC WD 2.0, the reference frame is a frame immediately preceding the current frame in display order. The frame index i of the current frame indicates the display order. When the frame index of the current frame is i, the frame index of the reference frame is (i-1). In one example, the current frame and the reference frame are in the same group of frames (GoF).
[0111] In one aspect, for example in MPEG V-DMC WD 2.0, the mesh encoding process starts with preprocessing. The preprocessing may convert the input dynamic mesh (denoted as M(i)) into a base mesh m(i) and a set of displacements d(i). The encoder may compress the base mesh m(i) and the displacements d(i) to generate a compressed code stream b(i).
[0112] like Figure 4 As shown, preprocessing may include mesh extraction, atlas parameterization, and subdivision surface fitting. Mesh extraction may use simplification techniques to extract the input mesh M(i) and generate an extracted mesh dm(i). The extracted mesh dm(i) may be re-parameterized. The generated mesh may be denoted as pm(i). Subdivision surface fitting may take the re-parameterized mesh pm(i) and the input mesh M(i) as input to generate a base mesh m(i) and a displacement set d(i).
[0113] For intra frames, the base mesh can be encoded using a static mesh codec, where the positions are encoded using spatial prediction. For inter frames, the base mesh can be encoded using motion field coding. The motion field of the base mesh can be defined as the difference between the vertex positions of the current frame and the vertex positions of the reference frame of the current frame, and motion field coding can utilize temporal prediction.
[0114] In the present disclosure, methods and systems for base mesh coding in mesh compression are provided. In one aspect, base mesh coding includes quantization of positions and motion fields in mesh compression. In one example, for example in MPEG V-DMC WD2.0, vertex positions of intra frames and motion fields of inter frames can utilize a workflow including bit depth quantization, prediction, and entropy coding. An example of the workflow is described in Figure 7. In bit depth quantization, the position (and / or motion field) can be quantized to a bit depth value specified by the encoder. For example, if the encoded bit depth value is 12, the position is quantized to an integer between 0 and 4095, so the quantized motion field can be the difference between two 12-bit integers. The quantized value (e.g., quantized position or quantized motion field) can be predicted. For example, the quantized position is predicted using a multi-parallelogram algorithm, and the quantized motion field is predicted using a neighborhood prediction algorithm. The prediction residual can be further compressed by entropy coding.
[0115] like Figure 7 As shown, the workflow (700) includes bit depth quantization (702), prediction (704), and entropy encoding (706). In bit depth quantization (702), the position (or position coordinate) of the vertex in the base mesh and the 2D texture coordinate are quantized into bit depth values. The bit depth value may be specified (or defined) by the encoder. The base mesh may include a subset of multiple vertices of the mesh. In one aspect, if the bit depth value of the position is n and the bit depth value of the 2D texture coordinate is m, the position is quantized to a value between 0 and 2. n -1, and quantizes the 2D texture coordinates to a value between 0 and 2 m -1, where n and m are positive integers. For example, if the bit depth value of the position is 12 and the bit depth value of the 2D texture coordinate is 13, the position is quantized to an integer between 0 and 4095, and the 2D texture coordinate is quantized to an integer between 0 and 8191. The quantized value may be predicted at prediction (704) by a prediction mode or a prediction algorithm (e.g., by a multi-parallelogram algorithm for the position of the vertex and a stretch prediction algorithm for the 2D texture coordinate). The prediction residual of the position of the vertex and the prediction residual of the 2D texture coordinate may be further compressed by entropy coding (706).
[0116] In one aspect, for example in MPEG V-DMC WD 2.0, the bit depth values of quantization are limited or constrained.In the present disclosure, position coding and / or motion field coding with fine quantization granularity is provided.
[0117] In the present disclosure, basic grid coding with fine quantization granularity is provided. Basic grid coding may include prediction, quantization, dequantization, and entropy coding. An example of basic grid coding (800) is shown in Figure 8 As shown in Figure 8As shown, the base mesh encoding (800) includes prediction (802), quantization (804), dequantization (808), and entropy encoding (806). In one example, the base mesh encoding (800) may also include bit depth quantization (not shown) located before the prediction (802). The bit depth quantization may be configured to quantize the positions of the vertices and 2D texture coordinates of the base mesh into bit depth values. At the prediction (802), when the base mesh encoding (800) includes bit depth quantization, the bit depth value generated at the time of bit depth quantization is predicted based on an algorithm or prediction mode (e.g., a multi-parallelogram algorithm for the positions of the vertices and a stretch prediction algorithm for the 2D texture coordinates). When the base mesh encoding (800) does not include bit depth quantization, the positions of the vertices and the 2D texture coordinates in the base mesh may be predicted based on a suitable prediction mode or algorithm. Therefore, the position prediction of the vertex, the 2D texture coordinate prediction, the position prediction residual of the position prediction, and the 2D texture coordinate prediction residual of the 2D texture coordinate prediction may be generated at the prediction (802). The position prediction residuals for the position prediction of the vertices and the 2D texture coordinate prediction residuals for the 2D texture coordinate prediction may be further quantized at quantization (804). The quantization (804) is configured to reduce the precision of the prediction residuals according to a quantization parameter (QP) (e.g., a quantization step value). Thus, less important information is discarded and significant compression is achieved by quantization (804). The quantized prediction residuals are further compressed by entropy coding (806).
[0118] Still reference Figure 8 , the quantized prediction residual obtained at quantization (804) may be provided to dequantization (808). Dequantization (808) may dequantize the quantized prediction residual to obtain the prediction residual generated at prediction (802). The prediction residual may be further provided to prediction (802). The position and 2D texture coordinates of the vertex may be reconstructed based on the prediction and the prediction residual. The reconstructed position and 2D texture coordinates of the vertex may be used as reference information for subsequent vertices.
[0119] In one aspect, the position of the current vertex is predicted by the positions of neighboring vertices (e.g., the positions of encoded vertices). In one aspect, the position of the current vertex is predicted by an algorithm such as multi-parallelogram prediction or any suitable position algorithm. In one aspect, the motion field of the current vertex is predicted based on one or more encoded motion fields by an algorithm such as a neighborhood prediction algorithm or any suitable motion field prediction algorithm.
[0120] The prediction residual is the difference between the true value (e.g., the true value of the position) and the predicted value (e.g., the predicted position), and the prediction residual can be quantized by a quantization step value. The quantization step value can be a positive integer, a positive rational number, or a positive real number.
[0121] The quantized prediction residual may be further compressed by entropy coding. The entropy coding may be fixed length coding, variable length coding, Huffman coding, arithmetic coding, or any other suitable entropy coding.
[0122] At the same time, the quantized prediction residual can be dequantized. For example, the quantized prediction residual can be dequantized at dequantization (808). A reconstructed value can be calculated by adding the dequantized prediction residual and the predicted value. The reconstructed value can be used as a reference for future vertices.
[0123] In the present disclosure, quantization of the position of the vertex and quantization of the motion field of the vertex are provided. On the one hand, the quantization step value of the motion field of the vertex and the quantization step value of the position of the vertex are dependent on each other. For example, the quantization step value of the motion field is determined by the quantization process (e.g., the quantization step value) of the position of the vertex.
[0124] In one aspect, the quantization step size of the motion field of a vertex depends on the quantization step size of the position of the vertex. For example, the quantization step size of the motion field is equal to the quantization step size of the position.
[0125] In one aspect, the quantization step value of the motion field of a vertex is equal to a multiple of the quantization step value of the position of the vertex, where the multiple can be a positive integer, a positive rational number, or a positive real number.
[0126] In one aspect, the quantization step value of the motion field of a vertex depends on a function associated with the quantization step value of the position of the vertex. For example, the quantization step value of the motion field is equal to a linear function of the quantization step value of the position of the vertex. The linear function may include the form f(x)=ax+b, where x is the quantization step value of the position, a is a non-negative real number, b is a real number, and f(x) is the quantization step value of the motion field.
[0127] In one aspect, the quantization step value of the motion field of a vertex is equal to a monotonically non-decreasing function of the quantization step value of the position of the vertex.
[0128] In one aspect, the quantization step value of the motion field of a vertex is equal to a monotonically non-decreasing function of the quantization error of the position of the vertex. The quantization error is the difference between the unquantized value and the quantized value. For example, the quantization error is the difference between the unquantized position prediction residual and the quantized position prediction residual.
[0129] In one aspect, a codestream syntax is provided for signaling a quantization step size value. The quantization step size value may be a fraction (e.g., a positive binary rational number). ), where the denominator of a positive binary rational number is a power of 2, the numerator m is a positive integer, and n is a non-negative integer. Binary rational numbers can be represented by signals using the integer pair (n, m)
[0130] In one aspect, the quantization step value is controlled at the sequence level. In a sequence, each intra frame may use the same quantization step value to quantize the position, and each inter frame may use the same quantization step value to quantize the motion field. A sequence parameter set including the quantization step value for the position and the motion field may be listed in a sequence header. An example of a syntax table for the quantization step value at the sequence level is shown in Table 1 below.
[0131] Table 1. Example of a syntax table for sequence-level quantization step values
[0132]
[0133]
[0134] As shown in Table 1, the parameter n of the positive binary rational number can be defined by a code stream syntax (e.g., qpLog2Denominator), which is the logarithm (log2) of the denominator of the quantization step value to the base 2. The parameter n may include values between 0 and 7. The parameter m may be defined by a code stream syntax (e.g., qpNumeratorMinusl) plus 1 and is defined as the value of the numerator of the quantization step value. The parameter m may include values between 1 and 256. The determination syntax element (e.g., motionQPEqualsPositionQP) is a 1-bit flag that signals (or indicates) whether the quantization step value of the position and the quantization step value of the motion field are the same. If motionQPEqualsPositionQP=1, the two quantization step values of the position and the motion field are the same. Otherwise, if motionQPEqualsPositionQP is not equal to 1, the quantization step value of the position is different from the quantization step value of the motion field.
[0135] On the other hand, the quantization step value is controlled at the frame level, wherein each frame may have a corresponding quantization step value for quantization of a position or a motion field, depending on whether the corresponding frame is an intra frame or an inter frame. For example, when the frame is an intra frame, each intra frame has a corresponding quantization step value for quantization of a position. When the frame is an inter frame, each inter frame has a corresponding quantization step value for quantization of a motion field. A picture parameter set may be listed in a frame header, the picture parameter set including quantization step values for a position and a motion field. An example of a syntax table of quantization step values at the frame level is provided in Table 2 as follows.
[0136] Table 2. Example of a syntax table for quantization step values at the frame level
[0137]
[0138]
[0139] As shown in Table 2, the parameter n of the positive binary rational number is defined by the codestream syntax (e.g., qpLog2Denominator), which is the log2 of the denominator of the quantization step value. The parameter n may include values between 0 and 7. The parameter m of the positive binary rational number is defined by the codestream syntax (e.g., qpNumeratorMinusl) plus 1. Therefore, the parameter m is defined as the value of the numerator of the quantization step value. The parameter m may include values between 1 and 256. The operation syntax (e.g., copyQP) is a 1-bit flag that signals whether the quantization step value of the current frame is copied from the reference frame of the current frame. If copyQP=1, the quantization step value of the current frame is copied from the quantization step value of the reference frame. Otherwise, if copyQP is not equal to 1, the quantization step value is not copied from the quantization step value of the reference frame.
[0140] It should be noted that at the level of Table 1 and Table 2, the code stream syntax (such as qpLog2Denominator and qpNumeratorMinus1) allows the quantization step value to be less than 1. For example, if qpLog2Denominator = 7, qpNumeratorMinus1 = 80, then the quantization step value is
[0141] In the present disclosure, methods and systems are provided for 2D texture coordinate encoding in mesh compression. In one example, such as in MPEG V-DMC WD 2.0, 2D texture coordinate encoding utilizes a workflow including bit depth quantization, prediction, and entropy coding. An example of the workflow is described in Figure 7 In bit depth quantization, the 2D texture coordinates of the vertex can be quantized based on the bit depth value specified by the encoder. When the bit depth value is n, the 2D texture coordinates can be quantized to a value between 0 and 2. n -1. For example, if the bit depth value of the 2D texture coordinate is 13, the 2D texture coordinate is quantized to an integer between 0 and 8191. The quantized value (e.g., the quantized 2D texture coordinate) may be further predicted by an algorithm for the 2D texture coordinate (e.g., a stretch prediction algorithm), and the prediction residual may be further compressed by entropy coding. In one example, such as in MPEG V-DMC WD 2.0, the bit depth value for quantization of the 2D texture coordinate is limited.
[0142] In the present disclosure, texture coordinate encoding with fine quantization granularity is provided. Texture coordinate encoding may include prediction, quantization, dequantization, and entropy encoding, such as Figure 8As shown. The 2D texture coordinates may be predicted by a coding algorithm (e.g., a stretch prediction algorithm), or may be directly encoded without prediction. In one aspect, different quantization and entropy schemes may be applied, depending on whether the 2D texture coordinates are predicted or directly encoded.
[0143] In one example, if the 2D texture coordinates of the vertices are directly encoded without prediction, quantization such as rounding quantization is applied to the quantization of the 2D texture coordinates. Entropy coding (eg, fixed length coding) may then be applied to encode the quantized 2D texture coordinates.
[0144] In one aspect, rounding quantization is configured to convert a number (e.g., a real number, a rational number, or an integer) into an integer. In one example, rounding quantization can convert a 2D texture coordinate into an integer. Rounding quantization can round a number (e.g., a 2D texture coordinate) to the nearest integer. The nearest integer can be higher (or greater than), lower (or less than), or equal to the actual value of the 2D texture coordinate. In one aspect, an offset can be added to the value (e.g., a 2D texture coordinate) before rounding quantization. For example, an offset is added to the 2D texture coordinate to generate an updated 2D texture coordinate. The updated 2D texture coordinate is further quantized by rounding quantization. In one example, the offset is a positive number, such as 0.5, 0.3333, etc. In one example, the offset is a negative number, such as -0.5, -0.3333, etc.
[0145] In one aspect, the integers resulting from the rounded quantization are encoded by entropy encoding (eg, by fixed length encoding).The fixed length encoded values may further be written into a codestream.
[0146] In one example, if the 2D texture coordinates are predicted by an encoding algorithm (e.g., a stretch prediction algorithm), then by quantization (e.g., Figure 8 The prediction residual is quantized (804) in quantization. In one example, the prediction residual is quantized by a fraction (e.g., by a binary rational number ) quantizes the prediction residual, where m and n are both integers, and m>0, n>=0. In the example of quantization of the prediction residual, the prediction residual is divided by (or multiplied by) And further round the result (e.g., the result of division or multiplication) to an integer. The result can be rounded to the nearest integer, and the nearest integer can be higher (or greater than), lower (or less than), or equal to the actual value of the result. In one aspect, an offset is added to the result (e.g., the result of division or multiplication) before rounding quantization. In one example, the offset can be a positive number, such as 0.5, 0.3333, etc. In one example, the offset can be a negative number, such as -0.5, -0.3333, etc.
[0147] In one aspect, the integers resulting from the rounded quantization are encoded by entropy encoding, eg, by variable length encoding (eg, Golomb encoding).
[0148] An example of syntax and semantics related to quantization of 2D texture coordinates is shown in Table 3.
[0149] Table 3. Example of syntax table related to quantization of 2D texture coordinates
[0150]
[0151] As shown in Table 3, the mesh attribute bit depth syntax (e.g., mesh_attribute_bit_depth_minus1[i]) plus 1 may specify the number of bits used to represent the component of the i-th attribute. The first mesh attribute quantization step size syntax (e.g., mesh_attribute_quantization_step_size_log2_denominator[i]) may indicate a binary rational number The parameter n may be defined as the log2 of the denominator of the quantization step size of the i-th attribute and may include values between 0 and 7.
[0152] The second mesh attribute quantization step size syntax (e.g., mesh_attribute_quantization_step_size_numerator[i]) may indicate a dyadic rational number The parameter m may be defined as the value of the numerator of the i-th attribute quantization step and may include values between 0 and 255.
[0153] Still referring to Table 3, the first flag syntax (e.g., mesh_attribute_per_face_flag[i]) may indicate whether the i-th attribute has a value defined for each face (or for each face). When the flag syntax is equal to 1, the flag syntax indicates that the i-th attribute has a value defined for each face. When the first flag syntax (e.g., mesh_attribute_per_face_flag[i]) is equal to 0, the first flag syntax indicates that the i-th attribute has a value defined for each vertex (or for each vertex).
[0154] The syntax elements in Table 3 include a second flag syntax (e.g., mesh_attribute_separate_index_flag[i]) to specify whether to append the i-th attribute as a specific index sequence. When the second flag syntax is equal to 1, the second flag (e.g., mesh_attribute_separate_index_flag[i]) specifies that the i-th attribute is appended as a specific index sequence. When mesh_attribute_separate_index_flag[i] is equal to 0, mesh_attribute_separate_index_flag[i] specifies that the i-th attribute index is copied from the position index or the index of another attribute. When the first flag syntax (e.g., mesh_attribute_per_face_flag[i]) is equal to 1, the second flag syntax (e.g., mesh_attribute_separate_index_flag[i]) is set to a default value of 0.
[0155] In Table 3, when the second flag syntax (e.g., mesh_attribute_separate_index_flag[i]) is equal to 0, the reference syntax (e.g., mesh_attribute_reference_index_plusl[i]) minus 1 may specify the reference index of the i-th attribute. When mesh_attribute_reference_index_plusl[i] minus 1 is equal to -1, the reference index is a position index. When mesh_attribute_reference_index_plusl[i] minus 1 is not equal to -1, the reference index may specify the attribute used as a reference for the i-th attribute.
[0156] Fig. 9 A flow chart outlining a process (900) according to an aspect of the present disclosure is shown. The process (900) may be used in a video decoder. In various aspects, the process (900) is performed by a processing circuit, such as a processing circuit that performs the functions of a video decoder (110), a processing circuit that performs the functions of a video decoder (210), etc. In some aspects, the process (900) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (900). The process starts at (S901) and proceeds to (S910).
[0157] At (S910), a code stream is received, the code stream including base mesh information of a base mesh. The base mesh includes a subset of a plurality of vertices of a mesh in a current mesh frame.
[0158] At (S920), a position prediction of a current vertex of the base mesh is determined. A motion field prediction of a current vertex of the base mesh is determined.
[0159] At (S930), a position prediction residual of a position prediction of the current vertex is determined based on the first quantization step size value. A motion field prediction residual of a motion field prediction of the current vertex is determined based on a second quantization step size value, wherein the second quantization step size value depends on the first quantization step size value.
[0160] At (S940), the position of the current vertex of the base mesh is reconstructed based on the position prediction and the position prediction residual. The motion field of the current vertex of the base mesh is reconstructed based on the motion field prediction and the motion field prediction residual.
[0161] In one aspect, to determine a position prediction residual, a quantized position prediction residual of a position prediction of a current vertex is determined based on a first entropy coding. The quantized position prediction residual of the position prediction of the current vertex is further dequantized based on a first quantization step value to determine a position prediction residual. In one aspect, to determine a motion field prediction residual, a quantized motion field prediction residual of a motion field prediction of a motion field of the current vertex is determined based on a second entropy coding. The quantized motion field prediction residual of the motion field prediction of the current vertex is dequantized based on a second quantization step value to determine a motion field prediction residual.
[0162] In one aspect, the second quantization step value is equal to a multiple of the first quantization step value.
[0163] In one aspect, the second quantization step value is equal to one of a linear function and a monotonically non-decreasing function of the first quantization step value.
[0164] In one aspect, the second quantization step size value is equal to a monotonically non-decreasing function of a quantization error associated with the position prediction of the current vertex. The quantization error is the difference between a quantized position prediction residual and an unquantized position prediction residual for the current vertex.
[0165] In one aspect, the first quantization step value is a positive binary rational number Definition, where m is a positive integer between 0 and 7, and n is a non-negative integer between 1 and 256.
[0166] In one aspect, when the first quantization step value and the second quantization step value are received at the sequence level, the first quantization step value and the second quantization step value are equal when a syntax element in the codestream indicates that the first quantization step value is equal to the second quantization step value.
[0167] On the one hand, when a first quantization step value and a second quantization step value are received at a frame level, when a syntax element in a code stream is a first value, the second quantization step value is equal to a quantization step value of a reference frame of a current grid frame, and when a syntax element in a code stream is a second value, the second quantization step value is equal to the first quantization step value.
[0168] In process (900), a 2D texture coordinate prediction of a 2D texture coordinate of a current vertex of a base mesh is determined. A 2D texture coordinate prediction residual of the 2D texture coordinate prediction of the current vertex is determined. Based on the 2D texture coordinate prediction and the 2D texture coordinate prediction residual, the 2D texture coordinate of the current vertex of the base mesh is reconstructed.
[0169] In one aspect, when the 2D texture coordinates of the current vertex are directly encoded, in order to determine the 2D texture coordinate prediction residual, a quantized 2D texture coordinate of the current vertex is determined, wherein the quantized 2D texture coordinate is encoded by fixed length coding. The quantized 2D texture coordinate of the current vertex is dequantized to obtain the 2D texture coordinate. The quantized 2D texture coordinate is obtained by rounding quantization, wherein the rounding quantization is configured to convert the 2D texture coordinate into an integer.
[0170] In one aspect, when predicting the 2D texture coordinates of the current vertex by the stretch prediction algorithm, in order to determine the 2D texture coordinate prediction residual, a quantized 2D texture coordinate prediction residual of the current vertex is determined, wherein the quantized 2D texture coordinate prediction residual is encoded by variable length coding. The quantized 2D texture coordinate prediction residual of the current vertex is dequantized to determine the 2D texture coordinate prediction residual. The quantized 2D texture coordinate prediction residual is obtained by rounding quantization, wherein the rounding quantization is configured to convert the 2D texture coordinate prediction residual into an integer.
[0171] On the one hand, rounding quantization is done by binary rational numbers Defined as follows, where m is a positive integer between 0 and 7, and n is a non-negative integer between 0 and 255. The quantized 2D texture coordinate prediction residual is obtained by rounding off a multiple of the 2D texture coordinate prediction residual and the reciprocal of a binary rational number.
[0172] Then, the process proceeds to (S999) and terminates.
[0173] The process (900) may be adapted as appropriate. Steps in the process (900) may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.
[0174] Fig.10A flow chart outlining a process (1000) according to an aspect of the present disclosure is shown. The process (1000) may be used in a video encoder. In various aspects, the process (1000) is performed by a processing circuit, such as a processing circuit that performs the functions of a video encoder (103), a processing circuit that performs the functions of a video encoder (303), etc. In some aspects, the process (1000) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (1000). The process starts at (S1001) and proceeds to (S1010).
[0175] At (S1010), a position prediction of a current vertex of a base mesh associated with a mesh in a current mesh frame is determined. A motion field prediction of a current vertex of the base mesh is determined.
[0176] At (S1020), the position prediction residual of the position prediction of the current vertex is quantized based on the first quantization step value to generate a quantized position prediction residual. The motion field prediction residual of the motion field prediction of the current vertex is quantized based on the second quantization step value to generate a quantized motion field prediction residual. The second quantization step value depends on the first quantization step value.
[0177] At (S1030), the quantized position prediction residual of the position prediction of the current vertex is entropy encoded. The quantized motion field prediction residual of the motion field prediction of the current vertex is entropy encoded.
[0178] In one aspect, the second quantization step value is equal to a multiple of the first quantization step value.
[0179] In one aspect, the second quantization step value is equal to one of a linear function and a monotonically non-decreasing function of the first quantization step value.
[0180] In the method, a 2D texture coordinate prediction of a 2D texture coordinate of a current vertex of a base mesh is determined. A 2D texture coordinate prediction residual of the 2D texture coordinate prediction of the current vertex is quantized to generate a quantized 2D texture coordinate prediction residual of the current vertex. The quantized 2D texture coordinate prediction residual of the current vertex is further entropy encoded.
[0181] In one aspect, when directly encoding the 2D texture coordinates of the current vertex, in order to quantize the 2D texture coordinate prediction residual, the 2D texture coordinates of the current vertex are quantized based on rounded quantization, wherein the rounded quantization is configured to convert the 2D texture coordinates into integers. The quantized 2D texture coordinates are further encoded based on fixed length coding.
[0182] On the one hand, when predicting the 2D texture coordinates of the current vertex by the stretch prediction algorithm, in order to determine the 2D texture coordinate prediction residual, the multiple of the 2D texture coordinate prediction residual and the binary rational number The 2D texture coordinate prediction residual of the current vertex is quantized by rounding the reciprocal of , where m is a positive integer between 0 and 7, and n is a non-negative integer between 0 and 255. The quantized 2D texture coordinate prediction residual is further encoded based on variable length coding.
[0183] In one aspect, in order to quantize the 2D texture coordinate prediction residual, an offset is added to the 2D texture coordinate prediction residual to generate an updated 2D texture coordinate prediction residual. Further, by quantizing the multiple of the updated 2D texture coordinate prediction residual and the binary rational number The updated 2D texture coordinate prediction residual of the current vertex is quantized by rounding the reciprocal of .
[0184] Then, the process proceeds to (S1099) and terminates.
[0185] The process (1000) may be adapted as appropriate. Steps in the process (1000) may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.
[0186] In one aspect, a method for processing dynamic mesh data includes processing a code stream of the dynamic mesh data according to a format rule. For example, the code stream may be a code stream decoded / encoded using any of the decoding methods and / or encoding methods described herein. The format rule may specify one or more constraints of the code stream and / or one or more processes to be performed by a decoder and / or encoder.
[0187] In one example, a code stream of dynamic mesh data is processed according to a format rule. The code stream includes base mesh information of a base mesh, and the base mesh includes a subset of multiple vertices of a mesh in a current mesh frame. The format rule specifies: determining (i) a position prediction of a current vertex of the base mesh and (ii) a motion field prediction of the current vertex of the base mesh. The format rule specifies: (i) determining a position prediction residual of the position prediction of the current vertex based on a first quantization step value, and (ii) determining a motion field prediction residual of the motion field prediction of the current vertex based on a second quantization step value, wherein the second quantization step value depends on the first quantization step value. The format rule specifies: (i) processing the position of the current vertex of the base mesh based on the position prediction and the position prediction residual, and (ii) processing the motion field of the current vertex of the base mesh based on the motion field prediction and the motion field prediction residual.
[0188] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig.11 A computer system (1100) suitable for implementing certain aspects of the disclosed subject matter is shown.
[0189] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to assembly, compilation, linking or similar mechanisms to create code comprising instructions that may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpretation, microcode execution, etc.
[0190] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.
[0191] Fig.11 The components of the computer system (1100) shown are examples and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing aspects of the present disclosure. Nor should the configuration of components be interpreted as having any dependency or requirement related to any one or combination of components shown in the exemplary aspects of the computer system (1100).
[0192] The computer system (1100) may include certain human interface input devices. Such human interface input devices may be responsive to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not depicted). Human interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, captured images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0193] The human-machine interface input device may include one or more of the following (only one of each is shown): keyboard (1101), mouse (1102), touchpad (1103), touch screen (1110), data gloves (not shown), joystick (1105), microphone (1106), scanner (1107), camera (1108).
[0194] The computer system (1100) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback of a touch screen (1110), a data glove (not shown), or a joystick (1105), but may also be a tactile feedback device that is not an input device), an audio output device (e.g., a speaker (1109), headphones (not depicted)), a visual output device (e.g., a screen (1110) including a CRT screen, an LCD screen, a plasma screen, an OLED screen, each with or without touch screen input capabilities, each with or without tactile feedback capabilities, some of which are capable of outputting two-dimensional visual output or output in excess of three dimensions through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted)), and a printer (not depicted).
[0195] The computer system (1100) may also include human-machine accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1120) with CD / DVD and other media (1121), thumb drives (1122), removable hard drives or solid-state drives (1123), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security software dogs (not depicted), etc.
[0196] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.
[0197] The computer system (1100) may also include an interface (1154) to one or more communication networks (1155). The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter attached to some common data port or peripheral bus (1149) (e.g., a USB port of the computer system (1100)); other network interfaces are typically integrated into the kernel of the computer system (1100) by attaching to a system bus as described below (e.g., connected to an Ethernet interface in a PC computer system or connected to a cellular network interface in a smartphone computer system). The computer system (1100) can use any of these networks to communicate with other entities. Such communications may be one-way receive only (e.g., broadcast television), one-way send only (e.g., CANBus connected to certain CANBus devices), or bidirectional, for example, using a LAN or WAN digital network to connect to other computer systems. Certain protocols and protocol stacks may be used on each of those networks and network interfaces as described above.
[0198] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the kernel ( 1140 ) of the computer system ( 1100 ).
[0199] The kernel (1140) may include one or more central processing units (CPUs) (1141), graphics processing units (GPUs) (1142), dedicated programmable processing units in the form of field programmable gate areas (FPGAs) (1143), hardware accelerators (1144) for certain tasks, graphics adapters (1150), etc. These devices, as well as read-only memory (ROM) (1145), random access memory (1146), internal mass storage (1147) such as internal non-user accessible hard drives, SSDs, etc., may be connected via a system bus (1148). In some computer systems, the system bus (1148) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the kernel's system bus (1148) or to the kernel's system bus (1148) via a peripheral bus (1149). In one example, a screen (1110) may be connected to a graphics adapter (1150). The architecture of the peripheral bus includes PCI, USB, etc.
[0200] The CPU (1141), GPU (1142), FPGA (1143) and accelerator (1144) can execute certain instructions, which can be combined to form the computer code mentioned above. The computer code can be stored in ROM (1145) or RAM (1146). Transition data can also be stored in RAM (1146), while permanent data can be stored, for example, in internal mass storage (1147). Fast storage and retrieval to any storage device can be performed by using a cache, which can be closely associated with one or more CPUs (1141), GPUs (1142), mass storage (1147), ROM (1145), RAM (1146), etc.
[0201] The computer readable medium may have thereon computer codes for performing various computer-implemented operations. The medium and computer codes may be those specially designed and constructed for the purposes of the present disclosure, or the medium and computer codes may be of a type well known and available to those skilled in the art of computer software.
[0202] As an example, and not by way of limitation, a computer system (1100) having an architecture, and in particular a kernel (1140), can provide functionality because one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage as described above, as well as some non-temporary kernel (1140) memories, such as a kernel internal mass storage (1147) or ROM (1145). Software implementing various aspects of the present disclosure can be stored in such devices and executed by the kernel (1140). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can cause the kernel (1140), in particular the processors therein (including CPUs, GPUs, FPGAs, etc.) to perform specific processes described herein or to perform specific parts of specific processes described herein, including defining data structures stored in RAM (1146) and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system may provide functionality due to logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1144)) that may replace or operate in conjunction with software to perform specific processes described herein or specific portions of specific processes described herein. Where appropriate, references to portions of software may include logic and vice versa. Where appropriate, references to portions of computer-readable media may include circuitry (e.g., integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.
[0203] As used in this disclosure, "at least one of..." or "one of..." is intended to include any one or combination of the listed elements. For example, reference to at least one of A, B, or C, at least one of A, B, and C, at least one of A, B, and / or C, and at least one of A to C is intended to include only A, only B, only C, or any combination thereof. Reference to one of A or B, and one of A and B is intended to include A or B or (A and B). Where applicable, use of "one of..." does not exclude any combination of the listed elements, such as when the elements are not mutually exclusive.
[0204] Although the present disclosure has described several examples of various aspects, there are changes, permutations, and various replacement equivalents that fall within the scope of the present disclosure. Therefore, it should be appreciated that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and therefore fall within the spirit and scope of the present disclosure.
Claims
1. A grid decoding method, comprising: receiving a code stream, the code stream comprising basic mesh information of a basic mesh, the basic mesh comprising a subset of a plurality of vertices of a mesh in a current mesh frame; determining (i) a position prediction of a current vertex of the base mesh and (ii) a motion field prediction of the current vertex of the base mesh; (i) determining a position prediction residual of the position prediction of the current vertex based on a first quantization step value, and (ii) determining a motion field prediction residual of the motion field prediction of the current vertex based on a second quantization step value, the second quantization step value being dependent on the first quantization step value; as well as (i) reconstructing the position of the current vertex of the base mesh based on the position prediction and the position prediction residual, and (ii) reconstructing the motion field of the current vertex of the base mesh based on the motion field prediction and the motion field prediction residual.
2. The method according to claim 1, wherein: Determining the position prediction residual includes: determining a quantized position prediction residual of the position prediction of the current vertex based on a first entropy coding, and Dequantizing the quantized position prediction residual of the position prediction of the current vertex based on the first quantization step value to determine the position prediction residual; and Determining the motion field prediction residual comprises: determining a quantized motion field prediction residual of the motion field prediction of the motion field of the current vertex based on a second entropy coding, and The quantized motion field prediction residual of the motion field prediction of the current vertex is dequantized based on the second quantization step value to determine the motion field prediction residual.
3. The method according to claim 2, wherein: The second quantization step value is equal to a multiple of the first quantization step value.
4. The method according to claim 2, wherein: The second quantization step value is equal to one of a linear function and a monotonically non-decreasing function of the first quantization step value.
5. The method according to claim 2, wherein: The second quantization step value is equal to a monotonically non-decreasing function of a quantization error associated with the position prediction of the current vertex, the quantization error being a difference between the quantized position prediction residual and an unquantized position prediction residual of the current vertex.
6. The method according to claim 2, wherein: The first quantization step value is a positive binary rational number Definition: m is a positive integer between 0 and 7, and n is a non-negative integer between 1 and 256.
7. The method according to claim 2, wherein: receiving, at a sequence level, the first quantization step value and the second quantization step value, and When the syntax element in the code stream indicates that the first quantization step value is equal to the second quantization step value, the first quantization step value is equal to the second quantization step value.
8. The method according to claim 2, wherein: receiving the first quantization step value and the second quantization step value at a frame level, When the syntax element in the code stream is a first value, the second quantization step value is equal to the quantization step value of the reference frame of the current grid frame, and When the syntax element in the code stream is a second value, the second quantization step value is equal to the first quantization step value.
9. The method according to claim 1, wherein: The method further comprises: Determining a 2D texture coordinate prediction of a two-dimensional 2D texture coordinate of the current vertex of the base mesh; Determining a 2D texture coordinate prediction residual of the 2D texture coordinate prediction of the current vertex; and The 2D texture coordinates of the current vertex of the base mesh are reconstructed based on the 2D texture coordinate prediction and the 2D texture coordinate prediction residual.
10. The method according to claim 9, wherein: When the 2D texture coordinates of the current vertex are directly encoded, determining the 2D texture coordinate prediction residual comprises: determining a quantized 2D texture coordinate of the current vertex, encoding the quantized 2D texture coordinate by fixed length coding, and The quantized 2D texture coordinates of the current vertex are dequantized to obtain the 2D texture coordinates, wherein the quantized 2D texture coordinates are obtained by rounding quantization, and the rounding quantization is configured to convert the 2D texture coordinates into integers.
11. The method according to claim 9, wherein: When predicting the 2D texture coordinates of the current vertex by a stretch prediction algorithm, determining the 2D texture coordinate prediction residual comprises: determining a quantized 2D texture coordinate prediction residual of the current vertex, encoding the quantized 2D texture coordinate prediction residual by variable length coding, and The quantized 2D texture coordinate prediction residual of the current vertex is dequantized to determine the 2D texture coordinate prediction residual, wherein the quantized 2D texture coordinate prediction residual is obtained by rounding quantization, and the rounding quantization is configured to convert the 2D texture coordinate prediction residual into an integer.
12. The method according to claim 11, wherein: The rounding quantization is performed by binary rational numbers Define, m is a positive integer between 0 and 7, n is a non-negative integer between 0 and 255, and The quantized 2D texture coordinate prediction residual is obtained by rounding a multiple of the 2D texture coordinate prediction residual and a reciprocal of the binary rational number.
13. A grid coding method, comprising: determining (i) a position prediction of a current vertex of a base mesh associated with a mesh in a current mesh frame and (ii) a motion field prediction of the current vertex of the base mesh; (i) quantizing a position prediction residual of the position prediction of the current vertex based on a first quantization step value to generate a quantized position prediction residual, and (ii) quantizing a motion field prediction residual of the motion field prediction of the current vertex based on a second quantization step value to generate a quantized motion field prediction residual, wherein the second quantization step value depends on the first quantization step value; as well as Entropy encoding is performed on (i) the quantized position prediction residual of the position prediction of the current vertex and (ii) the quantized motion field prediction residual of the motion field prediction of the current vertex.
14. The method according to claim 13, wherein: The second quantization step value is equal to a multiple of the first quantization step value.
15. The method according to claim 13, wherein: The second quantization step value is equal to one of a linear function and a monotonically non-decreasing function of the first quantization step value.
16. The method according to claim 13, wherein: The method further comprises: Determining a 2D texture coordinate prediction of a two-dimensional 2D texture coordinate of the current vertex of the base mesh; quantizing the 2D texture coordinate prediction residual of the 2D texture coordinate prediction of the current vertex to generate a quantized 2D texture coordinate prediction residual of the current vertex; and The quantized 2D texture coordinate prediction residual of the current vertex is entropy encoded.
17. The method of claim 16, wherein: When the 2D texture coordinates of the current vertex are directly encoded, quantizing the 2D texture coordinate prediction residual includes: quantizing the 2D texture coordinates of the current vertex based on rounded quantization, the rounded quantization configured to convert the 2D texture coordinates into integers; and The quantized 2D texture coordinates are encoded based on fixed length coding.
18. The method of claim 16, wherein: When predicting the 2D texture coordinates of the current vertex by a stretch prediction algorithm, determining the 2D texture coordinate prediction residual comprises: The multiple of the 2D texture coordinate prediction residual of the current vertex and the binary rational number quantize the 2D texture coordinate prediction residual by rounding the reciprocal of , m is a positive integer between 0 and 7, and n is a non-negative integer between 0 and 255; and The quantized 2D texture coordinate prediction residual is encoded based on variable length coding.
19. The method according to claim 18, wherein: Quantizing the 2D texture coordinate prediction residual further includes: adding an offset to the 2D texture coordinate prediction residual to generate an updated 2D texture coordinate prediction residual; and The multiple of the updated 2D texture coordinate prediction residual for the current vertex and the dyadic rational number The updated 2D texture coordinate prediction residual is quantized by rounding the reciprocal of .
20. A method for processing dynamic grid data, the method comprising: The code stream of the dynamic grid data is processed according to the format rule, wherein: The code stream includes base mesh information of a base mesh, the base mesh including a subset of a plurality of vertices of a mesh in a current mesh frame; and The format rules specify: determining (i) a position prediction of a current vertex of the base mesh and (ii) a motion field prediction of the current vertex of the base mesh; (i) determining a position prediction residual of the position prediction of the current vertex based on a first quantization step value, and (ii) determining a motion field prediction residual of the motion field prediction of the current vertex based on a second quantization step value, the second quantization step value being dependent on the first quantization step value; and (i) processing the position of the current vertex of the base mesh based on the position prediction and the position prediction residual, and (ii) processing the motion field of the current vertex of the base mesh based on the motion field prediction and the motion field prediction residual.