Uniform bit depth scaling

By using uniform bit depth scaling technology in 3D grid data processing, the vertex positions in the grid data are quantized, which solves the problems of large data volume and low transmission efficiency in the existing technology, and realizes efficient data compression and transmission.

CN120019660APending Publication Date: 2025-05-16TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004330.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-07-10
Filing Date
2024-07-12
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process and compress complex 3D grid data, especially during storage and transmission, resulting in large amounts of data and low transmission efficiency.

Method used

The uniform bit depth scaling technology is used to quantify the vertex positions in the grid data, and the positions of multiple vertices are converted into integers through the bit depth scaling function to realize data compression.

Benefits of technology

Through uniform bit depth scaling technology, the storage and transmission requirements of grid data are significantly reduced, data compression efficiency is improved, and data quality and accuracy are maintained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019660A_ABST
    Figure CN120019660A_ABST
Patent Text Reader

Abstract

And receiving a code stream, wherein the code stream comprises the basic grid information of the basic grid. The base grid includes a subset of the plurality of vertices of the grid in the current grid frame. A position of a current vertex of the base grid is determined based on a quantized position of the current vertex of the base grid, the quantized position of the current vertex being generated according to a bit depth scaling function. The bit depth scaling function is configured to convert a first subset of the positions of the plurality of vertices of the base grid to a first integer and to convert a second subset of the positions of the plurality of vertices of the base grid to a second integer. The total number of positions of the first subset is equal to the total number of positions of the second subset. The current vertex is reconstructed based on the determined position of the current vertex of the base grid.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Join by reference

[0002] This application claims the benefit of priority to U.S. Patent Application No. 18 / 769,315, filed on July 10, 2024, entitled “Uniform Bitdepth Scaling,” which claims the benefit of priority to U.S. Provisional Application No. 63 / 526,884, filed on July 14, 2023, entitled “Uniform Bitdepth Scaling.” The entire disclosures of these prior applications are incorporated herein by reference in their entirety. Technical Field

[0003] This disclosure describes aspects generally related to trellis coding. Background Art

[0004] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent that the work of the presently named inventors described in this background section and in various aspects of this specification was performed, it does not indicate that it qualifies as prior art at the time of filing, and it is never explicitly or implicitly admitted that it is prior art to the present disclosure.

[0005] Image / video compression can help transmit image / video data between different devices, storages, and networks with minimal quality degradation. In some examples, video codec techniques can compress video based on spatial redundancy and temporal redundancy. In one example, a video codec can use a technique called intra-frame prediction, which can compress an image based on spatial redundancy. For example, intra-frame prediction can use reference data from a current picture being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-frame prediction, which can compress an image based on temporal redundancy. For example, inter-frame prediction can predict samples in a current picture based on a picture previously reconstructed using motion compensation. Motion compensation can be indicated by a motion vector (MV).

[0006] Advances in three-dimensional (3D) capture, modeling, and rendering have facilitated 3D content across a variety of platforms and devices. For example, a baby's first steps can be filmed on one continent, and grandparents can see (and in some cases, interact with) and enjoy a fully immersive experience with the child on another continent. To achieve this sense of realism, models have become more complex, and a large amount of data is associated with the creation and use of these models. 3D meshes are widely used to represent such immersive content. Summary of the invention

[0007] Aspects of the present disclosure include code streams, methods, and apparatus for mesh processing.In some examples, the apparatus for mesh processing includes processing circuitry.

[0008] According to one aspect of the present disclosure, a mesh decoding method is provided. In the method, a code stream is received, the code stream including base mesh information of a base mesh. The base mesh includes a subset of multiple vertices of the mesh in a current mesh frame. The position of the current vertex of the base mesh is determined based on the quantized position of the current vertex of the base mesh, and the quantized position of the current vertex is generated according to a bit depth scaling function. The bit depth scaling function is configured to convert a first subset of the positions of the multiple vertices of the base mesh into a first integer, and convert a second subset of the positions of the multiple vertices of the base mesh into a second integer. The total number of positions of the first subset is equal to the total number of positions of the second subset. The current vertex is reconstructed based on the determined position of the current vertex of the base mesh.

[0009] In one aspect, the bit depth scaling function is based on the position bit depth m and the encoding bit depth n. The position bit depth m indicates that the position of the current vertex is between 0 and 2. m -1, the encoding bit depth n indicates that the position of the current vertex is quantized to a value between 0 and 2 n -1.

[0010] In one aspect, the bit depth scaling function is based on a value equal to the position bit depth minus the encoding bit depth.

[0011] In one aspect, the bit depth scaling function is based on an exponential term with a base of 2 and an exponent equal to the position bit depth minus the encoding bit depth.

[0012] In one aspect, the bit depth scaling function is configured to round the position of the current vertex to an integer.

[0013] In one aspect, the position bit depth is greater than the encoding bit depth.

[0014] In one aspect, the position bit depth is equal to or less than the encoding bit depth.

[0015] In one aspect, the bit depth scaling function is defined as Where x is the position of the current vertex, n is the encoding bit depth, and m is the position bit depth.

[0016] According to another aspect of the present disclosure, a mesh encoding method is provided. In the method, the position of a current vertex of a base mesh is quantized based on a bit depth scaling function to determine a quantized position of the current vertex. The base mesh includes a subset of multiple vertices of the mesh in the current mesh frame. The bit depth scaling function is configured to quantize a first subset of the positions of the multiple vertices of the base mesh into a first integer, and quantize a second subset of the positions of the multiple vertices of the base mesh into a second integer. The total number of positions of the first subset is equal to the total number of positions of the second subset. A position prediction of the quantized position of the current vertex is determined. A position prediction residual of the position prediction of the current vertex is encoded in a bitstream.

[0017] In one aspect, the bit depth scaling function is based on the position bit depth m and the encoding bit depth n. The position bit depth m indicates that the position of the current vertex is between 0 and 2. m -1, the encoding bit depth n indicates that the position of the current vertex is quantized to a value between 0 and 2 n -1.

[0018] In one aspect, the bit depth scaling function is based on a value equal to the position bit depth minus the encoding bit depth.

[0019] In one aspect, the bit depth scaling function is based on an exponential term with a base of 2 and an exponent equal to the position bit depth minus the encoding bit depth.

[0020] In one aspect, the bit depth scaling function is configured to round the position of the current vertex to an integer.

[0021] In one aspect, the position bit depth is greater than the encoding bit depth.

[0022] In one aspect, the position bit depth is equal to or less than the encoding bit depth.

[0023] In one aspect, the bit depth scaling function is defined as: Where x is the position of the current vertex, n is the encoding bit depth, and m is the position bit depth.

[0024] According to another aspect of the present disclosure, a method for processing mesh data is provided. In the method, a code stream of mesh data is processed according to a format rule. The code stream includes base mesh information of a base mesh. The base mesh includes a subset of multiple vertices of the mesh in a current mesh frame. The format rule specifies: the position of the current vertex of the base mesh is determined based on the quantized position of the current vertex of the base mesh, and the quantized position of the current vertex is generated according to a bit depth scaling function. The bit depth scaling function is configured to convert a first subset of the positions of the multiple vertices of the base mesh into a first integer, and convert a second subset of the positions of the multiple vertices of the base mesh into a second integer. The total number of positions of the first subset is equal to the total number of positions of the second subset. The format rule specifies: the current vertex is processed based on the determined position of the current vertex of the base mesh.

[0025] Aspects of the present disclosure also provide an apparatus for trellis coding. The apparatus for trellis coding comprises a processing circuit configured to implement any of the described methods for trellis coding.

[0026] Aspects of the present disclosure also provide an apparatus for trellis decoding. The apparatus for trellis decoding comprises a processing circuit configured to implement any of the described methods for trellis decoding.

[0027] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions which, when executed by a computer, cause the computer to perform any of the described methods for grid decoding, grid encoding, and grid data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Further features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0029] Figure 1 is a schematic illustration of an example of a block diagram of a communication system (100).

[0030] Figure 2 is a schematic illustration of an example of a block diagram of a decoder.

[0031] Figure 3 is a schematic illustration of an example of a block diagram of an encoder.

[0032] Figure 4 is a schematic illustration of an example of an encoding process (400) for mesh processing according to an aspect of the present disclosure.

[0033] Figure 5 is a schematic illustration of an example of a pre-processing step (500) according to an aspect of the present disclosure.

[0034] Figure 6 is a schematic illustration of a decoding process (600) for trellis processing according to an aspect of the present disclosure.

[0035] Figure 7 is a schematic illustration of an example of base grid encoding according to an aspect of the present disclosure.

[0036] Figure 8 A flow chart outlining a trellis decoding process according to some aspects of the present disclosure is shown.

[0037] Fig. 9 A flow chart outlining a trellis encoding process according to some aspects of the present disclosure is shown.

[0038] Fig.10 is a schematic illustration of a computer system according to an aspect of the present disclosure. DETAILED DESCRIPTION

[0039] Figure 1 A block diagram of a video processing system (100) in some examples is shown. The video processing system (100) is an example of an application of the disclosed subject matter, namely a video encoder and a video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., and the like.

[0040] The video processing system (100) includes an acquisition subsystem (113), which may include a video source (101). The video source (101) may include one or more images acquired by a camera and / or generated by a computer. For example, a digital camera creates, for example, an uncompressed video picture stream (102). In one example, the video picture stream (102) includes samples taken by the digital camera. The video picture stream (102), which is depicted as a thick line to emphasize the high amount of data compared to the encoded video data (104) (or the encoded video bitstream), can be processed by an electronic device (120), which includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. The encoded video data (104) (or encoded video bitstream), depicted as thin lines to emphasize the lower data volume compared to the video picture stream (102), can be stored on the streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1A client subsystem (106) and a client subsystem (108) in a streaming server (105) may access a copy (107) and a copy (109) of the encoded video data (104). The client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and creates an output video picture stream (111) that can be presented on a display (112) (e.g., a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), the video data (107), and the video data (109) (e.g., a video bitstream) may be encoded according to certain video encoding / compression standards. Examples of these standards include ITU-T H.265. In one example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of the VVC standard.

[0041] It should be noted that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may also include a video encoder (not shown).

[0042] Figure 2 An example of a block diagram of a video decoder (210) is shown. The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used to replace Figure 1 A video decoder (110) in an example of FIG.

[0043] The receiver (231) may receive one or more encoded video sequences, for example, included in a bitstream to be decoded by the video decoder (210). In one aspect, the encoded video sequences are received one at a time, wherein the decoding of each encoded video sequence is independent of the decoding of the other encoded video sequences. The encoded video sequence may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data and other data, for example, encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not depicted). The receiver (231) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory (215) may be located external to the video decoder (210) (not depicted). In still other applications, a buffer memory (not depicted) may be provided external to the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) may be provided internal to the video decoder (210) to, for example, handle playout timing. When the receiver (231) receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (215) may not be required, or the buffer memory (215) may be made smaller. For use on a traffic packet network such as the Internet, the buffer memory (215) may be required, and the buffer memory (215) may be relatively large, advantageously may have an adaptive size, and may be implemented at least partially in an operating system or similar element (not depicted) external to the video decoder (210).

[0044] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. The categories of symbols include information for managing the operation of the video decoder (210) and potential information for controlling a display device such as a display device (212) (e.g., a display screen) that is not part of the electronic device (230) but can be coupled to the electronic device (230), such as Figure 2As shown. The control information for the display device may be a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not depicted). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may be performed according to a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract a subgroup parameter set for at least one subgroup of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (220) may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the coded video sequence.

[0045] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).

[0046] Depending on the type of coded video picture or part of coded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of symbol (221) may involve multiple different units. Which units are involved and how they are involved may be controlled by subgroup control information parsed by parser (220) from the coded video sequence. For clarity, such subgroup control information flow between parser (220) and the multiple units below is not depicted.

[0047] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into a number of functional units as described below. In actual implementations operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purposes of describing the disclosed subject matter, the conceptual subdivision into the following multiple functional units is appropriate.

[0048] The first unit is a sealer / inverse transform unit (251). The sealer / inverse transform unit (251) receives quantized transform coefficients as symbols (221) from the parser (220) and control information including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit (251) may output a block including sample values, which may be input into an aggregator (255).

[0049] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses surrounding reconstructed information extracted from a current picture buffer (258) to generate a block of the same size and shape as the block being reconstructed. For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a per-sample basis.

[0050] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to a block that is inter-coded and potentially motion compensated. In this case, the motion compensated prediction unit (253) may access the reference picture memory (257) to extract samples for prediction. After the extracted samples are motion compensated according to the symbols (221) belonging to the block, these samples may be added to the output of the scaler / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) by the aggregator (255), thereby generating output sample information. The extraction of prediction samples by the motion compensated prediction unit (253) from the address in the reference picture memory (257) may be controlled by a motion vector, which may be provided to the motion compensated prediction unit (253) in the form of symbols (221), which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values ​​extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.

[0051] The output samples of the aggregator (255) may be employed by various loop filtering techniques in a loop filter unit (256). The video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video sequence (also referred to as a coded video bitstream) and available to the loop filter unit (256) as symbols (221) from the parser (220). The video compression may also be responsive to meta-information obtained during decoding of a coded picture or a previous (in decoding order) portion of the coded video sequence, and to previously reconstructed and loop filtered sample values.

[0052] The output of the loop filter unit (256) may be a sample stream that may be output to a display device (212) and stored in a reference picture memory (257) for future inter-picture prediction.

[0053] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed, and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) may become part of the reference picture memory (257), and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.

[0054] The video decoder (210) may perform decoding operations according to a predetermined video compression technology or standard such as ITU-T H.265. The encoded video sequence may conform to the syntax specified by the video compression technology or standard used in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as the only tools available for use under the profile. For compliance, it is also required that the complexity of the encoded video sequence is within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata of the HRD buffer management signaled in the encoded video sequence.

[0055] In one aspect, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0056] Figure 3 An example of a block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) may be used to replace Figure 1 A video encoder (103) in an example.

[0057] The video encoder (303) can be used to obtain the video source (301) (not Figure 3 In an example, the video source (301) is a part of the electronic device (320) to receive video samples, and the video source (301) can capture video images to be encoded by the video encoder (303). In another example, the video source (301) is a part of the electronic device (320).

[0058] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), the digital video sample stream may have any suitable bit depth (e.g. 8-bit, 10-bit, 12-bit, ...), any color space (e.g. BT.601 Y CrCB, RGB, ...) and any suitable sampling structure (e.g. Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The pictures themselves may be organized into a spatial pixel array, where each pixel may include one or more samples, depending on the sampling structure, color space, etc. used. The following description focuses on the samples.

[0059] According to one aspect, the video encoder (303) can encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other time constraints required. Implementing the appropriate encoding speed is a function of the controller (350). In some aspects, the controller (350) controls other functional units as described below and is functionally coupled to these other functional units. For clarity, the coupling is not depicted in the figure. The parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization techniques, etc.), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions related to the video encoder (303) optimized for a certain system design.

[0060] In some aspects, the video encoder (303) is configured to operate in a coding loop. As a simple description, in one example, the coding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Since the decoding of the symbol stream produces a bit-accurate result that is independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-accurately corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values ​​that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, such as due to channel errors) is also used in some related technologies.

[0061] The operation of the "local" decoder (333) may be similar to that described above in conjunction with Figure 2 The "remote" decoder of the video decoder (210) described in detail is identical. However, additional brief reference is made to Figure 2 , since the symbols are available and the entropy encoder (345) and the parser (220) can losslessly encode / decode the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (210) including the buffer memory (215) and the parser (220) may not be fully implemented in the local decoder (333).

[0062] On the one hand, except for the parsing / entropy decoding present in the decoder, the decoder technology is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on the decoder operation. The description of the encoder technology can be simplified because the encoder technology is mutually inverse to the decoder technology described comprehensively. In some areas, a more detailed description is provided below.

[0063] During operation, in some examples, the source encoder (330) may perform motion compensated predictive coding that predictively encodes an input picture by referencing one or more previously encoded pictures from a video sequence designated as "reference pictures." In this manner, the encoding engine (332) encodes the differences between pixel blocks of the input picture and pixel blocks of a reference picture that may be selected as a prediction reference for the input picture.

[0064] The local video decoder (333) may decode the encoded video data of the picture that may be designated as the reference picture based on the symbol created by the source encoder (330). The operation of the encoding engine (332) may advantageously be a lossy process. When the encoded video data may be decoded at the video decoder ( Figure 3 When the video sequence is decoded at a remote video decoder (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture memory (334). In this way, the video encoder (303) may store a copy of the reconstructed reference picture locally that has common content (absent transmission errors) with the reconstructed reference picture to be obtained by the remote video decoder.

[0065] The predictor (335) may perform a prediction search for the encoding engine (332). That is, for a new picture (or grid) to be encoded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may be used as appropriate prediction references for the new picture. The predictor (335) may operate on a pixel block by pixel block basis to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (335), it may be determined that the input picture may have prediction references taken from a plurality of reference pictures stored in the reference picture memory (334).

[0066] In one example, the mesh position quantization step size is defined based on a first parameter (e.g., mesh_position_quantization_step_size_log2_denominator) and a second parameter (e.g., mesh_position_quantization_step_size_numerator_minus1). In one example, mesh_position_quantization_step_size_log2_denominator is the logarithm (log2) of the denominator of the position quantization step size to the base 2, and its value is between 0 and 7 (including 0 and 7). In one example, mesh_position_quantization_step_size_numerator_minus1 plus 1 is the value of the numerator of the position quantization step size, and the value is between 1 and 256 (including 1 and 256). After the position encoding in the MEB, the reconstructed position value used for reference is scaled to the dynamic range of the inter-frame.

[0067] An example of the grid attribute encoding parameter syntax is as follows:

[0068]

[0069]

[0070] An example of the semantics of mesh attribute encoding parameters is defined by a parameter such as mesh_attribute_quantization_step_size_log2_denominator[i]. In one example, mesh_attribute_quantization_step_size_log2_denominator[i] is the log2 of the denominator of the i-th attribute quantization step size, and its value is between 0 and 7 (inclusive).

[0071] An example of a generic base grid sequence parameter set RBSP syntax is as follows:

[0072]

[0073]

[0074] An example of motion field quantization is defined by a first parameter (e.g., bmsps_inter_quantization_step_size_log2_denominator) and a second parameter (e.g., bmsps_inter_quantization_step_size_numerator_minus1). In one example, bmsps_inter_quantization_step_size_log2_denominator is the log2 of the denominator of the motion field quantization step size, and its value is between 0 and 7 (including 0 and 7). In one example, bmsps_inter_quantization_step_size_numerator_minus1 plus 1 is the value of the numerator of the motion field quantization step size, and the value is between 1 and 256 (including 1 and 256). After motion field encoding, the position value is reconstructed and can be used as a reference for future vertices. In one example, mesh attribute quantization is defined by a step size parameter (e.g., mesh_attribute_quantization_step_size_numerator[i]). In one example, mesh_attribute_quantization_step_size_numerator[i] is the value of the numerator of the i-th attribute quantization step size, which is between 0 and 255 (including 0 and 255). At the same time, the quantized prediction residual is dequantized. The reconstructed value is calculated by adding the dequantized prediction residual and the predicted value. The reconstructed value can be used as a reference for future vertices.

[0075] The controller (350) may manage encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.

[0076] The outputs of all the above functional units may be entropy encoded in an entropy encoder (345). The entropy encoder (345) applies lossless compression to the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into a coded video sequence.

[0077] The transmitter (340) may buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission over a communication channel (360), which may be a hardware / software link to a storage device that may store the encoded video data. The transmitter (340) may combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or ancillary data streams (source not shown).

[0078] The controller (350) may manage the operation of the video encoder (303). During encoding, the controller (350) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following picture types:

[0079] An intra picture (I picture) that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures.

[0080] Predictive pictures (P pictures), which can be encoded and decoded using intra prediction or inter prediction, which uses a motion vector and reference index to predict the sample values ​​of each block.

[0081] Bidirectional predictive pictures (B pictures), which can be encoded and decoded using intra prediction or inter prediction, which uses two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.

[0082] The source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined according to the coding allocation applied to the block's corresponding picture. For example, a block of an I picture may be non-predictively coded, or the block may be predictively coded (spatial prediction or intra prediction) with reference to an already coded block of the same picture. A block of a P picture may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. A block of a B picture may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.

[0083] The video encoder (303) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0084] In one aspect, the transmitter (340) may transmit additional data when transmitting the encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0085] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In one example, a particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.

[0086] In some aspects, a bidirectional prediction technique may be used for inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture that precede the current picture in the video in decoding order (but may be in the past and future, respectively, in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.

[0087] In addition, merge mode technology can be used for inter-picture prediction to improve coding efficiency.

[0088] According to some aspects of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks (e.g., polygonal blocks or triangle blocks). For example, according to the High-Efficiency Video Coding (HEVC) standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU includes three coding tree blocks (CTBs), which are a luminance CTB and two chrominance CTBs. Each CTU can be recursively divided into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be divided into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type for the CU, such as an inter prediction type or an intra prediction type. The CU is divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Typically, each PU includes a luma prediction block (PB) and two chroma PBs. On the one hand, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Taking the luma prediction block as an example, the prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0089] It should be noted that the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using any suitable technology. In one aspect, the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using one or more integrated circuits. In another aspect, the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using one or more processors executing software instructions.

[0090] Aspects of the present disclosure include techniques for uniform bit-depth scaling in mesh compression.

[0091] A mesh may include multiple polygons that describe the surface of a volumetric object. Each polygon of a mesh may be defined by the vertices of the corresponding polygon in a three-dimensional (3D) space and information about how the vertices are connected (which may be referred to as connectivity information). In some aspects, vertex attributes (e.g., color, normal, etc.) may be associated with vertices (or mesh vertices). Attributes (or vertex attributes) may also be associated with the surface of a mesh by utilizing mapping information that uses a two-dimensional (2D) attribute map to parameterize the mesh. This mapping may be described by a set of parameter coordinates (referred to as UV coordinates or texture coordinates) associated with mesh vertices. A 2D attribute map may be used to store high-resolution attribute information, such as textures, normals, displacements, etc. High-resolution attribute information may be used for various purposes, such as texture mapping and shading.

[0092] Dynamic mesh sequences may require a large amount of data because dynamic meshes may include a large amount of information that changes over time. Therefore, efficient compression techniques can be used to store and transmit such content. Mesh compression standards (e.g., Information and Communication (IC) mesh compression, MESHGRID, frame-based animated mesh compression (FAMC)) were previously developed by the Moving Picture Experts Group (MPEG) to handle dynamic meshes with constant connectivity, time-varying geometry and vertex attributes. However, these standards do not consider time-varying attribute graphs and connectivity information. DCC (Digital Content Creation) tools can generate such dynamic meshes. However, for volume acquisition technology, it may be challenging to generate dynamic meshes with constant connectivity, especially to generate dynamic meshes with constant connectivity under real-time constraints. Existing standards may not support this type of content (e.g., dynamic meshes with constant connectivity). The present disclosure includes aspects of a new mesh compression standard that can directly handle dynamic meshes with time-varying connectivity information and optional time-varying attribute graphs. Mesh compression can target lossy and lossless compression for various applications, such as real-time communication, storage, free viewpoint video, augmented reality (AR) and virtual reality (VR). Features such as random access and scalable / progressive coding can also be considered.

[0093] The mesh geometry information may include vertex connectivity information, 3D coordinates, 2D texture coordinates, etc. The 3D vertex coordinates and 2D texture coordinates occupy a large portion of the mesh geometry information. Therefore, it is necessary to compress the 3D vertex coordinates (also called vertex positions) and 2D texture coordinates to reduce the amount of data required to store and / or transmit the mesh geometry information.

[0094] Figure 4 An example of an encoding process (400) for mesh processing based on a related video codec (e.g., MPEG V-MeshTMv1.0) according to an aspect of the present disclosure is shown. Figure 4 As shown, the encoding process (400) may include a preprocessing step (400A) and an encoding step (400B). The preprocessing step (400A) may be configured to generate a base grid m(i) of the current frame and a displacement field d(i) of the current frame based on an input grid M(i) of the current frame, wherein the displacement field d(i) of the current frame includes a displacement vector. The encoding step (400B) may be configured to encode the base grid m(i), the displacement field d(i), and the texture information of the base grid m(i). The displacement field d(i) of the current frame may include a displacement vector. The index i may refer to the current frame. In one aspect, a mode decision method may be performed in the encoding process (400) to determine whether to apply inter-frame coding (also referred to as inter-frame prediction or inter-frame mode), intra-frame coding (also referred to as intra-frame prediction or intra-frame mode), etc. to the current frame. For example, the mode decision method may compare the cost of the intra-frame mode with the cost of the inter-frame mode, and determine the encoding mode of the base grid m(i) of the current frame based on which cost is smaller. In some examples, the base grid m(i) is encoded (e.g., encoded or decoded) using a skip mode. In one example, the skip mode is a special mode of the inter-frame mode. For example, the base grid m(i) can be intra-coded, or inter-coded, or encoded using a skip (SKIP) mode.

[0095] Still reference Figure 4, the preprocessing step (400A) may include a mesh extraction process (402), a parameterization process such as an atlas parameterization process (404), and a subdivision surface fitting process (406). The mesh extraction process (402) is configured to downsample the vertices of the input mesh M(i) to generate an extracted mesh dm(i) that may include a plurality of extracted (or downsampled) vertices. The number of the plurality of extracted vertices is less than the number of vertices of the input mesh M(i). The parameterization process such as the atlas parameterization process (404) is configured to map the extracted mesh dm(i) onto a planar domain, such as onto a UV atlas (or UV map), to generate a re-parameterized mesh pm(i). In one example, atlas parameterization can be performed based on a video processing tool (e.g., a UVAtlas tool). The subdivision surface fitting process (406) is configured to take the re-parameterized mesh pm(i) and the input mesh M(i) as input, and generate a base mesh m(i) and a displacement field d(i) including a displacement vector or a displacement set. In an example of the subdivision surface fitting process (406), pm(i) is subdivided using a subdivision scheme such as iterative interpolation to obtain a subdivided mesh. The iterative interpolation includes inserting a new point in the middle of each edge of the re-parameterized mesh pm(i) at each iteration. Any suitable subdivision scheme can be applied to subdivide pm(i). The displacement field d(i) is calculated by determining the nearest point on the surface of the input mesh M(i) for each vertex of the subdivided mesh.

[0096] Advantages of the subdivided mesh may include: the subdivided mesh has a subdivision structure that allows efficient compression while providing a reliable approximation to the input mesh. The improved compression efficiency may be obtained due to the following properties. The decimated mesh dm(i) may have a small number of vertices and may be encoded and transmitted using fewer bits than the input mesh M(i) or the number of subdivided meshes. Figure 4 , a base mesh m(i) can be generated based on the extracted mesh dm(i). In one example, the base mesh m(i) is the extracted mesh dm(i). Since the subdivided mesh can be generated based on the subdivision method, the subdivided mesh can be automatically generated by the decoder when the base mesh or the extracted mesh is decoded (for example, without using any information other than the subdivision scheme and the subdivision iteration count). On the decoder side, the displacement field d(i) can be generated by decoding the displacement vectors associated with the vertices of the subdivided mesh. In addition to enabling spatial / quality scalability, the subdivision structure enables efficient transformations such as wavelet decomposition, which can provide high compression performance.

[0097] For simplicity, the preprocessing step (400A) that can be applied to an input mesh (e.g., a 3D mesh) can be described using the preprocessing step (500) applied to a two-dimensional (2D) curve. The preprocessing step (400A) is similar to the preprocessing step (500), except that the 3D mesh can be replaced by a 2D curve.

[0098] Figure 5 An example of a preprocessing step (500) according to an aspect of the present disclosure is shown. Figure 5 As shown, an input 2D curve (represented by a 2D polyline) (502) can be downsampled to generate a base curve such as a polyline, which is referred to as a "decimated" curve (504). Then, a subdivision scheme can be applied to the decimated polyline (504) to generate a "decision" curve (506). In one example, the subdivision scheme can be an iterative interpolation scheme. The iterative interpolation scheme can include inserting a new point in the middle of each edge of the polyline (or decimated curve) (504) at each iteration. For example, a point (510) can be inserted in an edge (508) of the decimated curve (504). In one example, the edge (508) is located between a point (512) and a point (514). In addition, a point (522) can be added between the point (512) and the point (510), and a point (516) can be added between the point (510) and the point (514). Then, the decimal polyline (506) is deformed to generate a displacement curve (518). The displacement curve (518) may be a better approximation of the input curve (502) than the subdivided curve (506). For example, a displacement vector (e.g., (520)) is calculated for each vertex (e.g., (510)) of the subdivided curve (506) so that the shape of the displacement curve (518) is as close as possible to the shape of the input curve (502). An advantage of the subdivided curve (506) is that the subdivided curve (506) has a subdivision structure that allows for more efficient compression while providing a reliable approximation of the input curve (502).

[0099] The decimation curve (504) may have a smaller number of points and may be encoded and transmitted using a limited number of bits. Since the subdivision curve may be generated based on the subdivision scheme, the subdivision curve may be automatically generated by the decoder when decoding the base curve or the decimation curve (e.g., without using any information other than the subdivision method and the subdivision iteration count). The displacement curve may be generated by decoding the displacement vectors associated with the vertices of the subdivision curve. In addition to enabling spatial / quality scalability, the subdivision structure enables efficient transforms such as wavelet decomposition, which may provide high compression performance.

[0100] Still reference Figure 5In one example, the input mesh M(i) may include an input 2D curve (502). The base mesh m(i) may include a decimated curve (504), which is formed by downsampling the vertices of the input 2D curve (502). The displacement field dm(i) may include a plurality of displacement vectors, such as Figure 5 The displacement vector (520) is shown.

[0101] The encoding step (400B) may include base mesh encoding (408), displacement encoding (410), texture encoding (412), etc. The base mesh encoding (408) is configured to encode geometric information of a base mesh m(i) associated with a current frame. In intra-frame coding, the base mesh m(i) may first be quantized (e.g., quantized using uniform quantization) and then encoded, for example, by using a coding mode determined by a mode decision method. The coding mode may be an inter-frame mode, an intra-frame mode, a skip mode, etc. An encoder for intra-coding the base mesh m(i) may be referred to as a static mesh encoder. In inter-frame coding, a reference base mesh associated with a reference frame indicated by index j (e.g., a reconstructed quantized reference base mesh m'(j)) may be used to predict a base mesh m(i) associated with a current frame indicated by index i. The displacement encoding (410) is configured to encode a displacement field d(i) generated in the preprocessing step (400A). The displacement field d(i) may include a set of displacement vectors (or displacements) associated with vertices of a subdivided mesh. The texture encoding (412) is configured to encode attribute information of the base mesh m(i). The attribute information may include texture, normal, color, etc. The attribute information may be encoded based on a suitable codec (e.g., High Efficiency Video Coding (HEVC) or Next Generation Video Coding (VVC)).

[0102] On the one hand, reference Figure 4 , a mesh encoding process such as the encoding process (400) begins with preprocessing (e.g., preprocessing step (400A)). The preprocessing may convert an input mesh (e.g., an input dynamic mesh) M(i) into a base mesh m(i) and a displacement field d(i) including a displacement set (or a displacement vector set). The encoding step (400B) may compress the output from the preprocessing (e.g., m(i), d(i), etc.) and generate a compressed code stream b(i). The compressed code stream b(i) may include a compressed base mesh code stream, a compressed displacement field code stream, a compressed attribute code stream, etc.

[0103] Figure 6An example of a decoding process (600) for grid processing according to an aspect of the present disclosure is shown. The decoding process (600) may include a decoding step (605) and a post-processing step (610). A compressed code stream b(i) may be provided to the decoding step (605). In one example, for example for lossless transmission, the compressed code stream b(i) is the output b(i) from the encoding process (400). The decoding step (605) may extract various sub-code streams, such as a compressed base grid sub-stream, a compressed displacement field sub-stream, a compressed attribute sub-stream, etc. The decoding step (605) may decompress the sub-code streams to generate the following components: block (patch) metadata indicated by metadata (i) (metadata(i)), decoded base grid m" (i), decoded displacement field (including displacement) d" (i), decoded attribute map A" (i), etc.

[0104] On the one hand, the base grid substream may be provided to a grid decoder to generate a reconstructed quantized base grid m'(i). The decoded base grid (or reconstructed base grid) m"(i) may be obtained by applying inverse quantization to m'(i). The displacement field substream including encoded, packed and quantized wavelet coefficients may be decoded by a video and / or image decoder. Image unpacking and inverse quantization may be applied to the reconstructed, packed quantized wavelet coefficients to obtain unpacked and dequantized transform coefficients (e.g., wavelet coefficients). An inverse wavelet transform may be applied to the unpacked and dequantized wavelet coefficients to generate a decoded displacement field (or reconstructed displacement) d"(i).

[0105] The decoded components (e.g., including metadata(i), m"(i), d"(i), A"(i), etc.) can be provided to a post-processing step (610). A mesh (also referred to as a decoded / reconstructed mesh) M"(i) can be generated based on m"(i) and d"(i) by the post-processing step (610). In one example, the mesh M"(i) (also referred to as a reconstructed deformed mesh DM(i)) can be obtained by subdividing m"(i) using a subdivision scheme and applying reconstructed displacements d"(i) to the vertices of the subdivided mesh. In one example, DM(i) can include a displacement curve (518). In one example, when the encoding process (400), the decoding process (600), and the transmission are lossless, the mesh M"(i) can be the same as the input mesh M(i). When one of the encoding process (400), the decoding process (600), and the transmission is lossy, M"(i) is different from M(i). In various examples, the difference between M"(i) and M(i), if any, can be relatively small. In one example, a property graph A”(i) is also generated by a post-processing step (610).

[0106] In one aspect, the base grid may be intra-coded, inter-coded, or encoded using a SKIP mode, etc. In one example, the SKIP mode may be a special mode of the inter-mode, in which the base grid m(i) of the current frame indicated by index i is the same as the base grid m(j) of the reference frame indicated by index (also referred to as frame index) j. When the inter-mode is applied to encode the base grid in the current frame, the encoder may generate a predicted base grid for the current frame based on the reconstructed base grid of the reference frame. In one example, such as in MPEG V-DMC WD 2.0, the reference frame is a frame immediately preceding the current frame in display order. The frame index i of the current frame indicates the display order. When the frame index of the current frame is i, the frame index of the reference frame is (i-1). In one example, the current frame and the reference frame are in the same group of frames (GoF).

[0107] In one aspect, for example in MPEG V-DMC WD 2.0, the mesh encoding process starts with preprocessing. The preprocessing may convert the input dynamic mesh (denoted as M(i)) into a base mesh m(i) and a set of displacements d(i). The encoder may compress the base mesh m(i) and the displacements d(i) to generate a compressed code stream b(i).

[0108] like Figure 4 As shown, preprocessing may include mesh extraction, atlas parameterization, and subdivision surface fitting. Mesh extraction may use simplification techniques to extract the input mesh M(i) and generate an extracted mesh dm(i). The extracted mesh dm(i) may be re-parameterized. The generated mesh may be denoted as pm(i). Subdivision surface fitting may take the re-parameterized mesh pm(i) and the input mesh M(i) as input to generate a base mesh m(i) and a displacement set d(i).

[0109] For intra frames, the base mesh can be encoded using a static mesh codec, where the positions are encoded using spatial prediction. For inter frames, the base mesh can be encoded using motion field coding. The motion field of the base mesh can be defined as the difference between the vertex positions of the current frame and the vertex positions of the reference frame of the current frame, and motion field coding can utilize temporal prediction.

[0110] In the present disclosure, methods and systems are provided for uniform bit depth scaling in mesh compression. In one example, such as in MPEG V-DMC WD 2.0, vertex position coding can utilize a workflow including bit depth quantization, prediction, and entropy coding. An example of the workflow is described in Figure 7 In bit depth quantization, the position of the vertex can be quantized into a bit depth value specified by the encoder. For example, if the encoded bit depth value is 10, the vertex at the position is quantized into an integer between 0 and 1023.

[0111] like Figure 7 As shown, the workflow (700) includes bit depth quantization (702), prediction (704), and entropy encoding (706). In bit depth quantization (702), the position (or position coordinate) of the vertex in the base mesh and the 2D texture coordinate are quantized into bit depth values. The bit depth value may be specified (or defined) by the encoder. The base mesh may include a subset of multiple vertices of the mesh. In one aspect, if the bit depth value of the position is n and the bit depth value of the 2D texture coordinate is m, the position is quantized to a value between 0 and 2. n -1, and quantizes the 2D texture coordinates to a value between 0 and 2 m -1, where n and m are positive integers. For example, if the bit depth value of the position is 12 and the bit depth value of the 2D texture coordinate is 13, the position is quantized to an integer between 0 and 4095, and the 2D texture coordinate is quantized to an integer between 0 and 8191. The quantized value may be predicted at prediction (704) by a prediction mode or a prediction algorithm (e.g., by a multi-parallelogram algorithm for the position of the vertex and a stretch prediction algorithm for the 2D texture coordinate). The prediction residual of the position of the vertex and the prediction residual of the 2D texture coordinate may be further compressed by entropy coding (706).

[0112] In one aspect, such as in MPEG V-DMC WD 2.0, bit depth quantization is a stretch operation. If the source position has a bit depth of m and the encoded bit depth value is n, the stretch operation consists of converting the position of the vertex to the fractional The product can be further rounded to an integer.

[0113] However, this bit depth quantization based on the stretching operation may not be uniform. For example, when the bit depth of the pixel position value of the source grid (or base grid) is 13, the pixel position value can be an integer between 0 and 8191. If the encoded bit depth value is 11, the fraction is equal to Under stretch-based bit depth quantization, source mesh position values ​​0, 1, and 2 are quantized to 0. Source mesh position values ​​2727, 2728, 2729, 2730, and 2731 are quantized to 682. Thus, three position values ​​(0, 1, 2) are quantized to the same quantized value 0, and five position values ​​(2727, 2728, 2729, 2730, 2731) are quantized to the same quantized value 682. This is non-uniform quantization because the first group of three position values ​​corresponds to one quantized value, while the second group of five position values ​​corresponds to another quantized value. In uniform quantization, multiple groups of vertices of the source mesh can be quantized to integers. Each group of vertices can include the same total number of vertices and is further quantized to a corresponding integer.

[0114] In the present disclosure, various methods and / or systems are provided for uniform bit depth scaling in mesh compression. These methods and / or systems can be applied individually or in any combination. These methods and / or systems can be applied to mesh compression, or other applications of bit depth scaling, such as audio compression, audio processing, image compression, image processing, video compression, video processing, etc.

[0115] In the present disclosure, uniform bit depth scaling is provided. In one example, if the source bit depth (also referred to as position bit depth or source position bit depth) is m, and the target bit depth (also referred to as encoding bit depth) is n, then given a source (or position) x, where x is a non-negative integer of the source bit depth m, the uniform bit depth scaling (or uniform bit depth scaling function) s(x) is as shown in equation (1) below.

[0116]

[0117] As shown in equation (1), the uniform bit depth scaling s(x) is based on the position bit depth m and the encoding bit depth n. The position bit depth m indicates that the position x of the current vertex is between 0 and 2. m -1, the encoding bit depth n indicates that the position of the current vertex is quantized to a value between 0 and 2 n -1.

[0118] In one aspect, the uniform bit-depth scaling function is a rounding function configured to round the position of the current vertex to an integer.

[0119] In one aspect, the position bit depth m is greater than the encoding bit depth n.

[0120] In one aspect, the position bit depth m is equal to or less than the encoding bit depth n.

[0121] In one example, when the bit depth (or position bit depth) of the pixel position value of the source mesh is 13 and the encoding bit depth value is 11, using uniform bit depth scaling, the first group of four positions (0, 1, 2 and 3) are quantized to a quantized value of 0, the second group of four positions (2724, 2725, 2726 and 2727) are quantized to a quantized value of 681, and the third group of four positions (2728, 2729, 2730 and 2731) are quantized to a quantized value of 682. Therefore, based on the uniform bit depth scaling, uniform quantization is obtained. In uniform quantization, multiple groups of vertices of the source mesh (or base mesh) are quantized to integers. Each group of vertices is quantized to a corresponding integer and includes the same total number of vertices.

[0122] In one aspect, uniform bit depth scaling is used for both bit depth up-scaling and bit depth down-scaling. When the source bit depth m is less than the target bit depth n, the uniform bit depth scaling is configured to perform bit depth up-scaling. When the source bit depth m is greater than the target bit depth n, the uniform bit depth scaling is configured to perform bit depth down-scaling. When the source bit depth m is equal to the target bit depth n, no changes are made to the input value (or position value) x.

[0123] Figure 8 A flow chart outlining a process (800) according to an aspect of the present disclosure is shown. The process (800) may be used in a video decoder. In various aspects, the process (800) is performed by a processing circuit, such as a processing circuit that performs the functions of a video decoder (110), a processing circuit that performs the functions of a video decoder (210), etc. In some aspects, the process (800) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (800). The process starts at (S801) and proceeds to (S810).

[0124] At (S810), a code stream is received, the code stream including base mesh information of a base mesh. The base mesh includes a subset of a plurality of vertices of a mesh in a current mesh frame.

[0125] At (S820), a position of a current vertex of the base mesh is determined based on a quantized position of the current vertex of the base mesh, the quantized position of the current vertex being generated according to a bit depth scaling function. The bit depth scaling function is configured to convert a first subset of positions of a plurality of vertices of the base mesh into a first integer, and convert a second subset of positions of the plurality of vertices of the base mesh into a second integer. The total number of positions of the first subset is equal to the total number of positions of the second subset.

[0126] At (S830), the current vertex is reconstructed based on the determined position of the current vertex of the base mesh.

[0127] In one aspect, the bit depth scaling function is based on the position bit depth m and the encoding bit depth n. The position bit depth m indicates that the position of the current vertex is between 0 and 2. m -1, the encoding bit depth n indicates that the position of the current vertex is quantized to a value between 0 and 2 n -1.

[0128] In one aspect, the bit depth scaling function is based on a value equal to the position bit depth minus the encoding bit depth.

[0129] In one aspect, the bit depth scaling function is based on an exponential term with a base of 2 and an exponent equal to the position bit depth minus the encoding bit depth.

[0130] In one aspect, the bit depth scaling function is configured to round the position of the current vertex to an integer.

[0131] In one aspect, the position bit depth is greater than the encoding bit depth.

[0132] In one aspect, the position bit depth is equal to or less than the encoding bit depth.

[0133] In one aspect, the bit depth scaling function is defined as Where x is the position of the current vertex, n is the encoding bit depth, and m is the position bit depth.

[0134] Then, the process proceeds to (S899) and terminates.

[0135] The process (800) may be adapted as appropriate. Steps in the process (800) may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.

[0136] Fig. 9 A flow chart outlining a process (900) according to an aspect of the present disclosure is shown. The process (900) may be used in a video encoder. In various aspects, the process (900) is performed by a processing circuit, such as a processing circuit that performs the functions of a video encoder (103), a processing circuit that performs the functions of a video encoder (303), etc. In some aspects, the process (900) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (900). The process starts at (S901) and proceeds to (S910).

[0137] At (S910), a position of a current vertex of a base mesh is quantized based on a bit depth scaling function to determine a quantized position of the current vertex. The base mesh includes a subset of a plurality of vertices of a mesh in a current mesh frame. The bit depth scaling function is configured to quantize a first subset of positions of the plurality of vertices of the base mesh into a first integer, and quantize a second subset of positions of the plurality of vertices of the base mesh into a second integer. The total number of positions of the first subset is equal to the total number of positions of the second subset.

[0138] At (S920), a position prediction of the quantized position of the current vertex is determined.

[0139] At (S930), the position prediction residual of the position prediction of the current vertex is encoded in a bitstream.

[0140] In one aspect, the bit depth scaling function is based on the position bit depth m and the encoding bit depth n. The position bit depth m indicates that the position of the current vertex is between 0 and 2. m -1, the encoding bit depth n indicates that the position of the current vertex is quantized to a value between 0 and 2 n -1.

[0141] In one aspect, the bit depth scaling function is based on a value equal to the position bit depth minus the encoding bit depth.

[0142] In one aspect, the bit depth scaling function is based on an exponential term with a base of 2 and an exponent equal to the position bit depth minus the encoding bit depth.

[0143] In one aspect, the bit depth scaling function is configured to round the position of the current vertex to an integer.

[0144] In one aspect, the position bit depth is greater than the encoding bit depth.

[0145] In one aspect, the position bit depth is equal to or less than the encoding bit depth.

[0146] In one aspect, the bit depth scaling function is defined as: Where x is the position of the current vertex, n is the encoding bit depth, and m is the position bit depth.

[0147] Then, the process proceeds to (S999) and terminates.

[0148] The process (900) may be adapted as appropriate. Steps in the process (900) may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.

[0149] In one aspect, a method for processing mesh data includes processing a code stream of the mesh data according to a format rule. For example, the code stream may be a code stream decoded / encoded using any of the decoding methods and / or encoding methods described herein. The format rule may specify one or more constraints of the code stream and / or one or more processes to be performed by a decoder and / or encoder.

[0150] In one example, a codestream of mesh data is processed according to a format rule. The codestream includes base mesh information of a base mesh. The base mesh includes a subset of multiple vertices of the mesh in a current mesh frame. The format rule specifies: a position of a current vertex of the base mesh is determined based on a quantized position of the current vertex of the base mesh, and the quantized position of the current vertex is generated according to a bit depth scaling function. The bit depth scaling function is configured to convert a first subset of the positions of the multiple vertices of the base mesh into a first integer, and convert a second subset of the positions of the multiple vertices of the base mesh into a second integer. The total number of positions of the first subset is equal to the total number of positions of the second subset. The format rule specifies: a current vertex is processed based on the determined position of the current vertex of the base mesh.

[0151] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig.10 A computer system (1000) suitable for implementing certain aspects of the disclosed subject matter is shown.

[0152] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to assembly, compilation, linking or similar mechanisms to create code comprising instructions, which may be executed directly by one or more computer central processing units (CPU), graphics processing units (GPU), etc., or through interpretation, microcode execution, etc.

[0153] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.

[0154] Fig.10 The components of the computer system (1000) shown are examples and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing aspects of the present disclosure. Nor should the configuration of components be interpreted as having any dependency or requirement related to any one or combination of components shown in the exemplary aspects of the computer system (1000).

[0155] The computer system (1000) may include certain human-machine interface input devices. Such human-machine interface input devices may be responsive to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not depicted). Human-machine interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, captured images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0156] The human-machine interface input device may include one or more of the following (only one of each is shown): keyboard (1001), mouse (1002), touchpad (1003), touch screen (1010), data gloves (not shown), joystick (1005), microphone (1006), scanner (1007), camera (1008).

[0157] The computer system (1000) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback of a touch screen (1010), a data glove (not shown), or a joystick (1005), but may also be a tactile feedback device that is not an input device), an audio output device (e.g., a speaker (1009), headphones (not depicted)), a visual output device (e.g., a screen (1010) including a CRT screen, an LCD screen, a plasma screen, an OLED screen, each with or without touch screen input capabilities, each with or without tactile feedback capabilities, some of which are capable of outputting two-dimensional visual output or output in excess of three dimensions through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted)), and a printer (not depicted).

[0158] The computer system (1000) may also include human-machine accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1020) having CD / DVD and other media (1021), thumb drives (1022), removable hard drives or solid-state drives (1023), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security software dogs (not depicted), etc.

[0159] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.

[0160] The computer system (1000) may also include an interface (1054) to one or more communication networks (1055). The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter attached to some common data port or peripheral bus (1049) (e.g., a USB port of the computer system (1000)); other network interfaces are typically integrated into the kernel of the computer system (1000) by attaching to a system bus as described below (e.g., connected to an Ethernet interface in a PC computer system or connected to a cellular network interface in a smartphone computer system). The computer system (1000) can use any of these networks to communicate with other entities. Such communications may be one-way receive only (e.g., broadcast television), one-way send only (e.g., CANBus connected to certain CANBus devices), or bidirectional, for example, using a LAN or WAN digital network to connect to other computer systems. Certain protocols and protocol stacks may be used on each of those networks and network interfaces as described above.

[0161] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the kernel ( 1040 ) of the computer system ( 1000 ).

[0162] The kernel (1040) may include one or more central processing units (CPUs) (1041), graphics processing units (GPUs) (1042), dedicated programmable processing units in the form of field programmable gate areas (FPGAs) (1043), hardware accelerators (1044) for certain tasks, graphics adapters (1050), etc. These devices, as well as read-only memory (ROM) (1045), random access memory (1046), internal mass storage (1047) such as internal non-user accessible hard drives, SSDs, etc., may be connected via a system bus (1048). In some computer systems, the system bus (1048) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the kernel's system bus (1048) or to the kernel's system bus (1048) via a peripheral bus (1049). In one example, a screen (1010) may be connected to a graphics adapter (1050). The architecture of the peripheral bus includes PCI, USB, etc.

[0163] The CPU (1041), GPU (1042), FPGA (1043) and accelerator (1044) can execute certain instructions, which can be combined to form the computer code mentioned above. The computer code can be stored in ROM (1045) or RAM (1046). Transition data can also be stored in RAM (1046), while permanent data can be stored, for example, in internal mass storage (1047). Fast storage and retrieval to any storage device can be performed by using a cache, which can be closely associated with one or more CPUs (1041), GPUs (1042), mass storage (1047), ROM (1045), RAM (1046), etc.

[0164] The computer readable medium may have thereon computer codes for performing various computer-implemented operations. The medium and computer codes may be those specially designed and constructed for the purposes of the present disclosure, or the medium and computer codes may be of a type well known and available to those skilled in the art of computer software.

[0165] As an example, and not by way of limitation, a computer system (1000) having an architecture, particularly a kernel (1040), can provide functionality because one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage as described above, as well as some non-temporary kernel (1040) memories, such as a kernel internal mass storage (1047) or ROM (1045). Software implementing various aspects of the present disclosure can be stored in such devices and executed by the kernel (1040). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can cause the kernel (1040), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.) to perform specific processes described herein or to perform specific parts of specific processes described herein, including defining data structures stored in RAM (1046) and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system may provide functionality due to logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1044)) that may replace or operate in conjunction with software to perform specific processes described herein or specific portions of specific processes described herein. Where appropriate, references to software may include logic and vice versa. Where appropriate, references to computer-readable media may include circuitry (e.g., integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.

[0166] As used in this disclosure, "at least one of..." or "one of..." is intended to include any one or combination of the listed elements. For example, reference to at least one of A, B, or C, at least one of A, B, and C, at least one of A, B, and / or C, and at least one of A to C is intended to include only A, only B, only C, or any combination thereof. Reference to one of A or B, and one of A and B is intended to include A or B or (A and B). Where applicable, use of "one of..." does not exclude any combination of the listed elements, such as when the elements are not mutually exclusive.

[0167] Although the present disclosure has described several examples of various aspects, there are changes, permutations, and various replacement equivalents that fall within the scope of the present disclosure. Therefore, it should be appreciated that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and therefore fall within the spirit and scope of the present disclosure.

Claims

1. A grid decoding method, the method comprising: receiving a code stream, the code stream comprising basic mesh information of a basic mesh, the basic mesh comprising a subset of a plurality of vertices of a mesh in a current mesh frame; determining a position of a current vertex of the base mesh based on a quantized position of the current vertex of the base mesh, the quantized position of the current vertex generated according to a bit depth scaling function, the bit depth scaling function configured to convert a first subset of positions of a plurality of vertices of the base mesh into a first integer and convert a second subset of positions of the plurality of vertices of the base mesh into a second integer, the total number of positions of the first subset being equal to the total number of positions of the second subset; as well as The current vertex is reconstructed based on the determined position of the current vertex of the base mesh.

2. The method according to claim 1, wherein: The bit depth scaling function is based on the position bit depth m and the encoding bit depth n, The position bit depth m indicates that the position of the current vertex is between 0 and 2 m -1, and The encoding bit depth n indicates that the position of the current vertex is quantized to a value between 0 and 2. n -1.

3. The method according to claim 2, wherein: The bit depth scaling function is based on a value equal to the position bit depth minus the encoding bit depth.

4. The method according to claim 2, wherein: The bit depth scaling function is based on an exponential term with a base of 2 and an exponent equal to the position bit depth minus the encoding bit depth.

5. The method according to claim 2, wherein: The bit depth scaling function is configured to round the position of the current vertex to an integer.

6. The method according to claim 2, wherein: The position bit depth is greater than the encoding bit depth.

7. The method according to claim 2, wherein: The position bit depth is equal to or less than the encoding bit depth.

8. The method according to claim 1, wherein: The bit depth scaling function is defined as: x is the position of the current vertex, n is the encoding bit depth, and m is the position bit depth.

9. A grid coding method, comprising: quantizing a position of a current vertex of a base mesh based on a bit depth scaling function to determine a quantized position of the current vertex, the base mesh comprising a subset of a plurality of vertices of a mesh in a current mesh frame, the bit depth scaling function being configured to quantize a first subset of the positions of the plurality of vertices of the base mesh into a first integer and quantize a second subset of the positions of the plurality of vertices of the base mesh into a second integer, a total number of positions of the first subset being equal to a total number of positions of the second subset; determining a position prediction of the quantized position of the current vertex; as well as The position prediction residual of the position prediction of the current vertex is encoded in a bitstream.

10. The method according to claim 9, wherein: The bit depth scaling function is based on the position bit depth m and the encoding bit depth n, The position bit depth m indicates that the position of the current vertex is between 0 and 2 m -1, and The encoding bit depth n indicates that the position of the current vertex is quantized to a value between 0 and 2. n -1.

11. The method according to claim 10, wherein: The bit depth scaling function is based on a value equal to the position bit depth minus the encoding bit depth.

12. The method according to claim 10, wherein: The bit depth scaling function is based on an exponential term with a base of 2 and an exponent equal to the position bit depth minus the encoding bit depth.

13. The method according to claim 10, wherein: The bit depth scaling function is configured to round the position of the current vertex to an integer.

14. The method according to claim 10, wherein: The position bit depth is greater than the encoding bit depth.

15. The method according to claim 10, wherein: The position bit depth is equal to or less than the encoding bit depth.

16. The method according to claim 9, wherein: The bit depth scaling function is defined as: x is the position of the current vertex, n is the encoding bit depth, and m is the position bit depth.

17. A method for processing grid data, the method comprising: Processing the code stream of the grid data according to the format rule, wherein: The code stream includes base mesh information of a base mesh, the base mesh including a subset of a plurality of vertices of a mesh in a current mesh frame; and The format rules specify: determining a position of a current vertex of the base mesh based on a quantized position of the current vertex of the base mesh, the quantized position of the current vertex generated according to a bit depth scaling function, the bit depth scaling function configured to convert a first subset of positions of a plurality of vertices of the base mesh into a first integer and convert a second subset of positions of the plurality of vertices of the base mesh into a second integer, the total number of positions of the first subset being equal to the total number of positions of the second subset; and The current vertex is processed based on the determined position of the current vertex of the base mesh.

18. The method of claim 17, wherein: The bit depth scaling function is based on the position bit depth m and the encoding bit depth n, The position bit depth m indicates that the position of the current vertex is between 0 and 2 m -1, and The encoding bit depth n indicates that the position of the current vertex is quantized to a value between 0 and 2. n -1.

19. The method according to claim 18, wherein: The bit depth scaling function is based on an exponential term with a base of 2 and an exponent equal to the position bit depth minus the encoding bit depth.

20. The method according to claim 17, wherein: The bit depth scaling function is defined as: x is the position of the current vertex, n is the encoding bit depth, and m is the position bit depth.