Displacement coding and decoding in grid compression
By using processing circuits in the grid decoding device to receive and decode the base grid information and displacement information in the grid data, the problem of difficulty in efficiently compressing and decoding complex 3D grid data in the prior art is solved, and efficient processing of dynamic grids and time-changing data is realized.
Patent Information
- Application Number
- CN202480004281.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-23
- Filing Date
- 2024-04-24
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to efficiently compress and decode complex 3D grid data, especially when processing dynamic grids and time-varying attribute graphs and connectivity information.
An apparatus for grid decoding is adopted, the apparatus including a processing circuit to determine wavelet coefficients in a plurality of 2D blocks packaged into 2D images by receiving base grid information and displacement information in the code stream, and depackaging and displacement determination are performed based on the detailed levels of these wavelet coefficients.
It realizes efficient compression and decoding of complex 3D grid data, can process dynamic grids and time-changing attribute diagrams and connectivity information, and is suitable for applications such as real-time communication, storage and multi-platform rendering.
Smart Images

Figure CN119998840A_ABST
Abstract
Description
[0001] Incorporation by reference
[0002] This application claims the benefit of priority to U.S. Patent Application No. 18 / 644,014, filed on April 23, 2024, entitled “DISPLACEMENT CODING IN MESH COMPRESSION,” which claims the benefit of priority to U.S. Provisional Application No. 63 / 462,932, filed on April 28, 2023, entitled “Displacement Coding in Mesh Compression,” and U.S. Provisional Application No. 63 / 461,552, filed on April 24, 2023, entitled “Packing of Displacement Coding in Mesh Compression.” The entire disclosure of the prior application is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure describes aspects generally related to trellis coding. Background Art
[0004] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent that the work of the presently named inventors described in this background section and in various aspects of this specification was performed, it does not indicate that it qualifies as prior art at the time of filing, and it is never explicitly or implicitly admitted that it is prior art to the present disclosure.
[0005] Image / video compression can help transmit image / video data between different devices, storages, and networks with minimal quality degradation. In some examples, video codec techniques can compress video based on spatial redundancy and temporal redundancy. In one example, a video codec can use a technique called intra-frame prediction, which can compress an image based on spatial redundancy. For example, intra-frame prediction can use reference data from a current picture being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-frame prediction, which can compress an image based on temporal redundancy. For example, inter-frame prediction can be based on predicting samples in a current picture from a previously reconstructed picture using motion compensation. Motion compensation can be indicated by a motion vector (MV).
[0006] Advances in three-dimensional (3D) capture, modeling, and rendering have facilitated 3D content across a variety of platforms and devices. For example, capturing a baby's first steps on one continent, grandparents can see (and in some cases interact with) the child on another continent and enjoy a fully immersive experience with the child. To achieve this realism, models are becoming increasingly complex, and large amounts of data are associated with the creation and consumption of these models. 3D meshes are widely used to represent such immersive content. Summary of the invention
[0007] Aspects of the present disclosure include code streams, methods, and apparatus for mesh processing.In some examples, the apparatus for mesh processing includes processing circuitry.
[0008] According to one aspect of the present disclosure, a device for mesh decoding is provided. The device includes a processing circuit. The processing circuit is configured to receive a code stream, which includes base mesh information and displacement information of a mesh in a current mesh frame. The processing circuit is configured to determine multiple wavelet coefficients packed into multiple 2D blocks of a two-dimensional (2D) image based on the displacement information. The processing circuit is configured to unpack the packed multiple wavelet coefficients from the multiple 2D blocks based on the detail level of the multiple wavelet coefficients. The processing circuit is configured to determine multiple displacements of the mesh in the current mesh frame based on the unpacked multiple wavelet coefficients. The multiple displacements are associated with the mesh and the base mesh in the current mesh frame, and the base mesh includes a subset of multiple vertices of the mesh.
[0009] In one example, the processing circuit is configured to unpack a subset of the plurality of wavelet coefficients from each of the plurality of 2D blocks. The subset of the plurality of wavelet coefficients in each of the plurality of 2D blocks is from a corresponding level of detail in the levels of detail.
[0010] In one example, the processing circuit is configured to unpack a first subset of the plurality of wavelet coefficients from a first 2D block of the plurality of 2D blocks according to a first detail level in the detail levels. The processing circuit is configured to unpack a second subset of the plurality of wavelet coefficients from a second 2D block of the plurality of 2D blocks according to a second detail level in the detail levels. The first 2D block is at least one of a left block and an upper block of the second 2D block in the 2D image, and the first detail level is lower than the second detail level.
[0011] In one example, the first 2D block further includes one or more wavelet coefficients from the second level of detail, and the second 2D block does not include one or more wavelet coefficients from the first level of detail.
[0012] In one example, the plurality of wavelet coefficients are encoded according to intra mode. The processing circuit is configured to determine predicted values of the plurality of wavelet coefficients based on reconstructed wavelet coefficients of neighboring 2D blocks of the plurality of 2D blocks in the current grid frame.
[0013] In one example, the plurality of wavelet coefficients are encoded according to an inter-frame mode. The processing circuit is configured to (i) determine predicted values of the plurality of wavelet coefficients based on reconstructed wavelet coefficients of a reference grid frame from a current grid frame, and (ii) determine prediction residuals associated with the predicted values of the plurality of wavelet coefficients.
[0014] In one example, the plurality of wavelet coefficients are encoded in a copy mode. The processing circuit is configured to determine that the predicted values of the plurality of wavelet coefficients are reconstructed wavelet coefficients in a reference grid frame of the current grid frame.
[0015] In one example, the plurality of wavelet coefficients are encoded in a skip mode. The processing circuit is configured to determine, in a current grid frame, that the plurality of wavelet coefficients are zero.
[0016] In one aspect of the present disclosure, a method for mesh encoding is provided. In the method, a plurality of displacements of a mesh in a current mesh frame is determined. The plurality of displacements are associated with the mesh and a base mesh. The base mesh includes a subset of a plurality of vertices of the mesh. A wavelet transform is applied to the plurality of displacements to generate a plurality of wavelet coefficients. Based on a detail level of the plurality of wavelet coefficients, the plurality of wavelet coefficients are packed into a plurality of 2D blocks of a 2D image. The plurality of wavelet coefficients are encoded into a bitstream, wherein the bitstream includes displacement information associated with the plurality of wavelet coefficients and base mesh information associated with the base mesh.
[0017] In one example, to pack the plurality of wavelet coefficients, a subset of the plurality of wavelet coefficients is packed into each of the plurality of 2D blocks, wherein the subset of the plurality of wavelet coefficients in each of the plurality of 2D blocks is from a corresponding level of detail in the levels of detail.
[0018] In one example, to pack the plurality of wavelet coefficients, a first subset of the plurality of wavelet coefficients is packed into a first 2D block of the plurality of 2D blocks according to a first detail level in the detail levels. A second subset of the plurality of wavelet coefficients is packed into a second 2D block of the plurality of 2D blocks according to a second detail level in the detail levels. The first 2D block is at least one of a left block and an upper block of the second 2D block in the 2D image, and the first detail level is lower than the second detail level.
[0019] In one example, the first 2D block further includes one or more wavelet coefficients from the second level of detail, and the second 2D block does not include one or more wavelet coefficients from the first level of detail.
[0020] In one example, the plurality of wavelet coefficients are encoded according to an intra mode. To encode the plurality of wavelet coefficients, prediction values of the plurality of wavelet coefficients are determined based on reconstructed wavelet coefficients of neighboring 2D blocks of the plurality of 2D blocks in the current grid frame.
[0021] In one example, the plurality of wavelet coefficients are encoded according to an inter-frame mode. To encode the plurality of wavelet coefficients, prediction values of the plurality of wavelet coefficients are determined based on reconstructed wavelet coefficients of a reference grid frame from a current grid frame, and prediction residuals associated with the prediction values of the plurality of wavelet coefficients are determined.
[0022] In one example, the plurality of wavelet coefficients are encoded in a copy mode. To encode the plurality of wavelet coefficients, a prediction of the plurality of wavelet coefficients is determined as reconstructed wavelet coefficients in a reference grid frame of the current grid frame.
[0023] In one example, the plurality of wavelet coefficients are encoded in a skip mode. To encode the plurality of wavelet coefficients, in a current grid frame, the plurality of wavelet coefficients are determined to be zero.
[0024] In one aspect of the present disclosure, a method for mesh data processing is provided. In the method, a code stream of mesh data is processed according to format rules. In one example, the code stream includes base mesh information and displacement information of a mesh in a current mesh frame. The format rules specify: based on the displacement information, multiple wavelet coefficients in multiple 2D blocks packed into a 2D image are determined. The format rules specify: based on the detail level of the multiple wavelet coefficients, multiple wavelet coefficients packed into multiple 2D blocks are unpacked. The format rules specify: based on the unpacked multiple wavelet coefficients, multiple displacements of the mesh in the current mesh frame are determined. The multiple displacements are associated with the mesh and the base mesh in the current mesh frame. The base mesh includes a subset of multiple vertices of the mesh.
[0025] Aspects of the present disclosure also provide an apparatus for trellis coding. The apparatus for trellis coding comprises a processing circuit configured to implement any of the described methods for trellis coding.
[0026] Aspects of the present disclosure also provide a method for trellis decoding, which includes any method implemented by a device for trellis decoding.
[0027] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions which, when executed by a computer, cause the computer to perform any of the described methods for grid decoding / encoding. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0029] Figure 1 is a schematic illustration of an example of a block diagram of a communication system (100).
[0030] Figure 2is a schematic illustration of an example of a block diagram of a decoder.
[0031] Figure 3 is a schematic illustration of an example of a block diagram of an encoder.
[0032] Figure 4 is a schematic illustration of an example of an encoding process (400) for mesh processing according to an aspect of the present disclosure.
[0033] Figure 5 is a schematic illustration of an example of a pre-processing step (500) according to an aspect of the present disclosure.
[0034] Figure 6 is a schematic illustration of a decoding process (600) for trellis processing according to an aspect of the present disclosure.
[0035] Figure 7 is a schematic illustration of an example of a subdivision scheme according to an aspect of the present disclosure.
[0036] Figure 8 A flow chart outlining a trellis decoding process according to some aspects of the present disclosure is shown.
[0037] Fig. 9 A flow chart outlining a trellis encoding process according to some aspects of the present disclosure is shown.
[0038] Fig.10 is a schematic illustration of a computer system according to an aspect. DETAILED DESCRIPTION
[0039] Figure 1 A block diagram of a video processing system (100) in some examples is shown. The video processing system (100) is an example of an application of the disclosed subject matter, namely a video encoder and a video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., and the like.
[0040] The video processing system (100) includes an acquisition subsystem (113), which may include a video source (101). The video source (101) may include one or more images acquired by a camera and / or generated by a computer. For example, a digital camera creates, for example, an uncompressed video picture stream (102). In one example, the video picture stream (102) includes samples recorded by the digital camera. The video picture stream (102), which is depicted as a thick line to emphasize the high amount of data compared to the encoded video data (104) (or the encoded video bitstream), can be processed by an electronic device (120), which includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. The encoded video data (104) (or encoded video bitstream), depicted as thin lines to emphasize the lower amount of data compared to the video picture stream (102), can be stored on the streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1 A client subsystem (106) and a client subsystem (108) in a streaming server (105) may access a copy (107) and a copy (109) of the encoded video data (104). The client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and creates an output video picture stream (111) that can be presented on a display (112) (e.g., a display screen) or other presentation device (not depicted). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., a video bitstream) may be encoded according to certain video encoding / compression standards. Examples of such standards include ITU-T Recommendation H.265. In one example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.
[0041] It should be noted that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may also include a video encoder (not shown).
[0042] Figure 2 An example of a block diagram of a video decoder (210) is shown. The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used to replace Figure 1 A video decoder (110) in an example of FIG.
[0043] The receiver (231) may receive one or more encoded video sequences, for example, included in a bitstream to be decoded by the video decoder (210). In one aspect, the encoded video sequences are received one at a time, wherein the decoding of each encoded video sequence is independent of the decoding of the other encoded video sequences. The encoded video sequence may be received from a channel (201), which may be a hardware / software link leading to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data and other data, for example, encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not depicted). The receiver (231) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory (215) may be located outside the video decoder (210) (not depicted). In still other applications, a buffer memory (not depicted) may be provided outside the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) may be provided inside the video decoder (210) to, for example, handle playback timing. When the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (215) may not be required, or the buffer memory (215) may be made smaller. For use on a traffic packet network such as the Internet, the buffer memory (215) may be required, and the buffer memory (215) may be relatively large, advantageously may have an adaptive size, and may be implemented at least partially in an operating system or similar element (not depicted) external to the video decoder (210).
[0044] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (210) and potential information for controlling a presentation device such as a presentation device (212) (e.g., a display screen) that is not part of the electronic device (230) but can be coupled to the electronic device (230), such as Figure 2As shown. The control information for the rendering device may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not depicted). The parser (220) may parse / entropy decode the received coded video sequence. The encoding of the coded video sequence may be performed according to a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract a subgroup parameter set for at least one subgroup of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (220) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0045] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).
[0046] Depending on the type of coded video picture or part of coded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of symbol (221) may involve multiple different units. Which units are involved and how they are involved can be controlled by parser (220) through subgroup control information parsed from the coded video sequence. For clarity, such subgroup control information flow between parser (220) and the multiple units below is not depicted.
[0047] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into a number of functional units as described below. In actual implementations operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purposes of describing the disclosed subject matter, the conceptual subdivision into the following multiple functional units is appropriate.
[0048] The first unit is a sealer / inverse transform unit (251). The sealer / inverse transform unit (251) receives quantized transform coefficients as symbols (221) from the parser (220) and control information including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit (251) may output a block including sample values, which may be input into an aggregator (255).
[0049] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses surrounding reconstructed information extracted from a current picture buffer (258) to generate a block of the same size and shape as the block being reconstructed. For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a per-sample basis.
[0050] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to a block that is inter-coded and potentially motion compensated. In this case, the motion compensated prediction unit (253) may access the reference picture memory (257) to extract samples for prediction. After the extracted samples are motion compensated according to the symbols (221) belonging to the block, these samples may be added to the output of the scaler / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) by the aggregator (255), thereby generating output sample information. The extraction of prediction samples by the motion compensated prediction unit (253) from the address in the reference picture memory (257) may be controlled by a motion vector, which may be provided to the motion compensated prediction unit (253) in the form of symbols (221), which may have, for example, an X component, a Y component and a reference picture component. Motion compensation may also include interpolation of sample values extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.
[0051] The output samples of the aggregator (255) may be subjected to various loop filtering techniques in a loop filter unit (256). The video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video sequence (also referred to as the coded video bitstream) and available to the loop filter unit (256) as symbols (221) from the parser (220). The video compression may also be responsive to meta-information obtained during decoding of a coded picture or a previous (in decoding order) portion of the coded video sequence, and to previously reconstructed and loop filtered sample values.
[0052] The output of the loop filter unit (256) may be a sample stream that may be output to a rendering device (212) and stored in a reference picture memory (257) for future inter-picture prediction.
[0053] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed, and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) may become part of the reference picture memory (257), and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.
[0054] The video decoder (210) may perform decoding operations according to a predetermined video compression technology or standard, such as ITU-T H.265 Recommendation. The encoded video sequence may conform to the syntax specified by the video compression technology or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as the only tools available for use under the profile. For compliance, the complexity of the encoded video sequence may also be required to be within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata of the HRD buffer management signaled in the encoded video sequence.
[0055] In one aspect, the receiver (231) can receive additional (redundant) data when receiving the encoded video. The additional data can be included as part of the encoded video sequence. The additional data can be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data can take the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0056] Figure 3 An example of a block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used to replace Figure 1 A video encoder (103) in an example of FIG.
[0057] The video encoder (303) can be used to obtain the video source (301) (not Figure 3 In an example, the video source (301) is a part of the electronic device (320) to receive video samples, and the video source (301) can capture video images to be encoded by the video encoder (303). In another example, the video source (301) is a part of the electronic device (320).
[0058] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), the digital video sample stream may have any suitable bit depth (e.g. 8-bit, 10-bit, 12-bit, ...), any color space (e.g. BT.601 Y CrCB, RGB, ...) and any suitable sampling structure (e.g. Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The pictures themselves may be organized into a spatial pixel array, where each pixel may include one or more samples, depending on the sampling structure, color space, etc. used. The following description focuses on the samples.
[0059] According to one aspect, the video encoder (303) can encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other time constraints required. Implementing the appropriate encoding speed is a function of the controller (350). In some aspects, the controller (350) controls other functional units as described below and is functionally coupled to the other functional units described. For clarity, the coupling is not depicted in the figure. The parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization techniques, etc.), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions that are related to the video encoder (303) optimized for a certain system design.
[0060] In some aspects, the video encoder (303) is configured to operate in an encoding loop. As an oversimplified description, in one example, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how a (remote) decoder would also create sample data. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-accurately corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction portion of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, such as due to channel errors) is also used in some related technologies.
[0061] The operation of the "local" decoder (333) may be similar to that described above in conjunction with Figure 2 The "remote" decoder of the video decoder (210) described in detail is identical. However, additional brief reference is made to Figure 2 , since the symbols are available and the entropy encoder (345) and the parser (220) can losslessly encode / decode the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (210) including the buffer memory (215) and the parser (220) may not be fully implemented in the local decoder (333).
[0062] On the one hand, except for the parsing / entropy decoding present in the decoder, the decoder technology is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on the decoder operation. The description of the encoder technology can be simplified because the encoder technology is mutually inverse to the decoder technology described comprehensively. In some areas, a more detailed description is provided below.
[0063] During operation, in some examples, the source encoder (330) may perform motion compensated predictive coding that predictively encodes an input picture by referencing one or more previously encoded pictures from a video sequence designated as "reference pictures." In this manner, the encoding engine (332) encodes the differences between pixel blocks of the input picture and pixel blocks of a reference picture that may be selected as a prediction reference for the input picture.
[0064] The local video decoder (333) may decode the encoded video data of the picture that may be designated as the reference picture based on the symbol created by the source encoder (330). The operation of the encoding engine (332) may advantageously be a lossy process. When the encoded video data may be decoded at the video decoder ( Figure 3 When the video sequence is decoded at a remote video decoder (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture memory (334). In this way, the video encoder (303) may store a copy of the reconstructed reference picture locally that has common content (absent transmission errors) with the reconstructed reference picture to be obtained by the remote video decoder.
[0065] The predictor (335) may perform a prediction search for the encoding engine (332). That is, for a new picture to be encoded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may be used as appropriate prediction references for the new picture. The predictor (335) may operate on a pixel block by pixel block basis to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (335), the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (334).
[0066] The controller (350) may manage encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.
[0067] The outputs of all the above functional units may be entropy encoded in an entropy encoder (345). The entropy encoder (345) converts the symbols generated by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.
[0068] The transmitter (340) may buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission over a communication channel (360), which may be a hardware / software link to a storage device that may store the encoded video data. The transmitter (340) may combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or ancillary data streams (source not shown).
[0069] The controller (350) may manage the operation of the video encoder (303). During encoding, the controller (350) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following picture types:
[0070] An intra picture (I picture) that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh ("IDR") pictures.
[0071] Predictive pictures (P pictures), which can be encoded and decoded using intra prediction or inter prediction, which uses a motion vector and reference index to predict the sample values of each block.
[0072] Bidirectional predictive pictures (B pictures), which can be encoded and decoded using intra prediction or inter prediction, which uses two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.
[0073] The source picture may typically be spatially subdivided into blocks of samples (e.g. blocks of 4×4, 8×8, 4×8 or 16×16 samples) and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined by the coding allocation applied to the corresponding picture of the block. For example, blocks of an I picture may be non-predictively coded, or blocks of an I picture may be predictively coded (spatial prediction or intra prediction) with reference to already coded blocks of the same picture. Blocks of pixels of a P picture may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. Blocks of a B picture may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.
[0074] The video encoder (303) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.
[0075] In one aspect, the transmitter (340) may transmit additional data when transmitting the encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0076] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often shortened to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In one example, a particular picture being encoded / decoded is partitioned into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.
[0077] In some aspects, a bidirectional prediction technique may be used for inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture that precede the current picture in the video in decoding order (but may be in the past and future, respectively, in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.
[0078] In addition, merge mode technology can be used for inter-picture prediction to improve coding efficiency.
[0079] According to some aspects of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks (e.g., polygonal blocks or triangle blocks). For example, according to the High-Efficiency Video Coding (HEVC) standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU includes three coding tree blocks (CTBs), which are a luminance CTB and two chrominance CTBs. Each CTU can be recursively divided into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be divided into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type for the CU, such as an inter prediction type or an intra prediction type. The CU is divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Typically, each PU includes a luma prediction block (PB) and two chroma PBs. On the one hand, the prediction operation in the codec (encoding / decoding) is performed in units of prediction blocks. Using the luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luma values) for pixels, such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0080] It should be noted that the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using any suitable technology. In one aspect, the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using one or more integrated circuits. In another aspect, the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using one or more processors executing software instructions.
[0081] Aspects of the present disclosure include techniques for displacement encoding and decoding in mesh compression.
[0082] A mesh may include multiple polygons that describe the surface of a volumetric object. Each polygon of a mesh may be defined by the vertices of the corresponding polygon in a three-dimensional (3D) space and information about how the vertices are connected (which may be referred to as connectivity information). In some aspects, vertex attributes (e.g., color, normal, etc.) may be associated with vertices (or mesh vertices). Attributes (or vertex attributes) may also be associated with the surface of a mesh by utilizing mapping information that uses a two-dimensional (2D) attribute map to parameterize the mesh. This mapping may be described by a set of parameter coordinates (referred to as UV coordinates or texture coordinates) associated with mesh vertices. A 2D attribute map may be used to store high-resolution attribute information, such as textures, normals, displacements, etc. High-resolution attribute information may be used for various purposes, such as texture mapping and shading.
[0083] Dynamic mesh sequences may require a large amount of data because the dynamic mesh may include a large amount of information that changes over time. Therefore, efficient compression techniques can be used to store and transmit such content. Mesh compression standards (e.g., Information and Communication (IC) Mesh Compression, MESHGRID, Frame-based Animation Mesh Compression (FAMC)) were previously developed by the Moving Picture Experts Group (MPEG) to handle dynamic meshes with constant connectivity, time-varying geometry attributes, and vertex attributes. However, these standards may not take into account attribute graphs and connectivity information that change over time. DCC (Digital Content Creation) tools can generate such dynamic meshes. However, for volume acquisition technology, it may be challenging to generate dynamic meshes with constant connectivity, especially to generate dynamic meshes with constant connectivity under real-time constraints. Existing standards may not support this type of content (e.g., dynamic meshes with constant connectivity). The present disclosure includes various aspects for a new mesh compression standard that can directly handle dynamic meshes with connectivity information that changes over time and optional attribute graphs that change over time. Mesh compression can target lossy and lossless compression for various applications, such as real-time communication, storage, free viewpoint video, augmented reality (AR) and virtual reality (VR). Features such as random access and scalable / progressive codecs can also be considered.
[0084] Figure 4 An example of an encoding process (400) for mesh processing based on a related video codec (e.g., MPEG V-Mesh TMv1.0) according to an aspect of the present disclosure is shown. Figure 4As shown, the encoding process (400) may include a preprocessing step (400A) and an encoding step (400B). The preprocessing step (400A) may be configured to generate a base grid m(i) of the current frame and a displacement field d(i) of the current frame based on an input grid M(i) of the current frame, wherein the displacement field d(i) of the current frame includes a displacement vector. The encoding step (400B) may be configured to encode the base grid m(i), the displacement field d(i), and the texture information of the base grid m(i). The displacement field d(i) of the current frame may include a displacement vector. The index i may refer to the current frame. In one aspect, a mode decision method may be performed in the encoding process (400) to determine whether to apply inter-frame coding (also known as inter-frame prediction or inter-frame mode) or intra-frame coding (also known as intra-frame prediction or intra-frame mode), etc. to the current frame. For example, the mode decision method may compare the cost of the intra-frame mode with the cost of the inter-frame mode, and determine the encoding mode of the base grid m(i) of the current frame based on which cost is smaller. In some examples, the base grid m(i) is encoded or decoded using a skip mode. In one example, the skip mode is a special mode of the inter-frame mode. For example, the base grid m(i) can be intra-coded, inter-coded, or encoded using a skip (SKIP) mode.
[0085] Still reference Figure 4 , the preprocessing step (400A) may include a mesh extraction process (402), a parameterization process such as an atlas parameterization process (404), and a subdivision surface fitting process (406). The mesh extraction process (402) is configured to downsample the vertices of the input mesh M(i) to generate an extracted mesh dm(i) that may include a plurality of extracted (or downsampled) vertices. The number of the plurality of extracted vertices is less than the number of vertices of the input mesh M(i). The parameterization process such as the atlas parameterization process (404) is configured to map the extracted mesh dm(i) onto a planar domain, such as onto a UV atlas (or UV map), to generate a re-parameterized mesh pm(i). In one example, the atlas parameterization may be performed based on a video processing tool (e.g., a UV Atlas tool). The subdivision surface fitting process (406) is configured to take as input the re-parameterized mesh pm(i) and the input mesh M(i), and generate a base mesh m(i) and a displacement field d(i) including displacement vectors or displacement sets. In an example of the subdivision surface fitting process (406), pm(i) is subdivided using a subdivision scheme such as iterative interpolation to obtain a subdivided mesh. Iterative interpolation includes inserting a new point at the midpoint of each edge of the re-parameterized mesh pm(i) at each iteration. Any suitable subdivision scheme may be applied to subdivide pm(i). The displacement field d(i) is calculated by determining the closest point on the surface of the input mesh M(i) for each vertex of the subdivided mesh.
[0086] Advantages of the subdivided mesh may include: the subdivided mesh has a subdivision structure that allows efficient compression while providing a reliable approximation to the input mesh. The improved compression efficiency may be obtained due to the following properties. The decimated mesh dm(i) may have a small number of vertices and may be encoded and transmitted using a smaller number of bits than the input mesh M(i) or the subdivided mesh. Figure 4 , a base mesh m(i) can be generated based on the extracted mesh dm(i). In one example, the base mesh m(i) is the extracted mesh dm(i). Since the subdivided mesh can be generated based on the subdivision method, the subdivided mesh can be automatically generated by the decoder when decoding the base mesh or the extracted mesh (for example, without using any information other than the subdivision scheme and the subdivision iteration count). On the decoder side, the displacement field d(i) can be generated by decoding the displacement vectors associated with the vertices of the subdivided mesh. In addition to enabling spatial / quality scalability, the subdivision structure also enables efficient transformations such as wavelet decomposition, which can provide high compression performance.
[0087] For simplicity, the preprocessing step (400A) that can be applied to an input mesh (e.g., a 3D mesh) can be described using the preprocessing step (500) applied to a two-dimensional (2D) curve. The preprocessing step (400A) is similar to the preprocessing step (500), except that the 3D mesh can be replaced by a 2D curve.
[0088] Figure 5 An example of a preprocessing step (500) according to an aspect of the present disclosure is shown. Figure 5As shown, an input 2D curve (represented by a 2D polyline) (502) may be downsampled to generate a base curve such as a polyline, referred to as a "decimated" curve (504). A subdivision scheme may then be applied to the decimated polyline (504) to generate a "decision" curve (506). In one example, the subdivision scheme may be an iterative interpolation scheme. The iterative interpolation scheme may include inserting a new point at the midpoint of each edge of the polyline (or decimated curve) (504) at each iteration. For example, a point (510) may be inserted into an edge (508) of the decimated curve (504). In one example, the edge (508) is located between a point (512) and a point (514). Further, a point (522) may be added between the point (512) and the point (510), and a point (516) may be added between the point (510) and the point (514). The decimal polyline (506) may then be deformed to generate a displacement curve (518). The displacement curve (518) may be a better approximation of the input curve (502) than the subdivided curve (506). For example, a displacement vector (e.g., (520)) is calculated for each vertex (e.g., (510)) of the subdivided curve (506) so that the shape of the displacement curve (518) is as close as possible to the shape of the input curve (502). An advantage of the subdivided curve (506) is that the subdivided curve (506) has a subdivision structure that allows for more efficient compression while providing a reliable approximation of the input curve (502).
[0089] The decimated curve (504) may have a smaller number of points and may be encoded and transmitted using a limited number of bits. Since the subdivision curves may be generated based on the subdivision scheme, the subdivision curves may be automatically generated by the decoder when decoding the base curve or the decimated curve (e.g., without using any information other than the subdivision method and the subdivision iteration count). The displacement curves may be generated by decoding the displacement vectors associated with the subdivision curve vertices. In addition to enabling spatial / quality scalability, the subdivision structure may also enable efficient transforms such as wavelet decomposition, which may provide high compression performance.
[0090] Still reference Figure 5 In one example, the input mesh M(i) may include an input 2D curve (502). The base mesh m(i) may include a decimated curve (504) formed by downsampling the vertices of the input 2D curve (502). The displacement field dm(i) may include a plurality of displacement vectors, such as Figure 5 The displacement vector (520) is shown.
[0091] The encoding step (400B) may include base mesh encoding (408), displacement encoding (410), and texture encoding (412), etc. The base mesh encoding (408) is configured to encode geometric information of a base mesh m(i) associated with a current frame. In intra-frame coding, the base mesh m(i) may first be quantized (e.g., quantized using uniform quantization) and then encoded, for example, using a coding mode determined by a mode decision method. The coding mode may be an inter-frame mode, an intra-frame mode, or a skip mode, etc. An encoder for intra-coding the base mesh m(i) may be referred to as a static mesh encoder. In inter-frame coding, a reference base mesh associated with a reference frame indicated by index j (e.g., a reconstructed quantized reference base mesh m'(j)) may be used to predict a base mesh m(i) associated with a current frame indicated by index i. The displacement encoding (410) is configured to encode a displacement field d(i) generated in the preprocessing step (400A). The displacement field d(i) may include a set of displacement vectors (or displacements) associated with subdivided mesh vertices. The texture encoding (412) is configured to encode attribute information of the base mesh m(i). The attribute information may include texture, normal, and / or color, etc. The attribute information may be encoded based on a suitable codec (e.g., High Efficiency Video Coding (HEVC) or Next Generation Video Coding (VVC)).
[0092] On the one hand, reference Figure 4 , a mesh encoding process such as the encoding process (400) begins with preprocessing (e.g., preprocessing step (400A)). The preprocessing may convert an input mesh (e.g., an input dynamic mesh) M(i) into a base mesh m(i) and a displacement field d(i) including a displacement set (or a displacement vector set). The encoding step (400B) may compress the output from the preprocessing (e.g., m(i), and d(i), etc.) and generate a compressed code stream b(i). The compressed code stream b(i) may include a compressed base mesh code stream, a compressed displacement field code stream, and / or a compressed attribute code stream, etc.
[0093] Figure 6An example of a decoding process (600) for grid processing according to an aspect of the present disclosure is shown. The decoding process (600) may include a decoding step (605) and a post-processing step (610). A compressed code stream b(i) may be provided to the decoding step (605). In one example, for example for lossless transmission, the compressed code stream b(i) is the output b(i) from the encoding process (400). The decoding step (605) may extract various sub-code streams, such as a compressed base grid sub-stream, a compressed displacement field sub-stream, and / or a compressed attribute sub-stream, etc. The decoding step (605) may decompress the sub-code streams to generate the following components: patch metadata indicated by metadata (i), a decoded base grid m" (i), a decoded displacement field (including displacement) d" (i), and / or a decoded attribute map A" (i), etc.
[0094] On the one hand, the base grid substream can be provided to a grid decoder to generate a reconstructed quantized base grid m'(i). The decoded base grid (or reconstructed base grid) m"(i) can be obtained by applying inverse quantization to m'(i). The displacement field substream including the encoded packed and quantized wavelet coefficients can be decoded by a video and / or image decoder. Image unpacking and inverse quantization can be applied to the reconstructed packed and quantized wavelet coefficients to obtain unpacked and dequantized transform coefficients (e.g., wavelet coefficients). An inverse wavelet transform can be applied to the unpacked and dequantized wavelet coefficients to generate a decoded displacement field (or reconstructed displacement) d"(i).
[0095] The decoded components (e.g., including metadata(i), m”(i), d”(i), and / or A”(i), etc.) can be provided to a post-processing step (610). A mesh (also referred to as a decoded / reconstructed mesh) M”(i) can be generated by the post-processing step (610) based on m”(i) and d”(i). In one example, the mesh M”(i) (also referred to as a reconstructed deformed mesh DM(i)) can be obtained by subdividing m”(i) using a subdivision scheme and applying the reconstructed displacements d”(i) to the vertices of the subdivided mesh. In one example, DM(i) can include a displacement curve (518). In one example, when the encoding process (400), the decoding process (600), and the transmission are lossless, the mesh M”(i) can be the same as the input mesh M(i). When one of the encoding process (400), the decoding process (600), and the transmission is lossy, M”(i) is different from M(i). In various examples, the difference between M”(i) and M(i), if any, can be relatively small. In one example, a property graph A”(i) is also generated by a post-processing step (610).
[0096] In one aspect, the base grid may be intra-coded, inter-coded, or encoded using a skip (SKIP) mode, etc. In one example, the SKIP mode may be a special mode of the inter-frame mode, in which the base grid m(i) of the current frame indicated by the index i is the same as the base grid m(j) of the reference frame indicated by the index (also referred to as the frame index) j. When the inter-frame mode is applied to encode the base grid in the current frame, the encoder may generate a predicted base grid for the current frame based on the reconstructed base grid of the reference frame. In one example, such as in MPEG V-DMC WD 2.0, the reference frame is a frame immediately preceding the current frame in display order. The frame index i of the current frame indicates the display order. When the frame index of the current frame is i, the frame index of the reference frame is (i-1). In one example, the current frame and the reference frame are in the same group of frames (GoF).
[0097] In one aspect, for example in MPEG V-DMC WD 2.0, the mesh encoding process starts with preprocessing. The preprocessing may convert the input dynamic mesh (denoted as M(i)) into a base mesh m(i) and a set of displacements d(i). The encoder may compress the base mesh m(i) and the displacements d(i) to generate a compressed bitstream b(i).
[0098] Preprocessing may include mesh extraction, followed by atlas parameterization, and then subdivision surface fitting, which can be done in Figure 4 . Mesh extraction can use simplification techniques to extract the input mesh M(i) and produce an extracted mesh dm(i). The extracted mesh dm(i) can then be reparameterized. The generated mesh can be represented as pm(i). Subdivision surface fitting can take the reparameterized mesh pm(i) and the input mesh M(i) as input and produce a base mesh m(i) and a displacement set d(i).
[0099] An example of subdivision surface fitting is given in Figure 7 As shown in Figure 7 As shown, the re-parameterized mesh pm(i) (702) can be subdivided by applying a subdivision method (e.g., a midpoint subdivision scheme). The midpoint subdivision scheme can subdivide each triangle into 4 sub-triangles in each subdivision iteration, such as Figure 7 For example, in the initial iteration S 0 In the first iteration S, the re-parameterized mesh pm(i) (702) may include two triangles (701) and (703). 1 In the second iteration S, multiple vertices, such as vertex (704) and vertex (706), can be generated according to the midpoint subdivision scheme. Accordingly, triangle (701) is subdivided into 4 small triangles, and triangle (703) is also subdivided into 4 small triangles. 2In the example, multiple vertices, such as vertex (708) and vertex (710), may be generated according to the midpoint subdivision scheme. 1 Each triangle formed in can be 2 Subdivided into 4 triangles.
[0100] The displacement field d(i) may be computed by determining, for each vertex of the subdivided mesh, the closest point on the surface of the original (or input) mesh M(i).
[0101] Optionally, the encoder can encode a set of displacement vectors associated with the subdivided mesh vertices, referred to as the displacement field d(i). In one example, the reconstructed quantized base mesh m'(i) can be used to update the displacement field d(i) to generate an updated displacement field d'(i). A wavelet transform can then be applied to d'(i) and a set of wavelet coefficients can be generated. The wavelet coefficients can then be quantized and packed into a 2D image / video. The quantized and packed wavelet coefficients can be compressed using arithmetic coding, a conventional image / video encoder, or any other encoder.
[0102] The present disclosure includes various aspects of packing and / or displacement coding for displacement coding in grid compression. Packing may include packing of wavelet coefficients for displacement coding. Displacement coding may include encoding of wavelet coefficients. The various aspects described herein may be applied individually or in any combination. Further, these aspects are not limited to grid compression.
[0103] In some aspects, the packing of wavelet coefficients may be based on the level of detail of the wavelet coefficients. The level of detail may indicate the level (or degree) of mesh resolution of the mesh. For example, a high level of detail indicates a high resolution and a low level of detail indicates a low resolution. The levels of detail may include at least one low level of detail and at least one high level of detail. In one example, the levels of detail include a high level of detail, a medium level of detail, and a low level of detail. In one example, a higher level of detail (lower scale) captures high frequency wavelet coefficients, while a lower level of detail (higher scale) captures low frequency wavelet coefficients.
[0104] In some aspects, wavelet coefficients are packed into multiple 2D blocks of a 2D image / video based on the level of detail of the wavelet coefficients. The order of 2D blocks in a 2D image and / or which wavelet coefficients are included in each 2D block can be based on the level of detail of the wavelet coefficients.
[0105] When the wavelet coefficients are packed into 2D blocks, the one-dimensional structure of the displacement field can be converted into a two-dimensional structure, which includes the displacement values and the corresponding row / column numbers in the 2D blocks. In one example, sorting of the 2D blocks can be applied. For example, the 2D blocks can be sorted based on a sorting method (e.g., a zigzag order or a raster order). Then, the wavelet coefficients can be packed into the sorted (or serialized) 2D blocks according to the detail level.
[0106] In one example, a subset of the plurality of wavelet coefficients may be packed into each of the plurality of 2D blocks. The subset of the plurality of wavelet coefficients in each of the plurality of 2D blocks is from a corresponding level of detail in the levels of detail.
[0107] In one aspect, the packing of wavelet coefficients may be based on the detail level of the wavelet coefficients, wherein each 2D block may contain wavelet coefficients from only one detail level. In one example, the 2D block contains wavelet coefficients from only one of a low detail level and a high detail level. In one example, the 2D block contains wavelet coefficients from one of a plurality of low detail levels and a plurality of high detail levels.
[0108] In one aspect, the order or position of 2D blocks in a 2D image may be based on the detail levels of the wavelet coefficients. For example, a 2D block containing wavelet coefficients from a low detail level may be located before a 2D block containing wavelet coefficients from a high detail level. In one example, a 2D block containing wavelet coefficients from multiple low detail levels may be located before a 2D block containing wavelet coefficients from multiple high detail levels. The 2D blocks may also be sorted by multiple low detail levels and multiple high detail levels, such as in ascending order.
[0109] In one example, according to a first detail level in the detail levels, a first subset of the plurality of wavelet coefficients is packed into a first 2D block in the plurality of 2D blocks. According to a second detail level in the detail levels, a second subset of the plurality of wavelet coefficients is packed into a second 2D block in the plurality of 2D blocks. The first 2D block is at least one of a left block and an upper block of the second 2D block in the 2D image, and the first detail level is lower than the second detail level. In one example, the first detail level is a low detail level, and the second detail level is a high detail level.
[0110] In one aspect, the packing of wavelet coefficients may be based on the detail level of the wavelet coefficients. In one aspect, the order of the 2D blocks may be based on the detail level of the wavelet coefficients. For example, a 2D block containing wavelet coefficients from a low detail level may be located before a 2D block containing wavelet coefficients from a high detail level. Further, in some aspects, wavelet coefficients of more than one detail level may be included in a 2D block.
[0111] In one example, a 2D block containing wavelet coefficients from multiple low detail levels may be located before a 2D block containing wavelet coefficients from multiple high detail levels. In addition, in one aspect, wavelet coefficients from a high detail level (e.g., a small number of wavelet coefficients) may be packed into a 2D block containing wavelet coefficients from a low detail level. For example, for a 2D block including wavelet coefficients from more than one detail level, the number of wavelet coefficients from the high detail level is less than the number of wavelet coefficients from the low detail level.
[0112] In one aspect, wavelet coefficients from a low detail level may not be packed into a 2D block containing wavelet coefficients from a high detail level. For example, the first 2D block mentioned above may also include one or more wavelet coefficients from a second detail level (e.g., a high detail level), and the second 2D block mentioned above may not include one or more wavelet coefficients from the first detail level (e.g., a low detail level).
[0113] In certain aspects, the wavelet coefficients may be encoded using a prediction mode. Examples of prediction modes include intra mode, inter mode, copy mode, skip mode, and variations thereof.
[0114] In one aspect, the wavelet coefficients may be encoded using a prediction mode (e.g., inter-frame mode). In intra-frame mode, the wavelet coefficients of a frame may be encoded independently of other frames. The wavelet coefficients of a frame may be encoded by arithmetic coding, an image encoder, or any other suitable encoder.
[0115] In one aspect, the wavelet coefficients may be encoded using a prediction mode (e.g., inter-frame mode). In inter-frame mode, the wavelet coefficients of the current frame may be predicted based on the reconstructed wavelet coefficients of the reference frame of the current frame. The prediction residual of the predicted value may be further encoded by arithmetic coding, an image encoder, a video encoder, or any other suitable encoder.
[0116] In one aspect, the wavelet coefficients may be encoded using a prediction mode (eg, a copy mode). In the copy mode, the predicted wavelet coefficients of the current frame may be copied directly from the reconstructed wavelet coefficients of the reference frame of the current frame. The prediction residual may not be encoded.
[0117] In one aspect, the wavelet coefficients may be encoded using a prediction mode (eg, skip mode). In skip mode, the wavelet coefficients of the current frame may be set to zero. The prediction residual may not be encoded.
[0118] Figure 8A flow chart outlining a process (800) according to an aspect of the present disclosure is shown. The process (800) may be used in a device. The device may include a grid decoder (e.g., a dynamic grid decoder) and a video decoder. The video decoder is configured to, for example, decode a displacement component encoded using a video codec. In various aspects, the process (800) is performed by a processing circuit, such as a processing circuit that performs the functions of a video decoder (110), a processing circuit that performs the functions of a video decoder (210) and / or a grid decoder, etc. In some aspects, the process (800) is implemented in software instructions, so that when the processing circuit executes the software instructions, the processing circuit performs the process (800). The process starts at (S801) and proceeds to (S810).
[0119] At (S810), a code stream is received, the code stream including base grid information and displacement information of a grid in a current grid frame.
[0120] At (S820), a plurality of wavelet coefficients packed into a plurality of 2D blocks of a 2D image are determined based on the displacement information.
[0121] At (S830), a plurality of wavelet coefficients are unpacked from a plurality of 2D blocks based on a detail level of the plurality of wavelet coefficients.
[0122] At (S840), a plurality of displacements of a mesh in a current mesh frame is determined based on the unpacked plurality of wavelet coefficients. The plurality of displacements are associated with the mesh in the current mesh frame and a base mesh, and the base mesh includes a subset of a plurality of vertices of the mesh.
[0123] In one example, a subset of the plurality of wavelet coefficients is unpacked from each 2D block in the plurality of 2D blocks. The subset of the plurality of wavelet coefficients in each 2D block in the plurality of 2D blocks is from a corresponding level of detail in the levels of detail.
[0124] In one example, a first subset of the plurality of wavelet coefficients is unpacked from a first 2D block in the plurality of 2D blocks according to a first detail level in the detail levels. A second subset of the plurality of wavelet coefficients is unpacked from a second 2D block in the plurality of 2D blocks according to a second detail level in the detail levels. The first 2D block is at least one of a left block and an upper block of the second 2D block in the 2D image, and the first detail level is lower than the second detail level.
[0125] In one example, the first 2D block further includes one or more wavelet coefficients from the second level of detail, and the second 2D block does not include one or more wavelet coefficients from the first level of detail.
[0126] In one example, the plurality of wavelet coefficients are encoded according to an intra mode. Prediction values of the plurality of wavelet coefficients are determined based on reconstructed wavelet coefficients of neighboring 2D blocks of the plurality of 2D blocks in the current grid frame.
[0127] In one example, the plurality of wavelet coefficients are encoded according to an inter-frame mode. Prediction values of the plurality of wavelet coefficients are determined based on reconstructed wavelet coefficients of a reference grid frame from the current grid frame. Prediction residuals associated with the prediction values of the plurality of wavelet coefficients are determined.
[0128] In one example, the plurality of wavelet coefficients are encoded in a copy mode. The prediction values of the plurality of wavelet coefficients are determined to be reconstructed wavelet coefficients in a reference grid frame of the current grid frame.
[0129] In one example, the plurality of wavelet coefficients are encoded in a skip mode. In a current grid frame, the plurality of wavelet coefficients are determined to be zero.
[0130] Then, the process proceeds to (S899) and terminates.
[0131] The process (800) may be adapted as appropriate. One or more steps in the process (800) may be modified and / or omitted. Additional one or more steps may be added. Any suitable implementation order may be used.
[0132] Fig. 9 A flow chart outlining a process (900) according to an aspect of the present disclosure is shown. The process (900) may be used in a device. The device may include a grid encoder (e.g., a dynamic grid encoder) and a video encoder. In various aspects, the process (900) is performed by a processing circuit, such as a processing circuit that performs the functions of a video encoder (103), a processing circuit that performs the functions of a video encoder (303) and / or a grid encoder, etc. In some aspects, the process (900) is implemented in software instructions, so that when the processing circuit executes the software instructions, the processing circuit performs the process (900). The process starts at (S901) and proceeds to (S910).
[0133] At (S910), a plurality of displacements of a mesh in a current mesh frame is determined. The plurality of displacements are associated with the mesh and a base mesh. The base mesh includes a subset of a plurality of vertices of the mesh.
[0134] At (S920), wavelet transform is performed on the plurality of displacements to generate a plurality of wavelet coefficients.
[0135] At (S930), the plurality of wavelet coefficients are packed into a plurality of 2D blocks of the 2D image based on the detail levels of the plurality of wavelet coefficients.
[0136] At (S940), the plurality of wavelet coefficients are encoded into a code stream, wherein the code stream includes displacement information associated with the plurality of wavelet coefficients and base grid information associated with the base grid.
[0137] In one example, to pack the plurality of wavelet coefficients, a subset of the plurality of wavelet coefficients is packed into each of the plurality of 2D blocks, wherein the subset of the plurality of wavelet coefficients in each of the plurality of 2D blocks is from a corresponding level of detail in the levels of detail.
[0138] In one example, to pack the plurality of wavelet coefficients, a first subset of the plurality of wavelet coefficients is packed into a first 2D block of the plurality of 2D blocks according to a first detail level in the detail levels. A second subset of the plurality of wavelet coefficients is packed into a second 2D block of the plurality of 2D blocks according to a second detail level in the detail levels. The first 2D block is at least one of a left block and an upper block of the second 2D block in the 2D image, and the first detail level is lower than the second detail level.
[0139] In one example, the first 2D block further includes one or more wavelet coefficients from the second level of detail, and the second 2D block does not include one or more wavelet coefficients from the first level of detail.
[0140] In one example, the plurality of wavelet coefficients are encoded according to an intra mode. To encode the plurality of wavelet coefficients, prediction values of the plurality of wavelet coefficients are determined based on reconstructed wavelet coefficients of neighboring 2D blocks of the plurality of 2D blocks in the current grid frame.
[0141] In one example, the plurality of wavelet coefficients are encoded according to an inter-frame mode. To encode the plurality of wavelet coefficients, prediction values of the plurality of wavelet coefficients are determined based on reconstructed wavelet coefficients of a reference grid frame from a current grid frame, and prediction residuals associated with the prediction values of the plurality of wavelet coefficients are determined.
[0142] In one example, the plurality of wavelet coefficients are encoded in a copy mode. To encode the plurality of wavelet coefficients, prediction values of the plurality of wavelet coefficients are determined to be reconstructed wavelet coefficients in a reference grid frame of the current grid frame.
[0143] In one example, the plurality of wavelet coefficients are encoded in a skip mode. To encode the plurality of wavelet coefficients, in a current grid frame, the plurality of wavelet coefficients are determined to be zero.
[0144] Then, the process proceeds to (S999) and terminates.
[0145] The process (900) may be adapted as appropriate. One or more steps in the process (900) may be modified and / or omitted. Additional one or more steps may be added. Any suitable implementation order may be used.
[0146] In the present disclosure, a method for processing mesh data is provided. In the method, a code stream of mesh data is processed according to format rules. In one example, the code stream includes base mesh information and displacement information of a mesh in a current mesh frame. The format rule specifies: based on the displacement information, multiple wavelet coefficients in multiple 2D blocks packed into a 2D image are determined. The format rule specifies: based on the detail level of the multiple wavelet coefficients, multiple wavelet coefficients packed into multiple 2D blocks are unpacked. The format rule specifies: based on the unpacked multiple wavelet coefficients, multiple displacements of the mesh in the current mesh frame are determined. The multiple displacements are associated with the mesh and the base mesh in the current mesh frame. The base mesh includes a subset of multiple vertices of the mesh.
[0147] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig.10 A computer system (1000) suitable for implementing certain aspects of the disclosed subject matter is shown.
[0148] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to assembly, compilation, linking or similar mechanisms to create code comprising instructions that may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpretation, microcode execution, etc.
[0149] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.
[0150] Fig.10 The components of the computer system (1000) shown are examples and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing aspects of the present disclosure. Nor should the configuration of components be interpreted as having any dependency or requirement related to any one or combination of components shown in the exemplary aspects of the computer system (1000).
[0151] The computer system (1000) may include certain human-machine interface input devices. Such human-machine interface input devices may be responsive to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not depicted). Human-machine interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, captured images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0152] The human-machine interface input device may include one or more of the following (only one of each is shown): keyboard (1001), mouse (1002), touchpad (1003), touch screen (1010), data gloves (not shown), joystick (1005), microphone (1006), scanner (1007), camera (1008).
[0153] The computer system (1000) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback through a touch screen (1010), a data glove (not shown), or a joystick (1005), but may also be a tactile feedback device that is not an input device), audio output devices (e.g., speakers (1009), headphones (not depicted)), visual output devices (e.g., screens (1010) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities, some of which are capable of outputting two-dimensional visual outputs or outputs in more than three dimensions through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays and smoke boxes (not depicted), and printers (not depicted).
[0154] The computer system (1000) may also include human-machine accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1020) having CD / DVD and other media (1021), thumb drives (1022), removable hard drives or solid-state drives (1023), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security software dogs (not depicted), etc.
[0155] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.
[0156] The computer system (1000) may also include an interface (1054) to one or more communication networks (1055). The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter attached to some common data port or peripheral bus (1049) (e.g., a USB port of the computer system (1000)); other network interfaces are typically integrated into the kernel of the computer system (1000) by attaching to a system bus as described below (e.g., connected to an Ethernet interface in a PC computer system or connected to a cellular network interface in a smartphone computer system). The computer system (1000) can use any of these networks to communicate with other entities. Such communications may be one-way receive only (e.g., broadcast television), one-way send only (e.g., CANBus connected to certain CANBus devices), or bidirectional, for example, using a LAN or WAN digital network to connect to other computer systems. Certain protocols and protocol stacks may be used on each of those networks and network interfaces as described above.
[0157] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the kernel ( 1040 ) of the computer system ( 1000 ).
[0158] The kernel (1040) may include one or more central processing units (CPUs) (1041), graphics processing units (GPUs) (1042), dedicated programmable processing units in the form of field programmable gate areas (FPGAs) (1043), hardware accelerators (1044) for certain tasks, graphics adapters (1050), etc. These devices, as well as read-only memory (ROM) (1045), random access memory (1046), internal mass storage (1047) such as internal non-user accessible hard drives, SSDs, etc., may be connected via a system bus (1048). In some computer systems, the system bus (1048) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the kernel's system bus (1048) or to the kernel's system bus (1048) via a peripheral bus (1049). In one example, a screen (1010) may be connected to a graphics adapter (1050). The architecture of the peripheral bus includes PCI, USB, etc.
[0159] The CPU (1041), GPU (1042), FPGA (1043) and accelerator (1044) can execute certain instructions, which can be combined to form the computer code mentioned above. The computer code can be stored in ROM (1045) or RAM (1046). Transition data can also be stored in RAM (1046), while permanent data can be stored, for example, in internal mass storage (1047). Fast storage and retrieval to any storage device can be performed by using a cache, which can be closely associated with one or more CPUs (1041), GPUs (1042), mass storage (1047), ROM (1045), RAM (1046), etc.
[0160] The computer readable medium may have thereon computer codes for performing various computer-implemented operations. The medium and computer codes may be those specially designed and constructed for the purposes of the present disclosure, or the medium and computer codes may be of a type well known and available to those skilled in the art of computer software.
[0161] As an example, and not by way of limitation, a computer system (1000) having an architecture, particularly a kernel (1040), can provide functionality because one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage as described above, as well as some non-temporary kernel (1040) memories, such as a kernel internal mass storage (1047) or ROM (1045). Software implementing various aspects of the present disclosure can be stored in such devices and executed by the kernel (1040). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can cause the kernel (1040), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.) to perform specific processes described herein or to perform specific parts of specific processes described herein, including defining data structures stored in RAM (1046) and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system may provide functionality due to logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1044)) that may replace or operate in conjunction with software to perform specific processes described herein or specific portions of specific processes described herein. Where appropriate, references to portions of software may include logic and vice versa. Where appropriate, references to portions of computer-readable media may include circuitry (e.g., integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.
[0162] As used in this disclosure, "at least one of" or "one of" is intended to include any one or combination of the listed elements. For example, references to at least one of A, B, or C, at least one of A, B, and C, at least one of A, B, and / or C, and at least one of A to C are intended to include only A, only B, only C, or any combination thereof. References to one of A or B, and one of A and B are intended to include A or B or (A and B). Where applicable, the use of "one of" does not exclude any combination of the listed elements, such as when the elements are not mutually exclusive.
[0163] Although the present disclosure has described multiple examples of various aspects, there are changes, permutations, and various replacement equivalents that fall within the scope of the present disclosure. Therefore, it should be appreciated that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and therefore fall within the spirit and scope of the present disclosure.
Claims
1. A device for grid decoding, characterized in that: The device comprises: Processing circuitry for: receiving a code stream, wherein the code stream includes base grid information and displacement information of a grid in a current grid frame; Based on the displacement information, determining a plurality of wavelet coefficients packed into a plurality of 2D blocks of a two-dimensional (2D) image; unpacking the plurality of wavelet coefficients from the plurality of 2D blocks based on a level of detail of the plurality of wavelet coefficients; and A plurality of displacements of the mesh in the current mesh frame are determined based on the unpacked plurality of wavelet coefficients, the plurality of displacements being associated with the mesh in the current mesh frame and a base mesh including a subset of a plurality of vertices of the mesh.
2. The device according to claim 1, characterized in that The processing circuit is used for: A subset of the plurality of wavelet coefficients is unpacked from each 2D block in the plurality of 2D blocks, the subset of the plurality of wavelet coefficients in each 2D block in the plurality of 2D blocks being from a corresponding level of detail in the levels of detail.
3. The device according to claim 1 or 2, characterized in that: The processing circuit is used for: unpacking a first subset of the plurality of wavelet coefficients from a first 2D block in the plurality of 2D blocks according to a first level of detail in the levels of detail; and Unpacking a second subset of the plurality of wavelet coefficients from a second 2D block among the plurality of 2D blocks according to a second detail level among the detail levels, the first 2D block being at least one of a left block and an upper block of the second 2D block in the 2D image, the first detail level being lower than the second detail level.
4. The device according to claim 3, characterized in that The first 2D block also includes one or more wavelet coefficients from the second level of detail, and The second 2D block does not include one or more wavelet coefficients from the first level of detail.
5. The device according to any one of claims 1 to 4, characterized in that The plurality of wavelet coefficients are encoded according to an intra mode; and The processing circuit is used for: Prediction values of the plurality of wavelet coefficients are determined based on reconstructed wavelet coefficients of neighboring 2D blocks of the plurality of 2D blocks in the current grid frame.
6. The device according to any one of claims 1 to 4, characterized in that The plurality of wavelet coefficients are encoded according to an inter-mode; and The processing circuit is used for: (i) determining predicted values of the plurality of wavelet coefficients based on reconstructed wavelet coefficients of a reference grid frame from the current grid frame, and (ii) determining prediction residuals associated with the predicted values of the plurality of wavelet coefficients.
7. The device according to any one of claims 1 to 4, characterized in that The plurality of wavelet coefficients are encoded in a replication mode; and The processing circuit is used for: Determine the predicted values of the plurality of wavelet coefficients as reconstructed wavelet coefficients in a reference grid frame of the current grid frame.
8. The device according to any one of claims 1 to 4, characterized in that The plurality of wavelet coefficients are encoded in a skip mode; and The processing circuit is used for: In the current grid frame, the plurality of wavelet coefficients are determined to be zero.
9. A method for trellis coding, characterized in that: include: determining a plurality of displacements of a mesh in a current mesh frame, the plurality of displacements being associated with the mesh and a base mesh including a subset of a plurality of vertices of the mesh; performing a wavelet transform on the plurality of displacements to generate a plurality of wavelet coefficients; packing the plurality of wavelet coefficients into a plurality of two-dimensional (2D) blocks of a 2D image based on a level of detail of the plurality of wavelet coefficients; as well as The plurality of wavelet coefficients are encoded into a code stream, wherein the code stream includes displacement information associated with the plurality of wavelet coefficients and base grid information associated with the base grid.
10. The method according to claim 9, characterized in that The packing of the plurality of wavelet coefficients further comprises: A subset of the plurality of wavelet coefficients is packed into each of the plurality of 2D blocks, the subset of the plurality of wavelet coefficients in each of the plurality of 2D blocks being from a corresponding level of detail in the levels of detail.
11. The method according to claim 9 or 10, characterized in that: The packing of the plurality of wavelet coefficients further comprises: packing a first subset of the plurality of wavelet coefficients into a first 2D block in the plurality of 2D blocks according to a first level of detail in the levels of detail; and According to a second detail level among the detail levels, a second subset of the plurality of wavelet coefficients is packed into a second 2D block among the plurality of 2D blocks, the first 2D block being at least one of a left block and an upper block of the second 2D block in the 2D image, and the first detail level being lower than the second detail level.
12. The method according to claim 11, characterized in that The first 2D block also includes one or more wavelet coefficients from the second level of detail, and The second 2D block does not include one or more wavelet coefficients from the first level of detail.
13. The method according to any one of claims 9 to 12, characterized in that The plurality of wavelet coefficients are encoded according to an intra-frame mode; as well as The encoding of the plurality of wavelet coefficients comprises: determining predicted values of the plurality of wavelet coefficients based on reconstructed wavelet coefficients of neighboring 2D blocks of the plurality of 2D blocks in the current grid frame.
14. The method according to any one of claims 9 to 12, characterized in that The plurality of wavelet coefficients are encoded according to an inter-frame mode; as well as The encoding of the plurality of wavelet coefficients includes: (i) determining predicted values of the plurality of wavelet coefficients based on reconstructed wavelet coefficients of a reference grid frame from the current grid frame, and (ii) determining prediction residuals associated with the predicted values of the plurality of wavelet coefficients.
15. A method for grid data processing, characterized in that: The method comprises: Process the code stream of the grid data according to the format rules, where: The code stream includes base grid information and displacement information of the grid in the current grid frame; and The format rules specify: Based on the displacement information, determining a plurality of wavelet coefficients packed into a plurality of 2D blocks of a two-dimensional (2D) image; unpacking the plurality of wavelet coefficients packed into the plurality of 2D blocks based on the level of detail of the plurality of wavelet coefficients; and A plurality of displacements of the mesh in the current mesh frame are determined based on the unpacked plurality of wavelet coefficients, the plurality of displacements being associated with the mesh in the current mesh frame and a base mesh including a subset of a plurality of vertices of the mesh.