Inter coding in grid compression

By using inter-frame mode to encode the basic grid of the current frame in grid compression, and using the reference basic grid of the reference basic grid of the reference frame to reconstruct the basic grid of the current frame, the problem of difficulty in dealing with time-varying attribute diagrams and connection information in the prior art is solved, and efficient grid compression and transmission are achieved.

CN119999209APending Publication Date: 2025-05-13TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004278.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-23
Filing Date
2024-04-24
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process time-varying attribute diagrams and connection information in grid compression, especially under real-time constraints, and cannot meet the needs of complex volume acquisition technology.

Method used

The base grid of the current frame is encoded using inter-frame mode, by determining at least one reference frame, and reconstructing the base grid of the current frame using inter-frame mode based on the reference frame reference frame.

Benefits of technology

It realizes efficient compression of complex 3D mesh, reduces data volume, improves transmission and storage efficiency, and is suitable for real-time communication and other resource-constrained application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119999209A_ABST
    Figure CN119999209A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure include methods and apparatus of video decoding and video encoding, and methods of processing visual media data. The apparatus for video decoding includes a processing circuit configured to receive encoding information indicating that a current base grid of a current frame is encoded in an inter-frame mode and indicating at least one reference frame for encoding the current base grid of the current frame. The processing circuitry is configured to determine the at least one reference frame indicated by the encoding information, and reconstruct a current base grid of the current frame using the inter mode based on a respective reference base grid of each of the at least one reference frame.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Incorporated herein by reference

[0002] This application claims priority to U.S. patent application No. 18 / 644,020, filed on April 23, 2024, entitled “Inter-Frame Coding in Grid Compression,” which claims priority to U.S. Provisional Application No. 63 / 461,553, filed on April 24, 2023, entitled “Inter-Frame Coding in Grid Compression,” and U.S. Provisional Application No. 63 / 462,933, filed on April 28, 2023, entitled “Inter-Frame Coding in Grid Compression,” the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure describes embodiments that generally relate to mesh processing. Background Art

[0004] The background description provided herein is intended to generally present the context of the present disclosure. The extent to which the work of the presently named inventors described in the background section and in the examples of this specification is performed does not indicate that it is prior art at the time of filing of this disclosure, and it is never explicitly or implicitly admitted that it is prior art to the present disclosure.

[0005] Image / video compression can help transmit image / video data across different devices, storage devices, and networks while minimizing quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In an example, a video codec can use a technique called intra-frame prediction, which can compress an image based on spatial redundancy. For example, intra-frame prediction can use reference data from the current picture being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-frame prediction, which can compress an image based on temporal redundancy. For example, inter-frame prediction can predict samples in the current picture based on a previously reconstructed picture with motion compensation. Motion compensation can be indicated by a motion vector (MV).

[0006] Advances in three-dimensional (3D) capture, modeling, and rendering have facilitated 3D content across a variety of platforms and devices. For example, a baby’s first steps are captured on one continent, and grandparents on another can see (and in some cases interact with) and enjoy a fully immersive experience with their child. To achieve this level of realism, models have become more complex, and large amounts of data are associated with the creation and consumption of these models. 3D meshes are widely used to represent this immersive content. Summary of the invention

[0007] Various aspects of the present disclosure include methods and apparatus for mesh processing.

[0008] According to one aspect of the present disclosure, a grid decoding device includes a processing circuit. The processing circuit is configured to receive encoding information indicating that a current base grid of a current frame is encoded using an inter-frame mode and indicating at least one reference frame used to encode the current base grid of the current frame. The processing circuit is configured to determine the at least one reference frame indicated by the encoding information. The processing circuit is configured to reconstruct the current base grid of the current frame using the inter-frame mode based on a corresponding reference base grid of each reference frame in the at least one reference frame.

[0009] In one aspect, a display order of one of the at least one reference frames indicated by display index k precedes a display order of the current frame indicated by display index i. In an example, the one of the at least one reference frame and the current frame are in the same group of frames (GoF). In an example, the one of the at least one reference frame and the current frame are in different groups of frames.

[0010] In the example, k<(i-1).

[0011] In one aspect, a display order of one of the at least one reference frames indicated by display index k is after a display order of the current frame indicated by display index i. In an example, the one of the at least one reference frames and the current frame are in the same group of frames (GoF). In an example, the one of the at least one reference frames and the current frame are in different groups of frames.

[0012] In an example, the processing circuit is configured to determine a reference frame index indicated by the encoding information, and determine one of the at least one reference frame in the reference frame list of the current frame based on the reference frame index.

[0013] In an example, the at least one reference frame comprises two reference frames. The processing circuit is configured to determine a prediction of the current base grid based on corresponding reference base grids of the two reference frames, and to reconstruct the current base grid based on an average of the predictions of the current base grid.

[0014] In an example, the current base mesh includes at least two current vertices. For each of the reference base meshes, the prediction of the current base mesh based on the corresponding reference base mesh includes the prediction of the at least two current vertices. The encoding information indicates a motion field of the at least two current vertices, the motion field being associated with the corresponding reference base mesh and including a motion vector for each of the at least two current vertices. For each of the at least two current vertices, the processing circuit is configured to: determine a motion vector for the corresponding current vertex, and determine the prediction of the current vertex based on the motion vector of the corresponding current vertex and a corresponding reference vertex in the corresponding reference base mesh. In an example, for each of the at least two current vertices, the processing circuit is configured to determine the motion vector of the corresponding current vertex based on a motion vector predictor list and an index indicated by the encoding information.

[0015] In an example, the average of the predictions of the current base grid of the current frame is a weighted average of the predictions of the current base grid of the current frame.

[0016] In an example, the two reference frames are determined from two reference frame lists.

[0017] In one aspect, a grid encoding method includes: determining that a current base grid of a current frame is encoded using an inter-frame mode, determining at least one reference frame to encode the current base grid of the current frame, and encoding the current base grid of the current frame using the inter-frame mode based on a corresponding reference base grid of each reference frame in the at least one reference frame.

[0018] In one aspect, a display order of one of the at least one reference frames indicated by display index k precedes a display order of the current frame indicated by display index i. In one aspect, a display order of one of the at least one reference frames indicated by display index k follows a display order of the current frame indicated by display index i.

[0019] In one aspect, the at least one reference frame comprises two reference frames. The method comprises determining a prediction of the current base grid based on respective reference base grids of the two reference frames, and encoding the current base grid based on an average of the predictions of the current base grid.

[0020] In one aspect, a method for processing grid data is provided. In the method, a bitstream of the grid data is processed according to a format rule. The bitstream includes a first syntax element and a second syntax element, the first syntax element indicating that a current base grid of a current frame is encoded using an inter-frame mode, and the second syntax element indicating at least one reference frame used to encode the current base grid of the current frame. The format rule specifies that the at least one reference frame is determined based on the second syntax element, and the current base grid of the current frame is reconstructed based on the inter-frame mode and a corresponding reference base grid of each of the at least one reference frame. In an example, a display order of one of the at least one reference frames indicated by a display index k precedes a display order of the current frame indicated by a display index i.

[0021] Various aspects of the present disclosure also provide a trellis coding apparatus, wherein the trellis coding apparatus comprises a processing circuit, wherein the processing circuit is configured to implement any one of the trellis processing methods performed in an encoder.

[0022] Various aspects of the present disclosure also provide a grid processing method, which includes any method implemented by a grid processing device (eg, a decoder).

[0023] Various aspects of the present disclosure also provide a non-volatile computer-readable storage medium storing instructions, and when the instructions are executed by a computer, the computer is caused to perform any one of the grid processing methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Other features, properties and various advantages of the disclosed subject matter will become further apparent from the following detailed description and accompanying drawings, in which:

[0025] Figure 1 is a schematic diagram of an example of a block diagram of a communication system (100).

[0026] Figure 2 is a schematic diagram of an example block diagram of a decoder.

[0027] Figure 3 is a schematic diagram of an example of a block diagram of an encoder.

[0028] Figure 4 An example of an encoding process (400) for mesh processing according to an embodiment of the present disclosure is shown.

[0029] Figure 5 An example of a pre-processing step (500) according to an embodiment of the present disclosure is shown.

[0030] Figure 6 An example of a decoding process (600) for grid processing according to an embodiment of the present disclosure is shown.

[0031] Figure 7 A flow chart outlining a decoding process for trellis processing according to an embodiment of the present disclosure is shown.

[0032] Figure 8 A flow chart outlining an encoding process for mesh processing according to an embodiment of the present disclosure is shown.

[0033] Fig. 9 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION

[0034] Figure 1 A block diagram of a video processing system (100) in some examples is shown. The video processing system (100) is an example of an application for the disclosed subject matter, a video encoder and a video decoder in a streaming environment. The disclosed subject matter is equally applicable to other applications supporting images and / or video, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., and the like.

[0035] The video processing system (100) includes an acquisition subsystem (113), which may include a video source (101). The video source (101) may include at least one image captured by a camera and / or generated by a computer. For example, a digital camera may create an uncompressed video picture stream (102). In an embodiment, the video picture stream (102) includes samples taken by the digital camera. Compared to the encoded video data (104) (or the encoded video bitstream), the video picture stream (102) is depicted as a thick line to emphasize the high data volume of the video picture stream, and the video picture stream (102) can be processed by an electronic device (120), which includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. Compared to the video picture stream (102), the encoded video data (104) (or the encoded video bitstream) is depicted as a thin line to emphasize the lower amount of data of the encoded video data (104) (or the encoded video bitstream), which can be stored on the streaming server (105) for future use. At least one streaming client subsystem, such as Figure 1A client subsystem (106) and a client subsystem (108) in a streaming server (105) may access a streaming server (105) to retrieve a copy (107) and a copy (109) of the encoded video data (104). The client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and generates an output video picture stream (111) that can be presented on a display (112) (e.g., a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), the video data (107), and the video data (109) (e.g., a video bitstream) may be encoded according to certain video encoding / compression standards. Examples of such standards include ITU-T H.265. In an embodiment, the video coding standard under development is informally referred to as Versatile Video Coding (VVC), and the present application may be used in the context of the VVC standard.

[0036] It should be noted that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may also include a video encoder (not shown).

[0037] Figure 2 An example of a block diagram of a video decoder (210) is shown. The video decoder (210) may be provided in an electronic device (230). The electronic device (230) may include a receiver (231). The receiver (231) may include a receiving circuit, such as a network interface circuit. The video decoder (210) may be used to replace Figure 1 A video decoder (110) of an embodiment.

[0038] The receiver (231) may receive at least one encoded video sequence, for example, included in a bitstream, to be decoded by the video decoder (210). In one aspect, the encoded video sequences are received one at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequence may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data as well as other data, for example, encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not shown). The receiver (231) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other cases, the buffer memory (215) may be set outside the video decoder (210) (not shown). In other cases, a buffer memory (not shown) is set outside the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) may be configured inside the video decoder (210) to, for example, handle broadcast timing. When the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may not be necessary to configure the buffer memory (215), or the buffer memory may be made smaller. Of course, in order to use on a service packet network such as the Internet, a buffer memory (215) may also be required. The buffer memory may be relatively large and may have an adaptive size, and may be at least partially implemented in an operating system or a similar element (not shown) outside the video decoder (210).

[0039] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. The types of symbols include information for managing the operation of the video decoder (210) and potential information for controlling a display device such as a display device (212) (e.g., a display screen) that is not part of the electronic device (230) but can be coupled to the electronic device (230), such as Figure 2As shown in . The control information for the display device may be a parameter set fragment (not indicated) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (220) may parse / entropy decode the received coded video sequence. The encoding of the coded video sequence may be performed according to a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and the like. The parser (220) may extract a subgroup parameter set of at least one subgroup of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), and the like. The parser (220) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.

[0040] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215), thereby creating symbols (221).

[0041] Depending on the type of the coded video picture or a portion of the coded video picture (e.g., inter-frame and intra-frame pictures, inter-frame blocks and intra-frame blocks) and other factors, the reconstruction of the symbol (221) may involve at least two different units. Which units are involved and how they are involved may be controlled by subgroup control information parsed by the parser (220) from the coded video sequence. For the sake of brevity, such subgroup control information flow between the parser (220) and the at least two units below is not described.

[0042] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into several functional units as described below. In a practical embodiment operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purpose of describing the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.

[0043] The first unit is a sealer / inverse transform unit (251). The sealer / inverse transform unit (251) receives quantized transform coefficients as symbols (221) from the parser (220) and control information, including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit (251) can output a block including sample values, which can be input into an aggregator (255).

[0044] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of a current picture. Such predictive information may be provided by an intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates surrounding blocks of the same size and shape as the block being reconstructed using reconstructed information extracted from a current picture buffer (258). For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a per-sample basis.

[0045] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to an inter-frame coded and potentially motion compensated block. In this case, the motion compensated prediction unit (253) may access the reference picture memory (257) to extract samples for prediction. After the extracted samples are motion compensated according to the symbols (221), these samples may be added to the output of the scaler / inverse transform unit (251) (in this case referred to as residual samples or residual signals) by the aggregator (255) to generate output sample information. The acquisition of the predicted samples by the motion compensated prediction unit (253) from the address in the reference picture memory (257) may be controlled by a motion vector, and the motion vector is provided to the motion compensated prediction unit (253) in the form of the symbols (221), for example, including X, Y and reference picture components. Motion compensation may also include interpolation of sample values ​​extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.

[0046] The output samples of the aggregator (255) may be used by various loop filtering techniques in a loop filter unit (256). The video compression techniques may include in-loop filter techniques that are controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream) and that are available to the loop filter unit (256) as symbols (221) from the parser (220). However, in other embodiments, the video compression techniques may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or encoded video sequence, and to previously reconstructed and loop filtered sample values.

[0047] The output of the loop filter unit (256) may be a sample stream that may be output to a display device (212) and stored in a reference picture memory (257) for subsequent inter-picture prediction.

[0048] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed, and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) may become part of the reference picture memory (257), and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.

[0049] The video decoder (210) may perform decoding operations according to a predetermined video compression technology or standard (e.g., ITU-T H.265). The encoded video sequence may conform to the syntax specified by the video compression technology or standard used in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as the only tools available for use under the profile. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata of the HRD buffer management signaled in the encoded video sequence.

[0050] In an embodiment, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or signal noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0051] Figure 3 An example of a block diagram of a video encoder (303) is shown. The video encoder (303) is provided in an electronic device (320). The electronic device (320) includes a transmitter (340) (eg, a transmission circuit). The video encoder (303) may be used to replace Figure 1 A video encoder (103) in an embodiment.

[0052] The video encoder (303) can be used to obtain the video source (301) (not Figure 3 In another embodiment, the video source (301) is a part of the electronic device (320) to receive video samples, and the video source can collect video images to be encoded by the video encoder (303). In another embodiment, the video source (301) is a part of the electronic device (320).

[0053] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by a video encoder (303), wherein the digital video sample stream may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as at least two separate pictures that are given motion when viewed sequentially. The pictures themselves may be constructed as a spatial pixel array, wherein each pixel may include at least one sample depending on the sampling structure, color space, etc. used. The following description focuses on samples.

[0054] According to an embodiment, the video encoder (303) may encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other time constraints as required. Implementing an appropriate encoding speed is a function of the controller (350). In some embodiments, the controller (350) controls other functional units as described below and is functionally coupled to these units. For the sake of brevity, couplings are not shown in the figure. The parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology, etc.), picture size, group of pictures (group of pictures, GOP) layout, maximum motion vector search range, etc. The controller (350) can be used to have other suitable functions that are related to the video encoder (303) optimized for a certain system design.

[0055] In some embodiments, the video encoder (303) operates in a coding loop. As a simple description, in embodiments, the coding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-accurate between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values ​​that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, such as due to channel errors) is also used in some related technologies.

[0056] The operation of the "local" decoder (333) may be combined with, for example, Figure 2 The "remote" decoder described in detail for the video decoder (310) is identical. However, additional brief reference is made to Figure 2 , when symbols are available and the entropy encoder (345) and parser (220) are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and the parser (220), may not be fully implemented in the local decoder (333).

[0057] In an embodiment, in addition to the parsing / entropy decoding present in the decoder, the decoder technology is also present in the corresponding encoder in the same or substantially the same functional form. Therefore, the application focuses on the decoder operation. The description of the encoder technology can be simplified because the encoder technology is mutually inverse to the decoder technology described comprehensively. In some areas, a more detailed description is provided below.

[0058] During operation, in some embodiments, the source encoder (330) may perform motion compensated predictive coding. The motion compensated predictive coding predictively encodes an input picture with reference to at least one previously encoded picture from a video sequence designated as a "reference picture." In this manner, the encoding engine (332) encodes the difference between a pixel block of the input picture and a pixel block of a reference picture that may be selected as a prediction reference for the input picture.

[0059] The local video decoder (333) may decode the encoded video data of the picture that may be designated as the reference picture based on the symbol created by the source encoder (330). The operation of the encoding engine (332) may be a lossy process. When the encoded video data is available at the video decoder ( Figure 3 When the video sequence is decoded at a remote location (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture memory (334). In this way, the video encoder (303) may locally store a copy of the reconstructed reference picture that has common content (absent transmission errors) with the reconstructed reference picture to be obtained by the remote video decoder.

[0060] The predictor (335) may perform a prediction search for the encoding engine (332). That is, for a new picture to be encoded, the predictor (335) may search the reference picture memory (334) for sample data (as a reference pixel block) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may serve as a suitable prediction reference for the new picture. The predictor (335) may operate on a pixel block-by-pixel block basis to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (335), it may be determined that the input picture may have a prediction reference taken from at least two reference pictures stored in the reference picture memory (334).

[0061] The controller (350) may manage encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.

[0062] The outputs of all the above functional units may be entropy encoded in an entropy encoder (345). The entropy encoder (345) performs lossless compression on the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into a coded video sequence.

[0063] The transmitter (340) may buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission over a communication channel (360), which may be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (340) may combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or ancillary data streams (source not shown).

[0064] The controller (350) may manage the operation of the video encoder (303). During encoding, the controller (350) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following picture types:

[0065] Intra pictures (I pictures) can be encoded and decoded without using any other pictures in the sequence as a prediction source.Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures.

[0066] A predictive picture (P picture) may be encoded and decoded using intra prediction or inter prediction, which predicts sample values ​​of each block using at least one motion vector and a reference index.

[0067] Bidirectional predictive pictures (B pictures) can be encoded and decoded using intra prediction or inter prediction, which uses two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, at least two predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.

[0068] Those skilled in the art are aware of these variations of I pictures, P pictures and B pictures and their respective applications and features.

[0069] The source picture may typically be spatially subdivided into at least two blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined according to the coding allocation applied to the corresponding picture of the block. For example, blocks of I pictures may be non-predictively coded, or they may be predictively coded (spatial prediction or intra prediction) with reference to already coded blocks of the same picture. Pixel blocks of P pictures may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. Blocks of B pictures may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.

[0070] The video encoder (303) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0071] In an embodiment, the transmitter (340) may transmit additional data when transmitting the encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0072] The captured video may be taken as at least two source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlation in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In an embodiment, a particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case of using at least two reference pictures, the motion vector may have a third dimension that identifies the reference picture.

[0073] In some embodiments, bidirectional prediction techniques may be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be in the past and future in display order, respectively). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Specifically, the block may be predicted by a combination of the first reference block and the second reference block.

[0074] In addition, merge mode technology can be used in inter-picture prediction to improve coding efficiency.

[0075] According to some embodiments disclosed in the present application, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks, such as polygonal or triangular blocks. For example, according to the HEVC standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU includes three coding tree blocks (CTBs), which are a luminance CTB and two chrominance CTBs. Furthermore, each CTU can be split into at least one coding unit (CU) using a quadtree. For example, a 64×64 pixel CTU can be split into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In an embodiment, each CU is analyzed to determine the prediction type for the CU, such as an inter-prediction type or an intra-prediction type. In addition, depending on temporal and / or spatial predictability, the CU is split into at least one prediction unit (PU). Typically, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Taking the luminance prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, and the like.

[0076] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using any suitable technology. In an embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using at least one integrated circuit. In another embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using at least one processor that executes software instructions.

[0077] The present disclosure includes embodiments of methods and systems related to inter-frame coding in trellis compression.

[0078] A mesh may include several polygons that describe the surface of a volumetric object. Each polygon of a mesh may be defined by the vertices of the corresponding polygons in a three-dimensional (3D) space, and information about how the vertices are connected (which may be referred to as connection information). In some embodiments, vertex attributes such as color, normals, etc. may be associated with vertices (or mesh vertices). Using mapping information (which uses a two-dimensional (2D) attribute map to parameterize the mesh), attributes (or vertex attributes) may also be associated with the surface of the mesh. This mapping may be described by a set of parameter coordinates (referred to as UV coordinates or texture coordinates) associated with the mesh vertices. A 2D attribute map may be used to store high-resolution attribute information, such as textures, normals, displacements, etc. High-resolution attribute information may be used for various purposes, such as texture mapping and shading.

[0079] Dynamic mesh sequences may require a large amount of data because dynamic meshes may include a lot of information that varies over time. Therefore, efficient compression techniques can be used to store and transmit these contents. Mesh compression standards, such as Information and Communication (IC) Mesh Compression, MESHGRID, and Frame-based Animation Mesh Compression (FAMC), were previously developed by the Moving Picture Experts Group (MPEG) to address the problem of dynamic meshes with persistent connectivity, time-varying geometry, and vertex attributes. However, these standards may not consider time-varying attribute graphs and connectivity information. DCC (Digital Content Creation) tools can generate such dynamic meshes. However, it may be challenging for volumetric acquisition techniques to generate persistently connected dynamic meshes, especially under real-time constraints. Existing standards may not support this type of content (e.g., persistently connected dynamic meshes). A new mesh compression standard can be developed to directly handle dynamic meshes with time-varying connectivity information and optional time-varying attribute graphs. The new mesh compression standard can perform lossy and lossless compression for various applications, such as real-time communication, storage, free viewpoint video, augmented reality (AR), and virtual reality (VR). Features such as random access and scalable / progressive coding can also be considered.

[0080] Figure 4An example of an encoding process (400) for mesh processing based on a related video codec (such as MPEG V-Mesh TMv1.0) according to an embodiment of the present disclosure is shown. Figure 4 As shown, the encoding process (400) may include a preprocessing step (400A) and an encoding step (400B). The preprocessing step (400A) may be configured to generate a base grid m(i) of the current frame and a displacement field d(i) of the current frame according to an input grid M(i) of the current frame, the displacement field including a displacement vector. The encoding step (400B) may be configured to encode the base grid m(i), the displacement field d(i), and the texture information of the base grid m(i). The displacement field d(i) of the current frame may include a displacement vector. The index i may point to the current frame. In one aspect, a mode decision method may be performed in the encoding process (400) to determine whether to apply inter-frame coding (also known as inter-frame prediction or inter-frame mode), intra-frame coding (also known as intra-frame prediction or intra-frame mode), etc. to the current frame. For example, the mode decision method may compare the cost of the intra-frame mode and the cost of the inter-frame mode, and decide the encoding mode of the base grid m(i) of the current frame based on which cost is smaller. In some examples, skip mode is used to encode or decode (e.g., encode or decode) the base grid m(i). In an example, skip mode is a special mode of inter mode. For example, the base grid m(i) can be intra-coded, or inter-coded, or encoded using skip mode.

[0081] Still reference Figure 4, the preprocessing step (400A) may include a mesh extraction process (402), a parameterization process (such as an atlas parameterization process (404)), and a subdivision surface fitting process (406). The mesh extraction process (402) is configured to downsample the vertices of the input mesh M(i) to generate an extracted mesh dm(i), which may include at least two extracted (or downsampled) vertices. The number of at least two extracted vertices is less than the number of vertices of the input mesh M(i). The parameterization process, such as the atlas parameterization process (404), is configured to map the extracted mesh dm(i) to a planar domain, such as a UV atlas (or UV map), to generate a re-parameterized mesh pm(i). In an example, the atlas parameterization may be performed based on a video processing tool (such as a UVAtlas tool). The subdivision surface fitting process (406) is configured to take as input the re-parameterized mesh pm(i) and the input mesh M(i), and generate a base mesh m(i) and a displacement field d(i) including a displacement vector or a set of displacements. In an example of the subdivision surface fitting process (406), pm(i) is subdivided using a subdivision scheme (e.g., iterative interpolation) to obtain a subdivided mesh. Iterative interpolation includes inserting a new point in the middle of each edge of the re-parameterized mesh pm(i) at each iteration. Any suitable subdivision scheme can be applied to subdivide pm(i). The displacement field d(i) is calculated by determining the nearest point on the surface of the input mesh M(i) for each vertex of the subdivided mesh.

[0082] Advantages of subdividing meshes may include: the subdivided mesh has a subdivision structure that allows for efficient compression while providing a faithful approximation of the input mesh. The compression efficiency may be achieved due to the following properties. The decimated mesh dm(i) may have a small number of vertices and may be encoded and transmitted using fewer bits than the input mesh M(i) or the subdivided mesh. Figure 4 , a base mesh m(i) can be generated based on the extracted mesh dm(i). In the example, the base mesh m(i) is the extracted mesh dm(i). Since the subdivided mesh can be generated based on the subdivision method, the decoder can automatically generate the subdivided mesh when decoding the base mesh or the extracted mesh (for example, without using any information other than the subdivision scheme and the subdivision iteration count). On the decoder side, the displacement field d(i) can be generated by decoding the displacement vectors associated with the vertices of the subdivided mesh. In addition to allowing spatial / quality scalability, the subdivision structure can also implement efficient transforms such as wavelet decomposition, which can provide high compression performance.

[0083] For simplicity, the preprocessing step (500) applied to a two-dimensional (2D) curve may be used to illustrate the preprocessing step (400A) applicable to an input mesh (e.g., a 3D mesh). The preprocessing step (400A) and the preprocessing step (500) are similar except that the 3D mesh may be replaced by a 2D curve. Figure 5 An example of a preprocessing step (500) according to an embodiment of the present disclosure is shown. Figure 5 As shown, an input 2D curve (represented by a 2D polyline) (502) can be downsampled to generate a basic curve such as a polyline, referred to as a "decimate" curve (504). Then, a subdivision scheme can be applied to the decimate polyline (504) to generate a "decimate" curve (506). In an example, the subdivision scheme can be an iterative interpolation scheme. The iterative interpolation scheme can include inserting a new point in the middle of each edge of the polyline (or decimate curve) (504) at each iteration. For example, a point (510) can be inserted in an edge (508) of the decimate curve (504). In the example, the edge (508) is between points (512) and (514). In addition, a point (522) can be added between point (512) and point (510), and a point (516) can be added between point (510) and point (514). The decimate polyline (506) is then deformed to generate a displacement curve (518). The displacement curve (518) may be a closer approximation of the input curve (502) than the subdivided curve (506). For example, a displacement vector (e.g., (520)) is calculated for each vertex (e.g., (510)) of the subdivided curve (506) so that the shape of the displacement curve (518) is as close as possible to the shape of the input curve (502). An advantage of the subdivided curve (506) is that the subdivided curve (506) has a subdivision structure that allows for more efficient compression while providing a faithful approximation of the input curve (502).

[0084] The extracted curve (504) can have a small number of points and can be encoded and transmitted using a limited number of bits. Since the subdivision curves can be generated based on the subdivision scheme, when the base curve or the extracted curve is decoded, the decoder can automatically generate the subdivision curves (e.g., without using any information other than the subdivision method and the subdivision iteration count). The displacement curves are generated by decoding the displacement vectors associated with the subdivision curve vertices. In addition to allowing spatial / quality scalability, the subdivision structure can also enable efficient transforms such as wavelet decomposition, which can provide high compression performance.

[0085] Still reference Figure 5In an example, the input mesh M(i) may include an input 2D curve (502). The base mesh m(i) may include a decimated curve (504), which is formed by downsampling vertices of the input 2D curve (502). The displacement field dm(i) may include at least two displacement vectors, such as Figure 5 The displacement vector (520) is shown.

[0086] The encoding step (400B) may include base mesh encoding (408), displacement encoding (410), texture encoding (412), etc. The base mesh encoding (408) is configured to encode geometric information of a base mesh m(i) associated with a current frame. In intra-frame coding, the base mesh m(i) may be first quantized (e.g., using uniform quantization) and then encoded, for example, using a coding mode determined using a mode decision method. The coding mode may be an inter-frame mode, an intra-frame mode, a skip mode, etc. An encoder for intra-coding a base mesh m(i) may be referred to as a static mesh encoder. In inter-frame coding, a reference base mesh associated with a reference frame indicated by index j (e.g., a reconstructed quantized reference base mesh m'(j)) may be used to predict a base mesh m(i) associated with a current frame indicated by index i. The displacement encoding (410) is configured to encode a displacement field d(i) generated in the preprocessing step (400A). The displacement field d(i) may include a set of displacement vectors (or displacements) associated with subdivided mesh vertices. The texture encoding (412) is configured to encode attribute information of the base mesh m(i). The attribute information may include texture, normal, color, etc. The attribute information may be encoded based on a suitable codec, such as High Efficiency Video Coding (HEVC) or Universal Video Coding (VVC).

[0087] In one aspect, reference Figure 4 , a mesh encoding process such as encoding process (400) begins with preprocessing (e.g., preprocessing step (400A)). Preprocessing can convert an input mesh (e.g., an input dynamic mesh) M(i) together with a displacement field d(i) including a set of displacements (or a set of displacement vectors) into a base mesh m(i). Encoding step (400B) can compress the output from preprocessing (e.g., m(i), d(i), etc.) and generate a compressed bitstream b(i). The compressed bitstream b(i) can include a compressed base mesh bitstream, a compressed displacement field bitstream, a compressed attribute bitstream, etc.

[0088] Figure 6An example of a decoding process (600) for grid processing according to an embodiment of the present disclosure is shown. The decoding process (600) may include a decoding step (605) and a post-processing step (610). A compressed bitstream b(i) may be fed to the decoding step (605). In the example, for lossless transmission, the compressed bitstream b(i) is the output b(i) from the encoding process (400). The decoding step (605) may extract various sub-bitstreams, such as a compressed base grid substream, a compressed displacement field substream, a compressed attribute substream, and the like. The decoding step (605) may decompress the sub-bitstream to generate the following components: patch metadata indicated by metadata(i), a decoded base grid m"(i), a decoded displacement field (including displacement) d"(i), a decoded attribute map A"(i), and the like.

[0089] In one aspect, the base grid substream can be fed to a grid decoder to generate a reconstructed quantized base grid m'(i). The decoded base grid (or reconstructed base grid) m"(i) can be obtained by applying an inverse quantization to m'(i). The displacement field substream (including the encoded packed and quantized wavelet coefficients) can be decoded by a video and / or image decoder. Image unpacking and inverse quantization can be applied to the packed and quantized wavelet coefficients, which are reconstructed to obtain unpacked and unquantized transform coefficients (e.g., wavelet coefficients). An inverse wavelet transform can be applied to the unpacked and unquantized wavelet coefficients to generate a decoded displacement field (or reconstructed displacement) d"(i).

[0090] The decoded components (e.g., including metadata(i), m"(i), d"(i), A"(i), etc.) can be fed to a post-processing step (610). A mesh (also referred to as a decoded / reconstructed mesh) M"(i) can be generated by the post-processing step (610) based on m"(i) and d"(i). In an example, the mesh M"(i) (also referred to as a reconstructed deformed mesh DM(i)) can be obtained by subdividing m"(i) using a subdivision scheme and applying the reconstructed displacements d"(i) to the vertices of the subdivided mesh. In an example, DM(i) can include a displacement curve (518). In an example, when the encoding process (400), the decoding process (600), and the transmission are lossless, the mesh M"(i) can be the same as the input mesh M(i). When one of the encoding process (400), the decoding process (600), and the transmission is lossy, M"(i) is different from M(i). In various examples, the difference (if any) between M"(i) and M(i) is relatively small. In the example, the attribute graph A"(i) is also generated by the post-processing step (610).

[0091] In one aspect, the base grid may be intra-coded, inter-coded, or encoded using a skip (SKIP) mode, etc. In an example, the skip mode may be a special mode of the inter-frame mode, in which the base grid m(i) of the current frame indicated by index i is the same as the base grid m(j) of the reference frame indicated by index (also referred to as frame index) j. When the inter-frame mode is applied to encode the base grid in the current frame, the encoder may generate a predicted base grid for the current frame based on the reconstructed base grid of the reference frame. In an example, such as in MPEG V-DMC WD 2.0, the reference frame is a frame immediately preceding the current frame in display order. The frame index i of the current frame indicates the display order. When the frame index of the current frame is i, the frame index of the reference frame is (i-1). In an example, the current frame and the reference frame are in the same group of frames (GoF).

[0092] In one aspect, the GoF structure can specify the order of arrangement within and between frames. A GoF may include a group of frames that can be decoded independently. A GoF may include consecutive frames within a bitstream. A GoF may include an I frame that is intra-coded. An I frame may be encoded independently of other frames. In an example, an I frame is used as the starting point of a GoF. A GoF may include other frames (e.g., at least one P frame, at least one B frame, etc.) that may be predicted using at least one previously decoded frame. In an example, a GoF starts with an I frame, with P and B frames following the I frame, and specific reference constraints are applied to the frames in the GoF.

[0093] One aspect of the present disclosure describes methods, aspects, examples, and systems for inter-coding a base grid of a current frame in grid compression. In one aspect, inter-coding or inter-mode may be applied to encode a base grid of a current frame (also referred to as a current base grid). On the encoder side, at least one reference frame for encoding a current base grid of a current frame may be determined. The current base grid of the current frame may be encoded using an inter-mode based on a corresponding reference base grid of each reference frame in the at least one reference frame. On the decoder side, encoding information may be received by, for example, a decoder, the encoding information indicating that the current base grid of the current frame is encoded using an inter-mode. The encoding information may indicate at least one reference frame for encoding the current base grid of the current frame. The at least one reference frame indicated by the received encoding information may be determined by, for example, a processing circuit included in a decoder. The current base grid of the current frame may be reconstructed based on a corresponding reference base grid of each reference frame in the at least one reference frame, for example, by a processing circuit included in a decoder, and using an inter-mode.

[0094] In one aspect, a reference frame, such as one of at least one reference frame, may be a frame whose display order is before the current frame. If the display index of the current frame is i and the display index of the reference frame is k, then k < i. In one aspect, the display order of one of the at least one reference frame indicated by the display index k may be before the display order of the current frame indicated by the display index i. In an example, k < (i - 1).

[0095] In one aspect, one of the at least one reference frame (indicated by the display index k) and the current frame indicated by the display index i are in the same group of pictures (GoF), and k < i. In an example, a reference frame, such as one of at least one reference frame, may be a frame whose display order is before the current frame, and the reference frame and the current frame are in the same GoF. If the display index of the current frame is i and the display index of the reference frame is k, then k < i, and the two frames (i.e., the current frame and the reference frame) are in the same GoF.

[0096] In one aspect, one of the at least one reference frame indicated by the display index k and the current frame indicated by the display index i are in different groups of pictures (GoF), and k < i. In an example, a reference frame, such as one of at least one reference frame, may be a frame whose display order is before the current frame, and the reference frame and the current frame are in different GoFs. If the display index of the current frame is i and the display index of the reference frame is k, then k < i, and the two frames (i.e., the current frame and the reference frame) are in different GoFs.

[0097] In one aspect, the display order of one of the at least one reference frame indicated by the display index k may be after the display order of the current frame indicated by the display index i. In an example, a reference frame, such as one of the at least one reference frame indicated by the display index k, may be a frame whose display order is after the current frame. If the display index of the current frame is i and the display index of the reference frame is k, then k > i.

[0098] In one aspect, one of the at least one reference frame indicated by the display index k and the current frame indicated by the display index i are in the same group of pictures (GoF), and k > i. For example, the reference frame may be a frame whose display order is after the current frame, and the reference frame and the current frame are in the same GoF. If the display index of the current frame is i and the display index of the reference frame is k, then k > i, and the two frames (i.e., the reference frame and the current frame) are in the same GoF.

[0099] In one aspect, one of the at least one reference frame indicated by display index k and the current frame indicated by display index i are in different frame groups, and k>i. For example, the reference frame may be a frame whose display order is after the current frame, and the reference frame and the current frame are in different GoFs. If the display index of the current frame is i and the display index of the reference frame is k, then k>i, and the two frames (i.e., the reference frame and the current frame) are in different GoFs.

[0100] In one aspect, a reference frame (e.g., one of the at least one reference frame indicated by a display index k) may be selected from a reference frame list. A reference frame list may be created for each current frame (e.g., each current coded frame). In an example, a reference frame index indicated by the coding information may be determined, and one of the at least one reference frame may be determined from the reference frame list for the current frame based on the reference frame index.

[0101] In one aspect, at least one reference frame includes two reference frames, and two reference base meshes are associated with the two reference frames. The encoder can use the two reference frames to generate two predictions (also called predictors) of the current base mesh of the current frame, and use the average of the two predictions as the predicted current base mesh of the current frame. For each vertex of the current frame, two predicted vertices can be identified. The two predictors can then be averaged to generate a final predictor for the current base mesh. In an example, the encoder uses the weighted average of the two predictions as the predicted current base mesh of the current frame.

[0102] In an example, a prediction of the current base grid may be determined based on corresponding reference base grids of two reference frames, and the current base grid may be reconstructed based on an average of the predictions of the current base grid (e.g., two predictors). In an example, the average of the predictions of the current base grid of the current frame is a weighted average of the predictions of the current base grid of the current frame.

[0103] In one aspect, the current base mesh may include at least two current vertices.

[0104] In an example, for a reference base mesh indicated by index j (e.g., in each reference base mesh), a prediction of a current base mesh indicated by index i based on the corresponding reference base mesh may include predictions of at least two current vertices. The encoding information may indicate a motion field of the at least two current vertices. The motion field may be associated with the corresponding reference base mesh and include a motion vector for each of the at least two current vertices. For each of the at least two current vertices, a prediction of the current base mesh indicated by index i based on the corresponding reference base mesh may be determined. C (l, i) indicates the motion vector of each current vertex, and can be based on the motion vector of the corresponding current vertex and the corresponding reference vertex v in the corresponding reference base mesh indicated by index j.R (l,j) to determine the current vertex v C The index i can be the display index of the current frame. The index l of the current vertex can be a non-negative integer. The index j can be the display index of the reference frame m(j). R (l,j) can correspond to v C (l,i) and can share the same index l.

[0105] In an example, for each of the two predictors, a motion vector is signaled to indicate a displacement associated with each current vertex, such as a displacement from the current vertex to the predictor. In an example, at the encoder side, a first vertex is determined based on at least one of: (i) a current vertex v of the current base mesh indicated by index i C (l,i), and (ii) the current vertex v in the current base mesh C At least one adjacent vertex of (l,i). For example, the first vertex is the current vertex v C (l,i) and the current vertex v C Similarly, the second vertex is determined based on at least one of the following: (i) a reference vertex v of the reference base mesh indicated by index j R (l,j), and (ii) the reference vertex v in the reference base mesh indicated by index j R At least one adjacent vertex of (l,j). For example, the second vertex is the reference vertex v R (l,j) and the reference vertex v R In an example, the displacement is determined based on the first vertex and the second vertex. In another example, the displacement is in v C (l,i) and v R As described above, a motion vector indicating a displacement (e.g., a displacement between a first vertex and a second vertex) may be signaled. At the decoder side, the motion vector and the corresponding reference vertex v may be used to determine the displacement. R (l,j) to determine the current vertex v C (l,i) prediction, for example, the current vertex v C The prediction of (l,i) is v R (l,j) plus the shift indicated by the motion vector.

[0106] In one aspect, each of the two predictors associated with the two reference frames may have a motion vector predictor list. Each motion vector predictor list may include at least two candidate vectors. In an example, the number of at least two candidate vectors in the motion vector predictor list is at least 2. In an example, the number of at least two candidate vectors in the motion vector predictor list is greater than 2. The at least two candidate vectors may be determined from adjacent positions in the reference frame. An index of the motion vector predictor list may be signaled to indicate which candidate vector of the at least two candidate vectors may be used to generate the predictor. For example, based on v R (l,j) and the candidate vector selected using the index to determine the current vertex v C The predictor of (l,i).

[0107] In an example, at least two current vertices in a current base mesh share the same motion vector predictor list, and the at least two current vertices may have two motion vector predictor lists associated with two reference frames, respectively. Each current vertex may have an index indicating which candidate vector in the motion vector predictor list is used for the corresponding reference frame. A first current vertex may have a first index different from a second index of a second current vertex.

[0108] In another example, at least two current vertices in the current base mesh may have different motion vector predictor lists. A first current vertex may have a first motion vector predictor list that is different from a second motion vector predictor list of a second current vertex. In an example, the first current vertex may have a first index that is different from a second index of the second current vertex.

[0109] In an example, for each current vertex of the at least two current vertices, a motion vector of each current vertex may be determined based on a motion vector predictor list (eg, one of the motion vector predictor lists described above) and an index indicated by the encoding information. C The prediction of (l,i) can be based on the current vertex v C The motion vector of (l,i) and the corresponding reference vertex v in the corresponding reference base mesh R (l,j) to determine.

[0110] In one aspect, the two reference frames may be determined (eg, selected) from two lists of reference frames.

[0111] Figure 7A flow chart of an overview process (700) according to an embodiment of the present disclosure is shown. The process (700) may be used in an apparatus. The apparatus may include a grid decoder, such as a dynamic grid decoder and a video decoder. The video decoder is configured to, for example, decode a basic grid. In various embodiments, the process (700) is performed by a processing circuit, such as a processing circuit that performs the functions of a video decoder (110), a processing circuit that performs the functions of a video decoder (210), a grid decoder, and the like. In some embodiments, the process (700) is implemented as software instructions, so that when the processing circuit executes the software instructions, the processing circuit performs the process (700). The process starts at (S701) and proceeds to (S710).

[0112] At (S710), encoding information is received, the encoding information indicating that a current basic grid of a current frame is encoded in an inter-frame mode and indicating at least one reference frame used to encode the current basic grid of the current frame.

[0113] In one aspect, a display order of one of the at least one reference frames indicated by display index k precedes a display order of the current frame indicated by display index i. In an example, the one of the at least one reference frame and the current frame are in the same group of frames (GoF). In an example, the one of the at least one reference frame and the current frame are in different groups of frames.

[0114] In the example, k<(i-1).

[0115] In one aspect, a display order of one of the at least one reference frames indicated by display index k is after a display order of the current frame indicated by display index i. In an example, the one of the at least one reference frames and the current frame are in the same group of frames (GoF). In an example, the one of the at least one reference frames and the current frame are in different groups of frames.

[0116] At (S720), the at least one reference frame indicated by the encoding information is determined.

[0117] In an example, a reference frame index indicated by the received encoding information is determined. One of the at least one reference frame in the reference frame list of the current frame is determined based on the reference frame index.

[0118] At (S730), a current basic grid of the current frame is reconstructed using the inter-frame mode based on a corresponding reference basic grid of each of the at least one reference frame.

[0119] Then, the process proceeds to (S799) and ends.

[0120] The process (700) may be adjusted appropriately. At least one step in the process (700) may be modified and / or omitted. At least one additional step may be added. Any suitable order of implementation may be used.

[0121] In one aspect, the at least one reference frame comprises two reference frames. A prediction of the current base grid may be determined based on the respective reference base grids of the two reference frames; and the current base grid is reconstructed based on an average of the predictions of the current base grids. In an example, the average of the predictions of the current base grid of the current frame is a weighted average of the predictions of the current base grid of the current frame.

[0122] In an example, the two reference frames are determined from two reference frame lists.

[0123] In an example, the current base mesh comprises at least two current vertices. For each of the reference base meshes, the prediction of the current base mesh based on the corresponding reference base mesh comprises the prediction of the at least two current vertices. The encoding information indicates a motion field of the at least two current vertices, the motion field being associated with the corresponding reference base mesh and comprising a motion vector for each of the at least two current vertices.

[0124] In an example, for each current vertex of the at least two current vertices, a motion vector of the corresponding current vertex is determined. In an example, the motion vector of the corresponding current vertex is determined based on a motion vector predictor list and an index indicated by the encoding information. The prediction of the current vertex is determined based on the motion vector of the corresponding current vertex and a corresponding reference vertex in the corresponding reference base mesh.

[0125] The above method is applicable when the current base grid of the current frame is encoded in inter-frame mode. The above method is also applicable when the current base grid of the current frame is encoded in SKIP mode. When the current base grid is encoded in SKIP mode, the current base grid m(i) can be directly obtained using the reference base grid m(j) of at least one of the reference frames, for example m(i)=m(j). The above motion field includes a displacement vector of a zero motion vector.

[0126] Figure 8A flow chart of an overview process (800) according to an embodiment of the present disclosure is shown. The process (800) can be used in an apparatus. The apparatus may include a grid encoder, such as a dynamic grid encoder and a video encoder. The video encoder is configured to, for example, encode a basic grid. In various embodiments, the process (800) is performed by a processing circuit, such as a processing circuit that performs the functions of a video encoder (103), a processing circuit that performs the functions of a video decoder (303), a grid encoder, and the like. In some embodiments, the process (800) is implemented as software instructions, so that when the processing circuit executes the software instructions, the processing circuit performs the process (800). The process starts at (S801) and proceeds to (S810).

[0127] At (S810), it is determined that a current base grid of a current frame is encoded in an inter mode.

[0128] At (S820), at least one reference frame is determined to encode a current base grid of the current frame. In one aspect, a display order of one of the at least one reference frame indicated by display index k precedes a display order of the current frame indicated by display index i.

[0129] In one aspect, a display order of a reference frame among the at least one reference frame indicated by display index k follows a display order of the current frame indicated by display index i.

[0130] At (S830), a current basic grid of the current frame is encoded using the inter-frame mode based on a corresponding reference basic grid of each of the at least one reference frame.

[0131] In an example, the at least one reference frame comprises two reference frames. A prediction of the current base grid is determined based on respective reference base grids of the two reference frames. The current base grid is encoded based on an average of the predictions of the current base grid.

[0132] Then, the process proceeds to (S899) and ends.

[0133] The process (800) may be adjusted appropriately. At least one step in the process (800) may be modified and / or omitted. At least one additional step may be added. Any suitable order of implementation may be used.

[0134] In one aspect, a method for processing mesh data includes converting between a mesh data file and a bitstream of the mesh data according to a format rule. For example, the bitstream can be a bitstream decoded / encoded with any decoding and / or encoding method described herein. The format rule can specify at least one constraint of the bitstream and / or at least one process to be performed by a decoder and / or encoder.

[0135] In one aspect, the bitstream includes a first syntax element indicating that a current base grid of a current frame is encoded in an inter-frame mode, and a second syntax element indicating at least one reference frame used to encode the current base grid of the current frame. The format rule specifies that the at least one reference frame is determined based on the second syntax element, and the current base grid of the current frame is reconstructed based on the inter-frame mode and a corresponding reference base grid of each of the at least one reference frame.

[0136] The methods, embodiments and examples in the present disclosure may be used alone or in any order. For example, some embodiments and / or examples performed by a decoder may be performed by an encoder, and vice versa. Each method (or embodiment), encoder and decoder may be implemented by a processing circuit (e.g., at least one processor or at least one integrated circuit). In one example, at least one processor executes a program stored in a non-volatile computer-readable storage medium.

[0137] The above techniques may be implemented as computer software via computer-readable instructions and physically stored in at least one computer-readable storage medium. Fig. 9 A computer system (900) is shown that is suitable for implementing certain embodiments of the disclosed subject matter.

[0138] The computer software may be encoded in any suitable machine code or computer language, and may be assembled, compiled, linked, or the like to create a code comprising instructions, which may be directly executed by at least one computer central processing unit (CPU), graphics processing unit (GPU), or the like, or executed by decoding, microcode, or the like.

[0139] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablets, servers, smartphones, gaming devices, IoT devices, etc.

[0140] Fig. 9 The components shown for the computer system (900) are exemplary and are not intended to impose any limitations on the scope of use or functionality of computer software implementing the disclosed embodiments. Nor should the configuration of the components be interpreted as having any dependency or requirement on any one or combination of components shown in the exemplary embodiment of the computer system (900).

[0141] The computer system (900) may include certain human-computer interface input devices. Such human-computer interface input devices may respond to input from at least one human user through tactile input (e.g., keyboard input, sliding, data glove movement), audio input (e.g., sound, applause), visual input (e.g., gestures), and olfactory input (not shown). The human-computer interface device may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0142] The human-computer interface input device may include at least one of the following (only one of which is drawn): keyboard (901), mouse (902), touchpad (903), touch screen (910), data gloves (not shown), joystick (905), microphone (906), scanner (907), camera (908).

[0143] The computer system (900) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate at least one sense of a human user through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback through a touch screen (910), a data glove (not shown), or a joystick (905), but there may also be tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (909), headphones (not shown)), visual output devices (e.g., screens (910) including cathode ray tube screens, liquid crystal screens, plasma screens, organic light emitting diode screens, each of which has or does not have a touch screen input function, each of which has or does not have a tactile feedback function - some of which may output two-dimensional visual output or output of more than three dimensions through means such as stereoscopic image output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)) and printers (not shown).

[0144] The computer system (900) may also include human-accessible storage devices and their associated storage media, such as optical media including high-density read-only / rewritable optical disks (CD / DVD ROM / RW) (920) with CD / DVD or similar media (921), thumb drives (922), removable hard disk drives or solid state drives (923), traditional magnetic media such as tapes and floppy disks (not shown), special-purpose ROM / ASIC / PLD-based devices such as security software protectors (not shown), and the like.

[0145] Those skilled in the art should also understand that the term "computer-readable storage media" used in connection with the disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0146] The computer system (900) may also include an interface (954) to at least one communication network (955). For example, the network may be wireless, wired, or optical. The network may also be a local area network, a wide area network, a metropolitan area network, an in-vehicle network, an industrial network, a real-time network, a delay-tolerant network, and the like. The network also includes local area networks such as Ethernet, wireless local area networks, cellular networks (GSM, 3G, 4G, 5G, LTE, etc.), television wired or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), in-vehicle and industrial networks (including CANBus), and the like. Some networks typically require an external network interface adapter for connecting to some universal data port or peripheral bus (949) (e.g., a USB port of the computer system (900)); other systems are typically integrated into the core of the computer system (900) by connecting to a system bus as described below (e.g., an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smart phone computer system). By using any of these networks, the computer system (900) can communicate with other entities. The communication can be one-way, for receiving only (e.g., wireless television), one-way for sending only (e.g., CAN bus to certain CAN bus devices), or two-way, such as to other computer systems via a local or wide area digital network. Each of the above networks and network interfaces can use certain protocols and protocol stacks.

[0147] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces may be connected to the core (940) of the computer system (900).

[0148] The core (940) may include at least one central processing unit (CPU) (941), a graphics processing unit (GPU) (942), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (943), a hardware accelerator for a specific task (944), a graphics adapter (950), etc. These devices, as well as a read-only memory (ROM) (945), a random access memory (946), an internal mass storage (e.g., an internal non-user accessible hard disk drive, a solid state drive, etc.) (947), etc., may be connected via a system bus (948). In some computer systems, the system bus (948) may be accessed in the form of at least one physical plug so that it may be expanded by additional central processing units, graphics processing units, etc. Peripheral devices may be directly attached to the system bus (948) of the core, or connected via a peripheral bus (949). In an example, a screen (910) may be connected to a graphics adapter (950). The architecture of the peripheral bus includes a peripheral component interconnect (PCI), a universal serial bus (USB), etc.

[0149] The CPU (941), GPU (942), FPGA (943) and accelerator (944) can execute certain instructions, which can be combined to form the above-mentioned computer code. The computer code can be stored in ROM (945) or RAM (946). Transition data can also be stored in RAM (946), while permanent data can be stored in, for example, internal mass storage (947). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with at least one CPU (941), GPU (942), mass storage (947), ROM (945), RAM (946), etc.

[0150] The computer readable storage medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be specially designed and constructed for the purposes of the present disclosure, or may be the medium and code well known and available to those skilled in the art of computer software.

[0151] As an example and not a limitation, a computer system having an architecture (900), in particular a core (940), can provide the function of executing software contained in at least one tangible computer-readable medium as a processor (including a CPU, a GPU, an FPGA, an accelerator, etc.). Such a computer-readable medium can be a medium associated with the above-mentioned user-accessible mass storage, as well as a specific memory of the core (940) having non-volatility, such as a core internal mass storage (947) or a ROM (945). Software implementing various embodiments of the present disclosure can be stored in such a device and executed by the core (940). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can enable the core (940), in particular the processor therein (including a CPU, a GPU, an FPGA, etc.) to perform a specific process or a specific part of a specific process described herein, including defining a data structure stored in the RAM (946) and modifying such a data structure according to a software-defined process. Additionally or alternatively, the computer system may provide functionality hardwired in logic or otherwise contained in circuitry (e.g., accelerator (944)) that may operate in place of or in conjunction with software to perform specific processes or specific portions of specific processes described herein. Where appropriate, references to software may include logic and vice versa. Where appropriate, references to computer-readable media may include circuitry (e.g., an integrated circuit (IC)) storing execution software, circuitry containing execution logic, or both. The present disclosure includes any suitable combination of hardware and software.

[0152] As used in this disclosure, "at least one" or "one of" is intended to include any one or combination of the listed elements. For example, reference to at least one of A, B, or C, at least one of A, B, and C, at least one of A, B, and / or C, and at least one of A to C is intended to include only A, only B, only C, or any combination thereof. Reference to one of A or B and one of A and B is intended to include A or B or (A and B). Where applicable, the use of "one of" does not exclude any combination of the listed elements, such as when the elements are not mutually exclusive.

[0153] Although the present disclosure has described at least two exemplary embodiments, various changes, arrangements and various equivalent substitutions of the embodiments are within the scope of the present disclosure. Therefore, it should be understood that those skilled in the art can design a variety of systems and methods, which, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure.

Claims

1. A grid decoding device, characterized in that: The device comprises: The processing circuit is configured to: Receiving encoding information, the encoding information indicating that a current base grid of a current frame is encoded in an inter-frame mode and indicating at least one reference frame used to encode the current base grid of the current frame; determining the at least one reference frame indicated by the encoding information; and Based on a corresponding reference base grid of each of the at least one reference frame, a current base grid of the current frame is reconstructed using the inter-frame mode.

2. The device according to claim 1, characterized in that A display order of one of the at least one reference frame indicated by the display index k precedes a display order of the current frame indicated by the display index i.

3. The device according to claim 2, characterized in that k<(i-1).

4. The device according to claim 2, characterized in that The one reference frame of the at least one reference frame and the current frame are in the same frame group GoF.

5. The device according to claim 2, characterized in that The one reference frame of the at least one reference frame and the current frame are in different frame groups.

6. The device according to claim 1, characterized in that A display order of one of the at least one reference frame indicated by the display index k follows a display order of the current frame indicated by the display index i.

7. The device according to claim 1, characterized in that The at least one reference frame comprises two reference frames; and The processing circuit is configured to: determining a prediction of the current base grid based on the corresponding reference base grids of the two reference frames; and The current base grid is reconstructed based on the average value of the predictions of the current base grid.

8. The device according to claim 7, characterized in that The current base mesh includes at least two current vertices; and For each of the reference base meshes, A prediction of the current base mesh based on the corresponding reference base mesh includes a prediction of the at least two current vertices; The encoded information indicates a motion field of the at least two current vertices, the motion field being associated with the corresponding reference base mesh and comprising a motion vector for each of the at least two current vertices; as well as For each current vertex of the at least two current vertices, the processing circuit is configured to: Determine the motion vector corresponding to the current vertex; as well as A prediction of the current vertex is determined based on a motion vector of the corresponding current vertex and a corresponding reference vertex in the corresponding reference base mesh.

9. The device according to claim 7, characterized in that The current base mesh includes at least two current vertices; and For each of the reference base meshes, A prediction of the current base mesh based on the corresponding reference base mesh includes a prediction of the at least two current vertices; The encoded information indicates a motion field of the at least two current vertices, the motion field being associated with the corresponding reference base mesh and comprising a motion vector for each of the at least two current vertices; as well as For each current vertex of the at least two current vertices, the processing circuit is configured to: Determine a motion vector of a corresponding current vertex based on a motion vector predictor list and an index indicated by the encoding information; as well as A prediction of the current vertex is determined based on a motion vector of the corresponding current vertex and a corresponding reference vertex in the corresponding reference base mesh.

10. The device according to claim 7, characterized in that The average value of the predictions of the current base grid of the current frame is a weighted average value of the predictions of the current base grid of the current frame.

11. A method for trellis coding, characterized in that: The method comprises: Determining whether a current base grid of a current frame is encoded in inter-mode; determining at least one reference frame to encode a current base grid of the current frame; and A current base grid of the current frame is encoded using the inter-mode based on a corresponding reference base grid of each of the at least one reference frame.

12. The method according to claim 11, characterized in that A display order of one of the at least one reference frame indicated by the display index k precedes a display order of the current frame indicated by the display index i.

13. The method according to claim 11, characterized in that The at least one reference frame comprises two reference frames; and The method comprises: determining a prediction of the current base grid based on the corresponding reference base grids of the two reference frames; and The current base grid is encoded based on a predicted average value of the current base grid.

14. A method for processing grid data, characterized in that: The method comprises: The bit stream of the grid data is processed according to format rules, wherein The bitstream includes a first syntax element indicating that a current base grid of a current frame is encoded in an inter-frame mode and a second syntax element indicating at least one reference frame used to encode the current base grid of the current frame; and The format rule specifies that the at least one reference frame is determined based on the second syntax element, and a current base grid of the current frame is reconstructed based on the inter mode and a corresponding reference base grid of each of the at least one reference frame.

15. The method according to claim 14, characterized in that A display order of one of the at least one reference frame indicated by the display index k precedes a display order of the current frame indicated by the display index i.