Adaptive linear boost transform

Through the adaptive linear enhancement transformation method, the problem of difficulty in efficiently compressing and transmitting complex 3D grid data in the prior art is solved. Especially in the processing of dynamic grids, effective processing of attribute maps and connectivity information that change over time is realized, and data storage and transmission efficiency is improved.

CN120051801APending Publication Date: 2025-05-27TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004431.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-07-18
Filing Date
2024-07-19
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently compress and transmit complex 3D grid data, especially when processing dynamic grids, and existing standards fail to effectively process property graphs and connectivity information that vary over time.

Method used

Adaptive linear boost transformation method is adopted to determine the predicted value of the vertices in the grid frame through the prediction and update process and encode it in the code stream. This method calculates the updated and predicted values ​​of the vertex based on the value of the current vertex and the distance from adjacent vertices.

Benefits of technology

It realizes efficient compression and transmission of complex 3D grid data, especially in the processing of dynamic grids, which can effectively process attribute diagrams and connectivity information that change over time, improving the storage and transmission efficiency of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051801A_ABST
    Figure CN120051801A_ABST
Patent Text Reader

Abstract

In a mesh decoding method, a code stream is received, the code stream including prediction information of a plurality of vertices in a mesh frame. A predicted value of a current vertex of the plurality of vertices received in the code stream is determined based on (i) a value of the current vertex and (ii) each distance between the current vertex and one or more adjacent vertices of the current vertex in the grid frame. And reconstructing the current vertex based on the predicted value of the current vertex.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Join by reference

[0002] This application claims the benefit of priority to U.S. Patent Application No. 18 / 777,254, filed on July 18, 2024, entitled “Adaptive Linear Lifting Transform,” which claims the benefit of priority to U.S. Provisional Application No. 63 / 527,721, filed on July 19, 2023, entitled “Adaptive Linear Lifting Transform.” The entire disclosure of the prior application is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure describes aspects generally related to trellis coding. Background Art

[0004] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent that the work of the presently named inventors described in this background section and in various aspects of this specification was performed, it does not indicate that it qualifies as prior art at the time of filing, and it is never explicitly or implicitly admitted that it is prior art to the present disclosure.

[0005] Image / video compression can help transmit image / video data between different devices, storages, and networks with minimal quality degradation. In some examples, video codec techniques can compress video based on spatial redundancy and temporal redundancy. In one example, a video codec can use a technique called intra-frame prediction, which can compress an image based on spatial redundancy. For example, intra-frame prediction can use reference data from a current picture being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-frame prediction, which can compress an image based on temporal redundancy. For example, inter-frame prediction can predict samples in a current picture based on a picture previously reconstructed using motion compensation. Motion compensation can be indicated by a motion vector (MV).

[0006] Advances in three-dimensional (3D) capture, modeling, and rendering have facilitated 3D content across a variety of platforms and devices. For example, a baby's first steps can be filmed on one continent, and grandparents can see (and in some cases, interact with) and enjoy a fully immersive experience with the child on another continent. To achieve this sense of realism, models have become more complex, and a large amount of data is associated with the creation and use of these models. 3D meshes are widely used to represent such immersive content. Summary of the invention

[0007] Aspects of the present disclosure include code streams, methods, and apparatus for mesh processing.In some examples, the apparatus for mesh processing includes processing circuitry.

[0008] According to one aspect of the present disclosure, a mesh decoding method is provided. In the method, a code stream is received, the code stream including prediction information of multiple vertices in a mesh frame. Based on (i) the value of a current vertex among the multiple vertices received in the code stream and (ii) each distance between the current vertex and one or more adjacent vertices of the current vertex in the mesh frame, a prediction value of the current vertex is determined. The current vertex is reconstructed based on the prediction value of the current vertex.

[0009] In one aspect, the one or more neighboring vertices of the current vertex include a plurality of neighboring vertices. An updated value of the current vertex is determined. A predicted value of the current vertex is determined to be a sum of (i) the updated value of the current vertex and (ii) an average value of the plurality of neighboring vertices.

[0010] In one aspect, a ratio of each adjacent vertex in a plurality of adjacent vertices is determined. The ratio is equal to the value of the corresponding adjacent vertex divided by the distance between the current vertex and the corresponding adjacent vertex. A sum of the ratios of the plurality of adjacent vertices is determined. A product of (i) a constant and (ii) the sum of the ratios of the plurality of adjacent vertices is determined. An updated value for the current vertex is determined to be the value of the current vertex minus the product of (i) the constant and (ii) the sum of the ratios of the plurality of adjacent vertices.

[0011] In one aspect, the constant is equal to a scalar constant divided by a total number of the plurality of adjacent vertices.

[0012] In one aspect, the grid frame is a selected grid frame in a grid sequence.

[0013] In one aspect, a plurality of vertices are included in a selected area of ​​the mesh frame.

[0014] In one aspect, the plurality of vertices are included in a base layer of a mesh frame.

[0015] In one aspect, the plurality of vertices are included in an enhancement layer of the mesh frame.

[0016] In one aspect, a plurality of vertices are included in a selected frequency band of the mesh frame.

[0017] According to another aspect of the present disclosure, a mesh encoding method is provided. In the method, an initial prediction value of a current vertex is determined based on (i) a value of the current vertex among a plurality of vertices in a mesh frame and (ii) each value of one or more adjacent vertices of the current vertex in the mesh frame. An updated prediction value of the current vertex is determined based on (i) the initial prediction value of the current vertex and (ii) each distance between the current vertex and one or more adjacent vertices of the current vertex in the mesh frame. Based on the updated prediction value of the current vertex, the current vertex is encoded in a bitstream.

[0018] In one aspect, the one or more neighboring vertices of the current vertex include a plurality of neighboring vertices. An initial prediction value of the current vertex is determined as (i) a value of the current vertex minus (ii) an average of the plurality of neighboring vertices.

[0019] In one aspect, a ratio of each adjacent vertex in a plurality of adjacent vertices is determined. The ratio is equal to the value of the corresponding adjacent vertex divided by the distance between the current vertex and the corresponding adjacent vertex. A sum of the ratios of the plurality of adjacent vertices is determined. A product of (i) a constant and (ii) the sum of the ratios of the plurality of adjacent vertices is determined. An updated prediction value for the current vertex is determined to be the sum of the product of an initial prediction value for the current vertex and (i) the constant and (ii) the sum of the ratios of the plurality of adjacent vertices.

[0020] According to another aspect of the present disclosure, a method for processing mesh data is provided. In the method, a code stream of mesh data is processed according to a format rule. The code stream includes prediction information of multiple vertices in a mesh frame. The format rule specifies: based on (i) the value of the current vertex among the multiple vertices received in the code stream and (ii) each distance between the current vertex and one or more adjacent vertices of the current vertex in the mesh frame, determine the predicted value of the current vertex. The format rule specifies: processing the current vertex based on the predicted value of the current vertex.

[0021] Aspects of the present disclosure also provide an apparatus for trellis coding. The apparatus for trellis coding comprises a processing circuit configured to implement any of the described methods for trellis coding.

[0022] Aspects of the present disclosure also provide an apparatus for trellis decoding. The apparatus for trellis decoding comprises a processing circuit configured to implement any of the described methods for trellis decoding.

[0023] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions which, when executed by a computer, cause the computer to perform any of the described methods for grid decoding, grid encoding, and grid data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Further features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0025] Figure 1 is a schematic illustration of an example of a block diagram of a communication system (100).

[0026] Figure 2 is a schematic illustration of an example of a block diagram of a decoder.

[0027] Figure 3 is a schematic illustration of an example of a block diagram of an encoder.

[0028] Figure 4 is a schematic illustration of an example of an encoding process (400) for mesh processing according to an aspect of the present disclosure.

[0029] Figure 5 is a schematic illustration of an example of a pre-processing step (500) according to an aspect of the present disclosure.

[0030] Figure 6 is a schematic illustration of a decoding process (600) for trellis processing according to an aspect of the present disclosure.

[0031] Figure 7 is a schematic illustration of an example of a subdivision scheme according to an aspect of the present disclosure.

[0032] Figure 8 A flow chart outlining a trellis decoding process according to some aspects of the present disclosure is shown.

[0033] Fig. 9 A flow chart outlining a trellis encoding process according to some aspects of the present disclosure is shown.

[0034] Fig.10 is a schematic illustration of a computer system according to an aspect of the present disclosure. DETAILED DESCRIPTION

[0035] Figure 1 A block diagram of a video processing system (100) in some examples is shown. The video processing system (100) is an example of an application of the disclosed subject matter, namely a video encoder and a video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., and the like.

[0036] The video processing system (100) includes an acquisition subsystem (113), which may include a video source (101). The video source (101) may include one or more images acquired by a camera and / or generated by a computer. For example, a digital camera creates, for example, an uncompressed video picture stream (102). In one example, the video picture stream (102) includes samples taken by the digital camera. The video picture stream (102), which is depicted as a thick line to emphasize the high amount of data compared to the encoded video data (104) (or the encoded video bitstream), can be processed by an electronic device (120), which includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. The encoded video data (104) (or encoded video bitstream), depicted as thin lines to emphasize the lower amount of data compared to the video picture stream (102), can be stored on the streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1 A client subsystem (106) and a client subsystem (108) in a streaming server (105) may access a copy (107) and a copy (109) of the encoded video data (104). The client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and creates an output video picture stream (111) that can be presented on a display (112) (e.g., a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), the video data (107), and the video data (109) (e.g., a video bitstream) may be encoded according to certain video encoding / compression standards. Examples of these standards include ITU-T H.265. In one example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of the VVC standard.

[0037] It should be noted that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may also include a video encoder (not shown).

[0038] Figure 2An example of a block diagram of a video decoder (210) is shown. The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used to replace Figure 1 A video decoder (110) in an example of FIG.

[0039] The receiver (231) may receive one or more encoded video sequences, for example, included in a bitstream to be decoded by the video decoder (210). In one aspect, the encoded video sequences are received one at a time, wherein the decoding of each encoded video sequence is independent of the decoding of the other encoded video sequences. The encoded video sequence may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data and other data, for example, encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not depicted). The receiver (231) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory (215) may be located external to the video decoder (210) (not depicted). In still other applications, a buffer memory (not depicted) may be provided external to the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) may be provided internal to the video decoder (210) to, for example, handle playout timing. When the receiver (231) receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (215) may not be required, or the buffer memory (215) may be made smaller. For use on a traffic packet network such as the Internet, the buffer memory (215) may be required, and the buffer memory (215) may be relatively large, advantageously may have an adaptive size, and may be implemented at least partially in an operating system or similar element (not depicted) external to the video decoder (210).

[0040] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. The categories of symbols include information for managing the operation of the video decoder (210) and potential information for controlling a display device such as a display device (212) (e.g., a display screen) that is not part of the electronic device (230) but can be coupled to the electronic device (230), such as Figure 2As shown. The control information for the display device may be a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not depicted). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may be performed according to a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract a subgroup parameter set for at least one subgroup of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (220) may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the coded video sequence.

[0041] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).

[0042] Depending on the type of coded video picture or part of coded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of symbol (221) may involve multiple different units. Which units are involved and how they are involved may be controlled by subgroup control information parsed by parser (220) from the coded video sequence. For clarity, such subgroup control information flow between parser (220) and the multiple units below is not depicted.

[0043] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into a number of functional units as described below. In actual implementations operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purposes of describing the disclosed subject matter, the conceptual subdivision into the following multiple functional units is appropriate.

[0044] The first unit is a sealer / inverse transform unit (251). The sealer / inverse transform unit (251) receives quantized transform coefficients as symbols (221) from the parser (220) and control information including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit (251) may output a block including sample values, which may be input into an aggregator (255).

[0045] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses surrounding reconstructed information extracted from a current picture buffer (258) to generate a block of the same size and shape as the block being reconstructed. For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a per-sample basis.

[0046] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to a block that is inter-coded and potentially motion compensated. In this case, the motion compensated prediction unit (253) may access the reference picture memory (257) to extract samples for prediction. After the extracted samples are motion compensated according to the symbols (221) belonging to the block, these samples may be added to the output of the scaler / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) by the aggregator (255), thereby generating output sample information. The extraction of prediction samples by the motion compensated prediction unit (253) from the address in the reference picture memory (257) may be controlled by a motion vector, which may be provided to the motion compensated prediction unit (253) in the form of symbols (221), which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values ​​extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.

[0047] The output samples of the aggregator (255) may be employed by various loop filtering techniques in a loop filter unit (256). The video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video sequence (also referred to as a coded video bitstream) and available to the loop filter unit (256) as symbols (221) from the parser (220). The video compression may also be responsive to meta-information obtained during decoding of a coded picture or a previous (in decoding order) portion of the coded video sequence, and to previously reconstructed and loop filtered sample values.

[0048] The output of the loop filter unit (256) may be a sample stream that may be output to a display device (212) and stored in a reference picture memory (257) for future inter-picture prediction.

[0049] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed, and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) may become part of the reference picture memory (257), and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.

[0050] The video decoder (210) may perform decoding operations according to a predetermined video compression technology or standard such as ITU-T H.265. The encoded video sequence may conform to the syntax specified by the video compression technology or standard used in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as the only tools available for use under the profile. For compliance, it is also required that the complexity of the encoded video sequence is within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata of the HRD buffer management signaled in the encoded video sequence.

[0051] In one aspect, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0052] Figure 3 An example of a block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used to replace Figure 1 A video encoder (103) in an example.

[0053] The video encoder (303) can be used to obtain the video source (301) (not Figure 3 In an example, the video source (301) is a part of the electronic device (320) to receive video samples, and the video source (301) can capture video images to be encoded by the video encoder (303). In another example, the video source (301) is a part of the electronic device (320).

[0054] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), the digital video sample stream may have any suitable bit depth (e.g. 8-bit, 10-bit, 12-bit, ...), any color space (e.g. BT.601 Y CrCB, RGB, ...) and any suitable sampling structure (e.g. Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The pictures themselves may be organized into a spatial pixel array, where each pixel may include one or more samples, depending on the sampling structure, color space, etc. used. The following description focuses on the samples.

[0055] According to one aspect, the video encoder (303) can encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other time constraints required. Implementing the appropriate encoding speed is a function of the controller (350). In some aspects, the controller (350) controls other functional units as described below and is functionally coupled to these other functional units. For clarity, the coupling is not depicted in the figure. The parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization techniques, etc.), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions related to the video encoder (303) optimized for a certain system design.

[0056] In some aspects, the video encoder (303) is configured to operate in a coding loop. As a simple description, in one example, the coding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Since the decoding of the symbol stream produces a bit-accurate result that is independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-accurately corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values ​​that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, such as due to channel errors) is also used in some related technologies.

[0057] The operation of the "local" decoder (333) may be similar to that described above in conjunction with Figure 2 The "remote" decoder of the video decoder (210) described in detail is identical. However, additional brief reference is made to Figure 2 , since the symbols are available and the entropy encoder (345) and the parser (220) can losslessly encode / decode the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (210) including the buffer memory (215) and the parser (220) may not be fully implemented in the local decoder (333).

[0058] On the one hand, except for the parsing / entropy decoding present in the decoder, the decoder technology is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on the decoder operation. The description of the encoder technology can be simplified because the encoder technology is mutually inverse to the decoder technology described comprehensively. In some areas, a more detailed description is provided below.

[0059] During operation, in some examples, the source encoder (330) may perform motion compensated predictive coding that predictively encodes an input picture by referencing one or more previously encoded pictures from a video sequence designated as "reference pictures." In this manner, the encoding engine (332) encodes the differences between pixel blocks of the input picture and pixel blocks of a reference picture that may be selected as a prediction reference for the input picture.

[0060] The local video decoder (333) may decode the encoded video data of the picture that may be designated as the reference picture based on the symbol created by the source encoder (330). The operation of the encoding engine (332) may advantageously be a lossy process. When the encoded video data may be decoded at the video decoder ( Figure 3 When the video sequence is decoded at a remote video decoder (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture memory (334). In this way, the video encoder (303) may store a copy of the reconstructed reference picture locally that has common content (absent transmission errors) with the reconstructed reference picture to be obtained by the remote video decoder.

[0061] The predictor (335) may perform a prediction search for the encoding engine (332). That is, for a new picture (or grid) to be encoded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may be used as appropriate prediction references for the new picture. The predictor (335) may operate on a pixel block by pixel block basis to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (335), it may be determined that the input picture may have prediction references taken from a plurality of reference pictures stored in the reference picture memory (334).

[0062] In one example, the mesh position quantization step size is defined based on a first parameter (e.g., mesh_position_quantization_step_size_log2_denominator) and a second parameter (e.g., mesh_position_quantization_step_size_numerator_minus1). In one example, mesh_position_quantization_step_size_log2_denominator is the logarithm (log2) of the denominator of the position quantization step size to the base 2, and its value is between 0 and 7 (including 0 and 7). In one example, mesh_position_quantization_step_size_numerator_minus1 plus 1 is the value of the numerator of the position quantization step size, and the value is between 1 and 256 (including 1 and 256). After the position encoding in the MEB, the reconstructed position value used for reference is scaled to the dynamic range of the inter-frame.

[0063] An example of the grid attribute encoding parameter syntax is as follows:

[0064]

[0065]

[0066] An example of the semantics of mesh attribute encoding parameters is defined by a parameter such as mesh_attribute_quantization_step_size_log2_denominator[i]. In one example, mesh_attribute_quantization_step_size_log2_denominator[i] is the log2 of the denominator of the i-th attribute quantization step size, and its value is between 0 and 7 (inclusive).

[0067] An example of a generic base grid sequence parameter set RBSP syntax is as follows:

[0068]

[0069]

[0070] An example of motion field quantization is defined by a first parameter (e.g., bmsps_inter_quantization_step_size_log2_denominator) and a second parameter (e.g., bmsps_inter_quantization_step_size_numerator_minus1). In one example, bmsps_inter_quantization_step_size_log2_denominator is the log2 of the denominator of the motion field quantization step size, and its value is between 0 and 7 (including 0 and 7). In one example, bmsps_inter_quantization_step_size_numerator_minus1 plus 1 is the value of the numerator of the motion field quantization step size, and the value is between 1 and 256 (including 1 and 256). After motion field encoding, the position value is reconstructed and can be used as a reference for future vertices. In one example, mesh attribute quantization is defined by a step size parameter (e.g., mesh_attribute_quantization_step_size_numerator[i]). In one example, mesh_attribute_quantization_step_size_numerator[i] is the value of the numerator of the i-th attribute quantization step size, which is between 0 and 255 (including 0 and 255). At the same time, the quantized prediction residual is dequantized. The reconstructed value is calculated by adding the dequantized prediction residual and the predicted value. The reconstructed value can be used as a reference for future vertices.

[0071] The controller (350) may manage encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.

[0072] The outputs of all the above functional units may be entropy encoded in an entropy encoder (345). The entropy encoder (345) applies lossless compression to the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into a coded video sequence.

[0073] The transmitter (340) may buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission over a communication channel (360), which may be a hardware / software link to a storage device that may store the encoded video data. The transmitter (340) may combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or ancillary data streams (source not shown).

[0074] The controller (350) may manage the operation of the video encoder (303). During encoding, the controller (350) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following picture types:

[0075] Intra pictures (I pictures) that can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures.

[0076] Predictive pictures (P pictures), which can be encoded and decoded using intra prediction or inter prediction, which uses a motion vector and reference index to predict the sample values ​​of each block.

[0077] Bidirectional predictive pictures (B pictures), which can be encoded and decoded using intra prediction or inter prediction, which uses two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.

[0078] The source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined according to the coding allocation applied to the block's corresponding picture. For example, a block of an I picture may be non-predictively coded, or the block may be predictively coded (spatial prediction or intra prediction) with reference to an already coded block of the same picture. A block of a P picture may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. A block of a B picture may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.

[0079] The video encoder (303) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0080] In one aspect, the transmitter (340) may transmit additional data when transmitting the encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0081] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In one example, a particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.

[0082] In some aspects, a bidirectional prediction technique may be used for inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture that precede the current picture in the video in decoding order (but may be in the past and future, respectively, in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.

[0083] In addition, merge mode technology can be used for inter-picture prediction to improve coding efficiency.

[0084] According to some aspects of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks (e.g., polygonal blocks or triangle blocks). For example, according to the High-Efficiency Video Coding (HEVC) standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU includes three coding tree blocks (CTBs), which are a luminance CTB and two chrominance CTBs. Each CTU can be recursively divided into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be divided into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type for the CU, such as an inter prediction type or an intra prediction type. The CU is divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Typically, each PU includes a luma prediction block (PB) and two chroma PBs. On the one hand, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Taking the luma prediction block as an example, the prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0085] It should be noted that the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using any suitable technology. In one aspect, the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using one or more integrated circuits. In another aspect, the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using one or more processors executing software instructions.

[0086] Aspects of the present disclosure include methods and systems for adaptive linear transforms, such as adaptive linear lifting transforms.

[0087] A mesh may include multiple polygons that describe the surface of a volumetric object. Each polygon of a mesh may be defined by the vertices of the corresponding polygon in a three-dimensional (3D) space and information about how the vertices are connected (which may be referred to as connectivity information). In some aspects, vertex attributes (e.g., color, normal, etc.) may be associated with vertices (or mesh vertices). Attributes (or vertex attributes) may also be associated with the surface of a mesh by utilizing mapping information that uses a two-dimensional (2D) attribute map to parameterize the mesh. This mapping may be described by a set of parameter coordinates (referred to as UV coordinates or texture coordinates) associated with mesh vertices. A 2D attribute map may be used to store high-resolution attribute information, such as textures, normals, displacements, etc. High-resolution attribute information may be used for various purposes, such as texture mapping and shading.

[0088] Dynamic mesh sequences may require a large amount of data because dynamic meshes may include a large amount of information that changes over time. Therefore, efficient compression techniques can be used to store and transmit such content. Mesh compression standards (e.g., Information and Communication (IC) mesh compression, MESHGRID, frame-based animated mesh compression (FAMC)) were previously developed by the Moving Picture Experts Group (MPEG) to handle dynamic meshes with constant connectivity, time-varying geometry and vertex attributes. However, these standards do not consider time-varying attribute graphs and connectivity information. DCC (Digital Content Creation) tools can generate such dynamic meshes. However, for volume acquisition technology, it may be challenging to generate dynamic meshes with constant connectivity, especially to generate dynamic meshes with constant connectivity under real-time constraints. Existing standards may not support this type of content (e.g., dynamic meshes with constant connectivity). The present disclosure includes aspects of a new mesh compression standard that can directly handle dynamic meshes with time-varying connectivity information and optional time-varying attribute graphs. Mesh compression can target lossy and lossless compression for various applications, such as real-time communication, storage, free viewpoint video, augmented reality (AR) and virtual reality (VR). Features such as random access and scalable / progressive coding can also be considered.

[0089] The mesh geometry information may include vertex connectivity information, 3D coordinates, 2D texture coordinates, etc. The 3D vertex coordinates and 2D texture coordinates occupy a large portion of the mesh geometry information. Therefore, it is necessary to compress the 3D vertex coordinates (also called vertex positions) and 2D texture coordinates to reduce the amount of data required to store and / or transmit the mesh geometry information.

[0090] Figure 4 An example of an encoding process (400) for mesh processing based on a related video codec (e.g., MPEG V-Mesh TMv1.0) according to an aspect of the present disclosure is shown. Figure 4 As shown, the encoding process (400) may include a preprocessing step (400A) and an encoding step (400B). The preprocessing step (400A) may be configured to generate a base grid m(i) of the current frame and a displacement field d(i) of the current frame based on an input grid M(i) of the current frame, wherein the displacement field d(i) of the current frame includes a displacement vector. The encoding step (400B) may be configured to encode the base grid m(i), the displacement field d(i), and the texture information of the base grid m(i). The displacement field d(i) of the current frame may include a displacement vector. The index i may refer to the current frame. In one aspect, a mode decision method may be performed in the encoding process (400) to determine whether to apply inter-frame coding (also referred to as inter-frame prediction or inter-frame mode), intra-frame coding (also referred to as intra-frame prediction or intra-frame mode), etc. to the current frame. For example, the mode decision method may compare the cost of the intra-frame mode with the cost of the inter-frame mode, and determine the encoding mode of the base grid m(i) of the current frame based on which cost is smaller. In some examples, the base grid m(i) is encoded (e.g., encoded or decoded) using a skip mode. In one example, the skip mode is a special mode of the inter-frame mode. For example, the base grid m(i) can be intra-coded, or inter-coded, or encoded using a skip (SKIP) mode.

[0091] Still reference Figure 4, the preprocessing step (400A) may include a mesh extraction process (402), a parameterization process such as an atlas parameterization process (404), and a subdivision surface fitting process (406). The mesh extraction process (402) is configured to downsample the vertices of the input mesh M(i) to generate an extracted mesh dm(i) that may include a plurality of extracted (or downsampled) vertices. The number of the plurality of extracted vertices is less than the number of vertices of the input mesh M(i). The parameterization process such as the atlas parameterization process (404) is configured to map the extracted mesh dm(i) onto a planar domain, such as onto a UV atlas (or UV map), to generate a re-parameterized mesh pm(i). In one example, atlas parameterization can be performed based on a video processing tool (e.g., a UVAtlas tool). The subdivision surface fitting process (406) is configured to take the re-parameterized mesh pm(i) and the input mesh M(i) as input, and generate a base mesh m(i) and a displacement field d(i) including a displacement vector or a displacement set. In an example of the subdivision surface fitting process (406), pm(i) is subdivided using a subdivision scheme such as iterative interpolation to obtain a subdivided mesh. The iterative interpolation includes inserting a new point in the middle of each edge of the re-parameterized mesh pm(i) at each iteration. Any suitable subdivision scheme can be applied to subdivide pm(i). The displacement field d(i) is calculated by determining the nearest point on the surface of the input mesh M(i) for each vertex of the subdivided mesh.

[0092] Advantages of the subdivided mesh may include: the subdivided mesh has a subdivision structure that allows efficient compression while providing a reliable approximation to the input mesh. The improved compression efficiency may be obtained due to the following properties. The decimated mesh dm(i) may have a small number of vertices and may be encoded and transmitted using fewer bits than the input mesh M(i) or the number of subdivided meshes. Figure 4 , a base mesh m(i) can be generated based on the extracted mesh dm(i). In one example, the base mesh m(i) is the extracted mesh dm(i). Since the subdivided mesh can be generated based on the subdivision method, the subdivided mesh can be automatically generated by the decoder when the base mesh or the extracted mesh is decoded (for example, without using any information other than the subdivision scheme and the subdivision iteration count). On the decoder side, the displacement field d(i) can be generated by decoding the displacement vectors associated with the vertices of the subdivided mesh. In addition to enabling spatial / quality scalability, the subdivision structure enables efficient transformations such as wavelet decomposition, which can provide high compression performance.

[0093] For simplicity, the preprocessing step (400A) that can be applied to an input mesh (e.g., a 3D mesh) can be described using the preprocessing step (500) applied to a two-dimensional (2D) curve. The preprocessing step (400A) is similar to the preprocessing step (500), except that the 3D mesh can be replaced by a 2D curve.

[0094] Figure 5 An example of a preprocessing step (500) according to an aspect of the present disclosure is shown. Figure 5 As shown, an input 2D curve (represented by a 2D polyline) (502) can be downsampled to generate a base curve such as a polyline, which is referred to as a "decimated" curve (504). Then, a subdivision scheme can be applied to the decimated polyline (504) to generate a "decision" curve (506). In one example, the subdivision scheme can be an iterative interpolation scheme. The iterative interpolation scheme can include inserting a new point in the middle of each edge of the polyline (or decimated curve) (504) at each iteration. For example, a point (510) can be inserted in an edge (508) of the decimated curve (504). In one example, the edge (508) is located between a point (512) and a point (514). In addition, a point (522) can be added between the point (512) and the point (510), and a point (516) can be added between the point (510) and the point (514). Then, the decimal polyline (506) is deformed to generate a displacement curve (518). The displacement curve (518) may be a better approximation of the input curve (502) than the subdivided curve (506). For example, a displacement vector (e.g., (520)) is calculated for each vertex (e.g., (510)) of the subdivided curve (506) so that the shape of the displacement curve (518) is as close as possible to the shape of the input curve (502). An advantage of the subdivided curve (506) is that the subdivided curve (506) has a subdivision structure that allows for more efficient compression while providing a reliable approximation of the input curve (502).

[0095] The decimation curve (504) may have a smaller number of points and may be encoded and transmitted using a limited number of bits. Since the subdivision curve may be generated based on the subdivision scheme, the subdivision curve may be automatically generated by the decoder when decoding the base curve or the decimation curve (e.g., without using any information other than the subdivision method and the subdivision iteration count). The displacement curve may be generated by decoding the displacement vectors associated with the vertices of the subdivision curve. In addition to enabling spatial / quality scalability, the subdivision structure enables efficient transforms such as wavelet decomposition, which may provide high compression performance.

[0096] Still reference Figure 5In one example, the input mesh M(i) may include an input 2D curve (502). The base mesh m(i) may include a decimated curve (504), which is formed by downsampling the vertices of the input 2D curve (502). The displacement field dm(i) may include a plurality of displacement vectors, such as Figure 5 The displacement vector (520) is shown.

[0097] The encoding step (400B) may include base mesh encoding (408), displacement encoding (410), texture encoding (412), etc. The base mesh encoding (408) is configured to encode geometric information of a base mesh m(i) associated with a current frame. In intra-frame coding, the base mesh m(i) may first be quantized (e.g., quantized using uniform quantization) and then encoded, for example, by using a coding mode determined by a mode decision method. The coding mode may be an inter-frame mode, an intra-frame mode, a skip mode, etc. An encoder for intra-coding the base mesh m(i) may be referred to as a static mesh encoder. In inter-frame coding, a reference base mesh associated with a reference frame indicated by index j (e.g., a reconstructed quantized reference base mesh m'(j)) may be used to predict a base mesh m(i) associated with a current frame indicated by index i. The displacement encoding (410) is configured to encode a displacement field d(i) generated in the preprocessing step (400A). The displacement field d(i) may include a set of displacement vectors (or displacements) associated with vertices of a subdivided mesh. The texture encoding (412) is configured to encode attribute information of the base mesh m(i). The attribute information may include texture, normal, color, etc. The attribute information may be encoded based on a suitable codec (e.g., High Efficiency Video Coding (HEVC) or Next Generation Video Coding (VVC)).

[0098] On the one hand, reference Figure 4 , a mesh encoding process such as the encoding process (400) begins with preprocessing (e.g., preprocessing step (400A)). The preprocessing may convert an input mesh (e.g., an input dynamic mesh) M(i) into a base mesh m(i) and a displacement field d(i) including a displacement set (or a displacement vector set). The encoding step (400B) may compress the output from the preprocessing (e.g., m(i), d(i), etc.) and generate a compressed code stream b(i). The compressed code stream b(i) may include a compressed base mesh code stream, a compressed displacement field code stream, a compressed attribute code stream, etc.

[0099] Figure 6An example of a decoding process (600) for grid processing according to an aspect of the present disclosure is shown. The decoding process (600) may include a decoding step (605) and a post-processing step (610). A compressed code stream b(i) may be provided to the decoding step (605). In one example, for example for lossless transmission, the compressed code stream b(i) is the output b(i) from the encoding process (400). The decoding step (605) may extract various sub-code streams, such as a compressed base grid sub-stream, a compressed displacement field sub-stream, a compressed attribute sub-stream, etc. The decoding step (605) may decompress the sub-code streams to generate the following components: block (patch) metadata indicated by metadata (i) (metadata(i)), decoded base grid m" (i), decoded displacement field (including displacement) d" (i), decoded attribute map A" (i), etc.

[0100] On the one hand, the base grid substream may be provided to a grid decoder to generate a reconstructed quantized base grid m'(i). The decoded base grid (or reconstructed base grid) m"(i) may be obtained by applying inverse quantization to m'(i). The displacement field substream including encoded, packed and quantized wavelet coefficients may be decoded by a video and / or image decoder. Image unpacking and inverse quantization may be applied to the reconstructed, packed quantized wavelet coefficients to obtain unpacked and dequantized transform coefficients (e.g., wavelet coefficients). An inverse wavelet transform may be applied to the unpacked and dequantized wavelet coefficients to generate a decoded displacement field (or reconstructed displacement) d"(i).

[0101] The decoded components (e.g., including metadata(i), m"(i), d"(i), A"(i), etc.) can be provided to a post-processing step (610). A mesh (also referred to as a decoded / reconstructed mesh) M"(i) can be generated based on m"(i) and d"(i) by the post-processing step (610). In one example, the mesh M"(i) (also referred to as a reconstructed deformed mesh DM(i)) can be obtained by subdividing m"(i) using a subdivision scheme and applying reconstructed displacements d"(i) to the vertices of the subdivided mesh. In one example, DM(i) can include a displacement curve (518). In one example, when the encoding process (400), the decoding process (600), and the transmission are lossless, the mesh M"(i) can be the same as the input mesh M(i). When one of the encoding process (400), the decoding process (600), and the transmission is lossy, M"(i) is different from M(i). In various examples, the difference between M"(i) and M(i), if any, can be relatively small. In one example, a property graph A”(i) is also generated by a post-processing step (610).

[0102] In one aspect, the base grid may be intra-coded, inter-coded, or encoded using a SKIP mode, etc. In one example, the SKIP mode may be a special mode of the inter-mode, in which the base grid m(i) of the current frame indicated by index i is the same as the base grid m(j) of the reference frame indicated by index (also referred to as frame index) j. When the inter-mode is applied to encode the base grid in the current frame, the encoder may generate a predicted base grid for the current frame based on the reconstructed base grid of the reference frame. In one example, such as in MPEG V-DMC WD 2.0, the reference frame is a frame immediately preceding the current frame in display order. The frame index i of the current frame indicates the display order. When the frame index of the current frame is i, the frame index of the reference frame is (i-1). In one example, the current frame and the reference frame are in the same group of frames (GoF).

[0103] In one aspect, for example in MPEG V-DMC WD 2.0, the mesh encoding process starts with preprocessing. The preprocessing may convert the input dynamic mesh (denoted as M(i)) into a base mesh m(i) and a set of displacements d(i). The encoder may compress the base mesh m(i) and the displacements d(i) to generate a compressed code stream b(i).

[0104] like Figure 4 As shown, preprocessing may include mesh extraction, atlas parameterization, and subdivision surface fitting. Mesh extraction may use simplification techniques to extract the input mesh M(i) and generate an extracted mesh dm(i). The extracted mesh dm(i) may be re-parameterized. The generated mesh may be denoted as pm(i). Subdivision surface fitting may take the re-parameterized mesh pm(i) and the input mesh M(i) as input to generate a base mesh m(i) and a displacement set d(i).

[0105] In one aspect, pm(i) is subdivided. An example of a subdivision of pm(i) is Figure 7 As shown in Figure 7 As shown, the re-parameterized mesh pm(i) (702) can be subdivided by applying a subdivision method (e.g., a midpoint subdivision scheme). The midpoint subdivision scheme can subdivide each triangle into 4 sub-triangles in each subdivision iteration, such as Figure 7 For example, in the initial iteration S 0 In the first iteration S, the re-parameterized mesh pm(i) (702) may include two triangles (701) and (703). 1 In the second iteration S, multiple vertices, such as vertex (704) and vertex (706), can be generated according to the midpoint subdivision scheme. Accordingly, triangle (701) is subdivided into 4 small triangles, and triangle (703) is also subdivided into 4 small triangles. 2In the example, multiple vertices, such as vertex (708) and vertex (710), may be generated according to the midpoint subdivision scheme. 1 Each triangle formed in can be 2 Subdivided into 4 triangles.

[0106] The displacement field d(i) may be computed by determining, for each vertex of the subdivided mesh, the closest point on the surface of the original (or input) mesh M(i).

[0107] Optionally, the encoder can encode a set of displacement vectors associated with the subdivided mesh vertices, referred to as the displacement field d(i). In one example, the reconstructed quantized base mesh m'(i) can be used to update the displacement field d(i) to generate an updated displacement field d'(i). A wavelet transform can then be applied to d'(i) and a set of wavelet coefficients can be generated. The wavelet coefficients can then be quantized and packed into a 2D image / video. The quantized and packed wavelet coefficients can be compressed using arithmetic coding, a conventional image / video encoder, or any other encoder.

[0108] In one aspect, the wavelet transform (eg, in MPEG V-DMC 2.0) is a linear lifting transform.The linear lifting transform may include a prediction process and an update process.

[0109] In one example, the prediction process is defined in equation (1):

[0110]

[0111] Where v is the edge (v 1 ,v 2 ) is introduced (or located) in the middle of the signal. Signal(v), Signal(v 1 ) and Signal(v 2 ) indicate vertex v and vertex v respectively. 1 and vertex v 2 The value of the geometry / vertex attribute signal at .

[0112] In one example, the update process is defined in equation (2):

[0113]

[0114] where v * is the set of adjacent vertices of vertex v.

[0115] In the present disclosure, methods and systems for adaptive linear lifting transforms for vertex position or displacement vector encoding in mesh compression are provided. These methods and systems can be applied alone or in any combination. In addition, these methods and systems are not limited to mesh compression. These methods and systems can also be applied, for example, to audio processing, image processing, video processing, or signal processing.

[0116] In the present disclosure, an adaptive transformation is provided, such as an adaptive linear lifting transformation. The adaptive linear lifting transformation may include a prediction process and an update process. In one aspect, the weight in the update process is based on the distance between a vertex and its neighboring vertices.

[0117] On the one hand, for example at the encoder side, the prediction process of the adaptive linear lifting transform (or forward adaptive linear lifting transform) is defined in equation (3):

[0118]

[0119] Where v is located on the edge (v 1 ,v 2 ) is the vertex in the middle. Signal(v), Signal(v 1 ) and Signal(v 2 ) indicate vertex v and vertex v respectively. 1 and vertex v 2 In one example, Signal(v), Signal(v 1 ) and Signal(v 2 ) indicate vertex v and vertex v respectively. 1 and vertex v 2 The value of the geometry / vertex attribute signal (or geometry / vertex attribute information) at.

[0120] As shown in equation (3), for the prediction process of the adaptive linear lifting transform on the encoder side, the initial prediction value of the current vertex is determined as (i) the value of the current vertex (e.g., the initial value) minus (ii) the average value of multiple adjacent vertices. In the example of equation (3), the multiple adjacent vertices include two adjacent vertices v 1 and v 2 .

[0121] On the one hand, for example at the encoder side, the update process of the adaptive linear lifting transform is defined in equation (4):

[0122]

[0123] where v *is the set of adjacent vertices of vertex v, dist(v,w) is the distance function between vertex v and vertex w, and c is a scalar constant. The distance function provides the distance between vertex v and the set of adjacent vertices v * Signal(v) and Signal(w) are the values ​​of vertex v and vertex w, respectively, for example, the values ​​of the geometry / vertex attribute signals (or geometry / vertex attribute information) at vertex v and vertex w, respectively.

[0124] As shown in equation (4), multiple adjacent vertices (e.g., v * ). The ratio is equal to the value of the corresponding adjacent vertex (e.g., w) divided by the distance between the current vertex (e.g., v) and the corresponding adjacent vertex. The sum of the ratios of multiple adjacent vertices is determined. The product of (i) a constant (e.g., c) and (ii) the sum of the ratios of multiple adjacent vertices is determined. An updated prediction value for the current vertex is determined to be the sum of the product of the value of the current vertex (e.g., the initial prediction value obtained in equation (3)) and (i) the constant and (ii) the sum of the ratios of multiple adjacent vertices.

[0125] In the present disclosure, an inverse adaptive transform, such as an inverse adaptive linear lifting transform, is provided. The inverse adaptive linear lifting transform may include an updating process and a prediction process, wherein the updating process comes before the prediction process.

[0126] On the one hand, for example at the decoder side, the update process in the inverse adaptive transform is defined in equation (5):

[0127]

[0128] where v * is the set of adjacent vertices of vertex v, dist(v,w) is the distance function between vertex v and vertex w, and c is a scalar constant. The distance function provides the distance between vertex v and the set of adjacent vertices v * The distance between each vertex w in .

[0129] As shown in equation (5), determine the multiple adjacent vertices (e.g., v * ). The ratio is equal to the value of the corresponding adjacent vertex (e.g., w) divided by the distance between the current vertex (e.g., v) and the corresponding adjacent vertex (e.g., w). The sum of the ratios of multiple adjacent vertices is determined. The product of (i) a constant (e.g., c) and (ii) the sum of the ratios of multiple adjacent vertices is determined. An updated value for the current vertex is determined to be the value of the current vertex minus the product of (i) the constant and (ii) the sum of the ratios of multiple adjacent vertices. In one aspect, the value of the current vertex is the updated prediction value obtained in equation (4) and signaled to the decoder.

[0130] On the one hand, for example at the decoder side, the prediction process in the inverse adaptive linear lifting transform is defined in equation (6):

[0131]

[0132] Where v is located on the edge (v 1 ,v 2 ) is the vertex in the middle. Signal(v), Signal(v 1 ) and Signal(v 2 ) are vertex v and vertex v respectively. 1 and vertex v 2 The values ​​of, for example, vertex v, vertex v 1 and vertex v 2 The value of the geometry / vertex attribute signal (or geometry / vertex attribute information) at.

[0133] As shown in equation (6), for the prediction process of the inverse adaptive linear lifting transform on the decoder side, the prediction value of the current vertex is determined as (i) the value of the current vertex (e.g., the updated value obtained in equation (5)) minus (ii) the average value of multiple adjacent vertices. In the example of equation (6), the multiple adjacent vertices include two adjacent vertices.

[0134] On the one hand, the distance function shown in equation (4) and equation (5) is l 1 Distance (e.g., one-dimensional distance). In one aspect, the distance function is l 2 Distance (e.g., two-dimensional distance). In one aspect, the distance function is l 0 Distance, for two vertices, l 0 The distance function is equal to 2 (e.g., always equal to 2). In one aspect, the distance function is l p Distance (eg, p-dimensional distance), where p>=0.

[0135] On the one hand, the update process in the (forward) adaptive linear lifting transform is defined in Equation (7):

[0136]

[0137] where v * is the set of adjacent vertices of vertex v, dist(v,w) is the distance function between vertex v and vertex w, c is a scalar constant, |v * |v * The cardinality of the set v * Signal(v) and Signal(w) are the values ​​of vertex v and vertex w, respectively, for example, the values ​​of the geometry / vertex attribute signals (or geometry / vertex attribute information) at vertex v and vertex w, respectively.

[0138] On the one hand, the update process in the inverse adaptive linear lifting transform is defined in equation (8):

[0139]

[0140] where v * is the set of adjacent vertices of vertex v, dist(v,w) is the distance function between vertex v and vertex w, c is a scalar constant, |v * |v * The cardinality of the set v * The total number of vertices in .

[0141] In one aspect, an adaptive linear lifting transform is applied to selected (or pre-selected) mesh frames of a mesh sequence.

[0142] In one aspect, an adaptive linear lifting transform is applied to selected (or pre-selected) regions of the mesh frame.

[0143] In one aspect, a plurality of vertices are included in a base layer of a mesh frame. In one example, the base layer of a mesh frame is also referred to as a level of detail 0 video, which includes low frequency coefficients. In one example, the base layer is a base layer storing core geometry information of a mesh. The base layer may include vertices, edges, and faces of the mesh.

[0144] In one aspect, a plurality of vertices are included in an enhancement layer of a mesh frame. In one example, the enhancement layer of a mesh frame includes additional information or functionality relative to a base layer. For example, the enhancement layer may include texture coordinates, normals, and material properties (e.g., roughness, reflectivity, and emissivity).

[0145] In one aspect, an adaptive linear lifting transform is applied to selected (or pre-selected) frequency bands of the grid frames.

[0146] Figure 8 A flow chart outlining a process (800) according to an aspect of the present disclosure is shown. The process (800) may be used in a video decoder. In various aspects, the process (800) is performed by a processing circuit, such as a processing circuit that performs the functions of a video decoder (110), a processing circuit that performs the functions of a video decoder (210), etc. In some aspects, the process (800) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (800). The process starts at (S801) and proceeds to (S810).

[0147] At (S810), a code stream is received, the code stream including prediction information of a plurality of vertices in a mesh frame.

[0148] At (S820), a prediction value of the current vertex is determined based on (i) a value of the current vertex among a plurality of vertices received in the codestream and (ii) each distance between the current vertex and one or more neighboring vertices of the current vertex in the mesh frame.

[0149] At (S830), the current vertex is reconstructed based on the predicted value of the current vertex.

[0150] In one aspect, the one or more neighboring vertices of the current vertex include a plurality of neighboring vertices. An updated value of the current vertex is determined. A predicted value of the current vertex is determined to be a sum of (i) the updated value of the current vertex and (ii) an average value of the plurality of neighboring vertices.

[0151] In one aspect, a ratio of each adjacent vertex in a plurality of adjacent vertices is determined. The ratio is equal to the value of the corresponding adjacent vertex divided by the distance between the current vertex and the corresponding adjacent vertex. A sum of the ratios of the plurality of adjacent vertices is determined. A product of (i) a constant and (ii) the sum of the ratios of the plurality of adjacent vertices is determined. An updated value for the current vertex is determined to be the value of the current vertex minus the product of (i) the constant and (ii) the sum of the ratios of the plurality of adjacent vertices.

[0152] In one aspect, the constant is equal to a scalar constant divided by a total number of the plurality of adjacent vertices.

[0153] In one aspect, the grid frame is a selected grid frame in a grid sequence.

[0154] In one aspect, a plurality of vertices are included in a selected area of ​​the mesh frame.

[0155] In one aspect, the plurality of vertices are included in a base layer of a mesh frame.

[0156] In one aspect, the plurality of vertices are included in an enhancement layer of the mesh frame.

[0157] In one aspect, a plurality of vertices are included in a selected frequency band of the mesh frame.

[0158] Then, the process proceeds to (S899) and terminates.

[0159] The process (800) may be adapted as appropriate. Steps in the process (800) may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.

[0160] Fig. 9A flow chart outlining a process (900) according to an aspect of the present disclosure is shown. The process (900) may be used in a video encoder. In various aspects, the process (900) is performed by a processing circuit, such as a processing circuit that performs the functions of a video encoder (103), a processing circuit that performs the functions of a video encoder (303), etc. In some aspects, the process (900) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (900). The process starts at (S901) and proceeds to (S910).

[0161] At (S910), an initial prediction value of a current vertex is determined based on (i) a value of the current vertex among a plurality of vertices in a mesh frame and (ii) each value of one or more neighboring vertices of the current vertex in the mesh frame.

[0162] At (S920), an updated prediction value of the current vertex is determined based on (i) the initial prediction value of the current vertex and (ii) each distance between the current vertex and one or more neighboring vertices of the current vertex in the mesh frame.

[0163] At (S930), the current vertex is encoded in a code stream based on the updated prediction value of the current vertex.

[0164] In one aspect, the one or more neighboring vertices of the current vertex include a plurality of neighboring vertices. An initial prediction value of the current vertex is determined as (i) a value of the current vertex minus (ii) an average of the plurality of neighboring vertices.

[0165] In one aspect, a ratio of each adjacent vertex in a plurality of adjacent vertices is determined. The ratio is equal to the value of the corresponding adjacent vertex divided by the distance between the current vertex and the corresponding adjacent vertex. A sum of the ratios of the plurality of adjacent vertices is determined. A product of (i) a constant and (ii) the sum of the ratios of the plurality of adjacent vertices is determined. An updated prediction value for the current vertex is determined to be the sum of the product of an initial prediction value for the current vertex and (i) the constant and (ii) the sum of the ratios of the plurality of adjacent vertices.

[0166] Then, the process proceeds to (S999) and terminates.

[0167] The process (900) may be adapted as appropriate. Steps in the process (900) may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.

[0168] In one aspect, a method for processing mesh data includes processing a code stream of the mesh data according to a format rule. For example, the code stream may be a code stream decoded / encoded using any of the decoding methods and / or encoding methods described herein. The format rule may specify one or more constraints of the code stream and / or one or more processes to be performed by a decoder and / or encoder.

[0169] In one example, a code stream of mesh data is processed according to a format rule. The code stream includes prediction information for a plurality of vertices in a mesh frame. The format rule specifies that a predicted value of a current vertex is determined based on (i) a value of the current vertex in the plurality of vertices received in the code stream and (ii) each distance between the current vertex and one or more neighboring vertices of the current vertex in the mesh frame. The format rule specifies that a current vertex is processed based on the predicted value of the current vertex.

[0170] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig.10 A computer system (1000) suitable for implementing certain aspects of the disclosed subject matter is shown.

[0171] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to assembly, compilation, linking or similar mechanisms to create code comprising instructions that may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpretation, microcode execution, etc.

[0172] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.

[0173] Fig.10 The components of the computer system (1000) shown are examples and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing aspects of the present disclosure. Nor should the configuration of components be interpreted as having any dependency or requirement related to any one or combination of components shown in the exemplary aspects of the computer system (1000).

[0174] The computer system (1000) may include certain human-machine interface input devices. Such human-machine interface input devices may be responsive to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not depicted). Human-machine interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, captured images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0175] The human-machine interface input device may include one or more of the following (only one of each is shown): keyboard (1001), mouse (1002), touchpad (1003), touch screen (1010), data gloves (not shown), joystick (1005), microphone (1006), scanner (1007), camera (1008).

[0176] The computer system (1000) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback of a touch screen (1010), a data glove (not shown), or a joystick (1005), but may also be a tactile feedback device that is not an input device), audio output devices (e.g., speakers (1009), headphones (not depicted)), visual output devices (e.g., screens (1010) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities, some of which are capable of outputting two-dimensional visual outputs or outputs exceeding three dimensions through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted)), and printers (not depicted).

[0177] The computer system (1000) may also include human-machine accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1020) having CD / DVD and other media (1021), thumb drives (1022), removable hard drives or solid-state drives (1023), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security software dogs (not depicted), etc.

[0178] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.

[0179] The computer system (1000) may also include an interface (1054) to one or more communication networks (1055). The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter attached to some common data port or peripheral bus (1049) (e.g., a USB port of the computer system (1000)); other network interfaces are typically integrated into the kernel of the computer system (1000) by attaching to a system bus as described below (e.g., connected to an Ethernet interface in a PC computer system or connected to a cellular network interface in a smartphone computer system). The computer system (1000) can use any of these networks to communicate with other entities. Such communications may be one-way receive only (e.g., broadcast television), one-way send only (e.g., CANBus connected to certain CANBus devices), or bidirectional, for example, using a LAN or WAN digital network to connect to other computer systems. Certain protocols and protocol stacks may be used on each of those networks and network interfaces as described above.

[0180] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the kernel ( 1040 ) of the computer system ( 1000 ).

[0181] The kernel (1040) may include one or more central processing units (CPUs) (1041), graphics processing units (GPUs) (1042), dedicated programmable processing units in the form of field programmable gate areas (FPGAs) (1043), hardware accelerators (1044) for certain tasks, graphics adapters (1050), etc. These devices, as well as read-only memory (ROM) (1045), random access memory (1046), internal mass storage (1047) such as internal non-user accessible hard drives, SSDs, etc., may be connected via a system bus (1048). In some computer systems, the system bus (1048) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the kernel's system bus (1048) or to the kernel's system bus (1048) via a peripheral bus (1049). In one example, a screen (1010) may be connected to a graphics adapter (1050). The architecture of the peripheral bus includes PCI, USB, etc.

[0182] The CPU (1041), GPU (1042), FPGA (1043) and accelerator (1044) can execute certain instructions, which can be combined to form the computer code mentioned above. The computer code can be stored in ROM (1045) or RAM (1046). Transition data can also be stored in RAM (1046), while permanent data can be stored, for example, in internal mass storage (1047). Fast storage and retrieval to any storage device can be performed by using a cache, which can be closely associated with one or more CPUs (1041), GPUs (1042), mass storage (1047), ROM (1045), RAM (1046), etc.

[0183] The computer readable medium may have thereon computer codes for performing various computer-implemented operations. The medium and computer codes may be those specially designed and constructed for the purposes of the present disclosure, or the medium and computer codes may be of a type well known and available to those skilled in the art of computer software.

[0184] As an example, and not by way of limitation, a computer system (1000) having an architecture, particularly a kernel (1040), can provide functionality because one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage as described above, as well as some non-temporary kernel (1040) memories, such as a kernel internal mass storage (1047) or ROM (1045). Software implementing various aspects of the present disclosure can be stored in such devices and executed by the kernel (1040). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can cause the kernel (1040), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.) to perform specific processes described herein or to perform specific parts of specific processes described herein, including defining data structures stored in RAM (1046) and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system may provide functionality due to logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1044)) that may replace or operate in conjunction with software to perform specific processes described herein or specific portions of specific processes described herein. Where appropriate, references to software may include logic and vice versa. Where appropriate, references to computer-readable media may include circuitry (e.g., integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.

[0185] As used in this disclosure, "at least one of..." or "one of..." is intended to include any one or combination of the listed elements. For example, reference to at least one of A, B, or C, at least one of A, B, and C, at least one of A, B, and / or C, and at least one of A to C is intended to include only A, only B, only C, or any combination thereof. Reference to one of A or B, and one of A and B is intended to include A or B or (A and B). Where applicable, use of "one of..." does not exclude any combination of the listed elements, such as when the elements are not mutually exclusive.

[0186] Although the present disclosure has described several examples of various aspects, there are changes, permutations, and various replacement equivalents that fall within the scope of the present disclosure. Therefore, it should be appreciated that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and therefore fall within the spirit and scope of the present disclosure.

Claims

1. A grid decoding method, comprising: receiving a code stream, the code stream comprising prediction information of a plurality of vertices in a mesh frame; determining a predicted value of the current vertex based on (i) a value of a current vertex among the plurality of vertices received in the codestream and (ii) each distance between the current vertex and one or more neighboring vertices of the current vertex in the mesh frame; as well as The current vertex is reconstructed based on the predicted value of the current vertex.

2. The method according to claim 1, wherein: The one or more adjacent vertices of the current vertex include a plurality of adjacent vertices, and Determining the predicted value of the current vertex further includes: determining an updated value for the current vertex; and The predicted value of the current vertex is determined as a sum of (i) the updated value of the current vertex and (ii) an average value of the plurality of adjacent vertices.

3. The method according to claim 2, wherein: Determining the updated value of the current vertex includes: Determining a ratio of each adjacent vertex in the plurality of adjacent vertices, the ratio being equal to a value of the corresponding adjacent vertex divided by a distance between the current vertex and the corresponding adjacent vertex; determining a sum of ratios of the plurality of adjacent vertices; determining the product of (i) a constant and (ii) a sum of ratios of the plurality of adjacent vertices; and The update value of the current vertex is determined to be the value of the current vertex minus the product of (i) the constant and (ii) the sum of the ratios of the plurality of adjacent vertices.

4. The method according to claim 3, wherein: The constant is equal to a scalar constant divided by a total number of the plurality of adjacent vertices.

5. The method according to claim 1, wherein: The grid frame is a selected grid frame in a grid sequence.

6. The method according to claim 1, wherein: The plurality of vertices are included in a selected area of ​​the mesh frame.

7. The method according to claim 1, wherein: The plurality of vertices are included in a base layer of the mesh frame.

8. The method according to claim 1, wherein: The plurality of vertices are included in an enhancement layer of the mesh frame.

9. The method according to claim 1, wherein: The plurality of vertices are included in a selected frequency band of the mesh frame.

10. A grid coding method, comprising: Determine an initial predicted value of the current vertex based on (i) a value of the current vertex among a plurality of vertices in a grid frame and (ii) each value of one or more neighboring vertices of the current vertex in the grid frame; determining an updated prediction value of the current vertex based on (i) the initial prediction value of the current vertex and (ii) each distance between the current vertex and the one or more neighboring vertices of the current vertex in the mesh frame; as well as Based on the updated prediction value of the current vertex, the current vertex is encoded in a code stream.

11. The method according to claim 10, wherein: The one or more adjacent vertices of the current vertex include a plurality of adjacent vertices, and Determining the initial prediction value of the current vertex also includes: The initial prediction value of the current vertex is determined as (i) the value of the current vertex minus (ii) an average value of the plurality of adjacent vertices.

12. The method according to claim 11, wherein: Determining the updated predicted value of the current vertex includes: Determining a ratio of each adjacent vertex in the plurality of adjacent vertices, the ratio being equal to a value of the corresponding adjacent vertex divided by a distance between the current vertex and the corresponding adjacent vertex; determining a sum of ratios of the plurality of adjacent vertices; determining the product of (i) a constant and (ii) a sum of ratios of the plurality of adjacent vertices; and The updated prediction value of the current vertex is determined as the sum of the products of the initial prediction value of the current vertex and (i) the constant and (ii) the sum of the ratios of the plurality of adjacent vertices.

13. The method according to claim 12, wherein: The constant is equal to a scalar constant divided by a total number of the plurality of adjacent vertices.

14. The method according to claim 10, wherein: The grid frame is a selected grid frame in a grid sequence.

15. The method according to claim 10, wherein: The plurality of vertices are included in a selected area of ​​the mesh frame.

16. The method according to claim 10, wherein: The plurality of vertices are included in a base layer of the mesh frame.

17. The method according to claim 10, wherein: The plurality of vertices are included in an enhancement layer of the mesh frame.

18. The method according to claim 10, wherein: The plurality of vertices are included in a selected frequency band of the mesh frame.

19. A method for processing grid data, the method comprising: Processing the code stream of the grid data according to the format rule, wherein: The code stream includes prediction information of a plurality of vertices in a mesh frame; and The format rules specify: Determine a predicted value of the current vertex based on (i) a value of the current vertex among the plurality of vertices received in the codestream and (ii) each distance between the current vertex and one or more neighboring vertices of the current vertex in the mesh frame; and The current vertex is processed based on the predicted value of the current vertex.

20. The method of claim 19, wherein: The one or more adjacent vertices of the current vertex include a plurality of adjacent vertices, and The format rules specify: Determining an updated value of the current vertex; as well as The predicted value of the current vertex is determined as a sum of (i) the updated value of the current vertex and (ii) an average value of the plurality of adjacent vertices.