Location coding in grid compression

By adopting base mesh quantization technology in dynamic mesh coding, and using multi-parallelogram prediction and quantization step value, the compression problem of time-varying connection information and attribute mapping in dynamic mesh is solved, and efficient mesh vertex position compression and reconstruction is achieved, suitable for applications such as real-time communication and virtual reality.

CN120019417APending Publication Date: 2025-05-16TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004350.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-23
Filing Date
2024-05-24
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively compress and transmit time-varying connection information and attribute mapping in dynamic grid sequences, especially in applications such as real-time communication and virtual reality.

Method used

By using base grid quantization technology in mesh encoding, the vertex positions of dynamic mesh are compressed and reconstructed using multi-parallelogram prediction, quantization step value and inverse quantization prediction residual methods.

Benefits of technology

It realizes efficient compression and transmission of dynamic grids, improves the performance of applications such as real-time communication and virtual reality, and meets the needs of lossy and lossless compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005353484350000301
    Figure BDA0005353484350000301
  • Figure BDA0005353484350000302
    Figure BDA0005353484350000302
  • Figure BDA0005353484350000311
    Figure BDA0005353484350000311
Patent Text Reader

Abstract

A method and apparatus, the apparatus comprising computer code configured to cause one or more processors to obtain a mesh from a code stream, the mesh representing encoded volume data of at least one three-dimensional (Three-Dimensional, 3D) visual content, and the computer code configured to decode the encoded volume data based on a base mesh quantization of the mesh, wherein the base mesh quantization comprises predicting a current vertex by using an encoded vertex position through multi-parallelogram prediction, quantizing a prediction residual of the current vertex by quantizing a step size value, and quantizing the current vertex by adding an inverse quantized prediction residual of the quantized prediction residual to the multi-parallelogram prediction. The quantized prediction residual is compressed and the position of the current vertex is reconstructed as a reference for other vertexes.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Application No. 63 / 468,955 filed on May 25, 2023 and U.S. Application No. 18 / 672,164 filed on May 23, 2024, the disclosures of which are incorporated herein by reference in their entirety. Technical Field

[0003] The present application discloses a set of advanced video coding techniques, which include vertex grouping in mesh motion vector coding. Background Art

[0004] Advances in three-dimensional (3D) filming, 3D modeling, and 3D rendering technologies have facilitated the widespread use of 3D content on multiple platforms and devices. Today, a baby’s first steps can be filmed on one continent, and grandparents on another can see (and possibly interact with) and enjoy a fully immersive experience with the child. However, to achieve this realism, models are becoming increasingly complex, and the creation and consumption of these models is associated with large amounts of data. 3D meshes are widely used to represent this immersive content.

[0005] Dynamic mesh sequences may require a lot of data because it may consist of a large amount of information that changes over time. Therefore, effective compression techniques are needed to store and transmit such content. Previously, the Moving Picture Experts Group (MPEG) developed mesh compression standards IC, MESHGRID, and FAMC to handle dynamic meshes with constant connectivity, time-varying geometry, and vertex attributes. However, these standards do not consider time-varying attribute mapping and connectivity information. Digital Content Creation (DCC) tools usually generate such dynamic meshes. In contrast, volume acquisition techniques are particularly challenging to generate dynamic meshes with constant connectivity under real-time constraints. Existing standards do not support this type of content. MPEG plans to develop a new mesh compression standard to directly handle dynamic meshes with time-varying connectivity information and optional time-varying attribute mapping. The standard is aimed at lossy and lossless compression for various applications, such as real-time communication, storage, free-viewpoint video, augmented reality (AR) and virtual reality (VR). Features such as random access and scalable / progressive coding are also considered. Therefore, due to the above reasons, the industry urgently needs technical solutions to solve these problems arising in video coding technology. Summary of the invention

[0006] The present invention includes a method and an apparatus, the apparatus including a memory and one or more processors, the memory being configured to store computer program code, the one or more processors being configured to access the computer program code and operate according to the instructions of the computer program code. The computer program is configured to cause the processor to implement an acquisition code, the acquisition code being configured to cause the at least one processor to obtain a mesh from a bitstream, the mesh representing encoded volume data of at least one three-dimensional (3D) visual content, and the computer program is configured to cause the processor to decode the encoded volume data based on a base mesh quantization of the mesh, wherein the base mesh quantization includes: predicting a current vertex by using an encoded vertex position through multi-parallelogram prediction, quantizing a prediction residual by a quantization step value, and compressing the quantized prediction residual by adding an inverse quantized prediction residual of the quantized prediction residual to the multi-parallelogram prediction and reconstructing the position of the current vertex as a reference for other vertices.

[0007] According to aspects disclosed in the present application, base grid quantization may include: after quantizing the prediction residual and before reconstructing the position of the current vertex, dequantizing the quantized prediction residual.

[0008] According to aspects disclosed in the present application, base grid quantization may include: based on a vertex of the grid being determined as the first vertex to be encoded in the grid in time sequence, determining the position of the vertex of the grid to be directly encoded by quantizing the position of the vertex with a positive integer and entropy encoding the quantized position.

[0009] According to aspects disclosed in the present application, the quantization step value is a binary rational number, including a numerator m and a denominator, the denominator is 2 to the power of n, and the numerator m and the value n are both integers.

[0010] According to aspects disclosed herein, the numerator m may be a first integer greater than zero, and the value n may be a second integer greater than or equal to zero.

[0011] According to aspects disclosed herein, the quantization step value may be signaled by signaling the numerator m and the value n at least in the code stream.

[0012] According to various aspects disclosed in the present application, the quantization step value may be signaled in the code stream via a “mesh_position_quantization_step_size_log2_denominator” syntax and a “mesh_position_quantization_step_size_numerator_minus1” syntax. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Further features, characteristics and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0014] Figure 1 is a schematic diagram of aspects of a communication system according to one or more embodiments herein;

[0015] Figure 2 is a schematic diagram of aspects of a media processing system according to one or more embodiments herein;

[0016] Figure 3 is a schematic diagram of aspects of a decoder according to one or more embodiments herein;

[0017] Figure 4 is a schematic diagram of aspects of an encoder according to one or more embodiments herein;

[0018] Figure 5 is a schematic diagram of various aspects of an encoder side according to one or more embodiments herein;

[0019] Figure 6 is a schematic diagram of various aspects of a decoder side according to one or more embodiments herein;

[0020] Figure 7 is a schematic diagram of aspects of an encoder and decoder system based on media grid features according to one or more embodiments herein;

[0021] Figure 8 is a schematic diagram of aspects of mesh processing according to one or more embodiments herein;

[0022] Fig. 9 is a schematic diagram of aspects of mesh processing according to one or more embodiments herein;

[0023] Fig.10 is a flow chart illustrating aspects of mesh processing according to one or more embodiments herein;

[0024] Fig.11 is a flow chart illustrating aspects of mesh processing according to one or more embodiments herein;

[0025] Fig.12 is a flow chart illustrating aspects of mesh processing according to one or more embodiments herein;

[0026] Fig.13 is a schematic diagram of aspects of mesh processing according to one or more embodiments herein;

[0027] Fig.14 is a schematic diagram of aspects of mesh processing according to one or more embodiments herein;

[0028] Fig.15 is a schematic diagram of aspects of mesh processing according to one or more embodiments herein;

[0029] Fig.16 is a schematic diagram of aspects of mesh processing according to one or more embodiments herein;

[0030] Fig.17 is a schematic diagram of aspects of mesh processing according to one or more embodiments herein;

[0031] Fig.18 is a schematic diagram of aspects of mesh processing according to one or more embodiments herein;

[0032] Fig.19 is a schematic diagram of aspects of mesh processing according to one or more embodiments herein;

[0033] Fig. 20 is a schematic diagram of aspects of mesh processing according to one or more embodiments herein;

[0034] Fig.21 is a schematic diagram of aspects of mesh processing according to one or more embodiments herein;

[0035] Fig. 22 is a schematic diagram of aspects of mesh processing according to one or more embodiments herein;

[0036] Fig.23 is a schematic diagram of aspects of mesh processing according to one or more embodiments herein;

[0037] Fig.24 is a schematic diagram of aspects of mesh processing according to one or more embodiments herein;

[0038] Fig.25 is a schematic diagram of aspects of mesh processing according to one or more embodiments herein;

[0039] Fig.26 is a schematic diagram of aspects of mesh processing according to one or more embodiments herein; and

[0040] Fig. 27 is a schematic diagram of aspects of a system according to one or more embodiments. DETAILED DESCRIPTION

[0041] The suggested features discussed below may be used alone or in any combination in any order. In addition, the embodiments may be implemented by processing circuits (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0042] Figure 1 A simplified block diagram of a communication system 100 according to an embodiment disclosed in the present application is shown. The communication system 100 may include at least two terminals 102 and 103 interconnected by a network 105. For unidirectional transmission of data, a first terminal 103 may encode video data at a local location to transmit to another terminal 102 via the network 105. The second terminal 102 may receive the encoded video data of another terminal from the network 105, decode the encoded video data and display the recovered video data. Unidirectional data transmission is more common in applications such as media services.

[0043] Figure 1 A second pair of terminals 101 and 104 are shown, which are provided to support bidirectional transmission of encoded video, such as may occur during a video conference. For bidirectional transmission of data, each terminal 101 and 104 can encode video data collected at a local location for transmission to the other terminal via a network 105. Each terminal 101 and 104 can also receive encoded video data sent by another terminal, can decode the encoded video data, and can display the recovered video data on a local display device.

[0044] exist Figure 1 In the embodiment, terminals 101, 102, 103 and 104 can be shown as servers, personal computers and smart phones, but the principles disclosed in the present application are not limited thereto. The embodiments disclosed in the present application are applicable to laptop computers, tablet computers, media players and / or special video conferencing equipment. Network 105 represents any number of networks that transmit encoded video data between terminals 101, 102, 103 and 104, including, for example, wired and / or wireless communication networks. Communication network 105 can exchange data in circuit switching and / or packet switching channels. Representative networks include telecommunication networks, local area networks, wide area networks and / or the Internet. For the purpose of this discussion, unless explained below, the architecture and topology of network 105 may be insignificant for the operation disclosed in the present application.

[0045] As an example of the application of the subject matter disclosed in this application, Figure 2The placement of the video encoder and decoder in a streaming environment is shown. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., and the like.

[0046] The streaming system may include a capture subsystem 203, which may include a video source 201, such as a digital camera, for example, to create an uncompressed video sample stream 213. The sample stream 213 may be characterized as having a high data volume compared to an encoded video stream, and may be processed by an encoder 202 coupled to the video source 201, which may be, for example, a camera as discussed above. The encoder 202 may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. The encoded video stream 204 may be characterized as having a lower data volume compared to the sample stream, and may be stored on a streaming server 205 for future use. One or more streaming clients 212 and 207 may access the streaming server 205 to retrieve copies 208 and 206 of the encoded video stream 204. Client 212 may include a video decoder 211 that may decode an incoming copy of the encoded video stream 208 and create an output video sample stream 210 that may be presented on a display 209 or another presentation device (not depicted). In some streaming systems, video streams 204, 206, and 208 may be encoded according to certain video encoding / compression standards. Examples of these standards are mentioned above and further described herein.

[0047] Figure 3 It may be a functional block diagram of a video decoder 300 according to an embodiment of the present invention.

[0048] Receiver 302 may receive one or more coded video sequences to be decoded by decoder 300; in the same or another embodiment, one coded video sequence is received at a time, wherein the decoding of each coded video sequence is independent of the other coded video sequences. The coded video sequence may be received from channel 301, which may be a hardware / software link to a storage device storing the coded video data. Receiver 302 may receive the coded video data as well as other data, for example, receiver 302 may receive coded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not shown). Receiver 302 may separate the coded video sequence from the other data. To prevent network jitter, a buffer memory 303 may be coupled between receiver 302 and entropy decoder / parser 304 (hereinafter referred to as "parser"). When receiver 302 receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may not be necessary to configure buffer 303, or the buffer 303 may be made smaller. For use over traffic packet networks such as the Internet, a buffer 303 may also be required, which may be relatively large and may advantageously be adaptively sized.

[0049] The video decoder 300 may include a parser 304 to reconstruct symbols 313 from an entropy coded video sequence. The categories of these symbols include information for managing the operation of the decoder 300, and potential information for controlling a display device (e.g., display 312) that is not part of the decoder but may be coupled to the decoder. The control information for the (one or more) display devices may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser 304 may parse / entropy decode the received coded video sequence. The encoding of the coded video sequence may be based on a video coding technique or standard and may follow principles known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 304 may extract a subgroup parameter set for at least one subgroup of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. Subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CU), blocks, Transform Units (TU), Prediction Units (PU), etc. The entropy decoder / parser may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0050] The parser 304 may perform an entropy decoding / parsing operation on the video sequence received from the buffer 303, thereby creating the symbol 313. The parser 304 may receive the encoded data and selectively decode a specific symbol 313. In addition, the parser 304 may determine whether to provide the specific symbol 313 to the motion compensation prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.

[0051] Depending on the type of the coded video picture or a portion of the coded video picture (e.g., inter-frame and intra-frame pictures, inter-frame blocks and intra-frame blocks) and other factors, the reconstruction of the symbol 313 may involve multiple different units. Which units are involved and how they are involved can be controlled by the parser 304 through subgroup control information parsed from the coded video sequence. For the sake of brevity, such subgroup control information flow between the parser 304 and the multiple units below is not described.

[0052] In addition to the functional blocks already mentioned, decoder 300 may be conceptually subdivided into several functional units as described below. In actual implementations operating under commercial constraints, many of these units interact closely with each other and may be (at least partially) integrated with each other. However, for the purpose of describing the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.

[0053] The first unit is the sealer / inverse transform unit 305. The sealer / inverse transform unit 305 receives the quantized transform coefficients as symbols 313 from the parser 304, as well as control information, including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit 305 may output a block including sample values, which may be input into the aggregator 310.

[0054] In some cases, the output samples of the sealer / inverse transform unit 305 may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra-picture prediction unit 307. In some cases, the intra-picture prediction unit 307 generates a block of the same size and shape as the block being reconstructed using surrounding reconstructed information extracted from the current (partially reconstructed) picture 309. In some cases, the aggregator 310 adds the prediction information generated by the intra-prediction unit 307 to the output sample information provided by the sealer / inverse transform unit 305 on a per-sample basis.

[0055] In other cases, the output samples of the scaler / inverse transform unit 305 may belong to a block that is inter-coded and potentially motion compensated. In this case, the motion compensated prediction unit 306 may access the reference picture memory 308 to extract samples for prediction. After the extracted samples are motion compensated according to the symbol 313 associated with the block, these samples may be added to the output of the scaler / inverse transform unit (in this case referred to as residual samples or residual signals) by the aggregator 310 to generate output sample information. The motion compensation unit's retrieval of predicted samples from an address in the reference picture memory may be controlled by a motion vector, which is available to the motion compensated prediction unit in the form of the symbol 313, which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values ​​extracted from the reference picture memory when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and the like.

[0056] The output samples of aggregator 310 may be employed by various loop filtering techniques in loop filter unit 311. The video compression techniques may include in-loop filter techniques controlled by parameters included in the encoded video bitstream and available to loop filter unit 311 as symbols 313 from parser 304. However, the video compression techniques may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or encoded video sequence, and to previously reconstructed and loop filtered sample values.

[0057] The output of the loop filter unit 311 may be a sample stream that may be output to the display device 312 and stored in the reference picture memory 557 for subsequent inter-picture prediction.

[0058] Once certain coded pictures are fully reconstructed, they can be used as reference images for subsequent prediction. Once a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by parser 304), current reference picture 309 can become part of reference picture buffer 308, and new current picture memory can be reallocated before starting to reconstruct subsequent coded pictures.

[0059] The video decoder 300 may perform decoding operations according to a predetermined video compression technique recorded in a standard such as ITU-T Rec. H.265. The encoded video sequence may conform to the syntax specified by the video compression technique or standard used, and in the sense that the encoded video sequence follows the syntax of the video compression technique or standard, the encoded video sequence may conform to the syntax specified in the video compression technology document or standard and in particular in the profile in the video compression technology document or standard. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the level of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.

[0060] In an embodiment, the receiver 302 may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder 300 to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0061] Figure 4 It may be a functional block diagram of the video encoder 400 according to the embodiment disclosed in this application.

[0062] The encoder 400 may receive video samples from a video source 401 (which is not part of the encoder), which may capture video picture(s) to be encoded by the encoder 400 .

[0063] The video source 401 may provide a source video sequence in the form of a digital video sample stream to be encoded by the encoder (303), which may have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source 401 may be a storage device storing previously prepared videos. In a video conferencing system, the video source 401 may be a camera that collects local image information as a video sequence. The video data may be constructed as a plurality of separate pictures that impart motion when viewed sequentially. The pictures themselves may be constructed as a spatial pixel array, wherein each pixel may include one or more samples depending on the sampling structure, color space, etc. used. The relationship between pixels and samples may be easily understood by those skilled in the art. The following description focuses on samples.

[0064] According to an embodiment, the encoder 400 may encode and compress the pictures of the source video sequence into an encoded video sequence 410 in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller 402. The controller controls other functional units as described below and is functionally coupled to these units. For the sake of brevity, the coupling is not depicted. The parameters set by the controller may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology, etc.), picture size, picture group layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller 402 because these functions may be related to the video encoder 400 optimized for a specific system design.

[0065] Some video encoders operate in what is readily recognizable to those skilled in the art as a "coding loop". As an oversimplified description, the coding loop may consist of an encoding portion of an encoder 400 (hereinafter referred to as a "source encoder") (responsible for creating symbols based on an input picture to be encoded and (one or more) reference pictures) and a (local) decoder 406 embedded in the encoder 400, which reconstructs the symbols to create sample data in the same way as the (remote) decoder created the sample data (because any compression between the symbols and the encoded video code stream is lossless in the video compression techniques contemplated by the disclosed subject matter). The reconstructed sample stream is input into a reference picture memory 405. Since the decoding of the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the contents of the reference picture buffer are also bit-accurate between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction portion of the encoder are exactly the same sample values ​​that the decoder will "see" when using prediction during decoding. This basic principle of reference picture synchronization (and the drift that results when synchronization cannot be maintained, for example due to channel errors) is well known to those skilled in the art.

[0066] The operation of the "local" decoder 406 can be combined with the above Figure 3 The operation of the "remote" decoder 300 described in detail is identical. However, reference is also briefly made to Figure 4 , when symbols are available and the entropy encoder 408 and the parser 304 are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the decoder 300, including the channel 301, the receiver 302, the buffer 303 and the parser 304, may not be fully implemented in the local decoder 406.

[0067] At this point it can be observed that any decoder techniques other than parsing / entropy decoding present in the decoder must also be present in a substantially identical functional form in the corresponding encoder. The description of encoder techniques may be simplified because the encoder techniques are reciprocal to the decoder techniques described comprehensively. A more detailed description is only required in certain areas and is provided below.

[0068] As part of the operation of source encoder 403, source encoder 403 may perform motion compensated predictive coding that predictively encodes an input picture with reference to one or more previously encoded pictures from a video sequence designated as "reference frames." In this manner, encoding engine 407 encodes the differences between pixel blocks of an input frame and pixel blocks of (one or more) reference frames that may be selected as prediction reference (s) for the input frame.

[0069] The local video decoder 406 may decode the encoded video data of the frame that may be designated as the reference frame based on the symbol created by the source encoder 403. The operation of the encoding engine 407 may advantageously be a lossy process. When the encoded video data is available at the video decoder ( Figure 4 When decoded at a remote video decoder (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder 406 replicates the decoding process that may be performed by the video decoder on the reference frame, and the decoding process may cause the reconstructed reference frame to be stored in the reference picture cache 405. In this way, the encoder 400 may locally store a copy of the reconstructed reference frame that has common content (absent transmission errors) with the reconstructed reference frame to be obtained by the remote video decoder.

[0070] The predictor 404 may perform a prediction search for the encoding engine 407. That is, for a new frame to be encoded, the predictor 404 may search the reference picture memory 405 for sample data (as a candidate reference pixel block) or certain metadata, such as a reference picture motion vector, block shape, etc., that may serve as an appropriate prediction reference for the new picture. The predictor 404 may operate pixel-by-pixel based on the sample block to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor 404, it may be determined that the input picture may have a prediction reference obtained from a plurality of reference pictures stored in the reference picture memory 405.

[0071] The controller 402 may manage encoding operations of the video encoder 403 , including, for example, setting parameters and subgroup parameters for encoding video data.

[0072] The outputs of all the above functional units may be entropy encoded in the entropy encoder 408. The entropy encoder converts the symbols generated by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.

[0073] Transmitter 409 may buffer the encoded video sequence(s) created by entropy encoder 408 in preparation for transmission over communication channel 411, which may be a hardware / software link to a storage device where the encoded video data will be stored. Transmitter 409 may combine the encoded video data from video encoder 403 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0074] The controller 402 may manage the operation of the encoding device 400. During encoding, the controller 402 may assign a certain encoded picture type to each encoded picture, but this may affect the encoding technology that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following frame picture types:

[0075] Intra-pictures (I-pictures) can be pictures that can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra-pictures, including, for example, Independent Decoder Refresh Pictures. Those skilled in the art are aware of the variations of I-pictures and their corresponding applications and features.

[0076] A predictive picture (P picture), which may be a picture that can be encoded and decoded using intra prediction or inter prediction, uses at most one motion vector and a reference index to predict the sample values ​​of each block.

[0077] Bidirectional predictive pictures (B pictures), which can be pictures that can be encoded and decoded using intra prediction or inter prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.

[0078] The source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples) and coded block by block. The blocks may be predictively coded with reference to other (already coded) blocks, which are determined according to the coding allocation applied to the block's corresponding picture. For example, blocks of an I picture may be non-predictively coded, or may be predictively coded (spatial prediction or intra-prediction) with reference to already coded blocks of the same picture. Blocks of pixels of a P picture may be non-predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. Blocks of a B picture may be non-predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.

[0079] The encoder 400 (which may be a video encoder, for example) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T Rec. H.265. In the operation of the encoder 400, the encoder 400 may perform various compression operations, including predictive encoding operations that utilize temporal redundancy and spatial redundancy in an input video sequence. Therefore, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0080] In an embodiment, the transmitter 409 may transmit additional data when transmitting the encoded video. The source encoder 403 may include such data as part of the encoded video sequence. The additional data may include time / space / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.

[0081] Figure 5 A simplified block diagram workflow diagram 500 of an exemplary viewport-dependent process in the Omnidirectional Media Application Format (OMAF) standard is shown, which may allow 360-degree Virtual Reality (VR360) streaming as described in OMAF.

[0082] In acquisition block 501, video data A is acquired, where the video data is, for example, a plurality of image data and audio data at the same time instance when the image data can represent a scene in VR360. In processing block 503, image B at the same time instance is processed by one or more processing steps. i Processing includes stitching, mapping to a projected image relative to one or more virtual reality (VR) angles or other angles / (one or more) viewing angles, and packaging by region. Additionally, metadata may be created to indicate any such processed information and other information to aid in the transmission and rendering process.

[0083] Regarding data D, in the image encoding block 505, the projection picture is encoded into data E i And form a media file for viewport-independent streaming transmission. In the video encoding block 504, the video picture is encoded into data E v As a single-layer code stream. For example, regarding data B a The audio data may also be encoded into data E in the audio encoding block 502. a .

[0084] Data E a 、E v and E i 、The entire coded stream F iand / or F may be stored in a (Content Delivery Network, CDN) / cloud) server and may typically be fully transmitted to an OMAF player 520, for example, in a transmission box 507 or otherwise, and may be fully decoded by a decoder so that at least an area of ​​the decoded image corresponding to the current viewport is presented to the user in a display box 516 relative to various metadata from the head / eye tracking box 508, file playback, and orientation / viewport metadata (e.g., the angle at which the user may view through the VR image device relative to the viewport specification of the device). A notable feature of VR360 is that only a viewport may be displayed at any particular time, and such a feature may improve the performance of omnidirectional video systems by selective transmission based on the user's viewport (or any other criteria, such as recommended viewport timing metadata). For example, viewport-dependent transmission may be achieved by tile-based video encoding according to an exemplary embodiment.

[0085] As with the above-mentioned encoding frame, the OMAF player 520 according to the exemplary embodiment may similarly perform the encoding of the data F' and / or F'. i and metadata, and decapsulates one or more execution files / segments, and reverses one or more aspects of such encoding, decoding the audio data E' in the audio decoding block 510 i , decode the video data E' in the video decoding frame 513 v , decode the image data E' in the image decoding block 514 i , continue to process data B' in the audio rendering frame 511 a The audio rendering is performed, and the data D' is image rendered in the image rendering block 515 to output the display data A' in the VR360 format in the display block 516 according to various metadata such as orientation / viewport metadata. i , outputting audio data A' in the speaker / headphone frame 512 s Various metadata may affect the data decoding and rendering process according to various tracks, languages, qualities, views that may be selected by or for a user of the OMAF player 520, and it should be understood that the processing order described herein is presented for an exemplary embodiment and may be implemented in other orders according to other exemplary embodiments.

[0086] Figure 6A simplified block content flow chart 600 of (encoding) point cloud data is shown, and view position and angle dependent processing of point cloud data (hereinafter "V-PCC") is performed when acquiring / generating / (de)encoding / rendering / displaying 6 degrees of freedom (DoF) media. It should be understood that the described features can be used alone or in combination in any order, and that elements such as for encoding and decoding and other elements illustrated can be implemented by processing circuits (e.g., one or more processors or one or more integrated circuits), and according to exemplary embodiments, one or more processors can execute a program stored in a non-transitory computer readable medium.

[0087] Diagram 600 illustrates an exemplary embodiment for streaming encoded point cloud data in accordance with V-PCC.

[0088] At volume data acquisition block 601, a real-world visual scene or a computer-generated visual scene (or a combination of a real-world visual scene and a computer-generated visual scene) may be acquired by a set of camera devices or synthesized by a computer as volume data, and the volume data, which may be in any format, may be converted to a (quantized) point cloud data format by image processing at conversion to point cloud block 602. For example, according to an exemplary embodiment, data from the volume data may be region data, which is converted to points of a point cloud by extracting one or more values ​​described below from the volume data and any associated data into a desired point cloud format. According to an exemplary embodiment, the volume data may be a 3D data set of 2D images, such as a slice of a 2D projection of a 3D data set that may be projected from a 2D projection of the 3D data set. According to an exemplary embodiment, a point cloud data format includes representations of data points in one or more various spaces and can be used to represent volume data, and can provide improvements with respect to sampling and data compression, such as with respect to temporal redundancy, and point cloud data in, for example, an x-format, a y-format, a z-format, which represents a color value (e.g., RGB, etc.), brightness, intensity, etc. at each of multiple points of the cloud data, and can be used for progressive decoding, polygon meshing, direct rendering, octree 3D representation of 2D quadtree data.

[0089] When projected to the image frame 603, the acquired point cloud data may be projected onto a 2D image and encoded into an image / video picture using video-based point cloud coding (V-PCC). The projected point cloud data may consist of attributes, geometry, occupancy maps, and other metadata for point cloud data reconstruction, such as using a painter's algorithm, a ray casting algorithm, a (3D) binary space partitioning algorithm, etc.

[0090] On the other hand, in the scene generator box 609, the scene generator may generate some metadata to be used for rendering and displaying 6DoF media, for example, based on director intent or user preference. Such 6DoF media may include a 3D view of the scene from a rotation change on the 3D axes X, Y, and Z, similar to 360VR, in addition to allowing additional dimensions of moving forward / backward, up / down, and left / right in the virtual experience or at least based on the point cloud encoded data. The scene description metadata defines one or more scenes composed of the encoded point cloud data and other media data including VR360, light field, audio, etc., and the scene description metadata may be provided to one or more cloud servers and / or performed as follows. Figure 6 and file / segment encapsulation / decapsulation processing as shown in the related description.

[0091] Following the video encoding block 604 and the image encoding block 605 similar to the video and image encoding described above (and it will be appreciated that audio encoding may also be provided as described above), the file / segment encapsulation block 606 processes the encoded point cloud data so that the encoded point cloud data is combined into a media file for file playback or a sequence of initialization segments and media segments for streaming according to a particular media container file format (e.g., one or more video container formats, and for example using Dynamic Adaptive Streaming overHTTP (DASH) as described below), among other things, such descriptions represent exemplary embodiments. The file container may also include scene description metadata, such as from the scene generator block 1109, into the file or segment.

[0092] According to an exemplary embodiment, a file is packaged according to scene description metadata to include at least one view position and at least one or more angle views at one or more times at (one or more) view positions in 6DoF media, so that the file can be transmitted according to a request input by a user or creator. In addition, according to exemplary embodiments, a fragment of such a file may include one or more portions of such a file, such as a portion of 6DoF media that indicates a single viewpoint and angle at one or more times; however, these are merely exemplary embodiments and may change according to various conditions such as network, user, creator capabilities and input.

[0093] According to an exemplary embodiment, the point cloud data is divided into a plurality of 2D / 3D regions, which are independently encoded, for example, at one or more of the video encoding block 604 and the image encoding block 605. Each independently encoded partition of the point cloud data may then be packaged as a track in a file and / or segment in the file / segment packaging block 606. According to an exemplary embodiment, each point cloud track and / or metadata track may include some useful metadata for view position / angle related processing.

[0094] According to an exemplary embodiment, metadata useful for view position / angle related processing (contained in a file or fragment encapsulated with respect to a file / segment encapsulation box) includes one or more of the following: layout information of 2D / 3D partitions with indices, (dynamic) mapping information associating a 3D volume partition with one or more 2D partitions (e.g., any tile / tile group / slice / sub-picture), 3D position of each 3D partition in a 6DoF coordinate system, a representative view position / angle list, a selected view position / angle list corresponding to the 3D volume partition, an index of the 2D / 3D partition corresponding to the selected view position / angle, quality (grade) information of each 2D / 3D partition, and rendering information of each 2D / 3D partition that depends on, for example, each view position / angle. When such metadata is called upon request, for example by a user of a V-PCC player or by a content creator for a user of a V-PCC player, more efficient processing of specific portions of the 6DoF media required for such metadata may be allowed, so that the V-PCC player may provide an image focused on the portion of the 6DoF media with higher quality than other portions, rather than providing unused portions of the media.

[0095] Starting from the file / segment encapsulation box 606, a delivery mechanism (e.g., via DASH) can be used to deliver the file or one or more segments of the file directly to either the V-PCC player 625 and the cloud server, for example, in the cloud server box 607, the cloud server can extract one or more tracks and / or one or more specific 2D / 3D partitions from the file, and can merge multiple encoded point cloud data into one data.

[0096] Based on data such as the position / view tracking box 608, if (one or more) current view positions and angles are defined on the 6DoF coordinate system of the client system, the view position / angle metadata can be passed from the file / segment encapsulation box 606 in the cloud server box 607, or otherwise processed based on the file or segment already at the cloud server, so that the cloud server can extract the appropriate (one or more) partitions from the (one or more) storage files and merge them (if necessary) based on the metadata from, for example, a client system with a V-PCC player 625, and the extracted data can be transmitted to the client as a file or segment.

[0097] Regarding such data, in the file / segment decapsulation block 615, the file decapsulator processes the file or received fragment and extracts the encoded code stream and parses the metadata, and then in the video decoding box 610 and the image decoding box 611, the encoded point cloud data is decoded and reconstructed, and decoded and reconstructed into point cloud data in the point cloud reconstruction box 612, and the reconstructed point cloud data can be displayed in the display box 614 and / or may first be synthesized according to one or more various scene descriptions in the scene synthesis box 613 based on the scene description data of the scene generator box 609.

[0098] In view of the above, this exemplary V-PCC stream represents one or more of the following advantages over the V-PCC standard, including: one or more partitioning capabilities described for multiple 2D / 3D areas, the ability to combine compressed domains of coded 2D / 3D partitions into a single compliant coded video stream, and a stream extraction capability to extract the coded 2D / 3D of a coded picture into a compliant coded stream, wherein this V-PCC system support is further improved by including container formation for VVC streams to support mechanisms for including metadata that include one or more of the above-mentioned metadata.

[0099] In this case, according to an exemplary embodiment further described below, the term "mesh" refers to a combination of one or more polygons that describe the surface of a volumetric object. Each polygon is defined by its vertices in three-dimensional space and information about how the vertices are connected (referred to as connection information). Optionally, vertex attributes (such as color, normal, etc.) may be associated with mesh vertices. Attributes may also be associated with the surface of the mesh by mapping information that parameterizes the mesh using a 2D attribute map. Such a mapping may be described by a set of parameter coordinates (referred to as UV coordinates or texture coordinates) associated with the mesh vertices. A two-dimensional attribute map is used to store high-resolution attribute information, such as textures, normals, displacements, etc. According to an exemplary embodiment, such information may be used for various purposes, such as texture mapping and shading.

[0100] Nevertheless, dynamic mesh sequences may require a large amount of data because they may consist of a large amount of information that changes over time. For example, in contrast to a "static mesh" or "static mesh sequence", in which the information of the mesh may not change from one frame to another, a "dynamic mesh" or "dynamic mesh sequence" indicates motion in which the vertices represented by the mesh change from one frame to another. Therefore, effective compression techniques are needed to store and transmit such content. Previously, the Moving Picture Experts Group (MPEG) developed mesh compression standards IC, MESHGRID, and FAMC to handle dynamic meshes with constant connectivity and time-varying geometry and vertex attributes. However, these standards do not consider time-varying attribute graphs and connectivity information. Digital Content Creation (DCC) tools usually generate such dynamic meshes. In contrast, it is challenging for volume acquisition techniques to generate dynamic meshes with constant connectivity under real-time constraints. Existing standards do not support this type of content. According to exemplary embodiments herein, aspects of a new mesh compression standard are described to directly handle dynamic meshes with time-varying connectivity information and optionally time-varying property graphs, the standard targeting lossy and lossless compression for various applications such as real-time communications, storage, free viewpoint video, AR and VR. Features such as random access and scalable / progressive coding are also considered.

[0101] Figure 7An example framework 700 for dynamic mesh compression is shown, for example, for a method based on 2D atlas sampling. Each frame of the input mesh 701 can be preprocessed by a series of operations (e.g., tracing, re-meshing, parameterization, voxelization). Note that these operations can be encoder-only, which means that they may not be part of the decoding process, and this possibility can be signaled in the metadata by a flag (e.g., 0 for encoder-only and 1 for other cases). Afterwards, a mesh 702 with a 2D UV atlas can be obtained, where each vertex of the mesh has one or more associated UV coordinates on the 2D atlas. Then, by sampling on the 2D atlas, the mesh can be converted into multiple graphs including a geometric graph and an attribute graph. These 2D graphs can then be encoded by a video / image codec (e.g., HEVC, VVC, AV1, AVS3, etc.). On the decoder 703 side, the mesh can be reconstructed from the decoded 2D graph. Any post-processing and filtering can also be applied to the reconstructed mesh 704. It should be noted that for the purpose of 3D mesh reconstruction, other metadata can be signaled to the decoder side. Note that the graph boundary information (including the uv and xyz coordinates of the boundary vertices) can be predicted, quantized, and entropy encoded in the bitstream. The quantization step size can be configured at the encoder side to trade off between quality and bitrate.

[0102] In some embodiments, a 3D mesh may be segmented into segments (or blocks / graphs), and according to an exemplary embodiment, one or more 3D mesh segments may be considered a "3D mesh". Each segment consists of a set of connected vertices associated with its geometry, attributes, and connectivity information. Figure 8 As shown in the example 800 of volume data in FIG. 1 , a UV parameterization process 802 for mapping from 3D mesh segments to a 2D graph (e.g., to the 2D UV atlas block 702 described above) maps one or more mesh segments 801 to a 2D graph 803 in a 2D UV atlas 804. Each vertex (v n ) will specify a 2D UV coordinate in the 2D UV atlas. Note that the vertex (v n ) form connected components that are their 3D counterparts. The geometry, attributes, and connectivity information of each vertex may also be inherited from their 3D counterparts. For example, information may indicate that vertex v4 is directly connected to vertices v0, v5, v1, and v3, and similarly, information for each other vertex may be similarly indicated. In addition, according to an exemplary embodiment, such a 2D texture mesh will further indicate information, such as color information, on a block-by-block basis (such as by block of each triangle, for example, with v2, v5, and v3 as a "block").

[0103] For example, combined with Figure 8 For features of Example 800, see Fig. 9In the example 900, the 3D mesh segment 801 can also be mapped to multiple independent 2D graphs 901 and 902. In this case, one vertex in 3D can correspond to multiple vertices in the 2D UV atlas. Fig. 9 As shown, in a 2D UV atlas, the same 3D mesh segment is mapped to multiple 2D charts instead of Figure 8 For example, 3D vertices v1 and v4 have two 2D corresponding points v1, v1' and v4, v4' respectively. Therefore, a general 2D UV atlas for a 3D mesh can be represented by Fig.14 The diagram is composed of multiple graphs as shown, each of which may contain multiple (usually greater than or equal to 3) vertices associated with its 3D geometry, properties and connectivity information.

[0104] Fig. 9 Example 903 is shown, which shows a derived triangulation in a graph with boundary vertices B0, B1, B2, B3, B4, B5, B6, and B7. When such information is presented, any triangulation method can be applied to create connections between vertices (including boundary vertices and sampled vertices). For example, for each vertex, find the two closest vertices. Or for all vertices, continuously generate triangles until the minimum number of triangles is obtained after a certain number of attempts. As shown in example 903, there are various regular shaped repeating triangles that are usually closest to the boundary vertices and various strange shaped triangles, which have their own unique sizes and can be shared with other triangles or not shared with other triangles. The connection information can also be reconstructed by explicit signaling. According to an exemplary embodiment, if the polygon cannot be recovered by implicit rules, the encoder can signal the connection information in the code stream.

[0105] Boundary vertices B0, B1, B2, B3, B4, B5, B6, and B7 are defined in 2D UV space. The boundary can be determined by checking whether the boundary only appears in one triangle. According to an exemplary embodiment, the following information of the boundary vertex is important and should be signaled in the codestream: geometric information (e.g., 3D XYZ coordinates, even if currently in 2D UV parameter form) and 2D UV coordinates.

[0106] For the case where a boundary vertex in 3D corresponds to multiple vertices in the 2D UV atlas, such as Fig. 9 As shown, the mapping from 3D UVUZ to 2D UV can be a one-to-many mapping. Therefore, a UV-to-XYZ (or UV2XYZ) index can be signaled to indicate the mapping function. UV2XYZ can be an indexed 1D array that corresponds each 2D UV vertex to a 3D XYZ vertex.

[0107] According to an exemplary embodiment, in order to efficiently represent a mesh signal, a subset of mesh vertices and connection information between the subset of mesh vertices may be encoded first. In the original mesh, the connection between these vertices may not exist because these vertices are subsampled from the original mesh. There are different ways to represent the connection information between vertices, so such a subset is called a base mesh or base vertex.

[0108] According to an exemplary embodiment, a number of methods for dynamic mesh compression are implemented, and these methods are part of the edge-based vertex prediction framework described above, where a base mesh is first encoded, and then more additional vertices are predicted based on the connection information from the base mesh edges. It should be noted that these methods can be applied individually or in any combination.

[0109] For example, consider Fig.10 Example flow chart 1001 of vertex grouping for prediction mode. In S101, vertices in a mesh may be obtained, and in S102, the vertices in the mesh may be divided into different groups for prediction, for example, see Fig. 9. In one example, the partitioning is done using block / graph partitioning in S104. In another example, the partitioning S105 is done under each block / graph. The decision S103 whether to proceed to S104 or S105 may be signaled by a flag or the like. In the case of S105, several vertices of the same block / graph form a prediction group and will share the same prediction mode, while several other vertices of the same block / graph may use another prediction mode. In this article, the "prediction mode" may be considered as a specific mode that the decoder uses to predict the video content including the block, and the prediction mode may be clearly divided into an intra-frame prediction mode and an inter-frame prediction mode, and in each category, there may be different specific modes for the decoder to select from. According to an exemplary embodiment, each group ("prediction group") may share the same specific mode (e.g., an angle mode at a specific angle) or the same classification prediction mode (e.g., all intra-frame prediction modes, but may be predicted at different angles) according to an exemplary embodiment. Such grouping in S106 may be allocated at different levels by determining the corresponding number of vertices involved in each group. For example, according to an exemplary embodiment, every 64 vertices, 32 vertices, or 16 vertices following the scan order within the block / graph will be assigned the same prediction mode, and the other vertices may be assigned different prediction modes. For each group, the prediction mode may be an intra-frame prediction mode or an inter-frame prediction mode. This may be signaled or assigned. According to the example flowchart 1000, if a mesh frame or mesh slice is determined to be an intra-frame type in S107, for example, by checking whether a flag of the mesh frame or mesh slice indicates an intra-frame type, all vertex groups within the mesh frame or mesh slice should use the intra-frame prediction mode; otherwise, in S108, all vertices in each group may select an intra-frame prediction mode or an inter-frame prediction mode.

[0110] In addition, for a mesh vertex group using an intra-frame prediction mode, its vertices can only be predicted by using previously encoded vertices within the same sub-partition of the current mesh. Sometimes, according to an exemplary embodiment, the sub-partition may be the current mesh itself, and for a mesh vertex group using an inter-frame prediction mode, according to an exemplary embodiment, the vertices of the mesh vertex group can only be predicted by using previously encoded vertices from another mesh frame. Each of the above information may be determined and signaled by a flag or the like. The predicted features may occur in S110, and the results of the prediction and signaling may occur in S111.

[0111] According to an exemplary embodiment, for each vertex in a group of vertices in the example flowchart 1000 and the flowchart 1100 described below, after prediction, the residual will be a 3D displacement vector indicating the offset from the current vertex to its predicted vertex. The residual of a group of vertices needs to be further compressed. In one example, before entropy coding, the transformation in S111 and its signaling notification can be applied to the residual of the vertex group. The encoding of a group of displacement vectors can be processed by implementing the following method. For example, in one method, a group of displacement vectors, some displacement vectors or their components are correctly signaled to have only zero values. In another embodiment, a flag is signaled for each displacement vector, which indicates whether the vector has any non-zero vectors, and if not, the encoding of all components of the displacement vector can be skipped. In addition, in another embodiment, a flag is signaled for each group of displacement vectors, which indicates whether the group has any non-zero vectors, and if not, the encoding of all displacement vectors of the group can be skipped. Furthermore, in another embodiment, a flag is signaled for each component of a set of displacement vectors, indicating whether the component of the set has any non-zero vectors, and if not, encoding of the component of all displacement vectors of the set may be skipped. Furthermore, in another embodiment, there is a signaling to indicate that a set of displacement vectors or a component of a set of displacement vectors requires a transformation, and if not, the transformation may be skipped, and quantization / entropy coding may be applied directly to the set or the component of the set. Furthermore, in another embodiment, a flag may be signaled for each set of displacement vectors, indicating whether the set requires a transformation, and if not, transform encoding of all displacement vectors of the set may be skipped. Furthermore, in another embodiment, a flag is signaled for each component in a set of displacement vectors, indicating whether the component of the set requires a transformation, and if not, transform encoding of the component of all displacement vectors of the set may be skipped. The above-mentioned embodiments in this paragraph relating to the processing of vertex prediction residuals may also be combined and implemented in parallel on different blocks, respectively.

[0112] Fig.11An example flow chart 1100 is shown, in which a grid frame may be obtained in S121, which is encoded as a whole data unit, which means that there may be correlations between all vertices or attributes of the grid frame. Alternatively, depending on the determination in S122, the grid frame may be divided into smaller independent sub-partitions in S123, which are conceptually similar to slices or tiles in a 2D video or image. In S124, a prediction type may be assigned to the encoded grid frame or the encoded grid sub-partition. Possible prediction types include intra-frame coding type and inter-frame coding type. For the intra-frame coding type, only prediction based on the reconstructed part of the same frame or slice is allowed in S125. On the other hand, in addition to the grid intra-frame prediction, in S125, the inter-frame prediction type will allow prediction based on the previously encoded grid frame. In addition, the inter-frame prediction type can be classified into more sub-types, such as P type or B type. In the P type, only one predictor can be used for prediction, while in the B type, two predictors from two previously encoded grid frames can be used to generate the predictor. An example may be to take a weighted average of two predictors. When the grid frame is encoded as a whole, the frame may be considered as an intra-coded or inter-coded grid frame. In the case of an inter-grid frame, the P type or B type may be further identified by signaling. Alternatively, if the grid frame is encoded with further intra-frame partitioning, a prediction type is assigned to each sub-partition in S124. Each of the above information may be determined and signaled by a flag or the like, similar to Fig.10 The predicted features may appear in S126, and the result of the prediction and signaling notification may appear in S127.

[0113] Therefore, although dynamic mesh sequences may require a large amount of data due to containing a large amount of time-varying information, efficient compression techniques are still needed to store and transmit such content, and the features described in this article demonstrate this improved efficiency through the following improvements: at least allowing improved prediction of mesh vertex 3D positions by using previously decoded vertices in the same mesh frame (intra-frame prediction) or decoded vertices from previously encoded mesh frames (inter-frame prediction).

[0114] In addition, exemplary embodiments may generate displacement vectors of the third layer 1303 of the mesh based on one or more reconstructed vertices of previous layers (e.g., the second layer 1302 and the first layer 1301) of the third layer 1303 of the mesh. Assuming that the index of the second layer 1302 is T, a predictor of a vertex in the third layer 1303T+1 is generated based on at least the reconstructed vertices of the current layer or the second layer 1302. Examples of such layer-based prediction structures are as follows: Fig.13As shown in Example 1300, Example 1300 shows vertex prediction based on reconstruction: progressive vertex prediction using edge-based interpolation, where the predictor is generated based on previously decoded vertices rather than predictor vertices. The first layer 1301 can be a mesh bounded by a first polygon 1340, which has decoded vertices at its boundaries and interpolated vertices along the lines between those decoded vertices as its vertices. When progressive encoding proceeds from the first layer 1301 to the second layer 1302, an additional polygon 1341 can be formed by the displacement vectors from one of the interpolated vertices of the first layer to the additional vertices of the second layer 1302. Thus, the total number of vertices in the second layer 1302 can be greater than the total number of vertices in the first layer 1301. Similarly, when proceeding to the third layer 1303, the additional vertices of the second layer 1302 and the decoded vertices from the first layer 1301 can be encoded in a manner similar to that provided by the decoded vertices used when proceeding from the first layer 1301 to the second layer 1303; that is, multiple additional polygons can be formed. Note, refer to Fig.14 Example 1400 showing such progressive encoding, as opposed to Fig.13 Example 1400, where each of the additional polygons formed when proceeding from the first layer 1401 to the second layer 1403 and then to the third layer 1403 can be entirely within the polygon formed by the boundaries of the first layer 1401.

[0115] For such Examples 1300 and / or 1400, refer to the Fig.12 example flowchart 1200 according to an exemplary embodiment. Since the interpolated vertices on the current layer are predicted values, such values need to be reconstructed before the predictors for generating the vertices of the next layer. This is done by encoding the base mesh at S131, implementing vertex prediction at S132, and then adding the decoded displacement vectors of the current layer to the predictors of the vertices (e.g., of layer 1302) at S133. Then, the reconstructed vertices of this layer and all the decoded vertices of the previous layer(s) (e.g., checking the additional vertex values of these layers at S134) can be used to generate and signal the predictor vertices of the next layer 1303 at S135. This process can also be summarized as follows: Let P[t](Vi) denote the predictor of vertex Vi on layer t; Let R[t](Vi) denote the reconstructed vertex Vi on layer t; Let D[t](Vi) denote the displacement vector of vertex Vi on layer t; Let f(*) denote the predictor generator, and in particular, f(*) can be the average of two existing vertices. Then, for each layer t, according to the exemplary embodiment, there is the following:

[0116] P[t](Vi) = f(R[ss<t](Vj), R[mm<t](Vk)), where

[0117] Vj and Vk are the reconstructed vertices of the previous layer

[0118] R[t](Vi)=P[t](Vi)+D[t](Vi) - Equation (1)

[0119] Then, for all vertices in a mesh framework, all vertices are divided into layer zero (base mesh), layer one, layer two, and so on. Then, the reconstruction of vertices on a layer depends on the reconstruction of vertices on the previous layer. In the above, P, R, and D represent 3D vectors in the content represented by the 3D mesh, respectively. D is the decoded displacement vector, and quantization may or may not be applied to the vector.

[0120] According to an exemplary embodiment, vertex prediction using reconstructed vertices may be applied only to certain layers. For example, the zeroth layer and the first layer. For other layers, vertex prediction may still use adjacent predictor vertices without adding displacement vectors to other layers for reconstruction. Therefore, these other layers may be processed simultaneously without waiting for a previous layer to reconstruct these other layers. According to an exemplary embodiment, for each layer, whether reconstruction-based vertex prediction or predictor-based vertex prediction is selected may be signaled, or a layer (and its subsequent layers) that does not use reconstruction-based vertex prediction may be signaled.

[0121] For displacement vectors whose vertex predictors are generated by reconstructed vertices, quantization may be applied to these displacement vectors without further performing a transform, such as a wavelet transform, etc. For displacement vectors whose vertex predictors are generated by other predictor vertices, a transform may be required, and quantization may be applied to the transform coefficients of those displacement vectors.

[0122] Therefore, since dynamic mesh sequences may require a large amount of data, because dynamic mesh sequences may consist of a large amount of information that changes over time. Therefore, effective compression technology is needed to store and transmit such content. In the framework of the above-mentioned interpolation-based vertex prediction method, compressing the displacement vector is an important process, and this occupies a major part in the encoded bitstream, which is the focus of the disclosure of this application, and the features disclosed in this application alleviate this problem by providing such compression.

[0123] Furthermore, similar to the other examples described above, even with those embodiments, a dynamic mesh sequence may still require a large amount of data, as it may consist of a large amount of information that changes over time, and therefore, requires efficient compression techniques to store and transmit such content. In the framework of the above 2D atlas sampling method, important advantages can be achieved by inferring connectivity information from the sampled vertices plus the boundary vertices on the decoder side. This is a major part of the decoding process and is the focus of the other examples described below.

[0124] According to an exemplary embodiment, the connectivity information of the base grid may be inferred (derived) from the decoded boundary vertices and sampling vertices of each graph at the encoder and decoder sides.

[0125] Similar to the above, any triangulation method may be used to create connections between vertices (including boundary vertices and sample vertices).According to an exemplary embodiment, the connection type may be signaled in a high-level syntax (eg, sequence header, slice header).

[0126] As described above, the connection information can also be reconstructed by explicit signaling, such as a triangular mesh of an irregular shape. That is, if it is determined that the polygon cannot be restored by implicit rules, the encoder can signal the connection information in the code stream. And according to an exemplary embodiment, the overhead of such explicit signaling can be reduced according to the boundaries of the polygon.

[0127] According to an embodiment, only the connection information between boundary vertices and sampling positions is determined to be signaled, while the connection information between the sampling positions themselves is inferred.

[0128] Furthermore, in any embodiment, connectivity information may be signaled by prediction, such that only the difference in inferred connectivity (as a prediction) from one mesh to another may be signaled in the codestream.

[0129] It should be noted that according to an exemplary embodiment, the inferred triangle direction (e.g., each triangle is inferred in a clockwise or counterclockwise manner) may be signaled for all charts in a high-level syntax (e.g., sequence header, slice header, etc.), or may be fixed (assumed) by the encoder and decoder. The inferred triangle direction may also be signaled differently for each chart.

[0130] It is further noted that any reconstructed mesh may have different connectivity than the original mesh. For example, the original mesh may be a triangular mesh, while the reconstructed mesh may be a polygonal mesh (e.g., a quadrilateral mesh).

[0131] According to an exemplary embodiment, the connection information of any base vertex may not be signaled, instead, the same algorithm may be used on the encoder side and the decoder side to derive the edges between the base vertices. And according to an exemplary embodiment, the interpolation of the predicted vertices of the additional mesh vertices may be implemented based on the derived edges of the base mesh.

[0132] According to an exemplary embodiment, a flag may be used to signal whether the connection information of the base vertex is signaled or derived, and such a flag may be signaled at different levels of the codestream (eg, at a sequence level, a frame level, etc.).

[0133] According to an exemplary embodiment, the edges between the base vertices are first derived using the same algorithm on both the encoder and decoder sides. The derived edges are then compared to the original connections of the base mesh vertices, and the difference between the derived edges and the actual edges is signaled. Thus, after decoding the difference, the original connections of the base vertices can be restored.

[0134] In one example, for a derived edge, if it is determined to be erroneous when compared to an original edge, such information may be signaled in the codestream (by indicating the pair of vertices that form the edge); and for an original edge, if it is not derived, the original edge may be signaled in the codestream (by indicating the pair of vertices that form the edge). In addition, connections related to boundary edges and vertex interpolation involving boundary edges may be performed separately from internal vertices and edges.

[0135] Therefore, through the exemplary embodiments described herein, the above technical problems can be effectively improved through one or more of these technical solutions. For example, since a dynamic grid sequence may require a large amount of data, because a dynamic grid sequence may consist of a large amount of information that changes over time, the exemplary embodiments described herein at least represent an effective compression technology for storing and transmitting such content.

[0136] The above embodiments can also be applied to instance-based mesh coding, where the instance can be a mesh of an object or a mesh of a part of an object. Fig.15 The illustrative example 1500 of shows a grid example 1501, in which there are various instances 1502 (a grid representing a cup), instance 1503 (a grid representing a spoon), and instance 1504 (a grid representing a plate), and these instances can be separated and encoded independently. And each of instance 1501, instance 1502, instance 1503, and instance 1504 is shown separately in a bounding box, but it should be noted that instance 1501 can be considered to be shown as being delimited by a "grid-based bounding box", and instance 1502, instance 1503, and instance 1504 can be considered to be shown as being delimited by an "instance-based bounding box".

[0137] According to an embodiment, as in the MPEG V-DMC WD 2.0 standard, the mesh encoding process begins with preprocessing. The preprocessing converts the input dynamic mesh (denoted as M(i)) into a base mesh m(i) and a set of displacements d(i). The encoder compresses the new representation and generates a compressed code stream b(i).

[0138] According to an embodiment, as in Fig.16In the example 1600 of FIG. 1 , preprocessing includes mesh extraction 1601, followed by atlas parameterization 1602, and then subdivision surface fitting 1603. Mesh extraction 1601 uses a simplification technique to extract the input mesh M(i) and produce an extracted mesh dm(i). The extracted mesh dm(i) is then reparameterized. The resulting mesh is denoted pm(i). Subdivision surface fitting 1603 takes the reparameterized mesh pm(i) and the input mesh M(i) as input and produces a base mesh m(i) and a set of displacements d(i).

[0139] Depend on Fig.17 Example 1700 in shows lossless mesh compression. Duplicate geometry vertices are first merged. In this process, the generation of non-manifold structures should be avoided because the non-manifold structures will be confused with the original non-manifold structures and the processing overhead of the non-manifold structures is large. Before using Edgebreaker to encode the connection information, the non-manifold structure in the input mesh must be converted to a manifold structure. Then, the connection information of the manifold mesh is encoded by Edgebreaker. The geometric information and attribute information are encoded by using a predictive encoder. The texture map is compressed by using a video codec. On the decoder side, duplicate vertices and non-manifolds are losslessly reconstructed.

[0140] In lossless mesh compression, geometric compression consists of vertex position compression, 2D texture coordinate compression, and other geometric information (such as normals, etc.). And this article discloses an embodiment of a method and system for vertex position compression for lossless mesh compression. One of the basic ideas of vertex position compression according to the embodiments of this article is parallelogram prediction. The compression algorithm uses a triangle to introduce a new vertex from an edge. Fig.18 As shown in element 1801 of example 1800, the predicted position of the new vertex forms a parallelogram with two shared vertices of the opposite triangle and the third vertex. The multi-parallelogram prediction uses the average position given by two or more parallelogram predictions as much as possible. And element 1802 gives an example of a double parallelogram prediction.

[0141] The proposed methods can be used alone or in any order. In addition, each of the method (or embodiment), encoder and decoder can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0142] In the present application disclosure, many methods and systems for vertex position encoding in mesh compression are proposed. Note that these methods can be applied alone or in any form of combination. It should also be noted that these methods can be applied not only to dynamic meshes, but also to static meshes, where there is only one mesh frame or the mesh content does not change over time. In addition, the disclosed methods and systems are not limited to vertex position encoding. The disclosed methods and systems can also be applied to, for example, 2D texture coordinate encoding as a more general multi-prediction based scheme.

[0143] All triangles in a triangle mesh can be sorted. The order can be a traversal in the Edgebreaker algorithm or other algorithms. All vertices can be sorted. The order can be a traversal in the Edgebreaker algorithm or other algorithms. The order of triangles and vertices can be based on the same scheme (such as traversal) or based on different schemes.

[0144] In a triangle mesh, each triangle (also called a face) has three vertices. For two triangles that share an edge, we can apply parallelogram prediction to use one of the two opposite vertices as the predictor for the other of the two opposite vertices. Fig.18 As shown in example 1800, two triangles ABC and DBC share edge BC. Embodiments herein use the coordinates of A, B, and C (eg, assuming A, B, and C have been encoded) to predict the coordinates of D. The predicted coordinates D' will be

[0145] D' = B + C - A - Equation (2)

[0146] Thus, the four vertices (D', B, A, C) form a parallelogram. Because each vertex has a 3D coordinate, the above equation calculates each coordinate component-wise. For example, assuming that the subscripts x, y, and z represent 3D coordinates in xyz space, then

[0147] D x ' = B x + C x - A x - Equation (3)

[0148] D y ' = B y + C y - A y - Equation (4)

[0149] D z ' = B z + C z - A z - Equation (5)

[0150] If the position values ​​of A, B, and C have been encoded and can be used for prediction, the embodiments herein refer to triangle ABC as a prediction candidate for vertex D. Depending on the number of existing encoded vertices and shared edges, a vertex may contain zero, one, two, or more prediction candidates. If there is more than one prediction candidate, these prediction candidates are sorted based on the triangle order of the mesh.

[0151] According to an embodiment of the present invention, the vertex position encoding calculates the average of the available parallelogram predictions and uses the average as the prediction value. The associated prediction error (also called prediction residual) is encoded.

[0152] In addition, according to embodiments of the present invention, there is a position prediction. In a triangle mesh, all triangles are ordered. In addition, according to embodiments, all vertices are sorted.

[0153] Fig. 20 An example 2000 is shown according to an embodiment, in which a mesh position coding parameter syntax is shown and various syntaxes may be signaled, such as mesh_position_prediction_max_parallelograms_minus1 syntax, etc. mesh_position_prediction_max_parallelograms_minus1+1 specifies the maximum number of parallelograms used in mesh position prediction. According to an embodiment, mesh_position_prediction_max_parallelograms_minus1 should be in the range of 0 to 15.

[0154] According to an embodiment, there is a prediction for each vertex position attribute. For example, the input to the process can be a variable c that specifies the angular index to be used to predict the vertex position, and the output of the process can be indirect and modify the array VertextMarkingArray of clause I.9.2 and the array VertCoordValues ​​of clause I.9.1. According to an embodiment, the variable OppositeCornersArray is aliased O; the variable CornerToVertexArray is aliased V; the variable VertCoordValues ​​is aliased G; and the variable VertexMarkingArray is aliased MV. / / maxParallelograms: mesh_position_prediction_max_parallelograms_minus1 u(4). Not yet implemented in ref software using constants.

[0155] Let the variable maxParallelolograms (specifying the maximum number of parallelogram projections) and the variable v (specifying the index of the vertex position associated with c) be initialized as follows:

[0156] maxParallelograms=4

[0157] v=GetVertexIndex(CornerToVertexArray,c)

[0158] According to an embodiment, if MV[v] is strictly greater than 0, the vertex v has been predicted, and the process terminates and returns. Otherwise, the following applies:

[0159] / / Mark the vertices

[0160] MV[v]=1

[0161] / / Search for some parallelogram estimates around the corner vertices.

[0162] / / The triangle fan may not be complete,

[0163] / / But it is known that a vertex is always manifold, so each vertex has only one fan

[0164] / / In addition, due to the boundary, some diagonals may not be defined

[0165] So we use the OV accessor and test the result to filter out negative values.

[0166] According to an embodiment, let a 1D array predPos of size 3 (specifying the predicted position of the vertex associated with c, specifying the number of valid parallelograms found), and variables altC (specifying the corner index) and nextC (specifying the corner index) be initialized as follows:

[0167]

[0168] According to an embodiment, let the variable isBoundary specify whether nextC is on the boundary, let the variable count specify the number of valid parallelogram predictions found, let the variable startC specify the index of the extreme corner point of the fan, and initialize as follows:

[0169] isBoundary = (nextC! = c)

[0170] startC=altC

[0171] count=0

[0172] According to an embodiment, the variables prevV, oppoV, and nextV are set to 0, specifying the indices of the vertices associated with the previous corner point, the opposite corner point, and the next corner point, respectively. Therefore, the following applies:

[0173]

[0174]

[0175]

[0176] According to an embodiment, let variable b specify the index of the previous corner point on the boundary. Let variable bV specify the index of the vertex associated with b.

[0177]

[0178]

[0179] In addition, according to an embodiment, a method and system for encoding the vertex positions of a base mesh are disclosed. Fig.21 As shown in Example 2100, as in MPEG V-DMC WD 2.0, vertex position coding consists of bit depth quantization S2101, multi-parallelogram prediction S2102 and entropy coding S2103. In bit depth quantization S2101, the position of the vertex is quantized to a bit depth value specified by the encoder. For example, if the bit depth value is 12, the position is quantized to an integer between 0 and 4095, including 0 and 4095. The quantized position is then predicted by a multi-parallelogram algorithm in S2102, and the prediction residual is further compressed by entropy coding in S2103. According to the embodiments of this invention, although in MPEG V-DMC WD 2.0, the bit depth value of position quantization is limited, through the disclosure of this application, there is a vertex position coding system that is improved by having fine granularity quantization.

[0180] For example, Fig. 22 As shown in Example 2200, for improved vertex position coding with fine-grained quantization, multi-parallelogram prediction S2201, quantization S2201, inverse quantization S2205 and entropy coding S2204 are provided according to the embodiments disclosed in the present application.

[0181] Through multi-parallelogram prediction, the position of the current vertex is predicted by the encoded vertex positions.

[0182] The prediction residual is the difference between the true value and the predicted value of the position, and is quantized by the quantization step value. The quantization step value can be a positive integer, a positive rational number, or a positive real number.

[0183] The quantized prediction residual is further compressed by entropy coding, which can be fixed-length coding, variable-length coding, Huffman coding, or arithmetic coding.

[0184] At the same time, the quantized prediction residuals are dequantized. The reconstructed position of the current vertex is calculated by adding the prediction residuals to the multi-parallelogram predictions. The reconstructed position of the current vertex is used as a reference for future vertices.

[0185] If the vertex is the first vertex in a connected component of the mesh (to be encoded), no prediction is done. Instead, the position is encoded directly. The position is quantized by a quantization step value. The quantization step value can be a positive integer, a positive rational, or a positive real number. The quantized position is further compressed by entropy coding. Entropy coding can be fixed length coding, variable length coding, Huffman coding, arithmetic coding, etc. At the same time, the quantized position is dequantized to the reconstructed position of the current vertex and used as a reference for future vertices.

[0186] If the vertex is the second or third vertex (to be encoded) in the connected component of the mesh, no parallelogram prediction will be performed because there are not enough encoded vertices for parallelogram prediction. Instead, the previously encoded vertex positions will be used for prediction. Therefore, the encoded positions of the first vertex and the second vertex will be used for the prediction of the second vertex and the third vertex, respectively. The prediction residual is then calculated and quantized by the quantization step value. The quantization step value can be a positive integer, a positive rational number, or a positive real number. The quantized prediction residual is further compressed by entropy coding. Entropy coding can be fixed length coding, variable length coding, Huffman coding, or arithmetic coding, etc. At the same time, the quantized prediction residual is dequantized. The reconstructed position of the current vertex is calculated by adding the dequantized prediction residual to the predicted value of the current vertex. The reconstructed position of the current vertex will be used as a reference for future vertices.

[0187] According to the embodiments described above, an improvement in base grid quantization is proposed. The improvement provides fine granularity quantization for position, uv coordinates and dynamic fields in base grid coding. Experimental results show that under the disclosed C1 and C2 conditions, the BD rate (total grid bitstream) of D1 / D2 is reduced by at least 1.0%.

[0188] In V-DMC, the quantization for position, uv coordinates and motion field in base grid coding is bit depth quantization. Bit depth quantization is coarse and the choice of quantization step size is limited. In this paper, we propose an improvement to base grid quantization. This improvement provides fine-grained quantization for position, uv coordinates and motion field in base grid coding.

[0189] In mesh-based edge break (MEB) (MPEG EdgeBreaker), the position value of the vertex can be quantized with fine granularity by integers, rational numbers or real numbers. Therefore, the embodiment of this article uses the binary rational number qp=m / 2 n As the quantization step size, where m and n are both integers, and m>0, n>=0. Quantization step size schemes such as those in HEVC and VVC may also be used.

[0190] exist Fig.23 In example 2300, the quantization step size qp can be signaled in the MEB code stream. According to an embodiment, example 2300 represents a grid position coding parameter syntax according to an embodiment, which is semantically:

[0191] mesh_position_quantization_step_size_log2_denominator is the log2 of the denominator of the position quantization step size, and its value is between 0 and 7, including 0 and 7.

[0192] mesh_position_quantization_step_size_numerator_minus1+1 is the value of the numerator of the position quantization step size, and its value is between 1 and 256, including 1 and 256.

[0193] After position encoding in MEB, the reconstructed position values ​​used for reference are scaled to the inter-frame dynamic range.

[0194] For uv, quantization can be improved after uv prediction in MEB. The embodiments of this article use binary rational number quv as the quantization step size. Quantization step size schemes such as those in HEVC and VVC can also be used. The prediction residual is quantized by the quantization step size quv, and the quantized prediction residual is further compressed by entropy coding. Fig.24 Example 2400 of quantization step size quv is signaled in an MEB code stream, the example 2400 representing grid attribute coding parameter syntax,

[0195] mesh_attribute_quantization_step_size_log2_denominator[i] is semantically the log2 of the denominator of the i-th attribute quantization step size, and has a value between 0 and 7, inclusive.

[0196] mesh_attribute_quantization_step_size_numerator[i] is the value of the numerator of the i-th attribute quantization step size, and its value is between 0 and 255, including 0 and 255.

[0197] At the same time, the quantized prediction residual is dequantized. The reconstructed value is calculated by adding the dequantized prediction residual to the predicted value. The reconstructed value will be used as a reference for future vertices.

[0198] In addition, according to an embodiment, for dynamic field coding, the position values ​​of vertices can be quantized at a fine granularity. An example embodiment uses a binary rational number qm as a quantization step size. Quantization step size schemes such as those in HEVC and VVC can also be used.

[0199] The quantization step size qm may be signaled in the base grid code stream, such as in the syntax of the general base grid sequence parameter set RBSP. Fig.25 In the example 2500, it is semantically:

[0200] bmsps_inter_quantization_step_size_log2_denominator is the log2 of the denominator of the dynamic field quantization step size, and the value is between 0 and 7, including 0 and 7;

[0201] bmsps_inter_quantization_step_size_numerator_minus1+1 is the value of the numerator of the dynamic field quantization step size, and the value is between 1 and 256, including 1 and 256.

[0202] After the dynamic field encoding, the position values ​​are reconstructed and used as reference for future vertices.

[0203] The experimental results of this paper are Fig.26 2600. For example, position quantization result 2601 is shown, and 300 frame results are summarized in result 2601 for position quantization improvement. Under C1 condition, on average, the BD rate reduction of D1 / D2 / luminance of the total grid code stream is -0.2% / -0.3% / 0.0%, respectively. Under C2 condition, on average, the BD rate reduction of D1 / D2 / luminance of the total grid code stream is -0.1% / -0.1% / 0.0%, respectively. These results demonstrate BD rate reduction, and BD rate reduction can be further achieved by selecting an "optimized" quantization step size qp.

[0204] For UV quantization, see result 2602, where 300 frame results are summarized in result 2602 for UV quantization improvement. Under C1 condition, on average, the BD rate reduction of D1 / D2 / luminance of the total grid code stream is -1.1% / -1.1% / -0.3%, respectively. Under C2 condition, on average, the BD rate reduction of D1 / D2 / luminance of the total grid code stream is -0.8% / -0.8% / -0.1%, respectively. These results demonstrate BD rate reduction, and further BD rate reduction can be achieved by selecting an "optimized" quantization step size quv.

[0205] For combined position / UV quantization, see result 2603, where improved quantization is performed on position and UV coordinates using quantization steps qp and quv, respectively. The 300 frame results are summarized in result 2603. Under C1 conditions, on average, the BD rate reduction for D1 / D2 / luminance of the total grid code stream is -1.4% / -1.4% / -0.3%, respectively. Under C2 conditions, on average, the BD rate reduction for D1 / D2 / luminance of the total grid code stream is -0.9% / -0.9% / -0.1%, respectively. These results demonstrate BD rate reduction, and further BD rate reduction can be achieved by choosing "optimized" quantization steps qp and quv.

[0206] For dynamic field quantization, see result 2604, where 300 frame results are summarized for dynamic field quantization improvement. For dynamic field quantization testing, four sequences were evaluated under CTC conditions, namely solider / levi / mitch / thomas. Under C2 conditions, on average, the BD rate reduction for D1 / D2 / luminance of the total grid code stream was -0.2% / -0.2% / 0.0%, respectively. These results demonstrate BD rate reduction, and further BD rate reduction can be achieved by choosing an "optimized" quantization step size qm.

[0207] For combined position / UV / dynamic field quantization, see result 2605, where the results of fine quantization of position, UV coordinates and dynamic field are represented using quantization step sizes qp, quv and qm respectively. 300 frames of results are summarized in result 2605. Under C1 conditions, on average, the BD rate reduction of D1 / D2 / luminance of the total grid code stream is -1.4% / -1.4% / -0.3%, respectively. Under C2 conditions, on average, the BD rate reduction of D1 / D2 / luminance of the total grid code stream is -1.0% / -1.0% / -0.1%, respectively. These results demonstrate BD rate reduction, and BD rate reduction can be further achieved by selecting "optimized" quantization step sizes qp, quv and qm.

[0208] Therefore, through the embodiments of this invention, there are quantization improvements for base grid coding. The quantization improvements provide fine granularity quantization for position, uv coordinates and dynamic fields in base grid coding, and the BD rate of D1 / D2 (of the total grid code stream) is reduced by more than 1.0% under C1 and C2 conditions.

[0209] According to an embodiment, vertex position coding with fine granularity quantization is provided. As described above, the vertex position coding includes multi-parallelogram prediction, quantization, inverse quantization and entropy coding.

[0210] The proposed methods can be used alone or in any order. In addition, each of the method (or embodiment), encoder and decoder can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0211] In the present disclosure, many methods and systems for vertex position encoding in mesh compression are proposed. Note that these methods can be applied alone or in any combination. In addition, the disclosed methods and systems are not limited to vertex position encoding. The disclosed methods and systems can also be applied to, for example, 2D texture coordinate encoding, dynamic field encoding, etc.

[0212] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media, or implemented by one or more specially configured hardware processors. Fig. 27 A computer system 2700 suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0213] Computer software may be encoded using any suitable machine code or computer language, which may be assembled, compiled, linked or similarly constructed to create code comprising instructions that may be directly executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc., or executed through interpretation, microcode, etc.

[0214] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, IoT devices, etc.

[0215] Fig. 27The components shown for computer system 2700 are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing the embodiments disclosed herein. The configuration of components should also not be interpreted as having any dependency or requirement related to any one or combination of components shown in the exemplary embodiment of computer system 2700.

[0216] Computer system 2700 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users, for example, through: tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not depicted). Human-machine interface devices may also be used to capture certain media that are not necessarily directly related to a person's conscious input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video), etc.

[0217] The input human-machine interface device may include one or more of the following (only one of each is shown): keyboard 2701 , mouse 2702 , touch pad 2703 , touch screen 2710 , joystick 2705 , microphone 2706 , scanner 2708 , camera 2707 .

[0218] The computer system 2700 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback through a touch screen 2710 or a joystick 2705, but may also be a tactile feedback device that is not an input device), audio output devices (e.g., speakers 2709, headphones (not shown)), visual output devices (e.g., screens 2710 including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light-emitting diode (OLED) screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities, some of which are capable of outputting two-dimensional visual outputs or outputs of more than three dimensions through devices such as stereoscopic picture outputs, virtual reality glasses (not shown), holographic displays and smoke boxes (not shown), and printers (not shown).

[0219] The computer system 2700 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 2720 with CD / DVD 2711 or similar media, thumb drive 2722, removable hard drive or solid state drive 2723, traditional magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security software dogs (not shown), and the like.

[0220] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0221] The computer system 2700 may also include an interface 2799 to one or more communication networks 2798. For example, the network 2798 may be a wireless network, a wired network, an optical network. The network 2798 may also be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, and the like. Examples of the network 2798 include: a local area network such as Ethernet, a wireless LAN, a cellular network including Global System for Mobile communications (GSM), 3G, 4G, 5G, Long-Term Evolution (LTE), etc., a television wired or wireless wide area digital network including cable television, satellite television, and terrestrial broadcast television, a vehicle and industrial network including a CAN bus, and the like. Some networks 2798 typically require external network interface adapters connected to some general data ports or peripheral buses (2750 and 2751) (e.g., USB ports of computer system 2700); other network interfaces are typically integrated into the kernel of computer system 2700 by connecting to the system bus described below (e.g., connecting to an Ethernet interface in a PC computer system or connecting to a cellular network interface in a smartphone computer system). Computer system 2700 can use any of these networks 2798 to communicate with other entities. Such communications can be one-way receive only (e.g., broadcast television), one-way send only (e.g., CANBus to some CANBus devices), or two-way, such as connecting to other computer systems using a local area network or wide area network digital network. As described above, certain protocols and protocol stacks can be used on each of those networks and network interfaces.

[0222] The above-mentioned human-machine interface device, human-accessible storage device, and network interface may be attached to the kernel 2740 of the computer system 2700 .

[0223] The core 2740 may include one or more central processing units (CPUs) 2741, graphics processing units (GPUs) 2742, graphics adapters 2717, dedicated programmable processing units (FPGAs) in the form of field programmable gate areas (FPGAs) 2743, hardware accelerators 2744 for certain tasks, and the like. These devices may be connected through a system bus 2748 together with a read-only memory (ROM) 2745, a random access memory 2746, and an internal mass storage 2747 such as an internal non-user accessible hard drive, SSD, and the like. In some computer systems, the system bus 2748 may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, and the like. Peripheral devices may be directly attached to the core's system bus 2748, or connected to the core's system bus 2748 through a peripheral device bus 2749. The architecture of the peripheral bus includes PCI, USB, and the like.

[0224] The CPU 2441, GPU 2442, FPGA 2443, and accelerator 2744 may execute certain instructions, which may be combined to form the above-mentioned computer code. The computer code may be stored in ROM 2445 or RAM 2446. Transition data may also be stored in RAM 2446, while permanent data may be stored in, for example, internal mass storage 2747. Fast storage and retrieval of any memory device may be achieved by using a cache, which may be closely associated with one or more CPUs 2441, GPU 2442, mass storage 2747, ROM 2445, RAM 2446, etc.

[0225] The computer readable medium may have thereon computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes disclosed in this application, or the medium and computer code may be of a type well known and available to those skilled in the art of computer software.

[0226] As an example and not a limitation, a computer system having structure 2700, especially having kernel 2740, can provide functionality due to (one or more) processors (including CPU, GPU, FPGA, accelerator, etc.) executing software contained in one or more tangible computer-readable media. According to specific needs, the computer-readable medium may include one or more storage devices or chips. The software can enable the kernel 2740, especially the processor therein (including CPU, GPU, FPGA, etc.) to perform a specific process or a specific part of a specific process described in the disclosure of this application, including defining a data structure stored in RAM 2746 and modifying such data structure according to a process defined by the software. In addition or alternatively, the computer system can provide functionality due to logic hardwired or otherwise embodied in a circuit (e.g., accelerator 2744), which can replace the software or run with the software to perform a specific process or a specific part of a specific process described herein. Where appropriate, the part referring to the software may include logic, and vice versa. Where appropriate, the part referring to the computer-readable medium may include a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0227] Although the present application discloses a plurality of exemplary embodiments, there are modifications, substitutions and various replacement equivalents that fall within the scope of the present application discloses. Therefore, it should be understood that those skilled in the art will be able to design a variety of systems and methods, which, although not explicitly shown or described in the present application discloses, embody the principles disclosed in the present application and therefore fall within the spirit and scope of the present application discloses.

Claims

1. A method for video decoding, the method being executed by at least one processor, the method comprising: obtaining a mesh from a bitstream, the mesh representing encoded volume data of at least one three-dimensional (3D) visual content; as well as decoding the encoded volume data based on a base grid quantization of the grid, Among them, the base grid quantization includes: predicting the current vertex using the encoded vertex position through multi-parallelogram prediction, quantizing the prediction residual of the current vertex through a quantization step value, and compressing the quantized prediction residual and reconstructing the position of the current vertex by adding the inverse quantization prediction residual of the quantized prediction residual to the multi-parallelogram prediction for use as a reference for other vertices.

2. The video decoding method according to claim 1, wherein: The base grid quantization further includes: after quantizing the prediction residual and before reconstructing the position of the current vertex, dequantizing the quantized prediction residual.

3. The video decoding method according to claim 1, wherein: The base grid quantization includes: based on a vertex of the grid being determined as the first vertex of the grid to be encoded in time sequence, by quantizing the position of the vertex with a positive integer and entropy encoding the quantized position, determining to directly encode the position of the vertex of the grid.

4. The video decoding method according to claim 1, wherein: The quantization step value is a binary rational number, including a numerator m and a denominator, the denominator is 2 to the power of n, and the numerator m and the value n are both integers.

5. The video decoding method according to claim 4, in, The numerator m is a first integer greater than zero, and The value n is a second integer greater than or equal to zero.

6. The video decoding method according to claim 5, in, The quantization step value is signaled by at least signaling the numerator m and the value n in the code stream.

7. The video decoding method according to claim 4, in, The quantization step value is signaled in the codestream via the "mesh_position_quantization_step_size_log2_denominator" syntax and the "mesh_position_quantization_step_size_numerator_minus1" syntax.

8. A video decoding device, the device comprising: at least one memory configured to store computer program code; as well as At least one processor is configured to access the computer program code and operate according to the instructions of the computer program code, wherein the computer program code comprises: Obtaining code configured to cause the at least one processor to implement: obtaining a mesh from a bitstream, the mesh representing encoded volume data of at least one three-dimensional (3D) visual content; and decoding code configured to cause the at least one processor to decode the encoded volume data based on a base grid quantization of the grid, Among them, the base grid quantization includes: predicting the current vertex using the encoded vertex position through multi-parallelogram prediction, quantizing the prediction residual of the current vertex through a quantization step value, and compressing the quantized prediction residual and reconstructing the position of the current vertex by adding the inverse quantization prediction residual of the quantized prediction residual to the multi-parallelogram prediction for use as a reference for other vertices.

9. The video decoding apparatus according to claim 8, wherein: The base grid quantization further includes: after quantizing the prediction residual and before reconstructing the position of the current vertex, dequantizing the quantized prediction residual.

10. The video decoding apparatus according to claim 8, wherein: The base grid quantization includes: based on a vertex of the grid being determined as the first vertex of the grid to be encoded in time sequence, by quantizing the position of the vertex with a positive integer and entropy encoding the quantized position, determining to directly encode the position of the vertex of the grid.

11. The video decoding apparatus according to claim 8, wherein: The quantization step value is a binary rational number, including a numerator m and a denominator, the denominator is 2 to the power of n, and the numerator m and the value n are both integers.

12. The video decoding device according to claim 11, in, The numerator m is a first integer greater than zero, and The value n is a second integer greater than or equal to zero.

13. The video decoding apparatus according to claim 12, in, The quantization step value is signaled by at least signaling the numerator m and the value n in the code stream.

14. The video decoding apparatus according to claim 12, in, The quantization step value is signaled in the codestream via the "mesh_position_quantization_step_size_log2_denominator" syntax and the "mesh_position_quantization_step_size_numerator_minus1" syntax.

15. A non-transitory computer-readable medium storing a program, the program causing a computer to: obtaining a mesh from a bitstream, the mesh representing encoded volume data of at least one three-dimensional (3D) visual content; and decoding the encoded volume data based on a base grid quantization of the grid, in, The base grid quantization includes: predicting the current vertex using the encoded vertex position through multi-parallelogram prediction, quantizing the prediction residual through the quantization step value, and compressing the quantized prediction residual and reconstructing the position of the current vertex by adding the inverse quantization prediction residual of the quantized prediction residual to the multi-parallelogram prediction for reference to other vertices.

16. The non-transitory computer readable medium of claim 15, wherein: The base grid quantization further includes: after quantizing the prediction residual and before reconstructing the position of the current vertex, dequantizing the quantized prediction residual.

17. The non-transitory computer readable medium of claim 15, wherein: The base grid quantization includes: based on determining a vertex of the grid as the first vertex to be encoded in the grid in time sequence, determining to directly encode the position of the vertex of the grid by quantizing the position of the vertex with a positive integer and entropy encoding the quantized position.

18. The non-transitory computer readable medium of claim 15, wherein: The quantization step value is a binary rational number, including a numerator m and a denominator, the denominator is 2 to the power of n, and the numerator m and the value n are both integers.

19. The non-transitory computer readable medium of claim 18, in, The numerator m is a first integer greater than zero, and The value n is a second integer greater than or equal to zero.

20. The non-transitory computer readable medium of claim 19, in, The quantization step size value is signaled by signaling at least the numerator m and the value n in the codestream using a "mesh_position_quantization_step_size_log2_denominator" syntax and a "mesh_position_quantization_step_size_numerator_minus1" syntax.