Method, apparatus, and program for vertex position compression based on temporal prediction

The method for vertex position compression in dynamic meshes using temporal prediction addresses inefficiencies in existing standards by ordering vertices and generating prediction residuals, improving data compression and transmission efficiency.

JP7701560B2Active Publication Date: 2025-07-01TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024515865
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-03-28
Filing Date
2023-04-04
Publication Date
2025-07-01
Estimated Expiration
2043-04-04

AI Technical Summary

Technical Problem

Existing mesh compression standards do not effectively handle dynamic meshes with time-varying connectivity and attribute maps, leading to inefficiencies in data storage and transmission.

Method used

A method and system for vertex position compression using temporal prediction, which orders vertices based on the edgebreaker algorithm, determines adjacent estimation errors, and generates prediction residuals using inter-frame and intra-frame predictions to encode vertex positions efficiently.

Benefits of technology

This approach reduces the data required to represent dynamic meshes by effectively utilizing temporal prediction to minimize vertex position changes over time, enhancing data compression efficiency and transmission speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007701560000001
    Figure 0007701560000001
  • Figure 0007701560000002
    Figure 0007701560000002
  • Figure 0007701560000003
    Figure 0007701560000003
Patent Text Reader

Abstract

A plurality of adjacent vertices of a current vertex in a current frame of a mesh is determined. The current frame corresponds to the mesh at a first time point. Each of the plurality of adjacent vertices is connected to the current vertex through a respective edge in the mesh. A plurality of adjacent estimation errors of the plurality of adjacent vertices is determined. Each of the plurality of adjacent estimation errors indicates a difference between a reference vertex of a corresponding one of the plurality of adjacent vertices in a reference frame and the corresponding one of the plurality of adjacent vertices in the current frame. The reference frame corresponds to the mesh at a second time point. A prediction residual of the current vertex is determined based on the plurality of adjacent estimation errors of the plurality of adjacent vertices. Prediction information of the current vertex is generated based on the determined prediction residual.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Incorporation by Reference This application claims the benefit of priority to U.S. Patent Application No. 18 / 127,487, filed Mar. 28, 2023, entitled “Vertex Position Compression Based on Temporal Prediction,” which claims the benefit of priority to U.S. Provisional Application No. 63 / 345,824, filed May 25, 2022, entitled “Vertex Position Compression Based on Temporal Prediction.” The disclosure of the prior application is hereby incorporated by reference in its entirety.

[0002] Technical Field The present disclosure includes embodiments related to mesh processing.

Background Art

[0003] The background description provided herein is for the purpose of generally indicating the context of the present disclosure. The achievements of the inventors represented in this application, to the extent that they are not already described in this background section, are not to be recognized as prior art to the present disclosure, whether explicitly or implicitly, any more than aspects of this specification that may not qualify as prior art at the time of filing.

[0004] Advances in three-dimensional (3D) capture, modeling, and rendering have made 3D content ubiquitous across various platforms and devices. Today, it is possible to capture a baby's first steps on one continent and have the baby's grandparents on another continent view (and in some cases interact with) that child and enjoy a fully immersive experience. To achieve such realism, models have become increasingly sophisticated, and a significant amount of data is linked to the creation and consumption of those models. 3D meshes are widely used to represent such immersive content.

Summary of the Invention

Problems to be Solved by the Invention

[0005] Aspects of the present disclosure provide a method and apparatus for mesh processing. In some examples, the apparatus for mesh processing includes a processing circuit.

Means for Solving the Problems

[0006] According to an aspect of the present disclosure, a method for mesh processing executed in a video encoder is provided. In this method, in the current frame of the mesh in one of the two-dimensional model and the three-dimensional model, a plurality of adjacent vertices of the current vertex are determined. The current vertex and the plurality of adjacent vertices are included in the current frame and correspond to the mesh at a first time point. Each of the plurality of adjacent vertices is connected to the current vertex through each edge in the mesh. A plurality of adjacent estimation errors of the plurality of adjacent vertices of the current vertex are determined. Each of the plurality of adjacent estimation errors indicates the difference between the corresponding reference vertex of the plurality of adjacent vertices in the reference frame of the mesh and the corresponding one of the plurality of adjacent vertices in the current frame. The reference frame corresponds to the mesh at a second time point. Based on the plurality of adjacent estimation errors of the plurality of adjacent vertices, a prediction residual of the current vertex is determined. Based on the determined prediction residual of the current vertex, prediction information of the current vertex is generated.

[0007] In some embodiments, the corresponding reference vertex among the plurality of adjacent vertices is located at the same relative position as the current vertex in the current frame in the reference frame, and the reference frame and the current frame are generated at different time points.

[0008] In one example, an average adjacent estimation error of the plurality of adjacent estimation errors is determined. Based on the average adjacent estimation error, a prediction residual of the current vertex is determined.

[0009] In one example, to determine the prediction residual of the current vertex, an estimated error of the current vertex is determined based on the difference between the reference vertex of the current vertex and the current vertex. A prediction list of the current vertex is determined. The predictors of the prediction list include the estimated error of the current vertex, an average adjacent estimated error after the estimated error of the current vertex, and a plurality of adjacent estimated errors after the average adjacent estimated error. Each prediction index is associated with each of the predictors in the prediction list.

[0010] In some embodiments, the plurality of adjacent vertices are ordered based on an Edgebreaker algorithm in which the plurality of adjacent vertices are traversed in a spiraling triangle-spanning-tree order.

[0011] In one example, to determine the prediction residual of the current vertex, a difference between the average adjacent estimated error and the estimated error of the current vertex is determined. A difference between each of the plurality of adjacent estimated errors and the estimated error of the current vertex is also determined. An adjacent estimated error is selected from among the average adjacent estimated error and the plurality of adjacent estimated errors. The selected adjacent estimated error has the smallest difference. The prediction residual of the current vertex is determined as either (i) the estimated error of the current vertex or (ii) the selected adjacent estimated error.

[0012] In one example, to determine the prediction residual of the current vertex, according to inter-frame prediction, a prediction error is determined based on one of the plurality of adjacent vertices of the current vertex. The prediction error is compared with each predictor in the prediction list. The prediction residual is determined as the prediction error in response to the prediction error being smaller than the predictors in the prediction list.

[0013] In one example, to determine the prediction residual of the current vertex, the prediction residual is determined as a mixed estimated error. The mixed estimated error includes either (i) an average of the selected adjacent estimated error and the prediction error, and (ii) an average of the estimated error of the current vertex and the prediction error.

[0014] In one embodiment, the inter-frame prediction further includes one of the delta predictions based on the difference between the first predicted vertex and the current vertex, where the first predicted vertex includes said one of the plurality of adjacent vertices of the current vertex. The inter-frame prediction can also include a parallelogram prediction in which a second predicted vertex is determined based on the predicted triangles of the plurality of triangles of the mesh. The predicted triangle shares an edge with a triangle among the plurality of triangles including the current vertex. The second predicted vertex is on the opposite side of the shared edge, and the second predicted vertex and the predicted triangle form a parallelogram.

[0015] In one embodiment, the prediction information of the current vertex further includes a flag. The flag indicates that the prediction residual is one of (i) the prediction error of the current vertex or the selected adjacent estimation error based on the first value of the flag, (ii) the prediction error based on the second value of the flag, and (iii) the mixed estimation error based on the third value of the flag.

[0016] In one embodiment, the prediction information of the current vertex further includes index information indicating which predictor in the prediction list is the prediction residual based on the flag having the first value. The prediction information also includes prediction residual information. In one example, based on the flag having the first value, the prediction residual information is (i) the estimated error of the current vertex in response to the index information indicating that the prediction residual is the estimated error of the current vertex, and (ii) the selected adjacent estimated error in response to the index information indicating that the prediction residual is the selected adjacent estimated error. Based on the flag having the second value, it indicates the difference between the current vertex and the predicted vertex indicated by the prediction error. Based on the flag having the third value, the prediction residual information indicates the difference between the current vertex and the predicted vertex indicated by the mixed estimation error.

[0017] According to another aspect of the present disclosure, an apparatus is provided. The apparatus includes a processing circuit. The processing circuit can be configured to execute any of the described mesh processing methods.

[0018] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to execute any of the mesh processing methods described herein.

Brief Description of the Drawings

[0019] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

[0020]

Figure 1

[0021]

Figure 2

[0022]

Figure 3

[0023]

Figure 4

[0024]

Figure 5

[0025]

Figure 6A

[0026]

Figure 6B

[0027]

Figure 7

[0028]

Figure 8

[0029]

Figure 9

DETAILED DESCRIPTION OF THE INVENTION

[0030] FIG. 1 shows a block diagram of a video processing system (100) in some examples. The video processing system (100) is an example of the application of the disclosed subject matter and is a video encoder and a video decoder in a streaming environment. The disclosed subject matter is similarly applicable to other image and video-related applications, including, for example, the storage of compressed video on digital media such as video conferencing, digital TV, streaming services, CDs, DVDs, memory sticks, etc.

[0031] The video processing system (100) includes a capture subsystem (113) that can include a video source (101). The video source (101) can include one or more images captured by a camera and / or generated by a computer. For example, a digital camera can generate a stream of uncompressed video pictures (102). In one example, the stream of video pictures (102) includes samples taken by a digital camera. The stream of video pictures (102), shown in bold to emphasize that it has a large data volume compared to the encoded video data (104) (or encoded video bitstream), can be processed by an electronic device (120) that includes a video encoder (103) coupled to the video source (101). The video encoder (103) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (104) (or encoded video bitstream), shown as a thin line to emphasize that it has a small data volume compared to the stream of video pictures (102), can be stored in a streaming server (105) for future use. One or more streaming client subsystems, such as the client subsystems (106) and (108) of FIG. 1, can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) can include, for example, a video decoder (110) within an electronic device (130). The video decoder (110) decodes an incoming copy (107) of the encoded video data and generates an outgoing stream of video pictures (111) that can be rendered on a display (112) (such as a display screen) or other rendering device (not shown).In some streaming systems, encoded video data (104), (107), and (109) (e.g., video bitstreams) can be encoded according to certain video encoding / compression standards. Examples of these standards include ITU-T Recommendation H.265. In one example, a video encoding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0032] Note that electronic devices (120) and (130) can include other components (not shown). For example, electronic device (120) can include a video decoder (not shown), and electronic device (130) can also include a video encoder (not shown).

[0033] FIG. 2 shows an exemplary block diagram of a video decoder (210). The video decoder (210) can be included in an electronic device (230). The electronic device (230) can include a receiver (231). The receiver (231) can include a receiving circuit such as a network interface circuit. The video decoder (210) can be used in place of the video decoder (110) in the example of FIG. 1.

[0034] The receiver (231) can receive one or more encoded video sequences decoded by the video decoder (210). In certain embodiments, one encoded video sequence is received at a time, and the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequence can be received from a channel (201) that can be a hardware / software link to a storage device storing the encoded video data. The receiver (231) can receive the encoded video data together with other data, such as encoded audio data and / or auxiliary data streams, which can be transferred to their respective using entities (not shown). The receiver (231) can separate the encoded video sequence from other data. To handle network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter, "parser (220)"). In certain applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory may be outside the video decoder (210) (not shown). In still other applications, for example, to handle network jitter, there may be a buffer memory (not shown) outside the video decoder (210), and further, for example, to handle playback timing, there may be another buffer memory (215) inside the video decoder (210). If the receiver (231) receives data from a storage / transfer device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (215) may not be necessary or may be small. When used in a best-effort packet network such as the Internet, the buffer memory (215) may be necessary, may be relatively large, is advantageously of an adaptable size, and may be implemented at least in part in an operating system or similar element (not shown) outside the video decoder (210).

[0035] Video decoder (210) may include a parser (220) that reconstructs symbols (221) from an encoded video sequence. The categories of those symbols, as shown in FIG. 2, include information used to manage the operation of video decoder (210) and, potentially, information for controlling a rendering device (212) (such as a display screen) that is not an integral part of, but can be coupled to, electronic device (230). The control information for the rendering device may be in the form of a supplementary enhancement information (SEI) message or a video user utility information (VUI) parameter set fragment (not shown). Parser (220) can parse / entropy-decode the received encoded video sequence. The encoding of the encoded video sequence can follow a video encoding technology or standard and can follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity. Parser (220) can extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. Subgroups can include picture groups (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. Parser (220) can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0036] Parser (220) can perform an entropy-decode / parse operation on the video sequence received from buffer memory (215) to generate symbols (221).

[0037] The reconstruction of symbol (221) can involve multiple different units, depending on the type of the encoded video picture or a portion thereof (inter / intra picture, inter / intra block, etc.), and other factors. How each unit is involved can be controlled by subgroup control information parsed from the encoded video sequence by the parser (220). Such a flow of subgroup control information between the parser (220) and the following multiple units is not shown for clarity.

[0038] In addition to the function blocks already described, the video decoder (210) can conceptually be subdivided into several functional units as described below. In an actual implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate.

[0039] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives, as symbol (221) from the parser (220), the quantized transform coefficients and control information including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output a block including sample values that can be input to the aggregator (255).

[0040] In some cases, the output samples of the scaler / inverse transform unit (251) can belong to an intra-coded block. An intra-coded block is a block that does not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed part of the current picture. Such prediction information can be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses the surrounding already reconstructed information fetched from the current picture buffer (258) to generate a block of the same size and shape as the block being reconstructed. The current picture buffer (258) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (255) adds, in some cases, for each sample, the prediction information generated by the intra prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).

[0041] In other cases, the output samples of the scaler / inverse transform unit (251) can be associated with inter-coded and potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (253) can access the reference picture memory (257) to fetch the samples used for prediction. After motion-compensating the fetched samples according to the symbols (221) associated with the block, these samples can be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (in this case, called the residual samples or residual signal) to generate output sample information. The address in the reference picture memory (257) from which the motion compensation prediction unit (253) fetches the prediction samples can be controlled by the motion vector, and the motion vector can be available to the motion compensation prediction unit (253) in the form of symbols (221) that can have, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values fetched from the reference picture memory (257) when an exact motion vector of sub-samples is used, a motion vector prediction mechanism, and the like.

[0042] The output samples of the aggregator (255) can be applied with various loop filtering techniques within the loop filter unit (256). The video compression technique is controlled by the parameters included in the encoded video sequence (also called the encoded video bitstream), and can include in-loop filter techniques made available to the loop filter unit (256) as symbols (221) from the parser (220). Video compression can also respond to the meta information obtained during the decoding of the previous part (in decode order) of the encoded picture or encoded video sequence, and can also respond to the previously reconstructed and loop-filtered sample values.

[0043] The output of the loop filter unit (256) can be output to the rendering device (212) and can be a sample stream that can be stored in the reference picture memory (257) for use in future inter-picture prediction.

[0044] Once a certain coded picture is completely reconstructed, it can be used as a reference picture for future prediction. For example, when the coded picture corresponding to the current picture is completely reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before starting the reconstruction of the next coded picture.

[0045] The video decoder (210) can perform a decoding operation according to a predetermined video compression technique or standard, such as ITU-T Recommendation H.265. The encoded video sequence can conform to the syntax specified by the video compression technique or standard being used. This means that the encoded video sequence conforms to both the syntax of that video compression technique or standard and the profile described in that video compression technique or standard. Specifically, a profile can select certain tools from all the tools available in a video compression technique or standard as the only tools available under that profile. Also, as a requirement for compliance, the complexity of the encoded video sequence may be within the range defined by the level of the video compression technique or standard. In some cases, the level may restrict the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limitations set by the level can, in some cases, be further restricted through the virtual reference decoder (HRD) specifications and metadata signaled in the encoded video sequence for HRD buffer management.

[0046] In one embodiment, the receiver (231) can receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence(s). The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0047] Figure 3 shows an exemplary block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used in place of the video encoder (103) in the example of FIG. 1.

[0048] The video encoder (303) can receive video samples from a video source (301) (which is not part of the electronic device (320) in the example of FIG. 3) that can capture video images to be encoded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).

[0049] The video source (301) can provide a source video sequence to be encoded by the video encoder (303) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits...), any color space (e.g., BT.601 YCrCb, RGB,...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (301) may be a storage device that stores pre - prepared video. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that give motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, and each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. The following description focuses on the samples.

[0050] According to one embodiment, the video encoder (303) may encode and compress pictures of the source video sequence into an encoded video sequence (343) in real time or under any other optional time constraint conditions as needed. Enforcing an appropriate encoding speed is one function of the controller (350). In some embodiments, the controller (350) controls other functional units and is functionally coupled to other functional units as described below. The couplings are not shown for clarity. The parameters set by the controller (350) may include rate control related parameters (picture skip, quantizer, lambda value of rate-distortion optimization techniques, …), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other appropriate functions regarding the video encoder (303) optimized for a certain system design.

[0051] In some embodiments, the video encoder (303) is configured to operate in an encoding loop. As a radically simplified explanation, in one example, the encoding loop can include a source coder (330) (e.g., responsible for generating symbols such as a symbol stream based on an input picture and reference picture(s) to be encoded) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to generate sample data in the same way as a (remote) decoder does. The reconstructed sample stream (sample data) is input into the reference picture memory (334). Since the decoding of the symbol stream results in a bit-exact result independent of the decoder location (local or remote), the content in the reference picture memory (334) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as reference picture samples as what the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, e.g., due to channel errors) is also used in some related technologies.

[0052] The operation of the "local" decoder (333) can be the same as that of a "remote" decoder such as the video decoder (210) already described in detail in relation to FIG. 2. However, referring briefly to FIG. 2 as well, since symbols are available and the encoding / decoding of symbols into the encoded video sequence by the entropy encoder (345) and the parser (220) can be reversible, the entropy decoding part of the video decoder (210) including the buffer memory (215) and the parser (220) may not be fully implemented in the local decoder (333).

[0053] In one embodiment, decoder techniques other than the parse / entropy decoding that exists within the decoder exist in the corresponding encoder in the same or substantially the same functional form. Thus, the disclosed subject matter focuses on the operation of the decoder. The description of encoder techniques can be omitted since it is the reverse of the decoder techniques described comprehensively. In certain areas, more detailed descriptions are provided below.

[0054] During operation, in some examples, the source coder (330) can perform motion-compensated predictive coding, which predictively codes an input picture by referring to one or more previously encoded pictures from a video sequence designated as a "reference picture". In this way, the coding engine (332) codes the difference between a pixel block of the input picture and a pixel block of a reference picture that can be selected as a prediction reference for the input picture.

[0055] The local video decoder (333) can decode the encoded video data of a picture that can be designated as a reference picture based on the symbols generated by the source coder (330). The operation of the coding engine (332) can advantageously be a lossy process. If the encoded video data can be decoded by a video decoder (not shown in FIG. 3), the reconstructed video sequence can typically be a replica of the source video sequence with some errors. The local video decoder (333) can replicate the decoding process that can be performed on the reference picture by the video decoder and cause the reconstructed reference picture to be stored in the reference picture memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference picture that has the same content as the reconstructed reference picture that would be obtained by a remote video decoder (in the absence of transmission errors).

[0056] Predictor (335) can perform the prediction search of the encoding engine (332). That is, for a new picture to be encoded, the predictor (335) can search the reference picture memory (334) to obtain sample data (as candidate reference pixel blocks), or certain types of metadata that may serve as appropriate prediction references for the new picture, such as reference picture motion vectors, block shapes, etc. The predictor (335) can operate on a per sample block and pixel block basis to find an appropriate prediction reference. In some cases, the input picture can have prediction references derived from a plurality of reference pictures stored in the reference picture memory (334) as determined by the search results obtained by the predictor (335).

[0057] The controller (350) can manage the encoding operation of the source coder (330), including, for example, setting parameters and subgroup parameters used to encode video data.

[0058] The outputs of all the above-described functional units can undergo entropy encoding in the entropy encoder (345). The entropy encoder (345) converts the symbols generated by various functional units into an encoded video sequence by applying reversible compression to the symbols according to techniques such as Huffman coding, variable-length coding, arithmetic coding, etc.

[0059] The transmitter (340) can buffer the encoded video sequence generated by the entropy encoder (345) and prepare for transmission via a communication channel (360) that can be a hardware / software link to a storage device for storing the encoded video data. The transmitter (340) can merge the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0060] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a certain type of encoded picture type to each encoded picture, which can affect the encoding technology applicable to each picture. For example, a picture may often be assigned as one of the following picture types.

[0061] An intra picture (I picture) may be a picture that can be encoded and decoded without using other pictures in the sequence as a prediction source. Some video encoders allow various types of intra pictures, for example, including Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of these variations of I pictures, and their respective uses and characteristics.

[0062] A predicted picture (P picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction using at most one motion vector and a reference index to predict the sample values of each block.

[0063] A bi-directionally predicted picture (B picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction using at most two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi-predicted picture can use three or more reference pictures and associated metadata for the reconstruction of a single block.

[0064] The source picture is generally spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and may be encoded block by block. The blocks may be encoded predictively with reference to other (already encoded) blocks as determined by the encoding assignment applied to each block of the picture. For example, blocks of an I picture may be encoded non-predictively or may be encoded predictively with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be encoded predictively via spatial prediction or via temporal prediction with reference to one previously encoded reference picture. Blocks of a B picture may be encoded predictively via spatial prediction and via temporal prediction with reference to one or two previously encoded reference pictures.

[0065] The video encoder (303) can perform an encoding operation according to a predetermined video encoding technique or standard such as ITU-T Recommendation H.265. In that operation, the video encoder (303) can perform various compression operations including a predictive encoding operation that utilizes temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to a syntax specified by the video encoding technique or standard being used.

[0066] In one embodiment, the transmitter (340) can transmit additional data together with the encoded video. The source coder (330) can include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0067] Videos may be captured as a plurality of source pictures (video pictures) in a time sequence. Intra prediction (often abbreviated as intra prediction) utilizes spatial correlations within a given picture, while inter picture prediction utilizes (temporal or other) correlations between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. If a block within the current picture is similar to a reference block within a reference picture that has been previously encoded and is still buffered within the video, the block within the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension identifying the reference picture if multiple reference pictures are being used.

[0068] In some embodiments, in inter picture prediction, a dual prediction technique can be used. According to the dual prediction technique, two reference pictures such as a first reference picture and a second reference picture that both precede the current picture in the video in decoding order (however, the display order can be past and future respectively) are used. A block within the current picture can be encoded by a first motion vector pointing to a first reference block within the first reference picture and a second motion vector pointing to a second reference block within the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.

[0069] Furthermore, the merge mode technique can be used in inter picture prediction to improve encoding efficiency.

[0070] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks such as polygons or triangular blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and CTUs in a picture have the same size such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs) which are one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree divided into one or more coding units (CUs). For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type for the CU (such as an inter-prediction type or an intra-prediction type). A CU is divided into one or more prediction units (PUs) depending on temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In some embodiments, the prediction operation in encoding (coding) (encode / decode) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (such as luma values) for pixels such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels.

[0071] Note that the video encoders (103) and (303), and the video decoders (110) and (210) can be implemented using any suitable technology. In some embodiments, the video encoders (103) and (303), and the video decoders (110) and (210) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (103) and (303), and the video decoders (110) and (210) can be implemented using one or more processors that execute software instructions.

[0072] This disclosure includes embodiments related to methods and systems for mesh vertex position compression using temporal prediction.

[0073] A mesh can include a plurality of polygons that describe the surface of a volumetric object. Each polygon of the mesh can be defined by the vertices of the corresponding polygon in three-dimensional (3D) space and information on how those vertices are connected (which can be referred to as connectivity information). In some embodiments, vertex attributes such as color, normal, etc. can be associated with the mesh vertices. Attributes (or vertex attributes) can also be associated with the surface of the mesh by utilizing mapping information that parameterizes the mesh with a two-dimensional (2D) attribute map. Such a mapping can typically be described by a set of parametric coordinates called UV coordinates or texture coordinates associated with the mesh vertices. The 2D attribute map can be used to store high-resolution attribute information such as texture, normal, displacement, etc. Such information can be used for various purposes such as texture mapping and shading.

[0074] Since dynamic mesh sequences can contain a large amount of information that changes over time, they require a large amount of data. Therefore, an efficient compression technique is needed to store and transmit such content. Mesh compression standards such as IC, MESHGRID, and FAMC have been previously developed by MPEG to handle dynamic meshes with a certain connectivity, time-varying geometry, and vertex attributes. However, these standards may not consider time-varying attribute maps and connectivity information. DCC (Digital Content Creation) tools typically generate such dynamic meshes. However, generating a dynamic mesh with a certain connectivity can be difficult for volumetric acquisition techniques, especially under real-time constraints. This type of content (e.g., a dynamic mesh with a certain connectivity) may not be supported by existing standards. MPEG is planning to develop a new mesh compression standard to directly handle dynamic meshes with time-varying connectivity information and optionally time-varying attribute maps. The new mesh compression standard aims for irreversible and reversible compression for various applications such as real-time communication, storage, free viewpoint video, augmented reality (AR), and virtual reality (VR). Features such as random access and scalable / progressive encoding can also be considered.

[0075] Mesh geometry information can include vertex connectivity information, 3D coordinates, 2D texture coordinates, etc. The compression of vertex 3D coordinates, which can also be called vertex positions, can be important as the compression of vertex 3D coordinates can potentially consume a significant portion of the overall geometry-related data.

[0076] The dynamic mesh sequence M at time point t can be represented as M(t). If there is a mapping (or mapping operation) f from the vertex position of M(t) to the vertex position M(t0) at another time point, M(t) can be referred to as the frame to be position-tracked. Thus, M(t0) can be called the reference frame of the frame M(t), and the corresponding vertex in the reference frame can be called the reference vertex of the vertex in M(t). M(t) can correspond to the mesh at a first time point such as the current time point, and M(t0) can correspond to the mesh at a second time point such as the previous time point.

[0077] In the present disclosure, a method and / or system for vertex position compression are proposed. It should be noted that the method and / or system can be applied individually or in various combinations. Furthermore, the disclosed method and system are not limited to vertex position compression. The disclosed method and system can also be applied to, for example, two-dimensional (2D) texture coordinate compression or a more general time-based prediction-based scheme.

[0078] For a vertex V in the position-tracked frame M(t), the adjacent vertices of that vertex may be the vertices connected to V via edges, and these vertices are called the adjacent vertices (or neighboring vertices) of V. For example, as shown in FIG. 4, vertex A can have four adjacent vertices which are C, D, E, and B. Vertex E can have five adjacent vertices which are A, B, F, H, and D.

[0079] Figure 5 is a schematic diagram of an exemplary vertex position compression device (500) according to some embodiments of the present disclosure. As shown in Figure 5, the vertex position compression device (500) can include a vertex ordering module (502) configured to order the vertices of a mesh, a position prediction module (504) configured to calculate a predicted position of a vertex, a prediction index encoding module (506) configured to encode (or determine) a position prediction index (or a position prediction mode if necessary), and a prediction residual encoding module (508) configured to encode a position prediction residual. Note that the mesh processed by the vertex position compression device (500) can be a 3D mesh or a 2D mesh. For example, the mesh can be represented in a 3D coordinate system (e.g., x, y, z) or a 2D coordinate system (e.g., U and V).

[0080] The term "module" in the present disclosure can refer to a software module, a hardware module, or a combination thereof. A software module (e.g., a computer program) can be developed using a computer programming language. A hardware module can be implemented using a processing circuit and / or memory. Each module can be implemented using one or more processors (or a processor and memory). Similarly, a processor (or a processor and memory) can be used to implement one or more modules. Further, each module can be a part of an overall module that includes the functionality of that module.

[0081] In the present disclosure, the vertices of a frame (or frames) M(t) being position-tracked can be ordered. The order of the vertices in M(t) can be traced according to an edgebreaker algorithm or other partitioning algorithms.

[0082] Figures 6A and 6B show an exemplary ordering of triangles and vertices within frame M(t) based on the edge breaker algorithm. Figure 6A shows five exemplary patch configurations of the edge breaker algorithm. As shown in Figure 6A, V is the patch center vertex and T is the current triangle. The active gate (or current triangle) within each patch can be denoted as T. In patch C, complete triangles that fan out (or rotate) around V can be provided. In patch L, one or more missing triangles may be located to the left of the active gate T. In patch R, one or more missing triangles may be located to the right of the active gate T. In patch E, V is only adjacent to T. In patch S, one or more missing triangles may be located in positions other than to the left or right of the active gate T. Figure 6B shows an exemplary traversal of frame (600) that can order the triangles of frame (600) based on the traversal of the edge breaker algorithm. As shown in Figure 6B, the triangles within frame (600) can be traversed along a spiral triangle spanning tree. For example, the traversal can start from a triangle (602) of type C (or patch C). The traversal can proceed along a branch adjacent to the right edge of the triangle (e.g., (602)). The traversal can stop when it reaches a triangle of type E (e.g., (604)). According to the edge breaker algorithm, the triangles of frame (600) can be traversed (or ordered) in the sequence CRSRLECRRRLE as shown in Figure 6B. The vertices of each triangle of frame (600) can also be ordered based on the order of the triangles.

[0083] In the present disclosure, temporal prediction can be applied to predict vertices in the current frame based on reference vertices in a reference frame. The current frame can correspond to a mesh at a first time point such as the current time point. The reference frame can correspond to a mesh at a second time point such as a previous time point. In some embodiments, the connectivity of the vertices of the mesh in the current frame can be different from the connectivity of the vertices of the mesh in the reference frame. For a vertex V in the frame M(t) to be position-tracked (or the current frame), the position of the vertex V can be estimated by the position of a reference vertex f(V) in the reference frame (for example, M(t0), where t0 is a time point different from the time point t). Here, f is a mapping operation between M(t) and the reference frame. In some embodiments, the vertex V and the reference vertex f(V) in the reference frame are co-located. Thus, the reference vertex can have the same relative position as the vertex in the current frame M(t) within the reference frame. When the vertex V is predicted by the reference vertex f(V), an estimation error E can be determined as the difference between the position of V and the position of f(V) in Equation (1). E = V - f(V) Equation (1)

[0084] Since each vertex in the frame M(t) can have three-dimensional coordinates, the 3D coordinate components of the estimation error E can be given based on Equation (1). For example, assuming that the subscripts x, y, and z represent the 3D coordinates in the xyz space, the three-dimensional coordinate components of the estimation error E can be given by Equations (2) to (4). E x = V x - (f(V)) x Equation (2) E y = V y - (f(V)) y Equation (3) E z = V z - (f(V)) z Equation (4)

[0085] The estimation error E of vertex V can be predicted from the adjacent vertices (or neighboring vertices) of vertex V or determined in other ways. For the adjacent vertices (or neighboring vertices) of V, if the adjacent vertices are encoded and can be used for prediction, the estimation errors of the adjacent vertices can be applied to predict E.

[0086] Let V have N adjacent vertices (or neighboring vertices) V1, V2, …, V N that are encoded and can be used for prediction. In one example, the order of the adjacent vertices V1, V2, …, V N can be determined based on an edge breaker algorithm or other partitioning algorithms. For an adjacent vertex V i , the estimation error of the adjacent vertex V i can be determined as E i = V i - f(V i ) for i = 1, 2, …, N. f(V i ) can be the reference vertex for the adjacent vertex V i in the reference frame. E i can also be referred to as the neighboring estimation error related to vertex V, and each E i can be a prediction candidate for E. If duplicates are determined in the predicted value (or prediction candidates) E i , such duplicates can be removed from the list E1, E2, …, E N .

[0087] When N ≧ 2, multiple prediction candidates (or predicted values) E i are available. In one embodiment, the average of the predicted values E i can be used as an additional predictor and set as E0. For example, E0 can be defined by the following equation (5). E0 = (E1 + E2 + … + E N ) / N Equation (5) If the average E0 is equal to any of the predicted values E1, E2, …, E N , either the average predicted value or the duplicate predicted values can be removed.

[0088] On the encoder side, for each E i can be compared with E. Here, 0 ≤ i ≤ N. The estimation error can be selected from E0, E1, E2, …, E N The selected estimation error can have the minimum error (or minimum difference) between E and each predicted value (or predicted value or prediction candidate) E0, E1, E2, …, E N In some embodiments, the minimum error can be measured by the L 0 norm, the L 1 norm, the L 2 norm, or some other norm. For example, the L 0 norm can be determined as follows in Equation (6): L 0 norm = √{(E 0x - E x ) 2 + (E 0y - E y ) 2 + (E 0z - E z ) 2} Equation (6) Here, E x , E y , and E z are the coordinates of E in the xyz space, and E 0x , E 0y , and E 0z are the coordinates of E0 in the xyz space. In one example, the selected estimation error can be one of the E i corresponding to the minimum error. In one example, the selected estimation error can be E itself. For convenience, an index -1 can be used to represent E, such as E -1 = E. Thus, the selected estimation error can be stored (or identified) as an index whose value is between -1 and N. On the decoder side, the index of the selection indicating the selected estimation error is decoded, and the selected predictor (or the selected estimation error) can be recovered from the list of predictors E -1 , E0, E1, E2, …, E N .

[0089] In one embodiment, an upper limit M can be set. The upper limit M can indicate the number of prediction candidates that can be considered (or applied) in the prediction list. In one example, if N>M, then only the first M prediction candidates E1, E2, …, E M are considered for predicting E. Thus, E0 can be defined in the following equation (7): E0=(E1+E2+…+E M ) / M Equation (7) M can be an integer such as 4. Thus, a maximum of M (e.g., 4) prediction candidates can be considered. Note that the average prediction E0 can be arranged at different positions in the candidate list. For example, the average prediction E0' can be the first predictor in the prediction candidate list. For example, the average prediction E0' can be the last predictor in the prediction candidate list. For example, the average prediction E0' can be arranged among the predicted values E1', E2, …, E M .

[0090] Intra-frame prediction can be applied to predict the vertex position V. The intra-frame prediction can be, for example, delta prediction, parallelogram prediction, etc. In delta prediction, a predictor can be determined as an adjacent vertex of the vertex V. In parallelogram prediction, a predicted vertex can be determined based on predicted triangles of a plurality of triangles of the mesh. The predicted triangle shares an edge with a triangle among the plurality of triangles that includes the vertex V. The predicted vertex may be a vertex facing the shared edge. The predicted vertex and the predicted triangle can form a parallelogram.

[0091] The prediction error from the intra-frame prediction is the selected estimation error E ican be compared, where -1 ≦ i ≦ N. The prediction error of intra-frame prediction can indicate the true position of vertex V and the predicted position value based on intra-frame prediction. If the intra-frame prediction gives a smaller error (for example, the prediction error is smaller than the selected estimation error), the intra-frame prediction can be used and the intra prediction mode can be signaled. Otherwise, when the intra-frame prediction gives an error larger than the selected estimation error E i time prediction can be used and an inter prediction mode (for example, time prediction) can be signaled.

[0092] To predict the position of vertex V, hybrid prediction can be applied. For example, the average of intra-frame prediction and time prediction can be applied as the prediction of vertex position V.

[0093] In the present disclosure, the prediction index needs to be encoded only when there are multiple prediction candidates. For example, when multiple prediction candidates are available for the current vertex, the index indicating the selected prediction candidate can be encoded on the encoder side. On the decoder side, the decoder can determine the multiple prediction candidates in the same order as the encoder. The decoder can decode the encoded prediction index and reconstruct the selected prediction candidate based on the prediction index from the multiple prediction candidates.

[0094] In certain embodiments, it may not be necessary to encode the prediction index when no prediction candidates have been determined or when only one prediction candidate has been determined. Thus, the predicted value can be the current vertex itself or its only prediction candidate. In certain embodiments, the prediction index can always be encoded regardless of the number of available predictor candidates. When there are no available predictors, the value of the signaled index may not affect the decoding process on the decoder side.

[0095] When multiple prediction candidates are determined for the current vertex, such as N ≥ 1, the prediction index can be encoded using fixed-length coding. For example, when three prediction candidates are determined for the current vertex, four possible prediction indices, 0, 1, 2, and 3 (where 0 represents the average value of candidates 1, 2, and 3) may be required. Thus, a 2-bit binary number can be applied to represent each of the four prediction indices. It should be noted that different vertices of the triangular mesh can use different fixed lengths. For example, if there are seven prediction candidates for another vertex, that other vertex can use a 3-bit binary number for the representation of the prediction index. The output from the fixed-length coding can be further compressed by entropy coding such as arithmetic coding.

[0096] Alternatively, variable-length coding can also be used to encode the prediction index. For example, when four prediction candidates are determined for the current vertex, five prediction indices 0, 1, 2, 3, and 4 may be required. To represent the five prediction indices 0, 1, 2, 3, and 4, variable-length codes 0, 100, 101, 110, and 111 can be assigned respectively. Alternatively, variable-length codes 1, 01, 001, 0001, and 00001 can also be applied to represent the five prediction indices 0, 1, 2, 3, and 4. It should be noted that different vertices of the triangular mesh can use different variable lengths. The output from the variable-length coding can be further compressed by entropy coding such as arithmetic coding.

[0097] When an upper limit M that restricts the maximum number of prediction candidates is set, if there are more than M prediction candidates available for vertex V, only the first M prediction candidates may be used to predict vertex V. In this way, (M + 1) possible prediction indices can be encoded, where M indicates the first M prediction candidates and 1 indicates the average of the first M prediction candidates. When duplicates are identified in the predicted values, the duplicates can be removed. Thus, the possibility of prediction indices can also be reduced. The prediction indices can be encoded by fixed-length encoding, variable-length encoding, differential encoding, etc.

[0098] When in-frame prediction is also a prediction candidate, a 1-bit binary flag can be used to signal the position prediction mode. The position prediction mode can indicate whether inter-frame prediction (e.g., temporal prediction) is applied or in-frame prediction is applied.

[0099] When the average of in-frame prediction and inter-frame prediction is also a prediction candidate, a 3-symbol flag can be used to signal the position prediction mode. The position prediction mode can indicate which of inter-frame prediction, in-frame prediction, or the average of in-frame prediction and inter-frame prediction is applied.

[0100] In the present disclosure, the prediction residual of the vertex position can be encoded (or determined). When temporal prediction is applied, if the prediction index j is -1 (-1 means E itself), the prediction residual can be determined as E (e.g., V - f(V)). Thus, vertex V can be reconstructed as V = f(V)+E. Otherwise (e.g., the index is 0, or 1, 2, …, N), the prediction residual can be determined as E - E j as. Thus, E can be restored as E = E j + prediction residual (e.g., E - E j ), and vertex V can be further reconstructed as V = f(V)+(E j +(E - E j ).

[0101] When in-frame prediction is applied, the prediction residual can be determined as the difference between the true position value of vertex V and the predicted position value.

[0102] When an average of in-frame prediction and inter-frame prediction is applied, the prediction residual can be determined as the difference between the true position value of vertex V and the predicted position value.

[0103] In some embodiments, the prediction residual can be encoded using an encoding algorithm such as fixed-length coding, Golomb coding, arithmetic coding, etc. In some embodiments, the prediction residual can pass through a compacting transform such as a fast Fourier transform (FFT), discrete cosine transform (DCT), discrete sine transform (DST), discrete wavelet transform (DWT), etc. The output from the compacting transform can be encoded using an encoding algorithm such as fixed-length coding, Golomb coding, arithmetic coding, etc.

[0104] FIG. 7 shows a flowchart outlining a process (700) according to an embodiment of the present disclosure. The process (700) can be used in an encoder such as a video encoder. In various embodiments, the process (700) is executed by a processing circuit such as a processing circuit that executes the functions of video encoder (103), a processing circuit that executes the functions of video encoder (303), etc. In some embodiments, the process (700) is implemented in software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (700). The process starts at (S701) and proceeds to (S710).

[0105] (In S710), in the current frame of the mesh in either a two-dimensional model or a three-dimensional model, a plurality of adjacent vertices of the current vertex are determined. The current vertex and the plurality of adjacent vertices are included in the current frame and correspond to the mesh at a first point in time. Each of the plurality of adjacent vertices is connected to the current vertex through respective edges within the mesh.

[0106] (S720) In (S720), a plurality of adjacent estimation errors of a plurality of adjacent vertices of the current vertex are determined. Each of the plurality of adjacent estimation errors indicates a difference between a reference vertex of one corresponding adjacent vertex of the plurality of adjacent vertices in the reference frame of the mesh and the corresponding one adjacent vertex of the plurality of adjacent vertices in the current frame. The reference frame corresponds to the mesh at a second time point.

[0107] (S730) In (S730), a prediction residual of the current vertex is determined based on the plurality of adjacent estimation errors of the plurality of adjacent vertices.

[0108] (S740) In (S740), prediction information of the current vertex is generated based on the determined prediction residual of the current vertex.

[0109] In some embodiments, the reference vertex of one corresponding adjacent vertex of the plurality of adjacent vertices is located at the same relative position as the current vertex in the current frame in the reference frame, and the reference frame and the current frame are generated at different time points.

[0110] In one example, an average adjacent estimation error of the plurality of adjacent estimation errors is determined. The prediction residual of the current vertex is determined based on the average adjacent estimation error.

[0111] In one example, to determine the prediction residual of the current vertex, an estimation error of the current vertex is determined based on a difference between the reference vertex of the current vertex and the current vertex. A prediction list of the current vertex is determined. The predictors of the prediction list include the estimation error of the current vertex, the average adjacent estimation error following the estimation error of the current vertex, and the plurality of adjacent estimation errors following the average adjacent estimation error. Each predictor in the prediction list is associated with a respective prediction index.

[0112] In some embodiments, the plurality of adjacent vertices are ordered based on an edge breaker algorithm in which the plurality of adjacent vertices are traversed in a spiral triangular spanning tree order.

[0113] In one example, to determine the prediction residual of the current vertex, the difference between the average adjacent estimation error and the estimation error of the current vertex is determined. The differences between each of the plurality of adjacent estimation errors and the estimation error of the current vertex are also determined. An adjacent estimation error is selected from the average adjacent estimation error and the adjacent estimation errors from the plurality of adjacent estimation errors. The selected adjacent estimation error has the smallest difference. The prediction residual of the current vertex is determined as either (i) the estimation error of the current vertex or (ii) the selected adjacent estimation error.

[0114] In one example, to determine the prediction residual of the current vertex, according to inter-frame prediction, a prediction error is determined based on one of the plurality of adjacent vertices of the current vertex. The prediction error is compared with each predictor in the prediction list. The prediction residual is determined as the prediction error in response to the prediction error being smaller than the predictors in the prediction list.

[0115] In one example, to determine the prediction residual of the current vertex, the prediction residual is determined as a mixed estimation error. The mixed estimation error includes either (i) the average of the selected adjacent estimation error and the prediction error, and (ii) the average of the estimation error of the current vertex and the prediction error.

[0116] In certain embodiments, the inter-frame prediction further includes one of the following. Delta prediction based on the difference between the first predicted vertex and the current vertex. The first predicted vertex includes one of the plurality of adjacent vertices of the current vertex. The inter-frame prediction can also include parallelogram prediction in which a second predicted vertex is determined based on a predicted triangle among the plurality of triangles of the mesh. The predicted triangle shares an edge with a triangle among the plurality of triangles including the current vertex. The second predicted vertex is on the opposite side of the shared edge, and the second predicted vertex and the predicted triangle form a parallelogram.

[0117] In certain embodiments, the prediction information of the current vertex further includes a flag. The flag indicates that the prediction residual is one of (i) the estimation error of the current vertex or the selected adjacent estimation error based on the first value of the flag, (ii) the prediction error based on the second value of the flag, and (iii) the mixed estimation error based on the third value of the flag.

[0118] In one embodiment, the prediction information of the current vertex further includes index information indicating which predictor in the prediction list is the prediction residue based on the flag being the first value. The prediction information also includes prediction residue information. In one example, based on the flag being the first value, the prediction residue information is (i) the estimated error of the current vertex in response to the index information indicating that the prediction residue is the estimated error of the current vertex, and (ii) the selected adjacent estimated error in response to the index information indicating that the prediction residue is the selected adjacent estimated error. Based on the flag being the second value, the prediction residue information indicates the difference between the current vertex and the predicted vertex indicated by the prediction error. Based on the flag being the third value, the prediction residue information indicates the difference between the current vertex and the predicted vertex indicated by the mixed estimated error.

[0119] Then, the process proceeds to (S799) and ends.

[0120] The process (700) can be adapted as appropriate. The steps in the process (700) can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.

[0121] FIG. 8 shows a flowchart outlining a process (800) according to an embodiment of the present disclosure. The process (800) can be used in a decoder such as a video decoder. In various embodiments, the process (800) is executed by a processing circuit such as a processing circuit that executes the functions of the video decoder (110), a processing circuit that executes the functions of the video decoder (210), etc. In some embodiments, the process (800) is implemented in software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (800). The process starts at (S801) and proceeds to (S810).

[0122] In (S810), the encoding information of the current frame of the mesh in either the two-dimensional model or the three-dimensional model is received. The current frame includes a plurality of vertices and corresponds to the mesh at a first point in time. The encoding information includes the index information of the current vertex among the plurality of vertices. The index information indicates the prediction residual of the current vertex.

[0123] (S820) determines a plurality of adjacent vertices of the current vertex from the plurality of vertices, and each of the plurality of adjacent vertices is connected to the current vertex through its respective edge.

[0124] (S830) determines a plurality of adjacent estimation errors of the plurality of adjacent vertices of the current vertex. Each of the plurality of adjacent estimation errors indicates the difference between the reference vertex of the corresponding adjacent vertex among the plurality of adjacent vertices in the reference frame of the mesh and the corresponding adjacent vertex among the plurality of adjacent vertices in the current frame. The reference frame corresponds to the mesh at a second point in time.

[0125] (S840) determines the prediction residual of the current vertex based on (i) the index information and (ii) the plurality of adjacent estimation errors of the plurality of adjacent vertices.

[0126] (S850) reconstructs the current vertex based on the determined prediction residual of the current vertex.

[0127] Then, the process proceeds to (S899) and ends.

[0128] The process (800) can be appropriately adapted. The steps of the process (800) can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.

[0129] The above techniques can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media such as a non-transitory computer-readable storage medium. For example, FIG. 9 shows a computer system (900) suitable for implementing certain embodiments of the disclosed subject matter.

[0130] The computer software can be coded using any suitable machine code or computer language and can be applied with assembly, compilation, linking, or similar mechanisms to create code that includes instructions that can be executed directly by one or more processing circuits such as a computer central processing unit (CPU), a graphics processing unit (GPU), or through interpretation, microcode execution, etc.

[0131] The instructions can be executed on various types of computers or their components including, for example, a personal computer, a tablet computer, a server, a smartphone, a gaming device, an Internet-of-Things device, etc.

[0132] The components shown in FIG. 9 for the computer system (900) are exemplary in nature and are not intended to suggest any limitation as to the use or functionality scope of the computer software implementing embodiments of the present disclosure. The configuration of the components should not be construed as having any dependency or requirement regarding any one or combination of the components shown in the exemplary embodiment of the computer system (900).

[0133] The computer system (900) can include certain types of human interface input devices. Such human interface input devices can respond to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, movement of a data glove), voice input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not shown). Also, the human interface device can be used to capture certain types of media that are not necessarily directly related to conscious human input, such as voice (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., 2D video, 3D video including stereoscopic video).

[0134] The input human interface device may include one or more of a keyboard (901), a mouse (902), a trackpad (903), a touch screen (910), a data glove (not shown), a joystick (905), a microphone (906), a scanner (907), a camera (908) (only one of each is shown).

[0135] The computer system (900) may also include certain human interface output devices. Such human interface output devices may, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such human interface output devices may include the following. Tactile output devices (e.g., tactile feedback by a touch screen (910), a data glove (not shown), or a joystick (905); however, there may also be tactile feedback devices that do not function as input devices), audio output devices (e.g., speakers (909), headphones (not shown)), visual output devices (e.g., a screen (910) including a CRT screen, an LCD screen, a plasma screen, an OLED screen; each may or may not have a touch screen input function, each may or may not have a tactile feedback function, and some of them can output higher than three-dimensional output through means such as two-dimensional visual output or stereoscopic output; virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0136] The computer system (900) may also include human-accessible memory devices and associated media, such as optical media including a CD / DVD ROM / RW (920) together with a CD / DVD or similar media (921), a thumb drive (922), a removable hard drive or solid state drive (923), legacy magnetic media such as tapes and floppy disks (not shown), specialized ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.

[0137] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0138] The computer system (900) may also include an interface (954) to one or more communication networks (955). The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan area, in-vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include Ethernet®, wireless LAN, GSM, 3G, 4G, 5G, LTE, etc., cellular networks, cable TV, satellite TV, TV wired or wireless wide area digital networks including terrestrial broadcast TV, in-vehicle and industrial including CANBus, etc. Certain types of networks typically require an external network interface adapter attached to a certain type of general-purpose data port or peripheral bus (949) (such as a USB port of the computer system (900)). Others are typically integrated into the core of the computer system (900) by attachment to a system bus as described later (such as an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (900) can communicate with other entities. Such communication can be unidirectional, receive-only (such as broadcast TV), dedicated unidirectional transmission (such as CANbus to certain CANbus devices), or bidirectional to other computer systems using, for example, local or wide area digital networks. For each of the networks and network interfaces as described above, certain protocols and protocol stacks can be used.

[0139] The aforementioned human interface device, human-accessible storage device, and network interface may be attached to the core (940) of the computer system (900).

[0140] The core (940) may include one or more central processing units (CPUs) (941), a graphics processing unit (GPU) (942), a specialized programmable processing device in the form of a field programmable gate array (FPGA) (943), a hardware accelerator (944) for certain tasks, a graphics adapter (950), and the like. These devices may be connected through a system bus (948) together with a read-only memory (ROM) (945), a random access memory (946), an internal mass storage device such as an internal hard drive or a solid state drive (SSD) that is not accessible to users (947). In some computer systems, the system bus (948) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus (948) of the core or through a peripheral bus (949). In one example, a screen (910) can be connected to a graphics adapter (950). Architectures for peripheral buses include PCI, USB, and the like.

[0141] The CPU (941), GPU (942), FPGA (943), and accelerator (944) can execute certain instructions that can be combined to form the above-mentioned computer code. The computer code can be stored in the ROM (945) or the RAM (946). Temporary data can also be stored in the RAM (946), while persistent data can be stored, for example, in the internal mass storage device (947). By using a cache memory that can be closely associated with one or more CPUs (941), GPUs (942), the mass storage device (947), the ROM (945), the RAM (946), etc., fast storage and retrieval to any of the memory devices can be enabled.

[0142] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure, or they may be of the kind well known and available to those having skill in the computer software arts.

[0143] By way of example and not limitation, a computer system having an architecture (900), specifically a core (940), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied on one or more tangible computer-readable media. Such computer-readable media can be related to user-accessible mass storage as introduced above and certain types of storage of the core (940) having a non-transitory nature, such as a mass storage device (947) inside the core or a ROM (945). The software implementing various embodiments of the present disclosure can be stored on such a device and executed by the core (940). The computer-readable media can include one or more memory devices or chips according to specific needs. The software includes defining a data structure stored in a RAM (946) and modifying such a data structure according to a process defined by the software, and causing a specific process or a specific part described herein to be executed by the core (940) and specifically a processor (including a CPU, GPU, FPGA, etc.) therein. Additionally or alternatively, the computer system can provide functionality as a result of logic wired in a circuit (e.g., an accelerator (944)) or otherwise embodied, which can operate instead of or together with software for executing a specific process or a specific part of a specific process described herein. References to software include logic and, where appropriate, vice versa. References to computer-readable media can, where appropriate, include a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0144] Although several exemplary embodiments have been described, there are changes, substitutions, and various alternative equivalents within the scope of the present disclosure. Thus, those skilled in the art will be able to devise numerous systems and methods that embody the principles of the present disclosure and are thus within its spirit and scope, even though not explicitly shown or described herein.

Claims

1. A method of mesh processing executed by a video encoder, comprising: determining a plurality of adjacent vertices of a current vertex in a current frame of a mesh in either a two-dimensional model or a three-dimensional model, wherein the plurality of adjacent vertices and the current vertex are included in the current frame, the current frame corresponding to the mesh at a first time point, and each of the plurality of adjacent vertices is connected to the current vertex through respective edges in the mesh; determining a plurality of adjacent estimation errors of the plurality of adjacent vertices of the current vertex, wherein each of the plurality of adjacent estimation errors indicates a difference between a reference vertex of a corresponding one of the plurality of adjacent vertices in a reference frame of the mesh and the corresponding one of the plurality of adjacent vertices in the current frame, the reference frame corresponding to the mesh at a second time point; determining an average adjacent estimation error of the plurality of adjacent estimation errors; determining a prediction residual of the current vertex based on the average adjacent estimation error; generating prediction information of the current vertex based on the determined prediction residual of the current vertex. A method.

2. The method according to claim 1, wherein the reference vertex of the corresponding one of the plurality of adjacent vertices is located at the same relative position as the current vertex in the current frame in the reference frame, and the reference frame and the current frame are generated at different time points.

3. Determining the prediction residual of the current vertex comprises: determining an estimation error of the current vertex based on a difference between a reference vertex of the current vertex and the current vertex; further comprising determining a prediction list of the current vertex, wherein predictors of the prediction list include the estimation error of the current vertex, the average adjacent estimation error following the estimation error of the current vertex, and the plurality of adjacent estimation errors following the average adjacent estimation error, and respective prediction indices are associated with each predictor in the prediction list. The method according to claim 1.

4. The method according to claim 3, wherein the plurality of adjacent vertices are ordered based on an edge breaker algorithm in which the plurality of adjacent vertices are traversed in a spiral triangular spanning tree order.

5. Determining the prediction residual of the current vertex comprises: determining (i) the difference between the average adjacent estimation error and the estimation error of the current vertex, and (ii) the difference between each of the plurality of adjacent estimation errors and the estimation error of the current vertex; selecting an adjacent estimation error from the average adjacent estimation error and the plurality of adjacent estimation errors, wherein the selected adjacent estimation error has the smallest difference; further comprising determining the prediction residual of the current vertex as either (i) the estimation error of the current vertex or (ii) the selected adjacent estimation error; The method according to claim 3.

6. Determining the prediction residual of the current vertex: determining a prediction error based on one of the plurality of adjacent vertices of the current vertex according to inter-frame prediction; comparing the prediction error with each predictor in the prediction list; further comprising determining the prediction residual as the prediction error in response to the prediction error being smaller than the predictor in the prediction list. The method according to claim 5.

7. Determining the prediction residual of the current vertex: including determining the prediction residual as a mixed estimation error, the mixed estimation error including either (i) an average of the selected adjacent estimation error and the prediction error, and (ii) an average of the estimation error of the current vertex and the prediction error. The method according to claim 6.

8. The inter-frame prediction is: a delta prediction based on the difference between a first predicted vertex and the current vertex, the first predicted vertex including an adjacent vertex of one of the plurality of adjacent vertices of the current vertex; and a parallelogram prediction in which a second predicted vertex is determined based on a predicted triangle among the plurality of triangles of the mesh, the predicted triangle sharing an edge with a triangle among the plurality of triangles including the current vertex, the second predicted vertex being on the opposite side of the shared edge, and the second predicted vertex and the predicted triangle forming a parallelogram. The method according to claim 6.

9. The prediction information of the current vertex further includes a flag, the flag indicating that the prediction residual is one of (i) the estimation error of the current vertex or the selected adjacent estimation error based on a first value of the flag, (ii) the prediction error based on a second value of the flag, and (iii) the mixed estimation error based on a third value of the flag. The method according to claim 7.

10. The prediction information of the current vertex further includes index information indicating which predictor in the prediction list is the prediction residual based on the flag being the first value; prediction residual information, where the prediction residual information is: (i) based on the flag being the first value, in response to the index information indicating that the prediction residual is the estimation error of the current vertex, the estimation error of the current vertex, in response to the index information indicating that the prediction residual is the selected adjacent estimation error, indicating the selected adjacent estimation error; (ii) based on the flag being the second value, indicating the difference between the current vertex and the predicted vertex indicated by the prediction error: (iii) based on the flag being the third value, indicating the difference between the current vertex and the predicted vertex indicated by the mixed estimation error, The method according to claim 9. **Claim 11** An apparatus for mesh processing, the apparatus having a processing circuit configured to execute the method according to any one of claims 1 to 10. **Claim 12** A computer program for causing a computer to execute the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Encoding method for data streams representing time-varying graphic models

    JP2008516318A

  • A scalable compression method for time-consistent 3D mesh sequences.

    JP2010527523A