Grid decoding method and device, grid coding method and device, data processing method and device and storage medium
By selecting the prediction mode based on the prediction mode priority and accuracy in polygon mesh compression, the problem of low encoding accuracy and efficiency in the prior art is solved, and more efficient mesh compression is achieved.
Patent Information
- Application Number
- CN202411584953.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-29
- Filing Date
- 2024-11-07
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to effectively select prediction modes in polygon mesh compression, resulting in low encoding accuracy and efficiency.
通过接收网格中的多个属性的属性信息的码流,基于多个候选预测模式的预测模式优先级和精度,确定一个或多个预测模式,并基于这些预测模式对网格中的属性进行编码和解码。
The coding accuracy and efficiency of prediction residuals are improved, and the prediction mode selection during polygon mesh compression is optimized.
Smart Images

Figure HDA0005124062270000011 
Figure HDA0005124062270000021 
Figure HDA0005124062270000031
Abstract
Description
Cross - Reference to Related Applications
[0001] This application claims priority to U.S. Patent Application No. 18 / 820,160, filed on August 29, 2024, entitled "Prediction Mode Selection in Polygon Mesh Compression", which claims priority to U.S. Provisional Application No. 63 / 602,830, filed on November 27, 2023, entitled "Prediction Mode Selection in Polygon Mesh Compression". The entire content of the above - mentioned priority applications is hereby incorporated by reference into this application. Technical Field
[0002] This disclosure describes aspects generally related to mesh coding. Background Art
[0003] The description of the background art provided in this disclosure is intended to present the background of the disclosure as a whole. The work of the presently named inventors described in the background art section, and aspects of the specification that may not be prior art at the time of filing, are neither expressly nor impliedly admitted to be prior art against the disclosure.
[0004] Image / video compression helps to transmit image / video data between different devices, storage, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In an example, a video codec can use a technique called intra - prediction, which can compress an image based on spatial redundancy. For example, intra - prediction can perform sample prediction using reference data from the current image in reconstruction. In another example, a video codec can use a technique called inter - prediction, which can compress an image based on temporal redundancy. For example, inter - prediction can predict samples in the current image from a previously reconstructed image with motion compensation. Motion compensation can be indicated by a Motion Vector (MV).
[0005] Advances in three - dimensional (3D) capture, modeling, and rendering have contributed to the prevalence of 3D content on various platforms and devices. For example, a baby's first steps can be captured on one piece of land, and the baby's grandparents can watch (and in some cases interact) on another piece of land and enjoy a fully immersive experience with the child. To achieve this realism, models are becoming increasingly complex, and a large amount of data is associated with the creation and consumption of these models. 3D meshes are widely used to represent such immersive content. Summary of the Invention
[0006] Aspects of the present disclosure include bitstreams, methods, and apparatuses for mesh processing. In some examples, an apparatus for mesh processing includes processing circuitry.
[0007] According to one aspect of the present disclosure, a mesh decoding method is provided. In this method, a bitstream including attribute information of multiple attributes in a mesh is received. Based on one of the prediction mode priorities and prediction mode accuracies of one or more prediction modes among multiple candidate prediction modes, one or more prediction modes are determined from the multiple candidate prediction modes for a current attribute among the multiple attributes in the mesh. Based on the attribute information of the current attribute and the one or more prediction modes, a predicted value of the current attribute in the mesh is determined. The current attribute is reconstructed based on the predicted value of the current attribute.
[0008] According to another aspect of the present disclosure, a mesh encoding method is provided. In this method, multiple candidate prediction modes for multiple attributes in a mesh are determined. Based on one of the prediction mode priorities and prediction mode accuracies of one or more prediction modes among the multiple candidate prediction modes, one or more prediction modes are determined from the multiple candidate prediction modes for a current attribute among the multiple attributes in the mesh. The predicted value of the current attribute in the mesh is encoded based on the one or more prediction modes. Signal information indicating that one or more prediction modes are determined for the current attribute in the mesh from the multiple candidate prediction modes is encoded into the bitstream.
[0009] According to yet another aspect of the present disclosure, a mesh data processing method is provided. In this method, the bitstream of mesh data is processed according to format rules. The bitstream includes attribute information of multiple attributes in the mesh. The format rules specify that based on one of the prediction mode priorities and prediction mode accuracies of one or more prediction modes among multiple candidate prediction modes, one or more prediction modes are determined from the multiple candidate prediction modes for a current attribute among the multiple attributes in the mesh. The format rules specify that a predicted value of the current attribute in the mesh is determined based on the attribute information of the current attribute and the one or more prediction modes. The format rules specify that the current attribute is processed based on the predicted value of the current attribute.
[0010] Aspects of the present disclosure also provide a mesh encoding apparatus. The mesh encoding apparatus includes processing circuitry for implementing any of the described mesh encoding methods.
[0011] Aspects of the present disclosure also provide a mesh decoding apparatus. The mesh decoding apparatus includes processing circuitry for implementing any of the described mesh decoding methods.
[0012] Aspects of the present disclosure also provide a non - transitory computer - readable medium storing instructions. When the instructions are executed by a computer, the computer is caused to execute any of the described mesh decoding, mesh encoding, and mesh data processing methods.
[0013] The technical solutions of the present disclosure include methods and apparatuses for improving the selection of one or more prediction modes. The improvement includes at least one of coding accuracy or coding efficiency. In an example, a bitstream including attribute information of multiple attributes in a grid is received. Based on one of the prediction mode priorities and prediction mode accuracies of one or more prediction modes among multiple candidate prediction modes, one or more prediction modes are determined from the multiple candidate prediction modes for a current attribute among the multiple attributes in the grid. Based on the attribute information of the current attribute and the one or more prediction modes, a predicted value of the current attribute in the grid is determined. The current attribute is reconstructed based on the predicted value of the current attribute. For example, a method of determining one or more prediction modes based on one of the prediction mode priority and prediction mode accuracy is adopted to improve at least one of the coding accuracy or coding efficiency of the prediction residual. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0015] Figure 1 is an exemplary block diagram of a communication system (100).
[0016] Figure 2 is an exemplary block diagram of a decoder.
[0017] Figure 3 is an exemplary block diagram of an encoder.
[0018] Figure 4A is an exemplary diagram of cross-parallelogram prediction for grid processing according to one aspect of the present disclosure.
[0019] Figure 4B is an exemplary diagram of in-parallelogram prediction according to one aspect of the present disclosure.
[0020] Figure 5 shows a flowchart outlining a grid decoding process according to some aspects of the present disclosure.
[0021] Figure 6 shows a flowchart outlining a grid encoding process according to some aspects of the present disclosure.
[0022] Figure 7 is a schematic diagram of a computer system according to one aspect of the present disclosure. DETAILED DESCRIPTION
[0023] Figure 1A block diagram of a video processing system (100) in some examples is shown. The video processing system (100) is an example of an application for the disclosed subject matter, video encoders, and video decoders in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, Digital Video Discs (DVDs), memory sticks, etc.
[0024] The video processing system (100) includes an acquisition subsystem (113), and the acquisition subsystem includes a video source (101). The video source (101) includes one or more images captured by a camera and / or computer-generated. For example, a digital camera can create an uncompressed video picture stream (102). In one example, the video picture stream (102) includes samples taken by the digital camera. Compared with the encoded video data (104) (or encoded video bitstream), the video picture stream (102) is depicted as a thick line to emphasize the high data volume of the video picture stream. The video picture stream (102) can be processed by an electronic device (120), and the electronic device (120) includes a video encoder (103) coupled to the video source (101). The video encoder (103) can include hardware, software, or a combination of hardware and software to implement or carry out aspects of the disclosed subject matter described in more detail below. Compared with the video picture stream (102), the encoded video data (104) (or encoded video bitstream) is depicted as a thin line to emphasize the lower data volume of the encoded video data (104) (or encoded video bitstream (104)), which can be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1 the client subsystem (106) and the client subsystem (108) in, can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) can include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and produces an output video picture stream (111) that can be presented on a display (112) (such as a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), video data (107), and video data (109) (such as video bitstreams) can be encoded according to certain video coding / compression standards. Examples of these standards include ITU-T H.265. In an embodiment, a video coding standard that is being developed is informally referred to as VVC, and this application can be used in the context of the VVC standard.
[0025] Note that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may further include a video encoder (not shown).
[0026] Figure 2 An exemplary block diagram of a video decoder (210) is shown. The video decoder (210) may be provided in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used to replace Figure 1 the video decoder (110) in the example.
[0027] The receiver (231) may receive one or more encoded video sequences to be decoded by the video decoder (210), e.g., included in a bitstream. In one aspect, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data as well as other data, e.g., encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not labeled). The receiver (231) may separate the encoded video sequences from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other cases, the buffer memory (215) may be provided external to the video decoder (210) (not labeled). In other cases, a buffer memory (not labeled) is provided external to the video decoder (210) to, e.g., prevent network jitter, and another buffer memory (215) may be configured inside the video decoder (210) to, e.g., handle the playback timing. And when the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network, it may not be necessary to configure the buffer memory (215), or the buffer memory may be made smaller. Of course, for use on a service packet network such as the Internet, a buffer memory (215) may be required, which may be relatively large and may have an adaptive size, and may be implemented at least partially in the operating system or a similar element (not labeled) external to the video decoder (210).
[0028] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from an encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (210), as well as potential information for controlling a display device (212) (e.g., a display screen), etc., which is not a component of the electronic device (230), but can be coupled to the electronic device (230), as Figure 2 shown. The control information for the display device can be a parameter set segment (not labeled) of Supplementary Enhancement Information (SEI) or Video Usability Information (VUI). The parser (220) may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence can be performed according to video coding techniques or standards and can follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (220) may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. Subgroups may include Groups of Pictures (GOPs), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), and so on. The parser (220) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, MVs, and so on.
[0029] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).
[0030] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (221) may involve multiple different units. Which units are involved and the way they are involved can be controlled by subgroup control information parsed by the parser (220) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (220) and multiple units below are not described.
[0031] In addition to the functional blocks already mentioned, the video decoder (210) can conceptually be subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the functional units below.
[0032] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives the quantized transform coefficients as symbols (221) and control information from the parser (220), including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output blocks including sample values, and the sample values can be input into the aggregator (255).
[0033] In some cases, the output samples of the scaler / inverse transform unit (251) can belong to intra-coded blocks. An intra-coded block is a block that does not use predictive information from a previously reconstructed picture, but can use predictive information from a previously reconstructed part of the current picture. Such predictive information can be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates surrounding blocks of the same size and shape as the block being reconstructed using the reconstructed information extracted from the current picture buffer (258). For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the predictive information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) based on each sample.
[0034] In other cases, the output samples of the scaler / inverse transform unit (251) can belong to inter-coded and potentially motion-compensated blocks. In this case, the motion compensation prediction unit (253) can access the reference picture memory (257) to extract samples for prediction. After motion-compensating the extracted samples according to the symbols (221) belonging to the block, these samples can be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (which is called the residual sample or residual signal in this case), thereby generating output sample information. The address from which the motion compensation prediction unit (253) obtains the prediction samples from the reference picture memory (257) can be controlled by a motion vector, and the motion vector is in the form of the symbol (221) for use by the motion compensation prediction unit (253), and the symbol (221) includes, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, a motion vector prediction mechanism, etc.
[0035] The output samples of the aggregator (255) can be adopted by various loop filtering techniques in the loop filter unit (256). Video compression techniques can include in-loop filter techniques, which are controlled by parameters included in an encoded video sequence (also referred to as an encoded video bitstream), and the parameters can be used in the loop filter unit (256) as symbols (221) from the parser (220). Video compression can also respond to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or an encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0036] The output of the loop filter unit (256) can be a sample stream, which can be output to the display device (212) and stored in the reference picture memory (257) for subsequent inter-picture prediction.
[0037] Once fully reconstructed, some encoded pictures can be used as reference pictures for future prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before starting to reconstruct subsequent encoded pictures.
[0038] The video decoder (210) can perform decoding operations according to a predetermined video compression technique or standard such as ITU-T H.265. In the sense that the encoded video sequence conforms to the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard, the encoded video sequence can comply with the syntax specified by the video compression technique or standard used. Specifically, the profile can select certain tools from all the tools available in the video compression technique or standard as the only tools available under the profile. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0039] In one aspect, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (210) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, temporal, spatial, or signal noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.
[0040] Figure 3 An exemplary block diagram of a video encoder (303) is shown. The video encoder (303) is provided in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., transmission circuitry). The video encoder (303) may be used to replace Figure 1 the video encoder (103) in the example.
[0041] The video encoder (303) may receive video samples from a video source (301) (which is not Figure 3 part of the electronic device (320) in the example), and the video source may capture video images to be encoded by the video encoder (303). In another embodiment, the video source (301) is part of the electronic device (320).
[0042] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), and the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 Y CrCB, RGB,...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as spatial pixel arrays, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. The following focuses on the description of the samples.
[0043] According to one aspect, a video encoder (303) may encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other required time constraints. Implementing an appropriate encoding speed is a function of a controller (350). In some aspects, the controller (350) controls other functional units as described below and is functionally coupled to these units. For the sake of brevity, the couplings are not labeled in the figures. Parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, GOPs layout, maximum MV search range, etc. The controller (350) may be used for other suitable functions that relate to the video encoder (303) optimized for a certain system design.
[0044] In some aspects, the video encoder (303) operates in an encoding loop. As a simple description, in an embodiment, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data. The reconstructed sample stream (sample data) is input into a reference picture memory (334). Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory (334) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is also used in some related technologies.
[0045] The operation of the "local" decoder (333) may be the same as that of the "remote" decoder of the video decoder (310) described in detail above in connection with Figure 2 However, briefly referring additionally to Figure 2 , when symbols are available and the entropy encoder (345) and the parser (220) can encode / decode the symbols losslessly into an encoded video sequence, the entropy decoding part of the video decoder (210), including the buffer memory (215) and the parser (220), may not be fully implemented in the local decoder (333).
[0046] In one aspect, any decoder technology other than parsing / entropy decoding present in the decoder exists in the corresponding encoder in an identical or substantially identical functional form. Thus, this application focuses on decoder operations. The description of encoder technology can be simplified because encoder technology is inverse to the decoder technology described comprehensively. More detailed descriptions in certain areas are provided below.
[0047] During operation, in some embodiments, the source encoder (330) may perform motion-compensated predictive coding. Referencing one or more previously encoded pictures designated as "reference pictures" in the video sequence, the motion-compensated predictive coding performs predictive coding on the input picture. In this way, the coding engine (332) encodes the difference between the pixel blocks of the input picture and the pixel blocks of the reference picture, which can be selected as the prediction reference for the input picture.
[0048] The local video decoder (333) may decode the encoded video data of the picture that can be designated as a reference picture based on the symbols created by the source encoder (330). The operation of the coding engine (332) can be a lossy process. When the encoded video data can be decoded at a video decoder ( Figure 3 not shown), the reconstructed video sequence is typically a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that can be performed by the video decoder on the reference picture and can store the reconstructed reference picture in the reference picture cache (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference picture that has the same content (in the absence of transmission errors) as the reconstructed reference picture that will be obtained by the remote video decoder.
[0049] The predictor (335) may perform a prediction search for the coding engine (332). That is, for a new picture (or grid) to be encoded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can serve as an appropriate prediction reference for the new picture. The predictor (335) may operate block by block based on sample blocks to find a suitable prediction reference. In some cases, according to the search results obtained by the predictor (335), it can be determined that the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (334).
[0050] The controller (350) may manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding the video data.
[0051] The outputs of all the above functional units can be entropy encoded in an entropy encoder (345). The entropy encoder (345) losslessly compresses the symbols generated by various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.
[0052] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) to prepare for transmission over a communication channel (360), which can be a hardware / software link leading to a storage device that will store the encoded video data. The transmitter (340) can combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0053] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures can typically be assigned to any of the following picture types:
[0054] An intra picture (I picture), which can be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures.
[0055] A predictive picture (P picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses a motion vector and a reference index to predict the sample values of each block.
[0056] A bi - predictive picture (B picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses two motion vectors and reference indexes to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.
[0057] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be prediction-encoded with reference to other (already encoded) blocks, and the other blocks are determined according to the coding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non-prediction-encoded, or the blocks can be prediction-encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be prediction-encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be prediction-encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.
[0058] The video encoder (303) can perform encoding operations according to a predetermined video coding technique or standard such as ITU-T Rec. H.265. In operation, the video encoder (303) can perform various compression operations, including prediction coding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0059] In one aspect, the transmitter (340) can transmit additional data when transmitting the encoded video. The source encoder (330) can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0060] The captured video can be a plurality of source pictures (video pictures) in a time series. Intra picture prediction (often abbreviated as intra prediction) utilizes the spatial correlation within a given picture, while inter picture prediction utilizes the (temporal or other) correlation between pictures. In an embodiment, the particular picture being encoded / decoded is segmented into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.
[0061] In some aspects, bidirectional prediction techniques can be used in inter - picture prediction. According to the bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that are both before the current picture in decoding order (but may be past and future respectively in display order) in the video. A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.
[0062] In addition, merge mode techniques can be used in inter - picture prediction to improve coding efficiency.
[0063] According to some aspects of the present disclosure, predictions such as inter - picture prediction and intra - picture prediction are performed on a block - by - block basis, such as polygon blocks or triangle blocks. For example, according to the HEVC standard, pictures in a video picture sequence are segmented into Coding Tree Units (CTUs) for compression, and CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three Coding Tree Blocks (CTBs), which are one luma CTB and two chroma CTBs. Further, each CTU can be split into one or more Coding Units (CUs) in a quadtree manner. For example, a 64×64 - pixel CTU can be split into a 64×64 - pixel CU, or 4 32×32 - pixel CUs, or 16 16×16 - pixel CUs. In one aspect, each CU is analyzed to determine the prediction type for the CU, such as an inter - prediction type or an intra - prediction type. In addition, depending on temporal and / or spatial predictability, a CU is split into one or more Prediction Units (PUs). Generally, each PU includes a luma Prediction Block (PB) and two chroma PBs. In an embodiment, prediction operations in encoding (encoding / decoding) are performed on a prediction - block basis. Taking the luma prediction block as an example of the prediction block, the prediction block includes a matrix of pixel values (e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0064] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using any suitable technology. In one aspect, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using one or more integrated circuits. In another aspect, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using one or more processors that execute software instructions.
[0065] Aspects of the present disclosure include methods and systems for selecting prediction modes in polygon mesh compression.
[0066] A mesh includes a number of polygons that describe the surface of a volumetric object. Each polygon can be defined by vertices of the mesh in 3D space and information on how the vertices are connected (referred to as connectivity information). Vertex attributes (such as color, normal, displacement, etc.) can be associated with the mesh vertices. Attributes can also be associated with the surface of the mesh through mapping information that parameterizes the mesh using a two-dimensional (2D) attribute map. Such mapping can be described by a set of parametric coordinates (referred to as UV coordinates or texture coordinates) associated with the mesh vertices. The 2D attribute map can be used to store high-resolution attribute information such as texture, normal, displacement, etc. Such information can be used for various purposes such as texture mapping, shading, and mesh reconstruction, etc.
[0067] Mesh compression includes connectivity / topology encoding and value encoding for each attribute, and the value bitstream can be larger than the connectivity bitstream. To encode the values of position and UV attributes, a prediction scheme is typically used. For example, the value of each position or UV can be predicted by using a fixed value (e.g., zero or centroid), the previous position / UV value, the average of the last n positions / UVs, or parallelogram prediction. Parallelogram prediction can include cross-parallelogram prediction and in-parallelogram prediction. Then, entropy coding can be used to encode the prediction residuals.
[0068] In cross-parallelogram prediction, the position of a vertex can be predicted based on the positions of vertices in adjacent polygons. In one aspect, if the adjacent polygon is not planar or convex, cross-parallelogram prediction may result in poor prediction performance. In in-parallelogram prediction, the position of a vertex can be predicted based on vertices within the same polygon. Since polygons tend to be planar and convex, in-parallelogram prediction is generally more accurate. Figure 4A An example of cross-parallelogram prediction is shown in Figure 4B An example of in-parallelogram prediction is shown in
[0069] As Figure 4A shown, vertices (402), (404), and (406) form a polygon. Vertices (404), (406), and (408) form a polygon. Vertices (404), (408), (410), and (412) form a polygon. According to cross-parallelogram prediction, the position of vertex (412) can be predicted based on vertices in adjacent polygons (such as vertices (404), (406), and (408)). According to the parallelogram rule, the prediction (or predicted position) of vertex (412) is determined to be (414).
[0070] As Figure 4BAs shown, according to the in - parallelogram prediction, the position of vertex (410) can be predicted based on vertices in the same polygon (such as vertices (404), (412), and (408)). According to the parallelogram rule, the prediction (or predicted position) of vertex (410) is determined to be (416).
[0071] In polygon mesh compression, predictive coding can be used to encode attribute values (such as positions or texture coordinates), and these attribute values can be a large part of the total bitstream compared to the connectivity bitstream. Therefore, effective methods are needed to select prediction modes and encode the prediction residuals of attribute values.
[0072] In the present disclosure, methods for effectively selecting prediction modes and encoding the prediction residuals of attribute values are proposed and applied to polygon mesh compression. These methods can be applied individually or in any form of combination.
[0073] In one aspect, the attributes of a mesh can include the positions of vertices (e.g., 3D coordinates or UV coordinates), the normals perpendicular to the surface of the mesh at the vertices, the colors of the vertices or surfaces of the mesh, the texture coordinates mapping the vertices or surfaces to points on a texture image, the weights assigned to vertices, etc.
[0074] In one aspect, n prediction modes can be used to predict the value of an attribute (e.g., position or texture coordinate). When encoding one of the values of the attribute, m of the n prediction modes can be available.
[0075] In the present disclosure, various ways (or methods) can be applied to select one or more prediction modes and / or encode the prediction residuals based on one or more prediction modes.
[0076] In one aspect, the encoder and decoder agree on a set of predefined rules to select one or more prediction modes from multiple candidate prediction modes. For example, priorities are set for all prediction modes, and the highest - priority available prediction mode is selected. In an example, the priority of a prediction mode in a mesh depends on the accuracy of the prediction mode, the computational cost of the prediction mode, the interpretability of the prediction mode, the robustness of the prediction mode, or depends on a specific context or application.
[0077] In an example, for encoding the position value associated with a vertex of a mesh, the following priority order is applied: in - parallelogram prediction > cross - parallelogram prediction > delta coding. The attribute value (e.g., the position value) can also be predicted by using the average of the first k available predictions or the average of all m available predictions. Each prediction can be obtained based on the corresponding prediction mode.
[0078] In an example, delta encoding includes vertex sorting, difference calculation, delta encoding, and decoding / reconstruction. In vertex sorting, the vertices of a mesh are typically sorted in a specific manner, such as lexicographically by the coordinates of the vertices or based on the topological relationships of the vertices. In difference calculation, for each vertex except the first vertex, the coordinate difference (or delta) between the vertex and its previous neighbor is calculated. These deltas are typically smaller in magnitude than the absolute vertex positions, resulting in potential data compression. In delta encoding, the calculated deltas are then encoded using a suitable compression technique, such as entropy encoding or quantization. In decoding and reconstruction, the deltas are decoded and accumulated to reconstruct the original vertex positions.
[0079] In one aspect, an encoder may use all m available prediction modes to predict the value of an attribute and may select the most accurate mode. The accuracy may be defined, for example, as the norm of the prediction residual (e.g., L1 or L2 norm). The encoder then indicates the selected prediction mode and encodes the corresponding prediction residual (e.g., the prediction residual of the selected prediction mode).
[0080] In one aspect, a more accurate prediction may be obtained by combining m available predictions. For example, the encoder iterates through all combinations of the available predictions. In an example, when 3 prediction modes (e.g., modes 1 - 3) are available, the combinations of the 3 prediction modes include: the combination of modes 1 and 2, the combination of modes 1 and 3, the combination of modes 2 and 3, and the combination of modes 1 - 3. For each combination, the average of all predictions in the combination is calculated. The encoder then selects and signals the combination with the highest prediction accuracy and encodes the corresponding prediction residual. When the total number of combinations is too large, e.g., greater than a threshold, a limit on the maximum number of combinations is set.
[0081] In one aspect, instead of using prediction accuracy, the encoder may also use other metrics to select one or more prediction modes from the candidate (or available) prediction modes. In an example, the encoder compares the increase in the bitstream size of the encoded prediction residuals from the available prediction modes and then selects and signals the prediction mode with the smallest increase in bitstream size. In other words, the prediction mode with the smallest bitstream size may be selected and signaled.
[0082] In one aspect, signaling the prediction mode may also increase the bitstream size, especially for signaling combinations of prediction modes. Thus, the increase in bitstream size brought about by signaling each available prediction mode or combination of available prediction modes may be calculated. The prediction mode or combination of prediction modes with the overall smallest increase in bitstream size (e.g., increase in bitstream size from the prediction residuals + increase in bitstream size from the prediction mode (or from the combination of prediction modes)) is selected and signaled.
[0083] In one aspect, the calculation of the bitstream size increase may result in higher computational complexity and longer encoding time. Thus, other metrics that are easier to calculate the bitstream size increase (e.g., entropy) can be utilized to estimate the bitstream size increase. For example, the entropy coding is used to estimate the bitstream size increase.
[0084] Figure 5 A flowchart outlining a process (500) according to an aspect of the present disclosure is shown. The process (500) can be used in a video decoder. In various aspects, the process (500) is executed by a processing circuit, such as a processing circuit that performs the functions of the video decoder (110), a processing circuit that performs the functions of the video decoder (210), etc. In some aspects, the process (500) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the process (500). The process starts at (S501) and continues to (S510).
[0085] At operation (S510), a bitstream including attribute information of a plurality of attributes in a grid is received.
[0086] At operation (S520), one or more prediction modes are determined from a plurality of candidate prediction modes for a current attribute among the plurality of attributes in the grid based on one of the prediction mode priorities and prediction mode precisions of one or more of the plurality of candidate prediction modes.
[0087] At operation (S530), a predicted value of the current attribute in the grid is determined based on the attribute information of the current attribute and one or more prediction modes.
[0088] At operation (S540), the current attribute is reconstructed based on the predicted value of the current attribute.
[0089] Then, the process continues to (S599) and terminates.
[0090] The process (500) can be adjusted appropriately. Steps in the process (500) can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.
[0091] Figure 6 A flowchart outlining a process (600) according to an aspect of the present disclosure is shown. The process (600) can be used in a video encoder. In various aspects, the process (600) is executed by a processing circuit, such as a processing circuit that performs the functions of the video encoder (103), a processing circuit that performs the functions of the video encoder (303), etc. In some aspects, the process (600) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the process (600). The process starts at (S601) and continues to (S610).
[0092] At operation (S610), multiple candidate prediction patterns for multiple attributes in a grid are determined.
[0093] At operation (S620), based on one of the prediction pattern priority and the prediction pattern accuracy of one or more prediction patterns among the multiple candidate prediction patterns, one or more prediction patterns are determined from the multiple candidate prediction patterns for the current attribute among the multiple attributes in the grid.
[0094] At operation (S630), the predicted value of the current attribute in the grid is encoded based on the one or more prediction patterns.
[0095] At operation (S640), signal information is encoded into the bitstream, where the signal information indicates that one or more prediction patterns are determined for the current attribute in the grid from the multiple candidate prediction patterns.
[0096] Then, the process proceeds to (S699) and terminates.
[0097] The process (600) can be adjusted appropriately. Steps in the process (600) can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.
[0098] In one aspect, a method of processing grid data includes processing a bitstream of grid data according to format rules. For example, the bitstream can be a bitstream decoded / encoded in any decoding and / or encoding method described in the present disclosure. The format rules can specify one or more constraints of the bitstream and / or one or more processes performed by a decoder and / or an encoder.
[0099] In an example, a bitstream of grid data is processed according to format rules. The bitstream includes attribute information of multiple attributes in the grid. The format rules specify that based on one of the prediction pattern priority and the prediction pattern accuracy of one or more prediction patterns among the multiple candidate prediction patterns, one or more prediction patterns are determined from the multiple candidate prediction patterns for the current attribute among the multiple attributes in the grid. The format rules specify that a predicted value of the current attribute in the grid is determined based on the attribute information of the current attribute and the one or more prediction patterns. The format rules specify that the current attribute is processed based on the predicted value of the current attribute.
[0100] The above techniques can be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 7 A computer system (700) is shown that is suitable for implementing certain aspects of the disclosed subject matter.
[0101] Computer software can be encoded using any suitable machine code or computer language, and any suitable machine code or computer language can be combined, edited, linked, or similar mechanisms to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode, etc.
[0102] The instructions can be executed on various types of computers or their components, such as including personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0103] Figure 7 The components shown for the computer system (700) are exemplary and are not used to impose any limitations on the scope of use or functions of the computer software implementing various aspects of the present application. Nor should the configuration of the components be construed as having any dependency or requirement on any one component or combination thereof shown in the exemplary aspects of the computer system (700).
[0104] The computer system (700) may include certain human-machine interface input devices. Such human-machine interface input devices can respond to inputs from one or more human users through tactile inputs (such as keyboard input, swiping, data glove movement), audio inputs (such as sound, applause), visual inputs (such as gestures), olfactory inputs (not shown). The human-machine interface device can also be used to capture certain media, which does not have to be directly related to human conscious input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0105] The human-machine interface input device may include one or more of the following (only one is drawn): keyboard (701), mouse (702), touchpad (703), touch screen (710), data glove (not shown), joystick (705), microphone (706), scanner (707), camera (708).
[0106] The computer system (700) may also include certain human-machine interface output devices. Such human-machine interface output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback through a touch screen (710), a data glove (not shown), or a joystick (705), but there may also be tactile feedback devices that do not serve as input devices), audio output devices (e.g., speakers (709), headphones (not shown)), visual output devices (e.g., a screen (710) including a cathode ray tube (CRT) screen, a liquid crystal screen, a plasma screen, an organic light-emitting diode screen, each of which may or may not have touch screen input functionality, each of which may or may not have tactile feedback functionality - some of which may output two-dimensional visual output or output above three dimensions through means such as stereoscopic picture output; virtual reality glasses (not shown), holographic displays, and a smoke box (not shown)), and a printer (not shown).
[0107] The computer system (700) may also include human-accessible storage devices and their associated media, such as optical media including high-density read-only / rewritable compact discs with CD / DVD (CD / DVD ROM / RW) (720) or similar media (721), thumb drives (722), removable hard disk drives or solid state drives (723), traditional magnetic media such as tapes and floppy disks (not shown), dedicated devices based on ROM / Application-Specific Integrated Circuit (ASIC) / Programmable Logic Device (PLD) such as a security software protector (not shown), and so on.
[0108] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.
[0109] The computer system (700) may also include an interface (754) to one or more communication networks (755). The network may be wireless, wired, optical, for example. The network may also be a LAN, a wide area network, a metropolitan area network, a vehicular network, and an industrial network, a real-time network, a delay-tolerant network, and so on. Examples of the network also include Ethernet, wireless local area network, cellular networks (Global System for Mobile communications (GSM), 3G, 4G, 5G, Long-Term Evolution (LTE), etc.), LANs such as these, television cable or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), vehicular networks, and industrial networks (including Controller Area Network Bus (CANBus)), etc. Some networks typically require an external network interface adapter for connection to some general-purpose data port or peripheral bus (749) (e.g., the Universal Serial Bus (USB) port of the computer system (700)). Other systems are typically integrated into the core of the computer system (700) by connecting to the system bus described below (e.g., an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smart phone computer system). By using any of these networks, the computer system (700) can communicate with other entities. The communication may be unidirectional, only for receiving (e.g., wireless television), unidirectional only for sending (e.g., CAN bus to some CAN bus devices), or bidirectional (e.g., via a local or wide area digital network to other computer systems). Each of the above networks and network interfaces may use certain protocols and protocol stacks.
[0110] The above-mentioned human-machine interface device, human-accessible storage device, and network interface may be connected to the core (740) of the computer system (700).
[0111] The core (740) may include one or more CPUs (741), GPUs (742), dedicated programmable processing units in the form of Field Programmable Gate Areas (FGPA) (743), hardware accelerators for specific tasks (744), graphics adapters (750), etc. These devices, as well as read-only memory (ROM) (745), random access memory (746), internal mass storage (such as internal non-user-accessible hard disk drives, solid state drives, etc.) (747), etc., can be connected through a system bus (748). In some computer systems, the system bus (748) can be accessed in the form of one or more physical plugs for expansion with additional CPUs, GPUs, etc. Peripheral devices can be directly attached to the system bus (748) of the core or connected through a peripheral bus (749). In one example, a screen (710) can be connected to the graphics adapter (750). The architecture of the peripheral bus includes Peripheral Component Interconnect (PCI), USB, etc.
[0112] The CPU (741), GPUs (742), FPGA (743), and accelerator (744) can execute certain instructions that, when combined, can constitute the above computer code. The computer code can be stored in the ROM (745) or RAM (746). Transitional data can also be stored in the RAM (746), while permanent data can be stored in, for example, the internal mass storage (747). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (741), GPUs (742), mass storage (747), ROM (745), RAM (746), etc.
[0113] The computer-readable medium may have computer code for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this application or can be those well-known and available to those skilled in the field of computer software.
[0114] By way of example and not limitation, a computer system having an architecture (700), particularly a core (740), can function as a processor (including a CPU, GPU, FPGA, accelerator, etc.) to execute software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with the above-mentioned user-accessible mass storage, as well as specific memories of the non-volatile core (740), such as the core internal mass storage (747) or ROM (745). The software implementing various embodiments of the present application can be stored in such devices and executed by the core (740). Depending on specific needs, the computer-readable media can include one or more storage devices or chips. The software can cause the core (740), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific parts of specific processes described herein, including defining data structures stored in a Random Access Memory (RAM) (746) and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system can provide functions that are logically hardwired or otherwise included in a circuit (e.g., an accelerator (744)), which can operate instead of or in conjunction with the software to execute specific processes or specific parts of specific processes described herein. In appropriate cases, references to software can include logic, and vice versa. In appropriate cases, references to computer-readable media can include circuits (such as ICs) that store software for execution, circuits that contain execution logic, or both. The present application encompasses any suitable combination of hardware and software.
[0115] The use of "at least one" or "one of" in the present disclosure is intended to include any one or combination of the recited elements. For example, references to at least one of A, B, or C, at least one of A, B, and C, at least one of A, B, and / or C, and at least one of A through C are intended to mean including only A, only B, only C, or any combination thereof. References to one of A or B and one of A and B are intended to include A or B or (A and B). Where applicable, such as when the elements are not mutually exclusive, the use of "one of" does not exclude any combination of the recited elements.
[0116] Although the present application has described several examples of various aspects, various changes, permutations, and various equivalent replacements of the embodiments are within the scope of the present application. Therefore, it should be understood that those skilled in the art can design various systems and methods that, although not explicitly shown or described herein, embody the principles of the present application and are thus within the spirit and scope of the present application.
[0117] (1) A grid decoding method, the method comprising: receiving a bitstream including attribute information of a plurality of attributes in a grid; determining, for a current attribute among the plurality of attributes in the grid, the one or more prediction patterns from the plurality of candidate prediction patterns based on one of a prediction pattern priority and a prediction pattern accuracy of the one or more prediction patterns among the plurality of candidate prediction patterns; determining a predicted value of the current attribute in the grid based on the attribute information of the current attribute and the one or more prediction patterns; and reconstructing the current attribute based on the predicted value of the current attribute.
[0118] (2) The method according to feature (1), wherein determining the one or more prediction patterns further comprises: determining the one or more prediction patterns from the plurality of candidate prediction patterns based on a prediction pattern priority of the one or more prediction patterns among the plurality of candidate prediction patterns being higher than a prediction pattern priority of other candidate prediction patterns.
[0119] (3) The method according to feature (2), wherein the one or more prediction patterns include a plurality of prediction patterns; determining the predicted value of the current attribute further comprises: determining a plurality of candidate predicted values of the current attribute based on the plurality of prediction patterns, and determining the predicted value of the current attribute as an average value of the plurality of candidate predicted values.
[0120] (4) The method according to any one of features (1) to (3), wherein determining the one or more prediction patterns further comprises: determining the one or more prediction patterns as one of the plurality of candidate prediction patterns based on prediction pattern information written in the bitstream, the prediction pattern information indicating one candidate prediction pattern among the plurality of candidate prediction patterns, the one candidate prediction pattern corresponding to a minimum prediction residual among a plurality of prediction residuals corresponding to the plurality of candidate prediction patterns.
[0121] (5) The method according to any one of features (1) to (4), wherein determining the one or more prediction patterns further comprises: determining the one or more prediction patterns as one sub - combination among a plurality of sub - combinations of the plurality of candidate prediction patterns based on prediction pattern information written in the bitstream, the prediction pattern information indicating the sub - combination among the plurality of sub - combinations of the plurality of candidate prediction patterns, the sub - combination corresponding to a minimum average prediction residual among a plurality of average prediction residuals corresponding to the plurality of sub - combinations of the plurality of candidate prediction patterns.
[0122] (6) The method according to any one of features (1) to (5), wherein determining the one or more prediction patterns further includes: determining the one or more prediction patterns as one of the plurality of candidate prediction patterns based on prediction pattern information written in the bitstream, the prediction pattern information indicating that the prediction residual of one candidate prediction pattern among the plurality of candidate prediction patterns corresponds to the smallest bitstream size among the bitstream sizes of the prediction residuals of the plurality of candidate prediction patterns.
[0123] (7) The method according to any one of features (1) to (6), wherein determining the one or more prediction patterns further includes: determining the one or more prediction patterns as one of the plurality of candidate prediction patterns based on prediction information written in the bitstream, the prediction information indicating the one candidate prediction pattern among the plurality of candidate prediction patterns, the one candidate prediction pattern corresponding to the smallest total bitstream size among (i) the prediction residuals of the plurality of candidate prediction patterns and (ii) the total bitstream size of the prediction information of the plurality of candidate prediction patterns.
[0124] (8) The method according to any one of features (1) to (7), wherein determining the one or more prediction patterns further includes: determining the one or more prediction patterns as one sub - combination among the plurality of sub - combinations of the plurality of candidate prediction patterns based on prediction pattern information written in the bitstream, the prediction pattern information indicating the sub - combination among the plurality of sub - combinations of the plurality of candidate prediction patterns, the sub - combination corresponding to the smallest total bitstream size among (i) the prediction residuals of the plurality of sub - combinations of the plurality of candidate prediction patterns and (ii) the total bitstream size corresponding to the prediction information of the plurality of sub - combinations of the plurality of candidate prediction patterns.
[0125] (9) The method according to feature (6), wherein the bitstream size of the prediction residuals of the plurality of candidate prediction patterns is estimated by entropy coding.
[0126] (10) A trellis coding method, the method includes: determining a plurality of candidate prediction patterns for a plurality of attributes in a trellis; determining the one or more prediction patterns from the plurality of candidate prediction patterns for a current attribute among the plurality of attributes in the trellis based on one of a prediction pattern priority and a prediction pattern accuracy of the one or more prediction patterns among the plurality of candidate prediction patterns; encoding a predicted value of the current attribute in the trellis based on the one or more prediction patterns; and encoding signal information into a bitstream, wherein the signal information indicates the one or more prediction patterns determined for the current attribute in the trellis from the plurality of candidate prediction patterns.
[0127] (11) The method according to feature (10), wherein determining the one or more prediction patterns further includes: determining a prediction pattern priority for each candidate prediction pattern among the plurality of candidate prediction patterns; and determining the one or more prediction patterns from the plurality of candidate prediction patterns based on the prediction pattern priority of the one or more prediction patterns among the plurality of candidate prediction patterns being higher than the prediction pattern priorities of other candidate prediction patterns.
[0128] (12) The method according to feature (11), wherein the one or more prediction patterns include a plurality of prediction patterns; and encoding the predicted value of the current attribute further includes: determining a plurality of candidate predicted values of the current attribute based on the plurality of prediction patterns, and determining the predicted value of the current attribute as the average of the plurality of candidate predicted values.
[0129] (13) The method according to any one of features (10) to (12), wherein determining the one or more prediction patterns further includes: determining a prediction residual for each candidate prediction pattern among the plurality of candidate prediction patterns; and determining the one or more prediction patterns as a candidate prediction pattern among the plurality of candidate prediction patterns, the one candidate prediction pattern corresponding to the minimum prediction residual among the prediction residuals corresponding to the plurality of candidate prediction patterns.
[0130] (14) The method according to any one of features (10) to (13), wherein determining the one or more prediction patterns further includes: determining a plurality of sub - combinations of the plurality of candidate prediction patterns; determining an average prediction residual for each sub - combination among the plurality of sub - combinations of the plurality of candidate prediction patterns; and determining the one or more prediction patterns as a sub - combination among the plurality of sub - combinations of the plurality of candidate prediction patterns, the one sub - combination corresponding to the minimum average prediction residual among the average prediction residuals corresponding to the plurality of sub - combinations of the plurality of candidate prediction patterns.
[0131] (15) The method according to any one of features (10) to (14), wherein determining the one or more prediction patterns further includes: determining a bitstream size of the prediction residual for each candidate prediction pattern among the plurality of candidate prediction patterns; and determining the one or more prediction patterns as a candidate prediction pattern among the plurality of candidate prediction patterns such that the prediction residual of the one candidate prediction pattern among the plurality of candidate prediction patterns corresponds to the minimum bitstream size among the bitstream sizes of the prediction residuals of the plurality of candidate prediction patterns.
[0132] (16) The method according to any one of features (10) to (15), wherein determining the one or more prediction patterns further includes: determining (i) the prediction residual of each candidate prediction pattern among the plurality of candidate prediction patterns and (ii) the total bitstream size of the prediction information of the corresponding candidate prediction pattern; and determining the one or more prediction patterns as a candidate prediction pattern among the plurality of candidate prediction patterns, the one candidate prediction pattern corresponding to the smallest total bitstream size among (i) the prediction residuals of the plurality of candidate prediction patterns and (ii) the total bitstream size of the prediction information of the plurality of candidate prediction patterns.
[0133] (17) The method according to any one of features (10) to (16), wherein determining the one or more prediction patterns further includes: determining a plurality of sub - combinations of the plurality of candidate prediction patterns; determining (i) the prediction residual of each sub - combination among the plurality of sub - combinations of the plurality of candidate prediction patterns and (ii) the total bitstream size of the prediction information of the corresponding sub - combination of the plurality of candidate prediction patterns; and determining the one or more prediction patterns as a sub - combination among the plurality of sub - combinations of the plurality of candidate prediction patterns such that the one sub - combination among the plurality of sub - combinations of the plurality of candidate prediction patterns corresponds to the smallest total bitstream size among (i) the prediction residuals of the plurality of sub - combinations of the plurality of candidate prediction patterns and (ii) the total bitstream size of the prediction information corresponding to the plurality of sub - combinations of the plurality of candidate prediction patterns.
[0134] (18) The method according to feature (15), wherein determining the bitstream size of the prediction residual of each candidate prediction pattern among the plurality of candidate prediction patterns further includes: estimating the bitstream size of the prediction residual of the corresponding candidate prediction pattern among the plurality of candidate prediction patterns based on entropy coding.
[0135] (19) A method for processing grid data, the method including: processing the bitstream of the grid data according to format rules, wherein the bitstream includes attribute information of a plurality of attributes in the grid, and the format rules specify: determining the one or more prediction patterns from the plurality of candidate prediction patterns for a current attribute among the plurality of attributes in the grid based on one of a prediction pattern priority and a prediction pattern accuracy of one or more prediction patterns among the plurality of candidate prediction patterns; determining a predicted value of the current attribute in the grid based on the attribute information of the current attribute and the one or more prediction patterns; and processing the current attribute based on the predicted value of the current attribute.
[0136] (20) The method according to feature (19), wherein the formatting rule further specifies that: based on the prediction mode priority of one or more of the plurality of candidate prediction modes being higher than the prediction mode priority of other candidate prediction modes, the one or more prediction modes are determined from the plurality of candidate prediction modes.
[0137] (21) A trellis decoding apparatus, comprising a processing circuit configured to perform the method according to any one of features (1) to (9).
[0138] (22) A trellis encoding apparatus, comprising a processing circuit configured to perform the method according to any one of features (10) to (18).
[0139] (23) A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform the method according to any one of features (1) to (18).
Claims
1. A grid decoding method, the method comprising: receiving a code stream including attribute information of a plurality of attributes in a grid; determining the one or more prediction modes from the plurality of candidate prediction modes for a current attribute among the plurality of attributes in the grid based on one of a prediction mode priority and a prediction mode accuracy of the one or more prediction modes among the plurality of candidate prediction modes; determining a predicted value of the current attribute in the grid based on the attribute information of the current attribute and the one or more prediction modes; as well as The current attribute is reconstructed based on the predicted value of the current attribute.
2. The method according to claim 1, wherein: Determining the one or more prediction modes further comprises: Based on the prediction mode priorities of the one or more prediction modes among the multiple candidate prediction modes being higher than the prediction mode priorities of other candidate prediction modes, the one or more prediction modes are determined from the multiple candidate prediction modes.
3. The method according to claim 2, wherein: The one or more prediction modes include a plurality of prediction modes; Determining the predicted value of the current attribute further includes: determining a plurality of candidate prediction values of the current attribute based on the plurality of prediction modes, and The predicted value of the current attribute is determined as an average of the plurality of candidate predicted values.
4. The method according to claim 1, wherein: Determining the one or more prediction modes further comprises: The one or more prediction modes are determined as one of the multiple candidate prediction modes based on prediction mode information written in the code stream, the prediction mode information indicates the one candidate prediction mode among the multiple candidate prediction modes, and the one candidate prediction mode corresponds to the minimum prediction residual among the multiple prediction residuals corresponding to the multiple candidate prediction modes.
5. The method according to claim 1, wherein: Determining the one or more prediction modes further comprises: The one or more prediction modes are determined as one of multiple sub-combinations of the multiple candidate prediction modes based on prediction mode information written in the code stream, the prediction mode information indicates the sub-combination of the multiple sub-combinations of the multiple candidate prediction modes, and the sub-combination corresponds to the minimum average prediction residual among multiple average prediction residuals corresponding to the multiple sub-combinations of the multiple candidate prediction modes.
6. The method according to claim 1, wherein: Determining the one or more prediction modes further comprises: The one or more prediction modes are determined as one of the multiple candidate prediction modes based on prediction mode information written in the code stream, and the prediction mode information indicates that the prediction residual of the one candidate prediction mode among the multiple candidate prediction modes corresponds to the smallest code stream size among the code stream sizes of the prediction residuals of the multiple candidate prediction modes.
7. The method according to claim 1, wherein: Determining the one or more prediction modes further comprises: The one or more prediction modes are determined as one of the multiple candidate prediction modes based on prediction information written in the bitstream, the prediction information indicating the one candidate prediction mode among the multiple candidate prediction modes, and the one candidate prediction mode corresponds to the smallest total bitstream size between (i) prediction residuals of the multiple candidate prediction modes and (ii) total bitstream sizes of prediction information of the multiple candidate prediction modes.
8. The method according to claim 1, wherein: Determining the one or more prediction modes further comprises: The one or more prediction modes are determined as one of multiple sub-combinations of the multiple candidate prediction modes based on prediction mode information written in the codestream, the prediction mode information indicating the sub-combination of the multiple sub-combinations of the multiple candidate prediction modes, and the sub-combination corresponds to the smallest total codestream size among the total codestream sizes corresponding to (i) prediction residuals of the multiple sub-combinations of the multiple candidate prediction modes and (ii) prediction information of the multiple sub-combinations of the multiple candidate prediction modes.
9. A grid coding method, comprising: determining a plurality of candidate prediction modes for a plurality of attributes in a grid; determining the one or more prediction modes from the plurality of candidate prediction modes for a current attribute among the plurality of attributes in the grid based on one of a prediction mode priority and a prediction mode accuracy of the one or more prediction modes among the plurality of candidate prediction modes; encoding a predicted value of the current attribute in the grid based on the one or more prediction modes; as well as Signal information is encoded into a bitstream, wherein the signal information indicates that the one or more prediction modes are determined for the current attribute in the grid from the plurality of candidate prediction modes.
10. The method according to claim 9, wherein: Determining the one or more prediction modes further comprises: determining a prediction mode priority of each candidate prediction mode among the plurality of candidate prediction modes; and Based on the prediction mode priorities of the one or more prediction modes among the multiple candidate prediction modes being higher than the prediction mode priorities of other candidate prediction modes, the one or more prediction modes are determined from the multiple candidate prediction modes.
11. The method according to claim 9, wherein: Determining the one or more prediction modes further comprises: determining a prediction residual for each of the plurality of candidate prediction modes; and The one or more prediction modes are determined as one candidate prediction mode among the multiple candidate prediction modes, and the one candidate prediction mode corresponds to a minimum prediction residual among the prediction residuals corresponding to the multiple candidate prediction modes.
12. The method according to claim 9, wherein: Determining the one or more prediction modes further comprises: determining a plurality of subsets of the plurality of candidate prediction modes; determining an average prediction residual for each of the plurality of subsets of the plurality of candidate prediction modes; and The one or more prediction modes are determined as one of the multiple sub-combinations of the multiple candidate prediction modes, and the one sub-combination corresponds to a minimum average prediction residual among the average prediction residuals corresponding to the multiple sub-combinations of the multiple candidate prediction modes.
13. The method according to claim 9, wherein: Determining the one or more prediction modes further comprises: Determining a bitstream size of a prediction residual of each candidate prediction mode among the plurality of candidate prediction modes; and The one or more prediction modes are determined as one candidate prediction mode among the multiple candidate prediction modes, so that the prediction residual of the one candidate prediction mode among the multiple candidate prediction modes corresponds to the smallest bitstream size among the bitstream sizes of the prediction residuals of the multiple candidate prediction modes.
14. The method according to claim 9, wherein: Determining the one or more prediction modes further comprises: Determining (i) a prediction residual of each candidate prediction mode among the plurality of candidate prediction modes and (ii) a total bitstream size of prediction information of the corresponding candidate prediction mode; and The one or more prediction modes are determined as one of the multiple candidate prediction modes, and the one candidate prediction mode corresponds to the smallest total bitstream size between (i) the prediction residuals of the multiple candidate prediction modes and (ii) the total bitstream size of the prediction information of the multiple candidate prediction modes.
15. The method according to claim 9, wherein: Determining the one or more prediction modes further comprises: determining a plurality of subsets of the plurality of candidate prediction modes; determining (i) a prediction residual of each of the plurality of subsets of the plurality of candidate prediction modes and (ii) a total bitstream size of prediction information of the corresponding subsets of the plurality of candidate prediction modes; and The one or more prediction modes are determined as one of the multiple subcombinations of the multiple candidate prediction modes, so that the one of the multiple subcombinations of the multiple candidate prediction modes corresponds to the smallest total code stream size among the total code stream sizes corresponding to (i) the prediction residuals of the multiple subcombinations of the multiple candidate prediction modes and (ii) the prediction information of the multiple subcombinations of the multiple candidate prediction modes.
16. The method according to claim 13, wherein: Determining the bitstream size of the prediction residual of each candidate prediction mode among the multiple candidate prediction modes also includes: The code stream size of the prediction residual of the corresponding candidate prediction mode among the multiple candidate prediction modes is estimated based on entropy coding.
17. A method for processing grid data, the method comprising: Processing a code stream of the grid data according to a format rule, wherein the code stream includes attribute information of a plurality of attributes in the grid, The format rules specify: determining the one or more prediction modes from the plurality of candidate prediction modes for a current attribute among the plurality of attributes in the grid based on one of a prediction mode priority and a prediction mode accuracy of the one or more prediction modes among the plurality of candidate prediction modes; determining a predicted value of the current attribute in the grid based on the attribute information of the current attribute and the one or more prediction modes; and The current attribute is processed based on the predicted value of the current attribute.
18. A trellis decoding device, comprising a processing circuit, wherein the processing circuit is configured to execute the trellis decoding method according to any one of claims 1 to 8.
19. A grid coding device, comprising a processing circuit, wherein the processing circuit is configured to execute the grid coding method according to any one of claims 9 to 16.
20. A computer-readable storage medium storing a computer program, wherein the computer program is executed by a processor to execute the encoding method according to any one of claims 9 to 16 to form a code stream, and the code stream is stored in the computer-readable storage medium.