Grid encoding and decoding method, grid data processing method, medium and equipment
By identifying the context index of the mesh sector, the size of the sector is predicted based on its quadrilaterality and number of faces, thus solving the problem of low compression efficiency of polygon meshes in the prior art and achieving more efficient encoding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2026-03-10
AI Technical Summary
Existing video/mesh compression techniques are inefficient at predicting the face fan size of polygonal meshes, resulting in low encoding efficiency.
By identifying face sectors in the mesh, determining the context index of the face sector, predicting the size of the face sector based on whether the face is a quadrilateral face and whether the number of faces is greater than four, and using the context indicated by the context index for encoding and decoding, the encoding efficiency is improved.
It improves the accuracy of face fan size prediction in polygon mesh compression, thereby increasing coding efficiency.
Smart Images

Figure CN121644820A_ABST
Abstract
Description
Related applications
[0001] This application claims priority to U.S. Provisional Application No. 63 / 690,252, filed September 3, 2024, entitled "Accurate Prediction of Face Fan Size in Polygon Mesh Compression," and U.S. Application No. 19 / 232,762, filed June 9, 2025, entitled "Accurate Prediction of Face Fan Size in Polygon Mesh Compression," the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure describes various aspects relating generally to grid encoding and decoding, and in particular to methods for grid encoding and decoding, methods for processing grid data, media, and devices. Background Technology
[0003] The background description provided herein is for the purpose of generally presenting the context of this disclosure. The work of the currently attributed inventors is neither expressly nor impliedly acknowledged as prior art to this disclosure, with regard to the extent that such work is described in the background section and in aspects that may not conform to the prior art at the time of application.
[0004] Video / mesh compression can help transmit image / video data across different devices, storage devices, and networks with minimal quality degradation. In some examples, video (or mesh) codec techniques can compress video (or mesh) based on spatial and temporal redundancy. For instance, a video (or mesh) codec can use a technique called intra-frame prediction, which can compress images (or meshes) based on spatial redundancy. For example, intra-frame prediction can use reference data from the current image (or current mesh) in reconstruction for sample prediction. In another example, a video (or mesh) codec can use a technique called inter-frame prediction, which can compress images (or meshes) based on temporal redundancy. For example, inter-frame prediction can use motion compensation to predict samples in the current image (or current mesh) from a previously reconstructed image (or previously reconstructed mesh). Motion compensation can be indicated by motion vectors (MV). Summary of the Invention
[0005] This disclosure includes methods for grid encoding / decoding, methods for processing grid data, media, and apparatus. In some examples, an apparatus for grid encoding / decoding includes processing circuitry.
[0006] According to one aspect of this disclosure, a method for mesh decoding is provided. In this method, a bitstream comprising encoded information of a mesh is received. The mesh comprises face fans. Each face fan comprises at least two faces and a pivot vertex. The pivot vertex is the center point of the at least two faces, and the at least two faces are associated with the pivot vertex. A context index of the face fan is determined based on whether each of the at least two visited faces is a quadrilateral face and whether the total number of visited faces in the at least two faces is greater than four. The size of the face fan is decoded based on the context indicated by the context index.
[0007] According to another aspect of this disclosure, a method for mesh encoding is provided. In this method, face fans in a mesh are identified. A face fan comprises at least two faces and a pivot vertex, such that the pivot vertex is the center point of at least two faces, and at least two faces are associated with the pivot vertex. A context index of the face fan is determined based on whether each of the at least two visited faces is a quadrilateral face and whether the total number of visited faces in the at least two faces is greater than four. The size of the face fan is encoded based on the context indicated by the context index.
[0008] According to another aspect of this disclosure, a method for processing mesh data is provided. In this method, a bitstream of mesh data is processed according to format rules. The bitstream includes encoded information of the mesh. The mesh includes face fans. Each face fan includes at least two faces and a pivot vertex. The pivot vertex is the center point of the at least two faces, and the at least two faces are associated with the pivot vertex. The format rules instruct the determination of the context index of the face fan based on whether each of the at least two visited faces is a quadrilateral face and whether the total number of visited faces in the at least two faces is greater than four. The format rules instruct the processing of the face fan size based on the context indicated by the context index.
[0009] This disclosure also provides an apparatus for mesh decoding. The apparatus for mesh decoding includes processing circuitry configured to implement any of the described methods for mesh decoding.
[0010] This disclosure also provides an apparatus for trellis coding. The apparatus for trellis coding includes processing circuitry configured to implement any of the described methods for trellis coding. An electronic device is also provided, comprising: a processor; a memory; and at least a set of instructions stored in the memory and configured to be executed by the processor, which, when executed, implement the methods of this application. A method for storing a bitstream, performing trellis coding to generate a bitstream, and storing the bitstream is also provided. A method for transmitting a bitstream, performing trellis coding to generate a bitstream, and transmitting the bitstream is also provided.
[0011] This disclosure also provides a non-volatile computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the methods for grid decoding, encoding, and / or processing.
[0012] The technical solution disclosed herein includes a method and apparatus for predicting the size of face fans in a polygonal mesh based on a determined context index to improve encoding efficiency. In an example, face fans in the mesh are identified. Each face fan includes at least two faces and a pivot vertex. The pivot vertex is the center point of the at least two faces, and the at least two faces are associated faces connected to the pivot vertex. The context index of the face fan is determined based on whether each of the at least two visited faces is a quadrilateral face and whether the total number of visited faces in the at least two faces is greater than four. The size of the face fan is encoded based on the context indicated by the context index. Therefore, by predicting the size of the face fan associated with the pivot vertex based on a determined context index, encoding efficiency is improved. Attached Figure Description
[0013] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:
[0014] Figure 1 This is a schematic diagram of an example block diagram of a communication system.
[0015] Figure 2 This is a schematic diagram of an example block diagram of a decoder.
[0016] Figure 3 This is a schematic diagram of an example block diagram of an encoder.
[0017] Figure 4 An example of a topology configuration for compressing polygonal sector connectivity is shown, according to some aspects of this disclosure.
[0018] Figure 5 An example of a quadrilateral sector configuration according to some aspects of this disclosure is shown.
[0019] Figure 6 A flowchart outlining some aspects of the decoding method according to this disclosure is shown.
[0020] Figure 7 A flowchart outlining some aspects of the coding method according to this disclosure is shown.
[0021] Figure 8 It is a schematic diagram of a computer system based on one aspect. Detailed Implementation
[0022] Figure 1Block diagrams of some example video processing systems (100) are shown. The video processing system (100) is an example of an application of video encoders and video decoders in the disclosed subject matter and video processing environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital television, streaming media services, and storing compressed video on digital media including CDs, DVDs, Memory Sticks, etc.
[0023] The video processing system (100) may include an acquisition subsystem (113) that may include a video source (101) such as a digital camera, which creates an uncompressed video image stream (102). In an embodiment, the video image stream (102) includes samples captured by a digital camera. The video image stream (102) is depicted as a thick line to emphasize the high data volume of the video image stream compared to encoded video data (104) (or encoded video bitstream). The video image stream (102) may be processed by an electronic device (120) that includes a video encoder (mesh encoder) (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or carry out aspects of the disclosed subject matter as described in more detail below. Compared to the video image stream (102), the encoded video data (104) (or the encoded video bitstream (104)) is depicted as a thin line to emphasize the lower data volume of the encoded video data (104) (or the encoded video bitstream (104)), which can be stored on a streaming server (105) for future use. At least one streaming client subsystem, such as Figure 1 Client subsystems (106) and (108) can access a streaming server (105) to retrieve copies (107) and (109) of encoded video data (104). Client subsystem (106) may include, for example, a video decoder (mesh decoder) (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and produces an output video picture stream (111) that can be displayed on a display (112) (e.g., a screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), video data (107), and video data (109) (e.g., video streams) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T H.265. In embodiments, the video coding standard under development is informally referred to as Versatile Video Coding (VVC), and this application can be used in the context of the VVC standard.
[0024] It should be noted that electronic devices (120) and (130) may include other (not shown) components. For example, electronic device (120) may integrate a video decoder (not shown), and electronic device (130) may also integrate a video encoder (not shown).
[0025] Figure 2 This is an example block diagram of a video decoder (or mesh decoder) (210). The video decoder (210) may be disposed in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., receiving circuitry). The video decoder (210) may be used in place of... Figure 1 The video decoder (110) in the embodiment.
[0026] The receiver (231) can receive at least one encoded video sequence to be decoded by the video decoder (210), the at least one encoded video sequence being included in, for example, a bitstream. In one aspect, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequences can be received from a channel (201), which can be a hardware / software link to a storage device storing the encoded video data. The receiver (231) can receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that can be forwarded to their respective user entities (not indicated). The receiver (231) can separate the encoded video sequences from other data. To prevent network jitter, a buffer (215) can be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer (215) is part of the video decoder (210). In other cases, the buffer memory (215) may be located externally to the video decoder (210) (not shown). In other cases, a buffer memory (not shown) may be located externally to the video decoder (210) to prevent network jitter, for example, and another buffer memory (215) may be configured internally to handle broadcast timing, for example. When the receiver (231) receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network, a buffer memory (215) may not be necessary, or it may be made smaller. Of course, a buffer memory (215) may also be required for use on packet networks such as the Internet. This buffer memory may be relatively large and adaptive in size, and may be at least partially implemented in the operating system or a similar component (not shown) external to the video decoder (210).
[0027] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. These symbols may include information for managing the operation of the video decoder (210) and potential information for controlling a display device (212) (e.g., a display screen), which is not part of the electronic device (230) but may be coupled to it, such as... Figure 2 As shown in the figure. The control information for the display device may be a parameter set fragment (not shown) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (220) may parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract a subgroup parameter set of at least one subgroup of pixels in the subgroup of pixels in the encoded video sequence for use in the video decoder based on at least one parameter corresponding to a group. The subgroup may include Group of Pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (220) can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0028] The parser (220) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).
[0029] Depending on the type of encoded video frames or a subset of encoded video frames (e.g., inter-frame and intra-frame frames, inter-frame and intra-frame blocks) and other factors, the reconstruction of the symbol (221) may involve multiple different units. Which units are involved and how they are involved can be controlled by the subgroup control information parsed from the encoded video sequence by the parser (220). For brevity, the flow of such subgroup control information between the parser (220) and the various units described below is not described.
[0030] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with one another. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the functional units described below.
[0031] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives quantization transform coefficients as symbols (221) and control information from the parser (220), including the transform method used, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output a block containing sample values, which can be input into the aggregator (255).
[0032] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use predictive information from a previously reconstructed image, but can use predictive information from a previously reconstructed portion of the current image. Such predictive information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses reconstructed information extracted from the current picture buffer (258) to generate a surrounding block of the same size and shape as the block being reconstructed. For example, the current picture buffer (258) buffers a partially reconstructed current image and / or a fully reconstructed current image. In some cases, the aggregator (255) adds the predictive information generated by the intra-picture prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) based on each sample.
[0033] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to inter-frame coding and latent motion compensation blocks. In this case, the motion compensation prediction unit (253) can access the reference image memory (257) to extract samples for prediction. After motion compensation of the extracted samples according to the symbols (221), these samples can be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (referred to as residual samples or residual signals in this case) to generate output sample information. The motion compensation prediction unit (253) can obtain the predicted samples from the address in the reference image memory (257) under motion vector control, and the motion vector is available to the motion compensation prediction unit (253) in the form of the symbols (221), which, for example, include X, Y and reference image components. Motion compensation may also include interpolation of sample values extracted from the reference image memory (257) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.
[0034] The output samples of the aggregator (255) can be employed by various loop filtering techniques in the loop filter unit (256). Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video sequence (also referred to as the encoded video stream), and these parameters can be used as symbols (221) from the parser (220) in the loop filter unit (256). Video compression techniques may also respond to metadata obtained during decoding of a previous (in decoding order) portion of the encoded image or encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0035] The output of the loop filter unit (256) can be a sample stream, which can be output to a display device (212) and stored in a reference image memory (257) for subsequent inter-frame image prediction.
[0036] Once fully reconstructed, some of the encoded images can be used as reference images for future predictions. For example, once the encoded images corresponding to the current image have been fully reconstructed and the encoded images (by, for example, the parser (220)) are identified as reference images, the current image buffer (258) can become part of the reference image memory (257), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.
[0037] The video decoder (210) can perform decoding operations according to a predetermined video compression technique, such as that specified in the ITU-T H.265 standard. The encoded video sequence may conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the configuration file recorded in the video compression technique or standard. Specifically, the configuration file may select certain tools from all available tools in the video compression technique or standard as the only tools available under said configuration file. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference image size, etc. In some cases, the limitations set by the hierarchy can be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer, which is represented by signals in the encoded video sequence.
[0038] In one aspect, the receiver (231) may receive supplemental (redundant) data along with the encoded video. The supplemental data may be a portion of the encoded video sequence. The supplemental data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The supplemental data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.
[0039] Figure 3 This is an example block diagram of a video encoder (grid encoder) (303). The video encoder (303) is disposed in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., transmission circuitry). The video encoder (303) can be used in place of... Figure 1 The video encoder (103) in the embodiment.
[0040] The video encoder (303) can obtain data from the video source (301) (not) Figure 3 In one embodiment, a portion of the electronic device (320) receives video samples, the video source being capable of capturing video images to be encoded by a video encoder (303). In another embodiment, the video source (301) is a portion of the electronic device (320).
[0041] A video source (301) can provide a sequence of source videos encoded by a video encoder (303) in the form of a digital video sample stream, which can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601YCrCb, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (301) can be a storage device storing previously prepared video. In a video conferencing system, the video source (301) can be a camera that captures local image information as a video sequence. Video data can be provided as multiple individual pictures, which are given motion when viewed in sequence. The pictures themselves can be constructed as spatial pixel arrays, where each pixel can include at least one sample, depending on the sampling structure, color space, etc. used. The following focuses on describing the samples.
[0042] According to an embodiment, the video encoder (303) can encode and compress images of a source video sequence into an encoded video sequence (343) in real time or under any other required time constraints. Implementing an appropriate encoding rate is a function of the controller (350). In some embodiments, the controller (350) controls and is functionally coupled to other functional units described below. For simplicity, coupling is not shown in the figures. Parameters set by the controller (350) may include rate control related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) may be used with other suitable functions related to the video encoder (303) optimized for a particular system design.
[0043] In some embodiments, the video encoder (303) operates within an encoding loop. As a simplified description, in an embodiment, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder (333) embedded within the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data. The reconstructed sample stream (sample data) is input to a reference image memory (334). Since decoding of the symbol stream produces bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference image memory (334) also correspond bit-accurately between the local and remote encoders. In other words, the reference image samples "seen" by the encoder's prediction section are exactly the same sample values that the decoder will "see" during decoding when using the prediction. This fundamental principle of reference image synchronization (and the drift that occurs when synchronization cannot be maintained, for example, due to channel errors) is also used in some related techniques.
[0044] The operation of the “local” decoder (333) can be combined with, for example, the above-mentioned... Figure 2 The video decoder (210) is described in detail as the same as the "remote" decoder. However, a further brief reference is provided. Figure 2 When symbols are available and the entropy encoder (345) and parser (220) are able to encode / decode the symbols into an encoded video sequence without loss, the entropy decoding portion of the video decoder (210), including the buffer (215) and parser (220), may not be fully implemented in the local decoder (333).
[0045] In one aspect, decoder techniques other than parsing / entropy decoding, which exist in the decoder, also exist in the corresponding encoder in the same or substantially the same functional form. Therefore, this application focuses on decoder operation. The description of encoder techniques can be simplified because encoder techniques are inverses of the fully described decoder techniques. In certain areas, more detailed descriptions are provided below.
[0046] During operation, in some embodiments, the source encoder (330) may perform motion-compensated predictive coding. The motion-compensated predictive coding predictively encodes the input image with reference to at least one previously encoded image from the video sequence designated as a "reference image." In this manner, the encoding engine (332) encodes the differences between pixel blocks of the input image and pixel blocks of the reference image, which may be selected as a predictive reference for the input image.
[0047] The local video decoder (333) can decode encoded video data of a picture that can be designated as a reference picture, based on symbols created by the source encoder (330). The operation of the encoding engine (332) can be a lossy process. When the encoded video data can be decoded by the video decoder (333), Figure 3 When the source video sequence (not shown) is decoded, the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process, which can be performed by the video decoder on the reference image, and allows the reconstructed reference image to be stored in the reference image memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference image that shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.
[0048] The predictor (335) can perform a prediction search against the encoding engine (332). That is, for a new image to be encoded, the predictor (335) can search in the reference image memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new image. The predictor (335) can operate pixel-by-pixel based on the sample blocks to find suitable prediction references. In some cases, based on the search results obtained by the predictor (335), it can be determined that the input image may have prediction references obtained from multiple reference images stored in the reference image memory (334).
[0049] The controller (350) can manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.
[0050] The outputs of all the above-mentioned functional units can be entropy encoded in the entropy encoder (345). The entropy encoder (345) performs lossless compression on the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, and arithmetic coding, thereby converting the symbols into an encoded video sequence.
[0051] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission via a communication channel (360), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (340) can combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0052] The controller (350) manages the operation of the video encoder (303). During encoding, the controller (350) can assign a specific encoded image type to each encoded image, but this may affect the encoding techniques applicable to the corresponding images. For example, images can typically be assigned to any of the following image types:
[0053] Intra-frame pictures (I-pictures) are pictures that can be encoded and decoded without using any other pictures in the sequence as prediction sources. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures.
[0054] A predictive image (P-image) can be an image that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses a motion vector and a reference index to predict sample values for each block.
[0055] A bidirectional predictive image (B-image) can be an image that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses two motion vectors and a reference index to predict sample values for each block. Similarly, multiple predictive images can use more than two reference images and associated metadata to reconstruct a single block.
[0056] Source images are typically spatially subdivided into multiple sample blocks (e.g., 4×4, 8×8, 4×8, or 16×16 sample blocks), and encoded block by block. These blocks can be predictively coded with reference to other (already coded) blocks, determined based on the coding assignments of the corresponding images applied to the blocks. For example, blocks of an I-image can be non-predictively coded, or the blocks can be predictively coded (spatial or intra-frame prediction) with reference to already coded blocks of the same image. Pixel blocks of a P-image can be predictively coded with reference to a previously coded reference image via spatial or temporal prediction. Blocks of a B-image can be predictively coded with reference to one or two previously coded reference images via spatial or temporal prediction.
[0057] The video encoder (303) can perform encoding operations according to a predetermined video coding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (303) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0058] In this embodiment, the transmitter (340) may transmit additional data while transmitting encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0059] The acquired video can serve as multiple source images (video images) presented in a time series. Intra-frame image prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In an embodiment, a specific image being encoded / decoded is segmented into blocks, referred to as the current image. When a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. This motion vector points to the reference block in the reference image, and when multiple reference images are used, the motion vector may have a third dimension that identifies the reference image.
[0060] In some embodiments, bidirectional prediction techniques can be used in inter-frame image prediction. According to bidirectional prediction, two reference images are used, such as a first reference image and a second reference image, both preceding the current image in the video in decoding order (but possibly past and future in display order). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. Specifically, the block can be predicted using a combination of the first and second reference blocks.
[0061] In addition, merging mode techniques can be used in inter-frame image prediction to improve coding efficiency.
[0062] According to some embodiments disclosed in this application, predictions such as inter-frame image prediction and intra-frame image prediction are performed on a block-by-block basis. For example, according to the HEVC standard, images in a video image sequence are segmented into coding tree units (CTUs) for compression. The CTUs in the images have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Furthermore, each CTU can be further subdivided into at least one coding unit (CU) using a quadtree. For example, a 64×64 pixel CTU can be subdivided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In embodiments, each CU is analyzed to determine the prediction type used for the CU, such as inter-frame prediction or intra-frame prediction. Furthermore, depending on temporal and / or spatial predictability, the CU is divided into at least one prediction unit (PU). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In embodiments, prediction operations in encoding (encoding / decoding) are performed on a per-prediction-block basis. Taking a luma prediction block as an example, a prediction block includes a matrix of pixel values (e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0063] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using any suitable technology. In one aspect, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using at least one integrated circuit. In another aspect, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using at least one processor that executes software instructions.
[0064] Various aspects of this disclosure provide techniques for predicting face sector sizes in polygon mesh compression. In some aspects, methods are provided for more accurately predicting face sector sizes for polygon mesh compression.
[0065] Advances in 3D capture, modeling, and rendering have facilitated the ubiquitous presence of 3D content across numerous platforms and devices. Today, it's possible to take the first step in capturing images of a baby on one landmass, allowing the baby's grandparents to see (and potentially interact with) and enjoy a fully immersive experience with the child on another landmass. However, to achieve this realism, the models have become increasingly complex, and vast amounts of data are associated with their creation and consumption. 3D meshes are widely used to represent this immersive content.
[0066] A mesh can comprise several polygons describing the surface of a volumetric object. Each polygon can be defined by vertices in three-dimensional (3D) space and information about how these vertices are connected (called connectivity information). Vertex attributes (such as color, normals, displacement, etc.) can be associated with mesh vertices. Attributes can also be associated with the mesh's surface by utilizing mapping information parametrically assigned to the mesh using a two-dimensional (2D) attribute graph. This mapping can be described by a set of parametric coordinates, called UV coordinates or texture coordinates, associated with the mesh vertices. 2D attribute graphs can be used to store high-resolution attribute information, such as texture, normals, displacement, etc. This information can be used for various purposes, such as texture mapping, shading, and mesh reconstruction.
[0067] In the relevant examples, a polygonal fan algorithm is proposed to encode the connectivity of polygonal meshes. To encode connectivity at pivot vertices, the polygonal fan algorithm in the relevant examples first determines the polygonal fan. For each polygonal fan, entropy coding is used to encode the total number of faces and the degree of each face (e.g., the total number of associated vertices of the face). Then, by applying... Figure 4 One of the nine topological configurations C0 to C8 illustrated in the figure is used to encode the connectivity of polygonal sectors. For example, in configuration C0, a pivot vertex (502) may have a first polygonal sector (504) to be processed and a second polygonal sector (506) to be processed.
[0068] The polygonal fan algorithm in the relevant examples can be used to encode the connectivity of polygonal meshes. To encode the connectivity of a polygonal mesh, the size of each face fan encountered during traversal can be encoded. Therefore, it is desirable to improve the efficiency of encoding face fan sizes.
[0069] This disclosure includes methods for improving the coding efficiency of face fan size in polygonal fan algorithms. These methods can be applied individually or in any combination.
[0070] In the polygonal fan algorithm, a face fan is called a quadrilateral fan when all faces in the face fan are quadrilaterals. Figure 5 Examples of quadrilateral sector configurations in the current face sector (500) are shown. Figure 5 As shown, in configuration (502), there are no consecutively visited quadrilaterals, and the size of the current face fan (500) can be 4. In configuration (504), there is one visited quadrilateral (514), and the size of the current face fan (500) can be 3. In configuration (506), there are three consecutively visited quadrilaterals (516), (518), and (520), and the size of the current face fan (500) can be 1. In the examples, the face fan size indicates the size of the square or rectangle occupied by the fan. When two visited quadrilateral faces are around a pivot, if the two visited quadrilateral faces are consecutive, such as in configuration (508), the size of the current face fan (500) is expected to be two. Otherwise, if the two visited quadrilateral faces are not consecutive and touch at a single vertex, such as in configurations (510) and (512), the size of the current face fan is expected to be one.
[0071] Figure 5 The configuration shown assumes the pivot (or pivot vertex) has 4 associated quadrilateral faces, which may not always be valid. For example, as... Figure 4 As shown in configuration C8, when a pivot vertex (508) already has four visited faces (e.g., faces (510), (512), and (514)) and it has another face fan (516) to be encoded, then the pivot vertex (508) can have at least five associated faces, which violates the assumption that the pivot has four associated quadrilateral faces. In another case, this assumption is violated when some of the associated faces are not quadrilaterals. For these cases where the assumption of four associated quadrilateral faces is not satisfied, a method can be applied that utilizes the priority of the pivot vertex to select a context (e.g., an entropy coding context or a mathematical coding context) to encode the face fan size.
[0072] Various aspects of this disclosure include predicting the face fan size of a polygon mesh based on a context index determined according to the face fan configuration. For example, pseudocode is provided below to encode the face fan size:
[0073] As shown in the pseudocode, "encode_ffan_size_pivot_priority" corresponds to the method for encoding the fan size based on the context determined by using the priority of the pivot, and "encode_quad_fan_size" corresponds to the proposed method for encoding the fan size using the context determined according to the fan configuration.
[0074] Referring back to the pseudocode above, when it is determined that none of the visited faces are quadrilaterals according to "if(all_visited_faces_quad)", the face fan size is encoded using the pivot's priority according to "encode_ffan_size_pivot_priority(ffan_size)". For example, the context is selected based on the pivot's priority, and the face fan size is encoded based on the selected context.
[0075] When it is determined that all visited faces are quadrilaterals based on "if(all_visited_faces_quad)" and the visited face count is equal to or greater than 4 based on "if(visited_face_count<4)", the face fan size is encoded using the pivot's priority according to "encode_ffan_size_pivot_priority(ffan_size)". For example, the context is selected based on the pivot's priority, and the face fan size is encoded based on the selected context.
[0076] When it is determined that all visited faces are quadrilaterals based on "if(all_visited_faces_quad)", and the visited face count is less than 4 based on "if(visited_face_count<4)", the variable context_index is defined based on "int context_index". "int context_index" indicates that the variable "context_index" is an integer and refers to the context index. Furthermore, (i) when it is determined that the visited face count of the pivot vertex is equal to 2 according to "if(visited_face_count==2)", if it is determined that the two visited faces share an edge according to "context_index=two_visited_face_share_an_edge?2:4", then the context index is equal to 2; and if it is determined that the two visited faces do not share an edge according to "context_index=two_visited_face_share_an_edge?2:4", then the context index is equal to 4; (ii) when it is determined that the visited face count of the pivot vertex is not equal to 2 according to "if(visited_face_count==2)", then the context index is determined to be the visited face count according to "context_index=visited_face_count". Based on the determined context, the fan size is further encoded by subtracting one.
[0077] In one aspect, besides using a quad-fan configuration to select the context for encoding the fan size, this configuration can also be used to predict the face fan size, and the same context can be used to further encode the prediction residuals. For example, the predicted face fan size is obtained based on the quad-fan configuration. The context can be derived based on the configuration according to the pseudocode above. The prediction residuals can be determined based on the predicted values and the face fan size. The prediction residuals can be further encoded based on the context.
[0078] In one aspect, a quadrilateral sector configuration is applied to simultaneously predict the sector size and select the coding context, and the selected coding context is used to encode the prediction residuals.
[0079] Figure 6A flowchart outlining a method (600) according to one aspect of this disclosure is shown. The method (600) can be used in a mesh decoder. In various aspects, the method (600) is executed by processing circuitry, such as processing circuitry executing the functions of mesh decoder (110), processing circuitry executing the functions of mesh decoder (210), etc. In some aspects, the method (600) is implemented as software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes the method (600). The method begins at (S601) and proceeds to (S610).
[0080] At (S610), a bitstream including encoded information of a grid is received. The grid comprises face sectors. Each face sector comprises at least two faces and a pivot vertex. The pivot vertex is the center point of at least two faces, and at least two faces are associated with the pivot vertex.
[0081] At (S620), the context index of the face sector is determined based on whether each of the at least two visited faces is a quadrilateral face and whether the total number of visited faces in the at least two faces is greater than four.
[0082] At (S630), the size of the face fan is decoded based on the context indicated by the context index.
[0083] In one aspect, when at least one of the visited faces in at least two faces is not a quadrilateral face, the context index is determined based on the priority of the pivot vertex in the mesh.
[0084] In one aspect, when the total number of visited faces in at least two faces is greater than four, the context index is determined based on the priority of the pivot vertices in the mesh.
[0085] In one aspect, when the total number of visited faces is equal to two and the visited faces share an edge, the context index is determined as the first value.
[0086] In one aspect, when the total number of visited faces is equal to two and the visited faces do not share edges, the context index is determined to be the second value.
[0087] In one aspect, when the total number of visited faces is not equal to two and is less than four, the context index is determined as the total number of visited faces.
[0088] In one aspect, when each of at least two visited faces is a quadrilateral face and the total number of visited faces in at least two faces is less than four, the size value is decoded based on the context indicated by the context index. The size of the face sector is determined as the sum of the size value and one.
[0089] In one aspect, when at least one visited face is not a quadrilateral face or the total number of visited faces in at least two faces is greater than or equal to four, the size of the face fan is decoded based on the context indicated by the context index, which is determined based on the priority of the pivot vertices in the mesh.
[0090] In one aspect, the predicted size of a face fan is determined based on its configuration. The face fan configuration indicates whether each of the visited faces out of at least two is a quadrilateral and whether the total number of visited faces out of at least two is greater than or equal to four. The predicted residual of the face fan size is determined based on the context indicated by the context index. The face fan size is then determined based on the predicted value and the predicted residual.
[0091] Then, the method proceeds to (S699) and terminates.
[0092] The method (600) can be modified appropriately. At least one step in the method (600) can be modified and / or omitted. At least one additional step can be added. The order of any suitable implementation can be used.
[0093] Figure 7 A flowchart outlining a method (700) according to one aspect of this disclosure is shown. The method (700) can be used in a mesh encoder. In various aspects, the method (700) is executed by processing circuitry, such as processing circuitry that performs the functions of the mesh encoder (103), processing circuitry that performs the functions of the mesh encoder (303), etc. In some aspects, the method (700) is implemented as software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes the method (700). The method begins at (S701) and proceeds to (S710).
[0094] At (S710), a face fan is identified in the mesh, wherein the face fan comprises at least two faces and a pivot vertex such that the pivot vertex is the center point of at least two faces and at least two faces are associated with the pivot vertex.
[0095] At (S720), the context index of the face sector is determined based on whether each of the visited faces in at least two faces is a quadrilateral face and whether the total number of visited faces in at least two faces is greater than four.
[0096] At (S730), the size of the face fan is encoded based on the context indicated by the context index.
[0097] In one aspect, when at least one of the visited faces in at least two faces is not a quadrilateral face, the context index is determined based on the priority of the pivot vertex in the mesh.
[0098] In one aspect, when the total number of visited faces in at least two faces is greater than four, the context index is determined based on the priority of the pivot vertices in the mesh.
[0099] In one aspect, when the total number of visited faces is equal to two and the visited faces share an edge, the context index is determined as the first value.
[0100] In one aspect, when the total number of visited faces is equal to two and the visited faces do not share edges, the context index is determined to be the second value.
[0101] In one aspect, when the total number of visited faces is not equal to two and is less than four, the context index is determined as the total number of visited faces.
[0102] In one aspect, when each of the at least two visited faces is a quadrilateral face and the total number of visited faces in the at least two faces is less than four, the size of the face sector is encoded by minus one based on the context indicated by the context index.
[0103] In one aspect, when at least one visited face is not a quadrilateral face or the total number of visited faces out of at least two faces is greater than or equal to four, the size of the face fan is encoded based on the context indicated by the context index. The context index is determined based on the priority of pivot vertices in the mesh.
[0104] In one aspect, the predicted size of a face fan is determined based on its configuration. The face fan configuration indicates whether each of the visited faces out of at least two is a quadrilateral and whether the total number of visited faces out of at least two is greater than or equal to four. The predicted residual is determined based on the predicted size of the face fan. The predicted residual of the face fan size is encoded based on the context indicated by the context index.
[0105] Then, the method proceeds to (S799) and terminates.
[0106] The method (700) can be modified appropriately. At least one step in the method (700) can be modified and / or omitted. At least one additional step can be added. The order of any suitable implementation can be used.
[0107] In one aspect, a method for processing grid data includes processing the bitstream of the grid data according to format rules. For example, the bitstream can be a bitstream decoded / encoded using any of the decoding and / or encoding methods described herein. The format rules can indicate at least one constraint on the bitstream and / or at least one process to be performed by the decoder and / or encoder.
[0108] In the example, the bitstream includes encoded information about a mesh. This mesh comprises face fans. Each face fan consists of at least two faces and a pivot vertex. The pivot vertex is the center point of at least two faces, and at least two faces are associated with this pivot vertex. Format rules instruct the determination of the face fan's context index based on whether each of the at least two visited faces is a quadrilateral face and whether the total number of visited faces in the at least two faces is greater than four. Format rules instruct the processing of the face fan's size based on the context indicated by the context index.
[0109] The above-described technology can be implemented as computer software using computer-readable instructions and physically stored on at least one computer-readable medium. For example, Figure 8 A computer system (800) suitable for implementing certain aspects of the disclosed subject matter is shown.
[0110] Computer software can be encoded using any suitable machine code or computer language. These machine codes or languages can be assembled, compiled, linked, and so on to create code containing instructions that can be executed directly by at least one computer central processing unit (CPU), graphics processing unit (GPU), or through interpretation, microcode execution, etc.
[0111] The instructions can be executed on various types of computers or their components, including personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0112] Figure 8 The components shown for the computer system (800) are examples and are not intended to impose any limitation on the scope or functionality of computer software implementing various aspects of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement on any one or combination of the components shown in the example aspects of the computer system (800).
[0113] The computer system (800) may include certain human-computer interface input devices. Such human-computer interface input devices may respond to input from at least one human user through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). The human-computer interface device may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image cameras), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0114] The input human-machine interface device may include at least one of the following (only one of each is shown): keyboard (801), mouse (802), touchpad (803), touch screen (810), data glove (not shown), joystick (805), microphone (806), scanner (807), camera (808).
[0115] The computer system (800) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of at least one human user through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback from a touchscreen (810), a data glove (not shown), or a joystick (805), but may also include tactile feedback devices that are not input devices), audio output devices (e.g., speakers (809), headphones (not shown)), visual output devices (e.g., screens (810), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without tactile feedback capability, some of which may output two-dimensional or higher-dimensional visual outputs through stereoscopic output, etc.; virtual reality glasses (not shown), holographic displays, and smoke canisters (not shown)) and printers (not shown).
[0116] The computer system (800) may also include human-accessible storage devices and their associated media, such as optical media, including CD / DVD ROM / RW (820)(821) with media such as CD / DVD, thumb drives (822), removable hard disk drives or solid-state drives (823), conventional magnetic media such as magnetic tapes and floppy disks (not shown), devices based on dedicated ROM / ASIC / PLD such as security dongles (not shown), etc.
[0117] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include transmission media, carrier waves, or other transient signals.
[0118] The computer system (800) may also include an interface (854) to at least one communication network (855). For example, the network may be wireless, wired, or fiber optic. The network may also be a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a vehicle-mounted and industrial network, a real-time network, a latency-tolerant network, etc. Examples of networks include LANs such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital television networks including cable television, satellite television, and terrestrial broadcast television, and vehicle and industrial networks including CANBus, etc. Some networks typically require external network interface adapters to connect to certain general-purpose data ports or peripheral buses (849) (e.g., a USB port on the computer system (800)); other interfaces are typically integrated into the core of the computer system (800) by connecting to a system bus, as described below (e.g., an Ethernet interface connected to a PC computer system or a cellular network interface connected to a smartphone computer system). Using any of these networks, the computer system (800) can communicate with other entities. This communication can be unidirectional, receive-only (e.g., broadcast television), send-only (e.g., to a CANbus device), or bidirectional, such as to other computer systems using a local area network or a wide area digital network. As mentioned above, certain protocols and protocol stacks can be used for each of these networks and network interfaces.
[0119] The aforementioned human-machine interface device, human-accessible storage device, and network interface can be connected to the core (840) of the computer system (800).
[0120] The core (840) may include at least one central processing unit (CPU) (841), a graphics processing unit (GPU) (842), a dedicated programmable processing unit (843) in the form of a field-programmable gate array (FPGA), a hardware accelerator (844) for certain tasks, a graphics adapter (850), etc. These devices, as well as read-only memory (ROM) (845), random access memory (846), and internal mass storage (such as internal non-user-accessible hard disk drives, SSDs, etc.) (847), may be connected via a system bus (848). In some computer systems, the system bus (848) may be accessed as at least one physical connector to allow for expansion via additional CPUs, GPUs, etc. Peripheral devices may be directly connected to the core's system bus (848) or connected via a peripheral bus (849). In one example, a screen (810) may be connected to the graphics adapter (850). Peripheral bus architectures include PCI, USB, etc.
[0121] The CPU (841), GPU (842), FPGA (843), and accelerator (844) can execute certain instructions, which, when combined, constitute the aforementioned computer code. This computer code can be stored in ROM (845) or RAM (846). Transient data can also be stored in RAM (846), while permanent data can be stored, for example, in internal mass storage (847). By using a cache, fast storage and retrieval of any storage device can be achieved, and the cache can be closely associated with at least one CPU (841), GPU (842), mass storage (847), ROM (845), RAM (846), etc.
[0122] Computer-readable media may contain computer code for performing operations of various computer implementations. The media and computer code may be those specifically designed and constructed for the purposes of this disclosure, or they may be those well known and available to those skilled in the art of computer software.
[0123] By way of example and not limitation, a computer system having an architecture (800), particularly a core (840), can provide functionality by having a processor (including a CPU, GPU, FPGA, accelerator, etc.) execute software contained in at least one tangible computer-readable medium. Such a computer-readable medium may be a medium associated with a user-accessible mass storage as described above, and some memory of the core (840) having a non-transient nature, such as internal mass storage (847) or ROM (845). Software implementing various aspects of this disclosure may be stored in such a device and executed by the core (840). Depending on specific needs, the computer-readable medium may include at least one memory device or chip. The software may enable the core (840), particularly the processor therein (including a CPU, GPU, FPGA, etc.), to execute a specific process or a specific portion of a specific process described herein, including defining data structures stored in RAM (846) and modifying these data structures according to a software-defined process. In addition, or alternatively, a computer system may provide functionality as a result of logic hard-wired in a circuit (e.g., an accelerator (844)) or otherwise embodied, which may replace or operate with software to perform a particular process or a particular portion of a particular process described herein. References to software may include logic, and vice versa, where appropriate. References to computer-readable media may include circuitry (such as an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both, where appropriate. This disclosure includes any suitable combination of hardware and software.
[0124] As used in this disclosure, "at least one" or "one" is intended to include any one or a combination of the elements stated. For example, referring to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include only A, only B, only C, or any combination thereof. Mentioning one of A or B, or one of A and B, is intended to include either A or B or (A and B). Where applicable, the use of "one of" does not exclude any combination of the listed elements, for example, when the elements are not mutually exclusive.
[0125] While examples of several aspects are described in this disclosure, modifications, arrangements, and various alternative equivalents exist that fall within the scope of this disclosure. Therefore, it should be understood that those skilled in the art will be able to design numerous systems and methods that, although not expressly shown or described herein, embody the principles of this disclosure and are thus within its spirit and scope.
[0126] The above disclosure also includes the following features. These features can be combined in various ways, and are not limited to the combinations mentioned below.
[0127] (1) A method for mesh decoding, comprising: receiving a bitstream including encoded information of a mesh, the mesh including face fans, each face fan including at least two faces and a pivot vertex, the pivot vertex being the center point of the at least two faces and the at least two faces being associated with the pivot vertex; determining a context index of the face fan based on whether each visited face among the at least two faces is a quadrilateral face and whether the total number of visited faces among the at least two faces is greater than four; and decoding the size of the face fan based on a context indicated by the context index.
[0128] (2) Determining the context index according to the method of feature (1) further includes: determining the context index based on the priority of the pivot vertex in the mesh when at least one of the visited faces of the at least two faces is not the quadrilateral face.
[0129] (3) Determining the context index according to the method of feature (1) or (2) further includes: determining the context index based on the priority of the pivot vertex in the mesh when the total number of the visited faces in the at least two faces is greater than four.
[0130] (4) Determining the context index according to the method of features (1) to (3) further includes: determining the context index as a first value when the total number of visited faces is equal to two and the visited faces share edges.
[0131] (5) Determining the context index according to the method of features (1) to (4) further includes: determining the context index as a second value when the total number of visited faces is equal to two and the visited faces do not share edges.
[0132] (6) Determining the context index according to the method of features (1) to (5) further includes: determining the context index as the total number of visited faces when the total number of visited faces is not equal to two and is less than four.
[0133] (7) Decoding the size of the face fan according to the method of features (1) to (6) further includes: decoding the size value based on the context indicated by the context index when each of the at least two faces is the quadrilateral face and the total number of the at least two faces is less than four, and determining the size of the face fan as the sum of the size value and one.
[0134] (8) Decoding the size of the face fan according to the method of features (1) to (7) further includes: decoding the size of the face fan based on the context indicated by the context index when at least one visited face is not the quadrilateral face or the total number of visited faces of the at least two faces is greater than or equal to four, the context index being determined based on the priority of the pivot vertex in the mesh.
[0135] (9) Decoding the size of the face fan according to the method of features (1) to (8) further includes: determining a predicted value of the size of the face fan based on the configuration of the face fan, the configuration of the face fan indicating whether each of the visited faces of the at least two faces is a quadrilateral face and whether the total number of the visited faces of the at least two faces is greater than or equal to four; determining a predicted residual of the size of the face fan based on the context indicated by the context index; and determining the size of the face fan based on the predicted value and the predicted residual.
[0136] (10) A method for mesh encoding, comprising: identifying a face fan in a mesh, the face fan comprising at least two faces and a pivot vertex such that the pivot vertex is the center point of the at least two faces and the at least two faces are associated with the pivot vertex; determining a context index of the face fan based on whether each of the at least two faces being a quadrilateral face and whether the total number of the at least two faces being visited is greater than four; and encoding the size of the face fan based on a context indicated by the context index.
[0137] (11) Determining the context index according to the method of feature (10) further includes: determining the context index based on the priority of the pivot vertex in the mesh when at least one of the visited faces of the at least two faces is not the quadrilateral face.
[0138] (12) Determining the context index according to the method of feature (10) or (11) further includes: determining the context index based on the priority of the pivot vertex in the mesh when the total number of the visited faces in the at least two faces is greater than four.
[0139] (13) Determining the context index according to the method of features (10) to (12) further includes: determining the context index as a first value when the total number of the visited faces is equal to two and the visited faces share an edge.
[0140] (14) Determining the context index according to the method of features (10) to (13) further includes: determining the context index as a second value when the total number of visited faces is equal to two and the visited faces do not share edges.
[0141] (15) Determining the context index according to the method of features (10) to (14) further includes: determining the context index as the total number of visited faces when the total number of visited faces is not equal to two and is less than four.
[0142] (16) Encoding the size of the face fan according to the method of features (10) to (15) further includes: encoding the size of the face fan minus one based on the context indicated by the context index when each of the at least two faces is the quadrilateral face and the total number of the at least two faces is less than four.
[0143] (17) Encoding the size of the face fan according to the method of features (10) to (16) further includes: encoding the size of the face fan based on the context indicated by the context index when at least one visited face is not the quadrilateral face or the total number of visited faces of the at least two faces is greater than or equal to four, the context index being determined based on the priority of the pivot vertex in the mesh.
[0144] (18) Encoding the size of the face fan according to the method of features (10) to (17) further includes: determining a predicted value of the size of the face fan based on the configuration of the face fan, the configuration of the face fan indicating whether each of the visited faces of the at least two faces is a quadrilateral face and whether the total number of the visited faces of the at least two faces is greater than or equal to four; determining a prediction residual based on the predicted value of the size of the face fan; and encoding the prediction residual of the size of the face fan based on the context indicated by the context index.
[0145] (19) A method for processing mesh data, comprising: processing a bitstream of the mesh data according to a format rule, wherein: the bitstream includes encoded information of a mesh, the mesh includes face fans, each face fan including at least two faces and a pivot vertex, the pivot vertex being the center point of the at least two faces, and the at least two faces being associated with the pivot vertex; and the format rule instructs: determining a context index of the face fan based on whether each of the at least two faces is a quadrilateral face and whether the total number of the at least two faces is greater than four; and processing the size of the face fan based on the context indicated by the context index.
[0146] (20) According to the method of feature (19), the format rule indicates that when at least one of the visited faces of the at least two faces is not the quadrilateral face, the context index is determined based on the priority of the pivot vertex in the mesh.
[0147] (21) An apparatus for mesh decoding, comprising processing circuitry configured to perform the method of any one of features (1) to (9) and (19) to (20).
[0148] (22) A grid-encoded apparatus, comprising processing circuitry configured to perform a method of any one of features (10) to (18).
[0149] (23) A non-volatile computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform the method of any one of features (1) to (20).
Claims
1. A method of trellis decoding, characterized by, comprising: receiving a bitstream including encoded information of a mesh, the mesh including a face fan, the face fan including at least two faces and a pivot vertex, the pivot vertex being a center point of the at least two faces, and the at least two faces being associated to the pivot vertex; determining a context index of the face fan based on whether each visited face of the at least two faces is a quadrilateral face and whether a total number of the visited faces of the at least two faces is greater than four; and decoding a size of the face fan based on a context indicated by the context index.
2. The method of claim 1, wherein, Determining the context index further comprises: when at least one of the visited faces of the at least two faces is not the quadrilateral face, determining the context index based on a priority of the pivot vertex in the mesh.
3. The method of claim 1, wherein, Determining the context index further comprises: when the total number of the visited faces of the at least two faces is greater than four, determining the context index based on the priority of the pivot vertex in the mesh.
4. The method of claim 1, wherein, Determining the context index further comprises: when the total number of the visited faces is equal to two and the visited faces share an edge, determining the context index as a first value.
5. The method of claim 1, wherein, Determining the context index further comprises: when the total number of the visited faces is equal to two and the visited faces do not share an edge, determining the context index as a second value.
6. The method of claim 1, wherein, Determining the context index further comprises: when the total number of the visited faces is not equal to two and less than four, determining the context index as the total number of the visited faces.
7. The method according to any one of claims 1 to 6, characterized in that, Decoding the size of the face fan further comprises: when each visited face of the at least two faces is the quadrilateral face and the total number of the visited faces of the at least two faces is less than four, decoding a size value based on the context indicated by the context index, and determining the size of the face fan as a sum of the size value and one.
8. The method according to any one of claims 1 to 6, characterized in that, Decoding the size of the face fan further comprises: when at least one visited face is not the quadrilateral face or the total number of the visited faces of the at least two faces is greater than or equal to four, decoding the size of the face fan based on the context indicated by the context index, the context index being determined based on a priority of the pivot vertex in the mesh.
9. The method according to any one of claims 1 to 6, characterized in that, Decoding the size of the face fan further comprises: determining a prediction value of the size of the face fan based on a configuration of the face fan, the configuration of the face fan indicating whether each of the visited faces of the at least two faces is the quadrilateral face and whether the total number of the visited faces of the at least two faces is greater than or equal to four; determining a prediction residual of the size of the face fan based on the context indicated by the context index; and determining the size of the face fan based on the prediction value and the prediction residual.
10. A method of trellis encoding, characterized by, comprising: identifying a face fan in a mesh, the face fan including at least two faces and a pivot vertex, such that the pivot vertex is a center point of the at least two faces, and the at least two faces are associated to the pivot vertex; determining a context index of the face fan based on whether each of the visited faces of the at least two faces is a quadrilateral face and whether a total number of the visited faces of the at least two faces is greater than four; and encoding a size of the face fan based on a context indicated by the context index.
11. The method of claim 10, wherein, Determining the context index further includes: when at least one of the visited faces of the at least two faces is not the quadrilateral face, determining the context index based on a priority of the pivot vertex in the mesh.
12. The method of claim 10, wherein, Determining the context index further includes: when the total number of the visited faces of the at least two faces is greater than four, determining the context index based on the priority of the pivot vertex in the mesh.
13. The method of claim 10, wherein, Determining the context index further includes: when the total number of the visited faces is equal to two and the visited faces share an edge, determining the context index as a first value.
14. The method of claim 10, wherein, Determining the context index further includes: when the total number of the visited faces is equal to two and the visited faces do not share an edge, determining the context index as a second value.
15. The method of claim 10, wherein, Determining the context index further includes: when the total number of the visited faces is not equal to two and less than four, determining the context index as the total number of the visited faces.
16. The method according to any one of claims 10-15, characterized in that, Encoding the size of the face fan further includes: when each visited face of the at least two faces is the quadrilateral face and the total number of the visited faces of the at least two faces is less than four, encoding the size of the face fan minus one based on the context indicated by the context index.
17. The method according to any one of claims 10-15, characterized in that, Encoding the size of the face fan further includes: when at least one visited face is not the quadrilateral face or the total number of the visited faces of the at least two faces is greater than or equal to four, encoding the size of the face fan based on the context indicated by the context index, the context index being determined based on a priority of the pivot vertex in the mesh.
18. The method according to any one of claims 10-15, characterized in that, Encoding the size of the face fan further includes: determining a predicted value of the size of the face fan based on a configuration of the face fan, the configuration of the face fan indicating whether each of the visited faces of the at least two faces is the quadrilateral face and whether the total number of the visited faces of the at least two faces is greater than or equal to four; determining a prediction residual based on the predicted value of the size of the face fan; and encoding the prediction residual of the size of the face fan based on the context indicated by the context index.
19. A method of processing mesh data, characterized by, The method includes: processing a bitstream of the mesh data according to a format rule, wherein: the bitstream includes coded information of a mesh, the mesh including a face fan, the face fan including at least two faces and a pivot vertex, the pivot vertex being a center point of the at least two faces, and the at least two faces being associated to the pivot vertex; and the format rule indicates: determine a context index of the face fan based on whether each of the at least two faces that is visited is a quadrilateral face and whether a total number of the at least two faces that is visited is greater than four; and process a size of the face fan based on a context indicated by the context index.
20. The method of claim 19, wherein, The format rule indicates: when at least one of the visited faces of the at least two faces is not the quadrilateral face, determine the context index based on a priority of the pivot vertex in the mesh.
21. A method of storing a bitstream, characterized by, The method of encoding a mesh according to any one of claims 10-18 generates a bitstream, and stores the bitstream.
22. A method of transmitting a bitstream, the method comprising: The method of encoding a mesh according to any one of claims 10-18 generates a bitstream, and transmits the bitstream.
23. A non-transitory computer readable medium storing instructions, the instructions comprising: The instructions, when executed by a computer, cause the computer to perform the method according to any one of claims 1-22.
24. An electronic device, comprising: comprise: a processor; a memory; and at least one set of instructions stored in the memory and configured to be executed by the processor, when executed, implement the method according to any one of claims 1-22.