Grid processing method and device and computer readable storage medium
By receiving the code stream of encoded information of the mesh, it is determined whether the connectivity of the non-position attributes corresponds to at least the extreme case, and in extreme cases, the non-position attribute connection is determined based on the position connection of multiple 3D vertices, which solves the problem of inefficient encoding of non-position attribute connections of polygon mesh in the prior art, and realizes efficient encoding and decoding.
Patent Information
- Application Number
- CN202411513336.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-09-11
- Filing Date
- 2024-10-28
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to efficiently encode and decode the connectivity of non-positional properties in polygon mesh in three-dimensional space, especially in extreme cases, resulting in inefficient data transmission and storage.
By receiving the code stream of encoded information of the grid, it is determined whether the connectivity of the non-position attributes corresponds to at least the extreme case, and in the extreme case, the non-position attribute connection is determined based on the position connection of the multiple 3D vertices, thereby achieving efficient encoding and decoding.
It realizes efficient encoding and decoding of the non-position attribute connectivity of polygon mesh, reducing resource consumption for data transmission and storage, and improving coding efficiency.
Smart Images

Figure CN119991832A_ABST
Abstract
Description
Incorporation by reference
[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 547,965, filed on November 9, 2023, entitled “Zero Byte Coding of Attribute Connectivity in Polygon Meshes,” which is incorporated herein by reference in its entirety. Technical Field
[0002] The present disclosure describes technologies generally related to grid coding, and more particularly, to a grid processing method, apparatus, and computer-readable storage medium. Background Art
[0003] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent that the work of the presently named inventors described in this background section and in various aspects of this specification was performed, it does not indicate that it qualifies as prior art at the time of filing, and it is never explicitly or implicitly admitted that it is prior art to the present disclosure.
[0004] Various technologies have been developed to capture and represent the world in three-dimensional (3D) space, such as objects in the world, environments in the world, etc. 3D representations of the world enable more immersive forms of interaction and communication. For example, technological developments in 3D media processing (e.g., advances in three-dimensional (3D) capture, 3D modeling, and 3D rendering) have promoted the ubiquity of 3D media content on several platforms and devices. In one example, a baby's first steps can be filmed on one continent, and media technology can enable grandparents to watch (perhaps interact) and enjoy an immersive experience with the baby on another continent. According to one aspect of the present disclosure, in order to improve the immersive experience, 3D models have become increasingly complex, and the creation and consumption of 3D models occupy a large amount of data resources, such as data storage and data transmission resources. In some examples, a 3D mesh can be used as a 3D representation of the world. Summary of the invention
[0005] Aspects of the present disclosure include code streams, methods, and apparatus for trellis encoding / decoding.In some examples, an apparatus for trellis encoding / decoding includes processing circuitry.
[0006] Some aspects of the present disclosure provide a method for mesh processing. The method includes receiving a code stream of encoded information of a mesh, the mesh including multiple 3D vertices and at least non-positional attributes in a three-dimensional (3D) space. The encoded information includes positional connectivity of multiple 3D vertices in the 3D space. The method also includes: determining whether the encoded information of the mesh indicates that the non-positional attribute connectivity of the non-positional attribute corresponds to at least an extreme case; and when the non-positional attribute connectivity corresponds to an extreme case, determining the non-positional attribute connectivity based on the positional connectivity of the multiple 3D vertices.
[0007] Some aspects of the present disclosure also provide a method for mesh processing. The method includes: for a mesh including a plurality of 3D vertices in a three-dimensional (3D) space and at least a non-position attribute, determining whether non-position attribute connectivity of the non-position attribute corresponds to at least an extreme case; and encoding a signal into a bitstream of encoded information of the mesh, the signal indicating whether the non-position attribute connectivity of the non-position attribute corresponds to an extreme case.
[0008] Some aspects of the present disclosure provide a method for processing mesh data, the method comprising processing a code stream of the mesh data according to a format rule. The code stream comprises encoded information of the mesh. The mesh comprises a plurality of 3D vertices and at least a non-positional attribute in a three-dimensional (3D) space, and the encoded information comprises positional connectivity of the plurality of 3D vertices in the 3D space. The format rule specifies: determining whether the encoded information of the mesh indicates that the non-positional attribute connectivity of the non-positional attribute corresponds to at least an extreme case; and when the non-positional attribute connectivity corresponds to an extreme case, determining the non-positional attribute connectivity according to the positional connectivity of the plurality of 3D vertices.
[0009] Aspects of the present disclosure also provide an apparatus for mesh processing. The apparatus for mesh processing comprises a processing circuit configured to implement any of the methods for mesh processing described.
[0010] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions which, when executed by a computer, cause the computer to perform any of the described methods for grid processing.
[0011] Aspects of the present disclosure provide more efficient techniques for encoding connectivity of non-positional attributes such as UV coordinates, normals, etc., with high coding efficiency for polygonal mesh compression. For some special cases (also referred to as extreme cases), the connectivity of non-positional attributes of a mesh (also referred to as non-positional attribute connectivity) can be derived from the connectivity of vertices (3D vertices) of the mesh in 3D space (also referred to as positional connectivity). BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Further features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0013] Figure 1 A block diagram of a streaming system in some examples is shown.
[0014] Figure 2 is a schematic illustration of an example of a block diagram of a video decoder.
[0015] Figure 3 is a schematic illustration of an example of a block diagram of a video encoder.
[0016] Figure 4 An example of an encoding process for mesh processing according to an aspect of the present disclosure is shown.
[0017] Figure 5 An example of a decoding process for trellis processing according to another aspect of the present disclosure is shown.
[0018] Figure 6 A diagram illustrating the mapping of a 3D mesh to a 2D atlas in some examples is shown.
[0019] Figure 7 Examples of maps of grids are shown in further examples.
[0020] Figure 8 A flow chart outlining a decoding process according to some aspects of the present disclosure is shown.
[0021] Fig. 9 A flowchart outlining an encoding process according to further aspects of the present disclosure is shown.
[0022] Fig.10 is a schematic illustration of a computer system according to an aspect. DETAILED DESCRIPTION
[0023] Aspects of the present disclosure provide techniques in the field of grid processing.
[0024] A mesh (also referred to as a mesh model) includes a number of polygons (also referred to as faces) that describe the surface of a volumetric object. Each polygon may be defined by vertices in a three-dimensional (3D) space and information about how the vertices are connected (referred to as connectivity information). In some examples, the mesh also includes vertex attributes associated with the mesh vertices, such as color, normal, displacement, and the like. In addition, in some examples, the mesh can include attributes associated with the surface of the mesh using mapping information that uses a two-dimensional (2D) attribute map to parameterize the mesh. This mapping is typically described by a set of parameter coordinates (referred to as UV coordinates or texture coordinates) associated with the mesh vertices. 2D attribute maps are used to store high-resolution attribute information, such as textures, normals, displacements, and the like. 2D attribute maps can be used for various purposes, such as texture mapping, shading, mesh reconstruction, and the like.
[0025] Figure 1 A block diagram of a streaming system (100) in some examples is shown. The streaming system (100) is an example of an application of the disclosed subject matter, namely a mesh encoder and a mesh decoder located in a streaming environment. The disclosed subject matter is equally applicable to other mesh-enabled applications including, for example, video conferencing, 3D TV, streaming services, storage of compressed 3D data on digital media including CDs, DVDs, memory sticks, etc., and the like.
[0026] The streaming system (100) includes an acquisition subsystem (113), which may include a 3D source (101), such as a light detection and ranging (LIDAR) system, a 3D camera, a 3D scanner, a graphics generation component, etc., that creates an uncompressed 3D data stream (102). In one example, the 3D data stream (102) includes samples recorded by a 3D camera system. The 3D data stream (102), which is depicted as a thick line to emphasize the high amount of data compared to the encoded 3D data (104) (or encoded bitstream), can be processed by an electronic device (120), which includes a 3D encoder (103) coupled to the 3D source (101). The 3D encoder (103) may include hardware, software, or a combination of hardware and software to implement or implement aspects of the disclosed subject matter as described in more detail below. The encoded 3D data (104) (or the encoded bitstream), depicted as thin lines to emphasize the lower data volume compared to the 3D data stream (102), may be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1The client subsystem (106) and the client subsystem (108) in the streaming server (105) can access the streaming server (105) to retrieve the copy (107) and the copy (109) of the encoded 3D data (104). The client subsystem (106) can include, for example, a 3D decoder (110) in the electronic device (130). The 3D decoder (110) decodes the incoming copy (107) of the encoded 3D data and creates an output stream (111) of a 3D representation that can be presented on a display (112) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded 3D data (104), (107) and (109) (e.g., video streams) can be encoded according to certain 3D encoding / compression standards (e.g., mesh encoding / compression standards, etc.).
[0027] It should be noted that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a 3D decoder (not shown), and the electronic device (130) may further include a 3D encoder (not shown).
[0028] It should also be noted that in some examples, the 3D encoder and / or 3D decoder may use 2D encoding / decoder technology.For example, the 3D encoder and / or 3D decoder may include a video decoder or a video encoder.
[0029] Figure 2 An example of a block diagram of a video decoder (210) is shown. The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used to Figure 1 A 3D decoder (110) in the example of FIG.
[0030] The receiver (231) may receive one or more encoded video sequences, for example, included in a bitstream to be decoded by the video decoder (210). In one aspect, the encoded video sequences are received one at a time, wherein the decoding of each encoded video sequence is independent of the decoding of the other encoded video sequences. The encoded video sequence may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not shown). The receiver (231) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory (215) may be provided outside the video decoder (210) (not shown). In still other applications, a buffer memory (not shown) may be provided outside the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) may be provided inside the video decoder (210) to, for example, handle playback timing. When the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may not be necessary to configure the buffer memory (215), or the buffer memory (215) may be made smaller. For use on a traffic packet network such as the Internet, the buffer memory (215) may be required. The buffer memory (215) may be relatively large, advantageously may have an adaptive size, and may be at least partially implemented in an operating system or similar element (not shown) external to the video decoder (210).
[0031] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. The categories of symbols include information for managing the operation of the video decoder (210) and potential information for controlling a presentation device such as a presentation device (212) (e.g., a display screen) that is not part of the electronic device (230) but can be coupled to the electronic device (230), such as Figure 2As shown. The control information for the presentation device may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not indicated). The parser (220) may parse / entropy decode the received coded video sequence. The encoding of the coded video sequence may be performed according to a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract a subgroup parameter set for at least one subgroup of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (220) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0032] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).
[0033] Depending on the type of coded video picture or part of coded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of symbol (221) may involve multiple different units. Which units are involved and how they are involved may be controlled by subgroup control information parsed by parser (220) from the coded video sequence. For the sake of brevity, such subgroup control information flow between parser (220) and the multiple units below is not described.
[0034] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into a number of functional units as described below. In actual implementations operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, the conceptual subdivision into a number of functional units is appropriate in the following.
[0035] The first unit is a sealer / inverse transform unit (251). The sealer / inverse transform unit (251) receives quantized transform coefficients as symbols (221) from the parser (220) and control information including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit (251) may output a block including sample values, which may be input into an aggregator (255).
[0036] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use predictive information from a previously reconstructed picture, but can use predictive information from a previously reconstructed portion of a current picture. Such predictive information may be provided by an intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses surrounding reconstructed information extracted from a current picture buffer (258) to generate a block of the same size and shape as the block being reconstructed. For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a per-sample basis.
[0037] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to a block that is inter-coded and potentially motion compensated. In this case, the motion compensated prediction unit (253) may access the reference picture memory (257) to extract samples for prediction. After the extracted samples are motion compensated according to the symbols (221) belonging to the block, these samples may be added to the output of the scaler / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) by the aggregator (255) to generate output sample information. The extraction of prediction samples by the motion compensated prediction unit (253) from the address in the reference picture memory (257) may be controlled by a motion vector, which may be provided to the motion compensated prediction unit (253) in the form of symbols (221), which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.
[0038] The output samples of the aggregator (255) may be used by various loop filtering techniques in a loop filter unit (256). The video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video sequence (also referred to as the coded video bitstream) and available to the loop filter unit (256) as symbols (221) from the parser (220). The video compression may also be responsive to meta-information obtained during decoding of a coded picture or a previous (in decoding order) portion of the coded video sequence, and to previously reconstructed and loop filtered sample values.
[0039] The output of the loop filter unit (256) may be a sample stream that may be output to a rendering device (212) and stored in a reference picture memory (257) for subsequent inter-picture prediction.
[0040] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed, and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) may become part of the reference picture memory (257), and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.
[0041] The video decoder (210) may perform decoding operations according to a predetermined video compression technology or standard, such as ITU-T H.265 Recommendation. The encoded video sequence may conform to the syntax specified by the video compression technology or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as the only tools available for use under the profile. For compliance, the complexity of the encoded video sequence may also be required to be within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata of the HRD buffer management signaled in the encoded video sequence.
[0042] In one aspect, a receiver (231) can receive additional (redundant) data when receiving an encoded video. The additional data can be included as part of the encoded video sequence. The additional data can be used by a video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data can take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0043] Figure 3 An example of a block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) may be used Figure 1 3D encoder (103) in the example.
[0044] The video encoder (303) can be used to obtain the video source (301) (not Figure 3 In an example, the video source (301) is a part of the electronic device (320) to receive the video samples, and the video source (301) can obtain the video image to be encoded by the video encoder (303). In another example, the video source (301) is a part of the electronic device (320).
[0045] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The pictures themselves may be constructed as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. The following description focuses on the samples.
[0046] According to one aspect, the video encoder (303) can encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other time constraints required. Implementing the appropriate encoding speed is a function of the controller (350). In some aspects, the controller (350) controls other functional units as described below and is functionally coupled to the other functional units described. For the sake of brevity, the coupling is not depicted in the figure. The parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization techniques, etc.), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions related to the video encoder (303) optimized for a certain system design.
[0047] In some aspects, the video encoder (303) is configured to operate in an encoding loop. As a simplified description, in one example, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Since the decoding of the symbol stream produces a bit-accurate result that is independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-accurately corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, such as due to channel errors) is also used in some related technologies.
[0048] The operation of the "local" decoder (333) may be similar to that described above in conjunction with Figure 2 The "remote" decoder of the video decoder (210) described in detail is identical. However, additional brief reference is made to Figure 2 , since the symbols are available and the entropy encoder (345) and the parser (220) can losslessly encode / decode the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (210) including the buffer memory (215) and the parser (220) may not be fully implemented in the local decoder (333).
[0049] On the one hand, decoder techniques other than parsing / entropy decoding present in a decoder are present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on decoder operation. The description of encoder techniques can be simplified because encoder techniques are mutually inverse to the decoder techniques described comprehensively. In some areas, a more detailed description is provided below.
[0050] During operation, in some examples, the source encoder (330) may perform motion compensated predictive coding that predictively encodes an input picture by referencing one or more previously encoded pictures from a video sequence designated as "reference pictures." In this manner, the encoding engine (332) encodes the differences between pixel blocks of the input picture and pixel blocks of a reference picture that may be selected as a prediction reference for the input picture.
[0051] The local video decoder (333) may decode the encoded video data of the picture that may be designated as the reference picture based on the symbol created by the source encoder (330). The operation of the encoding engine (332) may advantageously be a lossy process. When the encoded video data may be decoded at the video decoder ( Figure 3 When the video sequence is decoded at a remote video decoder (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture memory (334). In this way, the video encoder (303) may store a copy of the reconstructed reference picture locally that has common content (absent transmission errors) with the reconstructed reference picture to be obtained by the remote video decoder.
[0052] The predictor (335) may perform a prediction search for the encoding engine (332). That is, for a new picture to be encoded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may be used as appropriate prediction references for the new picture. The predictor (335) may operate on a pixel block by pixel block basis to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (335), the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (334).
[0053] The controller (350) may manage encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.
[0054] The outputs of all the above functional units may be entropy encoded in an entropy encoder (345). The entropy encoder (345) applies lossless compression to the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into a coded video sequence.
[0055] The transmitter (340) may buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission over a communication channel (360), which may be a hardware / software link to a storage device that may store the encoded video data. The transmitter (340) may combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or ancillary data streams (source not shown).
[0056] The controller (350) may manage the operation of the video encoder (303). During encoding, the controller (350) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following picture types:
[0057] Intra pictures (I pictures) that can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures.
[0058] Predictive pictures (P pictures), which can be encoded and decoded using intra prediction or inter prediction, which uses motion vectors and reference indices to predict sample values for each block.
[0059] Bidirectional predictive pictures (B pictures), which can be encoded and decoded using intra prediction or inter prediction, which uses two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.
[0060] The source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined by the coding allocation applied to the corresponding picture of the block. For example, blocks of an I picture may be non-predictively coded, or blocks of an I picture may be predictively coded (spatial prediction or intra prediction) with reference to already coded blocks of the same picture. Blocks of pixels of a P picture may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. Blocks of a B picture may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.
[0061] The video encoder (303) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.
[0062] In one aspect, the transmitter (340) may transmit additional data when transmitting the encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0063] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In one example, a particular picture being encoded / decoded is partitioned into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.
[0064] In some aspects, a bidirectional prediction technique may be used for inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be in the past and future, respectively, in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.
[0065] In addition, merge mode technology can be used for inter-picture prediction to improve coding efficiency.
[0066] According to some aspects of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the High-Efficiency Video Coding (HEVC) standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU includes three coding tree blocks (CTBs), which are a luminance CTB and two chrominance CTBs. Each CTU can be recursively divided into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be divided into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type for the CU, such as an inter-prediction type or an intra-prediction type. The CU is divided into one or more prediction units (PUs) based on temporal and / or spatial predictability. Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In one aspect, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Using the luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luma values) for pixels, such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0067] It should be noted that the encoder (103) and encoder (303) and decoder (110) and decoder (210) may be implemented using any suitable technology. In one aspect, the encoder (103) and encoder (303) and decoder (110) and decoder (210) may be implemented using one or more integrated circuits. In another aspect, the encoder (103) and encoder (303) and decoder (110) and decoder (210) may be implemented using one or more processors executing software instructions.
[0068] In some examples, the 3D data includes a mesh model, the 3D encoder (103) may include a mesh encoder, and the 3D decoder (110) may include a mesh decoder.
[0069] According to one aspect of the present disclosure, a dynamic mesh is a mesh in which at least one of the components (geometric information, connectivity information, mapping information, vertex attributes, and attribute mapping) changes over time. A dynamic mesh may be described by a sequence of meshes (also referred to as mesh frames). In some examples, a mesh frame in a dynamic mesh may be a representation of the surface of an object at different times, and each mesh frame is a representation of the surface of an object at a specific time (also referred to as a time instance). A dynamic mesh may require a large amount of data because the dynamic mesh may include a large amount of information that changes over time. Compression techniques for meshes enable efficient storage and transmission of media content in a mesh representation.
[0070] Dynamic mesh sequences can require a lot of data because dynamic meshes can include a lot of information that changes over time. Therefore, efficient compression techniques can be used to store and transmit such content.
[0071] Figure 4 An example of an encoding process (400) for mesh processing according to an aspect of the present disclosure is shown. Figure 4As shown, the encoding process (400) includes a preprocessing step (410) and an encoding step (420). The preprocessing step (410) is configured to generate a base grid m(i) of the current frame and a displacement field d(i) of the current frame based on an input grid M(i) of the current frame, and the displacement field d(i) of the current frame includes a displacement vector. The encoding step (420) is configured to encode the base grid m(i), the displacement field d(i), and the texture information of the base grid m(i). The displacement field d(i) of the current frame includes a displacement vector. The index i is used to refer to the current frame. In one aspect, a mode decision method can be performed in the encoding process (400) to determine whether to apply inter-frame coding (also referred to as inter-frame prediction or inter-frame mode), intra-frame coding (also referred to as intra-frame prediction or intra-frame mode), etc. to the current frame. For example, the mode decision method can compare the cost of the intra-frame mode with the cost of the inter-frame mode, and determine the encoding mode of the base grid m(i) of the current frame based on which of the two has a smaller cost. In some examples, the base grid m(i) is encoded using a SKIP mode. In one example, the SKIP mode is a special mode of the INTER mode. For example, the base grid m(i) may be intra-coded, inter-coded, or encoded using a SKIP mode.
[0072] Still reference Figure 4 , the preprocessing step (410) may include a mesh extraction process (412), a parameterization process such as an atlas parameterization process (414), and a subdivision surface fitting process (416). The mesh extraction process (412) is configured to downsample the vertices of the input mesh M(i) to generate an extracted mesh dm(i) that may include a plurality of extracted (or downsampled) vertices. In an example, the number of the plurality of extracted vertices is less than the number of vertices of the input mesh M(i). The parameterization process such as the atlas parameterization process (414) is configured to map the extracted mesh dm(i) onto a planar domain, such as onto a UV atlas (or UV map), to generate a re-parameterized mesh pm(i). In an example, the atlas parameterization may be performed based on a video processing tool (e.g., a UV Atlas tool). The subdivision surface fitting process (416) is configured to take as input the re-parameterized mesh pm(i) and the input mesh M(i), and generate a base mesh m(i) and a displacement field d(i) including displacement vectors or displacement sets. In an example of the subdivision surface fitting process (416), pm(i) is subdivided using a subdivision scheme such as iterative interpolation to obtain a subdivided mesh. Iterative interpolation includes inserting a new point in the middle of each edge of the re-parameterized mesh pm(i) at each iteration. Any suitable subdivision scheme can be applied to subdivide pm(i). The displacement field d(i) is calculated by determining the closest point on the surface of the input mesh M(i) for each vertex of the subdivided mesh.
[0073] Advantages of the subdivided mesh may include: the subdivided mesh has a subdivision structure that can be efficiently compressed while providing a reliable approximation of the input mesh. The improved compression efficiency can be obtained due to the following properties. The decimated mesh dm(i) may have a small number of vertices and may be encoded and transmitted using a smaller number of bits than the input mesh M(i) or the subdivided mesh. Figure 4 , a base mesh m(i) can be generated based on the extracted mesh dm(i). In one example, the base mesh m(i) is the extracted mesh dm(i). Since the subdivided mesh can be generated based on the subdivision method, the subdivided mesh can be automatically generated by the decoder when the base mesh or the extracted mesh is decoded (for example, without using any information other than the subdivision scheme and the subdivision iteration count). On the decoder side, the displacement field d(i) can be generated by decoding the displacement vectors associated with the vertices of the subdivided mesh. In addition to enabling spatial / quality scalability, the subdivision structure enables efficient transformations such as wavelet decomposition, which can provide high compression performance.
[0074] exist Figure 4 In the example of , the encoding step (420) includes base mesh encoding (422), displacement encoding (424), texture encoding (426), etc. The base mesh encoding (422) is configured to encode geometric information of a base mesh m(i) associated with a current frame. In intra-frame coding, the base mesh m(i) may first be quantized (e.g., quantized using uniform quantization) and then encoded, for example, by using a coding mode determined by a mode decision method. The coding mode may be an inter-frame mode, an intra-frame mode, a skip mode, etc. An encoder for intra-coding a base mesh m(i) may be referred to as a static mesh encoder. In inter-frame coding, a reference base mesh associated with a reference frame indicated by index j (e.g., a reconstructed quantized reference base mesh m'(j)) may be used to predict a base mesh m(i) associated with a current frame indicated by index i. The displacement encoding (424) is configured to encode a displacement field d(i) generated in the preprocessing step (410). The displacement field d(i) may include a set of displacement vectors (or displacements) associated with subdivided mesh vertices. The texture encoding (426) is configured to encode attribute information of the base mesh m(i). The attribute information may include texture, normal, and / or color, etc. The attribute information may be encoded based on a suitable codec (e.g., High Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC)).
[0075] On the one hand, reference Figure 4, a mesh encoding process such as encoding process (420) begins with preprocessing (e.g., preprocessing step (410)). Preprocessing can convert an input mesh (e.g., an input dynamic mesh) M(i) into a base mesh m(i) and a displacement field d(i) including a displacement set (or a displacement vector set). The encoding step (420) can compress the output from the preprocessing (e.g., m(i), d(i), etc.) and generate a compressed code stream b(i). The compressed code stream b(i) can include a compressed base mesh code stream, a compressed displacement field code stream, and / or a compressed attribute code stream, etc.
[0076] Figure 5 An example of a decoding process (500) for grid processing according to an aspect of the present disclosure is shown. The decoding process (500) may include a decoding step (510) and a post-processing step (520). A compressed code stream b(i) may be provided to the decoding step (510). In one example, for example for lossless transmission, the compressed code stream b(i) is the output b(i) from the encoding process (400). The decoding step (510) may extract various sub-code streams, such as a compressed base grid sub-stream, a compressed displacement field sub-stream, and / or a compressed attribute sub-stream, etc. The decoding step (510) may decompress the sub-code streams to generate the following components: patch metadata indicated by metadata (i), a decoded base grid m" (i), a decoded displacement field (including displacement) d" (i), and / or a decoded attribute map A" (i), etc.
[0077] On the one hand, the base grid substream can be provided to a grid decoder to generate a reconstructed quantized base grid m'(i). The decoded base grid (or reconstructed base grid) m"(i) can be obtained by applying inverse quantization to m'(i). The displacement field substream including the encoded packed and quantized wavelet coefficients can be decoded by a video and / or image decoder. Image unpacking and inverse quantization can be applied to the reconstructed packed and quantized wavelet coefficients to obtain unpacked and dequantized transform coefficients (e.g., wavelet coefficients). An inverse wavelet transform can be applied to the unpacked and dequantized wavelet coefficients to generate a decoded displacement field (or reconstructed displacement) d"(i).
[0078] The decoded components (e.g., including metadata (i), m" (i), d" (i), A" (i), etc.) can be provided to a post-processing step (520). A mesh (also referred to as a decoded / reconstructed mesh) M" (i) can be generated based on m" (i) and d" (i) by the post-processing step (520). In one example, the mesh M" (i) (also referred to as a reconstructed deformed mesh DM(i)) can be obtained by subdividing m" (i) using a subdivision scheme and applying the reconstructed displacements d" (i) to the vertices of the subdivided mesh. In one example, DM(i) may include a displacement curve. In one example, when the encoding process (400), the decoding process (500), and the transmission are lossless, the mesh M" (i) may be the same as the input mesh M(i). If one of the encoding process (400), the decoding process (500), and the transmission is lossy, then M" (i) is different from M(i). In various examples, the difference between M" (i) and M(i), if any, may be relatively small. In one example, an attribute map A”(i) is also generated by a post-processing step (520).
[0079] In some examples, the mesh may also include attributes associated with the vertices, such as color, normal, etc. The attributes may be associated with the surface of the mesh by utilizing mapping information, which parameterizes the mesh using a 2D attribute map. The mapping information is typically described by a set of parameter coordinates (referred to as UV coordinates or texture coordinates) associated with the mesh vertices. The 2D attribute map (referred to as a texture map in some examples) is used to store high-resolution attribute information, such as texture, normal, displacement, etc. Such information may be used for various purposes, such as texture mapping and shading.
[0080] In some embodiments, the mesh may include other information referred to as geometric information, connectivity information, mapping information, vertex attributes, and attribute maps. In some examples, the geometric information is described by a 3D position set associated with the vertices of the mesh. In one example, (x, y, z) coordinates can be used to describe the 3D position of the vertex, and are also referred to as 3D coordinates. In some examples, connectivity information includes a vertex index set describing how to connect vertices to create a 3D surface. In some examples, mapping information describes how to map the mesh surface to a 2D area of a plane. In one example, mapping information is described together with UV parameters / texture coordinate (u, v) sets associated with mesh vertices and connectivity information. In some examples, vertex attributes include scalar or vector attribute values associated with mesh vertices. In some examples, attribute maps include attributes associated with mesh surfaces and stored as 2D images / videos. In one example, the mapping between a video (e.g., a 2D image / video) and a mesh surface is defined by mapping information.
[0081] According to one aspect of the present disclosure, some techniques referred to as UV mapping or mesh parameterization are used to map the surface of a mesh in a 3D domain to a 2D domain. In some examples, the mesh is cut into patches (also referred to as patch components) in the 3D domain. A patch is a continuous subset of the mesh whose boundaries are formed by border edges. A border edge of a patch is an edge that belongs to only one polygon in the patch and is not shared by two adjacent polygons in the patch. In some examples, the vertices of the border edges in a patch are referred to as the border vertices of the patch, and the non-border vertices in the patch may be referred to as the internal vertices of the patch.
[0082] According to an aspect of the present disclosure, in some examples, the patches are parameterized as 2D shapes (also referred to as UV patches, 2D patches, or UV charts). In some examples, the 2D shapes can be packaged (e.g., oriented and placed) into a map also referred to as a UV atlas. In some examples, the map can also be processed using 2D image or video processing techniques.
[0083] In one example, the UV mapping technique generates a UV atlas (also referred to as a UV map) and one or more texture atlases (also referred to as texture maps) corresponding to a piece of a 3D mesh in 2D. The UV atlas includes assigning 3D vertices of a 3D mesh to 2D points in a 2D domain (e.g., a rectangle). The UV atlas is a mapping between the coordinates of a 3D surface and the coordinates of a 2D domain. In one example, a point at a 2D coordinate (u, v) in the UV atlas has a value formed by the coordinates (x, y, z) of the vertex in the 3D domain. In one example, the texture atlas includes color information of the 3D mesh. For example, a point at a 2D coordinate (u, v) in the texture atlas (having a 3D value (x, y, z) in the UV atlas) has a color that specifies the color attribute of the point at (x, y, z) in the 3D domain. In some examples, the coordinates (x, y, z) in the 3D domain are called 3D coordinates or xyz coordinates, and the 2D coordinates (u, v) are called uv coordinates or UV coordinates.
[0084] According to some aspects of the present disclosure, mesh compression may be performed by representing the mesh using one or more 2D maps (also referred to as 2D atlases in some examples) and then encoding the 2D maps using an image or video codec. Different techniques may be used to generate the 2D maps.
[0085] Figure 6 A diagram illustrating the mapping of a 3D grid (610) to a 2D atlas (620) in some examples is shown. Figure 6In the example of , the 3D mesh (610) includes four vertices 1 to 4 that form four patches A to D. Each patch has a set of vertices and associated attribute information. For example, patch A is formed by vertices 1, 2, 3 connected into a triangle; patch B is formed by vertices 1, 3, 4 connected into a triangle; patch C is formed by vertices 1, 2, 4 connected into a triangle; and patch D is formed by vertices 2, 3, 4 connected into a triangle. In some examples, vertices 1, 2, 3, 4 may have corresponding attributes, and the triangle formed by vertices 1, 2, 3, 4 may have corresponding attributes.
[0086] In one example, patches A, B, C, and D in 3D are mapped to a 2D domain, such as a 2D atlas (620) also referred to as a UV atlas (620) or a map (620). For example, patch A is mapped to a 2D shape (also referred to as a UV patch) A' in the map (620), patch B is mapped to a 2D shape (also referred to as a UV patch) B' in the map (620), patch C is mapped to a 2D shape (also referred to as a UV patch) C' in the map (620), and patch D is mapped to a 2D shape (also referred to as a UV patch) D' in the map (620). In some examples, coordinates in the 3D domain are referred to as (x, y, z) coordinates, and coordinates in the 2D domain (e.g., the map (620)) are referred to as UV coordinates. Vertices in a 3D mesh may have corresponding UV coordinates in the map (620).
[0087] The map (620) may be a geometry map having geometry information, or may be a texture map having color, normal, texture or other attribute information, or may be an occupancy map having occupancy information.
[0088] Although in Figure 6 In the examples of , each patch is represented by a triangle, but it should be noted that a patch may include any suitable number of vertices connected to form a continuous subset of the mesh. In some examples, the vertices in a patch are connected into triangles. It should be noted that other suitable shapes may be used to connect the vertices in a patch.
[0089] In one example, the geometric information of the vertex may be stored in a 2D geometric map. For example, the 2D geometric map stores the (x, y, z) coordinates of the sampling point at the corresponding point in the 2D geometric map. For example, a point at the (u, v) position in the 2D geometric map has a vector value of 3 components, which correspond to the x, y, and z values of the corresponding sampling point in the 3D grid.
[0090] According to one aspect of the present disclosure, the area in the map may not be completely occupied. Figure 6In the decoded image, the area outside the 2D shapes A', B', C', D' is undefined. After decoding, the sample values of the area outside the 2D shapes A', B', C', D' can be discarded. In some cases, an occupancy map is used to store some additional information for each pixel, such as a binary value used to identify whether the pixel belongs to a patch or is undefined.
[0091] Figure 7 An example of a map (700) of a mesh in some examples is shown. The map (700) is a UV map (also called a UV atlas) that includes texture / UV coordinates and UV connectivity. The map (700) includes a plurality of UV charts that may correspond to patches of a mesh in a 3D domain, the UV charts including UV coordinates of points (also called UV vertices) and connectivity of the points.
[0092] In some examples, when a mesh includes non-positional attributes (e.g., color, normals, displacement, texture, UV / texture coordinates) (e.g., Figure 7 ), the mesh codec needs to encode connectivity that is not a locality property. For example, Figure 7 As shown, the polygon mesh includes UV / texture coordinates, and the mesh codec needs to encode the corresponding connectivity of the UV chart (hereinafter referred to as UV connectivity), such as the connectivity of the UV vertices in the UV chart. It should be noted that in some examples, different non-position attributes may have corresponding connectivity.
[0093] In a related example, the connectivity of the non-positional attribute is encoded into a separate grid using a direct coding method, and the correspondence between the corners of the non-positional attribute and the 3D position can be represented by a signal. However, representing the correspondence between the corners of the non-positional attribute and the 3D position using a signal requires a large number of bits, and there is a large redundancy between the positional connectivity and the non-positional attribute connectivity (connectivity of the non-positional attribute), so encoding the non-positional attribute connectivity using a direct coding method may be inefficient.
[0094] In another related example, a seam edge is signaled. When an edge between two 3D vertices in a 3D space is split into two edges in a space for non-positional attributes (also referred to as a non-positional attribute space, such as a 2D UV chart), the edge in the 3D space is called a seam edge. In a related example, the non-positional attribute connectivity can be derived based on the information of the seam edge and the positional connectivity, thereby avoiding the angle correspondence between the angles of the non-positional attributes and the 3D vertices that are signaled.
[0095] According to some aspects of the present disclosure, signaling seam edges is also not efficient enough because the number of edges is much larger than the number of vertices and faces. Some aspects of the present disclosure provide more efficient techniques to encode the connectivity of non-positional attributes such as UV coordinates, normals, etc. for polygonal mesh compression.
[0096] Some aspects of the present disclosure provide techniques for encoding the connectivity of non-positional attributes (e.g., texture coordinates, normals, colors, etc.) of polygonal meshes with high coding efficiency (e.g., without using any additional bytes). In some examples, for some special cases (also referred to as extreme cases), the connectivity of the non-positional attributes of the mesh (also referred to as non-positional attribute connectivity) can be derived based on the connectivity of the vertices (3D vertices) of the mesh in 3D space (also referred to as positional connectivity). The encoder can determine whether the connectivity of the non-positional attributes of the mesh belongs to one or more extreme cases, and include the determined extreme case information in the bitstream of the encoded information of the mesh. When the decoder side knows that the connectivity of the non-positional attributes of the mesh belongs to an extreme case, the decoder can determine the connectivity of the non-positional attributes based on the connectivity of the 3D vertices of the mesh. It should be noted that the various disclosed techniques can be applied individually or in any combination to encode the connectivity of polygonal mesh attributes.
[0097] According to one aspect of the present disclosure, for polygon meshes, the correspondence between positional vertices (also referred to as spatial vertices, 3D vertices) and non-positional attribute vertices is generally a one-to-one mapping or a one-to-many mapping. For example, any non-positional attribute vertex (e.g., UV vertex) cannot be shared by different spatial vertices. In other words, the connectivity of non-positional attributes can be obtained by cutting the positional connectivity, and the gluing of positional vertices / edges is not allowed.
[0098] In various examples, connectivity of non-positional attributes can be obtained by cutting positional connectivity without gluing positional vertices / edges, and the number of vertices of non-positional attributes can carry information related to connectivity of non-positional attributes. Since cutting positional connectivity increases the number of non-positional attribute vertices compared to positional vertices, there are three cases related to cutting positional connectivity. These three cases are called the first case, the second case, and the third case.
[0099] In the first case, the number of vertices of non-position attributes (also called non-position attribute vertices, attribute vertices) is equal to the number of position vertices (vertices in 3D space, also called 3D vertices), then no cut is applied to the position connectivity, so the connectivity of the attribute (also called non-position attribute connectivity) and the position connectivity are the same. In some examples, the first case is called the first extreme case of no cut.
[0100] In the second case, the number of vertices of the non-positional attribute is equal to the total face degree. In one example, the total face degree is the sum of the number of vertices associated with each face. In one example, the face degree of a vertex is defined as the number of faces associated with the vertex, and the total face degree is the sum of the face degrees of the vertices. In the second case, the cut is applied to each non-boundary edge of the positional connectivity, so all faces of the non-positional attribute are disconnected / separated. In some examples, the second case is referred to as the second extreme case of cutting each edge.
[0101] The third case is between the first and second cases described above, for example, between the first extreme case of no cutting and the second extreme case of cutting each edge. In the third case, in some examples, the execution may signal how to cut the location connectivity into the attribute connectivity.
[0102] According to some aspects of the present disclosure, the first case and the second case are not mutually exclusive. For example, when all faces in positional connectivity are disconnected, there are no edges to cut, so positional connectivity and non-positional connectivity can have the same number of vertices, and the number of vertices in positional connectivity and non-positional connectivity are both equal to the total face degree. In this case, the positional attribute (3D position) and the non-positional attribute share the same connectivity, both are disconnected faces.
[0103] According to some aspects of the present disclosure, for a first extreme case and a second extreme case, two bits may be used to signal the connectivity of a non-positional attribute. For example, a first bit of the two bits may indicate whether a first case is true (the positional attribute and the non-positional attribute have the same number of vertices), and a second bit of the two bits may indicate whether a second case is true (the number of vertices of the non-positional attribute is equal to the total facet degree).
[0104] In some examples, instead of using an additional byte to signal the connectivity of the two extreme cases, the two bits can be part of other information signaled. For example, to signal the quantization bit (denoted as QT) of the UV coordinate, six bits in the byte (e.g., the first 6 bits in the byte) can be used to signal QT, and then the remaining 2 bits in the byte can be used to signal the connectivity of the two extreme cases. In this way, the use of any additional bytes to signal the connectivity of non-positional attributes can be avoided.
[0105] Figure 8A flow chart outlining a process (800) according to an aspect of the present disclosure is shown. The process (800) may be used for a trellis decoder. In various aspects, the process (800) is performed by a processing circuit, such as a processing circuit that performs the functions of a 3D decoder (110), etc. In some aspects, the process (800) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (800). The process starts at (S801) and proceeds to (S810).
[0106] At (S810), a code stream of encoded information of a mesh is received, the mesh comprising a plurality of 3D vertices in a three-dimensional (3D) space and at least non-positional attributes, the encoded information comprising positional connectivity of the plurality of 3D vertices in the 3D space.
[0107] At (S820), it is determined whether the encoded information of the mesh indicates that the non-positional attribute connectivity of the non-positional attribute corresponds to at least an extreme case.
[0108] At (S830), when the non-position attribute connectivity corresponds to an extreme case, the non-position attribute connectivity is determined according to the position connectivity of the plurality of 3D vertices.
[0109] In some examples, a non-position attribute connectivity corresponding to a first extreme case is determined. In the first extreme case, the non-position attribute connectivity matches the position connectivity. In one example, when the position connectivity is appropriately determined, the non-position attribute connectivity is the same as the position connectivity.
[0110] In some examples, a non-position attribute connectivity corresponding to a second extreme case is determined. In the second extreme case, the faces of the non-position attribute are not connected. In one example, when the position connectivity is appropriately determined, the non-position attribute connectivity can be obtained by disconnecting all faces.
[0111] In some examples, a signal is decoded from at least the encoded information of the grid. The signal indicates whether the non-positional attribute connectivity of the non-positional attribute corresponds to at least an extreme case.
[0112] In some examples, a first bit and a second bit are decoded from the encoded information of the grid, the first bit indicating whether the non-positional attribute connectivity of the non-positional attribute corresponds to a first extreme case, in which the non-positional attribute connectivity matches the positional connectivity; the second bit indicating whether the non-positional attribute connectivity of the non-positional attribute corresponds to a second extreme case, in which the faces of the non-positional attribute are not connected. For example, a byte is decoded from the encoded information of the grid, the byte including 6 bits indicating the quantization parameter, the first bit, and the second bit.
[0113] It should be noted that the non-positional attribute may be any suitable non-positional attribute, such as color, normal, displacement, texture, UV coordinates, etc.
[0114] Then, the process proceeds to (S899) and terminates.
[0115] The process (800) may be adapted as appropriate. Steps in the process (800) may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.
[0116] Fig. 9 A flow chart outlining a process (900) according to an aspect of the present disclosure is shown. The process (900) may be used for a trellis encoder. In various aspects, the process (900) is performed by a processing circuit, such as a processing circuit that performs the functions of a 3D encoder (103), etc. In some aspects, the process (900) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (900). The process starts at (S901) and proceeds to (S910).
[0117] At (S910), for a mesh including a plurality of 3D vertices in a three-dimensional (3D) space and at least non-positional attributes, it is determined whether non-positional attribute connectivity of the non-positional attributes corresponds to at least an extreme case.
[0118] At (S920), a signal indicating whether the non-positional attribute connectivity of the non-positional attribute corresponds to an extreme case is encoded into a code stream of the encoded information of the grid.
[0119] In some examples, a first number of non-position attribute vertices representing non-position attributes in the 2D map is checked to see if it is equal to a second number of the plurality of 3D vertices. When the first number is equal to the second number, the non-position attribute connectivity is determined to be a first extreme case, in which the non-position attribute connectivity matches the position connectivity of the plurality of 3D vertices in the 3D space.
[0120] In some examples, a first number of non-position attribute vertices representing the non-position attribute in the 2D map is checked to see if it is equal to a total number of face degrees of the mesh, the total number of face degrees being the sum of the number of vertices associated with each face of the mesh. When the first number is equal to the total number of face degrees, the non-position attribute connectivity is determined to be a second extreme case, in which the faces of the non-position attribute are not connected.
[0121] In some examples, a first bit and a second bit are encoded into a bitstream of encoded information of a mesh. The first bit indicates whether the non-positional attribute connectivity of the non-positional attribute corresponds to a first extreme case, in which the non-positional attribute connectivity matches the positional connectivity of multiple 3D vertices in a 3D space; the second bit indicates whether the non-positional attribute connectivity of the non-positional attribute corresponds to a second extreme case, in which the faces of the non-positional attribute are not connected.
[0122] In some examples, a byte is encoded into a bitstream of encoded information of the grid. The byte includes 6 bits indicating a quantization parameter, a first bit and a second bit.
[0123] In some examples, the non-positional attributes include at least one of color, normal, displacement, texture, and UV coordinates.
[0124] Then, the process proceeds to (S999) and terminates.
[0125] The process (900) may be adapted as appropriate. Steps in the process (900) may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.
[0126] According to one aspect of the present disclosure, a method for processing a mesh is provided. In the method, conversion between a mesh file and a code stream of a compressed mesh is performed according to a format rule. For example, the code stream may be a code stream decoded / encoded in any decoding and / or encoding method described herein. The format rule may specify one or more constraints of the code stream and / or one or more processes to be performed by a decoder and / or encoder.
[0127] In one example, a code stream includes encoded information of a mesh, the mesh including a plurality of 3D vertices and at least a non-positional attribute in a three-dimensional (3D) space, the encoded information including positional connectivity of the plurality of 3D vertices in the 3D space. The format rule specifies: determining whether the encoded information of the mesh indicates that the non-positional attribute connectivity corresponds to at least an extreme case; and determining the non-positional attribute connectivity according to the positional connectivity of the plurality of 3D vertices when the non-positional attribute connectivity corresponds to an extreme case.
[0128] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig.10 A computer system (1000) suitable for implementing certain aspects of the disclosed subject matter is shown.
[0129] Computer software may be encoded using any suitable machine code or computer language that may be subjected to assembly, compilation, linking or similar mechanisms to create code comprising instructions that may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpretation, microcode execution, etc.
[0130] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.
[0131] Fig.10 The components of the computer system (1000) shown are examples and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing aspects of the present disclosure. Nor should the configuration of components be interpreted as having any dependency or requirement related to any one or combination of components shown in the exemplary aspects of the computer system (1000).
[0132] The computer system (1000) may include certain human-machine interface input devices. Such human-machine interface input devices may be responsive to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not shown). Human-machine interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, captured images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0133] The human-machine interface input device may include one or more of the following (only one of each is shown): keyboard (1001), mouse (1002), touchpad (1003), touch screen (1010), data gloves (not shown), joystick (1005), microphone (1006), scanner (1007), camera (1008).
[0134] The computer system (1000) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback of a touch screen (1010), a data glove (not shown), or a joystick (1005), but may also be a tactile feedback device that is not an input device), audio output devices (e.g., speakers (1009), headphones (not shown)), visual output devices (e.g., screens (1010) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities, some of which are capable of outputting two-dimensional visual outputs or outputs exceeding three dimensions through devices such as stereoscopic image output, virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown), and printers (not shown).
[0135] The computer system (1000) may also include human-computer accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1020) having CD / DVD and other media (1021), thumb drives (1022), removable hard disk drives or solid-state drives (1023), traditional magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security software dogs (not shown), etc.
[0136] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.
[0137] The computer system (1000) may also include an interface (1054) to one or more communication networks (1055). The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter attached to some common data port or peripheral bus (1049) (e.g., a USB port of the computer system (1000)); as described below, other network interfaces are typically integrated into the kernel of the computer system (1000) by attaching to the system bus (e.g., connected to an Ethernet interface in a PC computer system or connected to a cellular network interface in a smartphone computer system). The computer system (1000) can use any of these networks to communicate with other entities. Such communications may be one-way receive only (e.g., broadcast television), one-way send only (e.g., CANBus connected to certain CANBus devices), or bidirectional, for example, using a LAN or WAN digital network to connect to other computer systems. As described above, certain protocols and protocol stacks may be used on each of those networks and network interfaces.
[0138] The above-mentioned human-machine interface device, human-machine accessible storage device and network interface may be attached to the kernel (1040) of the computer system (1000).
[0139] The core (1040) may include one or more central processing units (CPUs) (1041), graphics processing units (GPUs) (1042), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (1043), hardware accelerators (1044) for certain tasks, graphics adapters (1050), etc. These devices, as well as read-only memory (ROM) (1045), random access memory (1046), internal mass storage (1047) such as internal non-user accessible hard drives, SSDs, etc., may be connected via a system bus (1048). In some computer systems, the system bus (1048) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (1048) or to the core's system bus (1048) via a peripheral bus (1049). In one example, a touch screen (1010) may be connected to a graphics adapter (1050). The architecture of the peripheral bus includes PCI, USB, etc.
[0140] The CPU (1041), GPU (1042), FPGA (1043) and accelerator (1044) can execute certain instructions, which can be combined to form the above-mentioned computer code. The computer code can be stored in ROM (1045) or RAM (1046). Transitional data can also be stored in RAM (1046), while permanent data can be stored, for example, in internal mass storage (1047). Fast storage and retrieval to any storage device can be performed by using a cache, which can be closely associated with the following: one or more CPUs (1041), GPUs (1042), mass storage (1047), ROM (1045), RAM (1046), etc.
[0141] The computer readable medium may have thereon computer codes for performing various computer-implemented operations. The medium and computer codes may be those specially designed and constructed for the purposes of the present disclosure, or the medium and computer codes may be of a type well known and available to those skilled in the art of computer software.
[0142] As an example, and not by way of limitation, a computer system (1000) having an architecture, particularly a kernel (1040), may provide functionality due to one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be media associated with user-accessible mass storage as described above, as well as certain non-temporary kernel (1040) memories, such as kernel internal mass storage (1047) or ROM (1045). Software implementing various aspects of the present disclosure may be stored in such devices and executed by the kernel (1040). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software may cause the kernel (1040), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform specific processes or specific parts of specific processes described herein, including defining data structures stored in RAM (1046) and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system may provide functionality due to logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1044)) that may replace software or operate in conjunction with software to perform specific processes or specific portions of specific processes described herein. Where appropriate, references to portions of software may include logic and vice versa. Where appropriate, references to portions of computer-readable media may include circuitry (e.g., integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.
[0143] As used in this disclosure, "at least one of" or "one of" is intended to include any one or combination of the listed elements. For example, reference to at least one of A, B, and C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include only A, only B, only C, or any combination thereof. Reference to one of A and B, and one of A and B is intended to include A or B or (A and B). Where applicable, the use of "one of" does not exclude any combination of the listed elements, such as when the elements are not mutually exclusive.
[0144] Although the present disclosure has described several examples of various aspects, there are changes, permutations, and various replacement equivalents that fall within the scope of the present disclosure. Therefore, it should be appreciated that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and therefore fall within the spirit and scope of the present disclosure.
[0145] The above disclosure also includes the features mentioned below. Features can be combined in various ways and are not limited to the combinations mentioned below.
[0146] (1) A method for mesh processing, the method comprising: receiving a code stream of encoded information of a mesh, the mesh comprising a plurality of 3D vertices and at least non-positional attributes in a three-dimensional (3D) space, the encoded information comprising positional connectivity of the plurality of 3D vertices in the 3D space; determining whether the encoded information of the mesh indicates that the non-positional attribute connectivity of the non-positional attributes at least corresponds to an extreme case; and when the non-positional attribute connectivity corresponds to an extreme case, determining the non-positional attribute connectivity based on the positional connectivity of the plurality of 3D vertices.
[0147] (2) According to the method described in feature (1), the method also includes: determining that the non-location attribute connectivity corresponds to a first extreme case, in which the non-location attribute connectivity matches the location connectivity.
[0148] (3) According to the method described in any one of features (1) to (2), the method includes: determining that the non-position attribute connectivity corresponds to a second extreme case, in which the faces of the non-position attribute are disconnected.
[0149] (4) A method according to any one of features (1) to (3), the method comprising: decoding a signal from at least the encoded information of the grid, the signal indicating whether the non-location attribute connectivity of the non-location attribute corresponds to at least an extreme case.
[0150] (5) According to the method described in any one of features (1) to (4), the method includes: decoding a first bit and a second bit from the encoded information of the grid, the first bit indicating whether the non-positional attribute connectivity of the non-positional attribute corresponds to a first extreme case, in which the non-positional attribute connectivity matches the positional connectivity; the second bit indicating whether the non-positional attribute connectivity of the non-positional attribute corresponds to a second extreme case, in which the faces of the non-positional attribute are disconnected.
[0151] (6) According to the method described in any one of features (1) to (5), the method includes: decoding a byte from the encoded information of the grid, the byte including 6 bits indicating a quantization parameter, a first bit and a second bit.
[0152] (7) A method according to any one of features (1) to (6), wherein the non-position attribute includes at least one of color, normal, displacement, texture, and UV coordinates.
[0153] (8) A method for mesh processing, the method comprising: for a mesh including multiple 3D vertices in a three-dimensional (3D) space and at least non-positional attributes, determining whether non-positional attribute connectivity of the non-positional attributes corresponds to at least an extreme case; and encoding a signal indicating whether the non-positional attribute connectivity of the non-positional attributes corresponds to an extreme case into a code stream of encoded information of the mesh.
[0154] (9) According to the method described in feature (8), the method also includes: checking whether a first number of non-position attribute vertices used to represent non-position attributes in the 2D map is equal to a second number of multiple 3D vertices; and when the first number is equal to the second number, determining that the non-position attribute connectivity corresponds to a first extreme case, in which the non-position attribute connectivity matches the position connectivity of multiple 3D vertices in the 3D space.
[0155] (10) According to the method described in any one of features (8) to (9), the method further includes: checking whether a first number of non-positional attribute vertices representing non-positional attributes in the 2D map is equal to the total number of face degrees of the mesh, the total number of face degrees being the sum of the number of vertices associated with each face of the mesh; and when the first number is equal to the total number of face degrees, determining that the non-positional attribute connectivity corresponds to a second extreme case, in which the faces of the non-positional attributes are disconnected.
[0156] (11) According to the method described in any one of features (8) to (10), the method further includes: encoding a first bit and a second bit into a code stream of encoded information of the mesh, the first bit indicating whether the non-positional attribute connectivity of the non-positional attribute corresponds to a first extreme case, in which the non-positional attribute connectivity matches the positional connectivity of multiple 3D vertices in 3D space; the second bit indicating whether the non-positional attribute connectivity of the non-positional attribute corresponds to a second extreme case, in which the faces of the non-positional attribute are disconnected.
[0157] (12) According to the method described in any one of features (8) to (11), the method further includes: encoding a byte into a code stream of encoded information of the grid, the byte including 6 bits indicating a quantization parameter, a first bit and a second bit.
[0158] (13) A method according to any one of features (8) to (12), wherein the non-position attributes include at least one of color, normal, displacement, texture, and UV coordinates.
[0159] (14) A method for processing mesh data, the method comprising processing a code stream of the mesh data according to a format rule. The code stream comprises encoded information of the mesh, the encoded information comprising positional connectivity of a plurality of 3D vertices of the mesh in a three-dimensional (3D) space, and a correspondence between the plurality of 3D vertices and UV vertices in a UV space of the mesh. The format rule specifies: determining whether the encoded information of the mesh indicates that non-positional attribute connectivity of a non-positional attribute at least corresponds to an extreme case; and when the non-positional attribute connectivity corresponds to an extreme case, determining the non-positional attribute connectivity according to the positional connectivity of the plurality of 3D vertices.
[0160] (15) A method according to feature (14), wherein the format rule further specifies: determining non-positional attribute connectivity corresponding to a first extreme case, in which the non-positional attribute connectivity matches the positional connectivity.
[0161] (16) A method according to any one of features (14) to (15), wherein the format rule further specifies: determining the connectivity of non-positional attributes corresponding to a second extreme case, in which the faces of the non-positional attributes are disconnected.
[0162] (17) A method according to any one of features (14) to (16), wherein the format rule further specifies: decoding a signal from at least the encoded information of the grid, the signal indicating whether the non-positional attribute connectivity of the non-positional attribute corresponds to at least an extreme case.
[0163] (18) A method according to any one of features (14) to (17), wherein the format rule further specifies: determining a first bit and a second bit from the encoded information of the grid, the first bit indicating whether the non-positional attribute connectivity of the non-positional attribute corresponds to a first extreme case, in which the non-positional attribute connectivity matches the positional connectivity; the second bit indicating whether the non-positional attribute connectivity of the non-positional attribute corresponds to a second extreme case, in which the faces of the non-positional attribute are disconnected.
[0164] (19) A method according to any one of features (14) to (18), wherein the format rule further specifies: decoding a byte from the encoded information of the grid, the byte comprising 6 bits indicating a quantization parameter, a first bit and a second bit.
[0165] (20) A method according to any one of features (14) to (19), wherein the non-position attributes include at least one of color, normal, displacement, texture, and UV coordinates.
[0166] (21) An apparatus for grid processing, comprising a processing circuit configured to perform the method of any one of features (1) to (7).
[0167] (22) An apparatus for grid processing, comprising a processing circuit configured to perform the method of any one of features (8) to (13).
[0168] (23) A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform the method described in any one of features (1) to (19).
Claims
1. A method for grid processing, characterized in that: include: receiving a bitstream of encoded information of a mesh, the mesh comprising a plurality of 3D vertices in a three-dimensional 3D space and at least non-positional attributes, the encoded information comprising positional connectivity of the plurality of 3D vertices in the 3D space; determining whether the encoded information of the grid indicates that the non-positional attribute connectivity of the non-positional attribute corresponds to at least an extreme case; as well as When the non-position attribute connectivity corresponds to the extreme case, the non-position attribute connectivity is determined according to the position connectivity of the plurality of 3D vertices.
2. The method according to claim 1, characterized in that Also includes: Determining the non-position attribute connectivity corresponds to a first extreme case, in which the non-position attribute connectivity matches the position connectivity.
3. The method according to claim 1, characterized in that Also includes: Determining the non-positional attribute connectivity corresponds to a second extreme case, in which the faces of the non-positional attribute are not connected.
4. The method according to claim 1, characterized in that: Also includes: A signal is decoded from at least the encoded information of the grid, the signal indicating whether the non-positional attribute connectivity of the non-positional attribute corresponds to at least the extreme case.
5. The method according to claim 4, characterized in that Also includes: A first bit and a second bit are decoded from the encoded information of the grid, the first bit indicating whether the non-positional attribute connectivity of the non-positional attribute corresponds to a first extreme case, in which the non-positional attribute connectivity matches the positional connectivity; the second bit indicating whether the non-positional attribute connectivity of the non-positional attribute corresponds to a second extreme case, in which the faces of the non-positional attribute are disconnected.
6. The method according to claim 5, characterized in that Also includes: A byte is decoded from the encoded information of the grid, the byte including 6 bits indicating a quantization parameter, the first bit, and the second bit.
7. The method according to any one of claims 1 to 6, characterized in that The non-positional attributes include at least one of color, normal, displacement, texture, and UV coordinates.
8. A method for grid processing, characterized in that: include: For a mesh including a plurality of 3D vertices in a three-dimensional 3D space and at least a non-position attribute, determining whether a non-position attribute connectivity of the non-position attribute corresponds to at least an extreme case; as well as A signal indicating whether the non-positional attribute connectivity of the non-positional attribute corresponds to the extreme case is encoded into a code stream of the encoded information of the grid.
9. The method according to claim 8, characterized in that Also includes: Checking whether a first number of non-position attribute vertices representing the non-position attribute in the two-dimensional (2D) map is equal to a second number of the plurality of 3D vertices; as well as When the first number is equal to the second number, determining the non-position attribute connectivity corresponds to a first extreme case in which the non-position attribute connectivity matches the position connectivity of the plurality of 3D vertices in the 3D space.
10. The method according to claim 8, characterized in that Also includes: Checking whether a first number of non-position attribute vertices representing the non-position attribute in the 2D map is equal to a total number of face degrees of the mesh, the total number of face degrees being a sum of the number of vertices associated with each face of the mesh; as well as When the first number is equal to the total number of the facet degrees, determining the non-positional attribute connectivity corresponds to a second extreme case, in which the faces of the non-positional attribute are not connected.
11. The method according to claim 8, characterized in that Also includes: A first bit and a second bit are encoded into the code stream of the encoded information of the mesh, the first bit indicating whether the non-position attribute connectivity of the non-position attribute corresponds to a first extreme case, in which the non-position attribute connectivity matches the position connectivity of the multiple 3D vertices in the 3D space; the second bit indicates whether the non-position attribute connectivity of the non-position attribute corresponds to a second extreme case, in which the faces of the non-position attribute are disconnected.
12. The method according to claim 11, characterized in that Also includes: A byte is encoded into the code stream of the encoded information of the grid, the byte including 6 bits indicating a quantization parameter, the first bit and the second bit.
13. A method for processing grid data, characterized in that: The method comprises: Process the code stream of the grid data according to the format rules, where: The code stream includes encoded information of a mesh, the mesh including a plurality of 3D vertices in a three-dimensional 3D space and at least non-positional attributes, the encoded information including positional connectivity of the plurality of 3D vertices in the 3D space; and The format rules specify: determining whether the encoded information of the grid indicates that the non-positional attribute connectivity of the non-positional attribute corresponds to at least an extreme case; and When the non-position attribute connectivity corresponds to the extreme case, the non-position attribute connectivity is determined according to the position connectivity of the plurality of 3D vertices.
14. A device for grid processing, characterized in that: The method comprises a processing module configured to perform the method of any one of claims 1 to 7, the method of any one of claims 8 to 12, or the method of claim 13.
15. A non-transitory computer-readable storage medium, characterized in that: Instructions are stored, and when the instructions are executed by at least one processor, the at least one processor executes the method of any one of claims 1 to 7, the method of any one of claims 8 to 12, or the method of claim 13.
16. A method for processing a video stream, characterized in that: The video code stream is generated according to the method described in any one of claims 8 to 12, or is decoded based on the method described in any one of claims 1 to 7.