Coding and decoding method, code stream transmission method, codec and storage medium
By extracting and reconstructing the geometric and attribute information of point clouds through point cloud encoding and decoding, the problem of insufficient entropy encoding and decoding performance in the existing AI-PCC standard is solved, and more efficient point cloud data compression and transmission are achieved.
Patent Information
- Application Number
- CN202511260673.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-09-06
- Filing Date
- 2025-09-04
- Publication Date
- 2026-03-10
AI Technical Summary
The existing AI-PCC (AI-based learning) standard for point cloud encoding lacks comprehensive entropy encoding and decoding processes, resulting in low entropy encoding and decoding performance for point clouds.
An encoding and decoding method is provided that reconstructs the geometry and attribute values of a point cloud by determining the syntax element identifier information of the point cloud's geometry and attribute information, thereby improving the entropy encoding and decoding performance.
By extracting and reconstructing the geometric and attribute information of point clouds, the entropy encoding and decoding performance of point clouds is improved, enabling more efficient point cloud data compression and transmission.
Smart Images

Figure CN121644805A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 691,627, filed on September 6, 2024, entitled “METHODAND SYSTEM OF CARRIAGE FOR AUXILIARY INFORMATION FOR LEARNING-BASEDSPARSECONVOLUTION POINT CLOUD CODING”. Technical Field
[0003] This application relates to the field of point cloud encoding and decoding technology, and in particular to an encoding and decoding method, a method for transmitting code streams, an encoder and decoder, and a storage medium. Background Technology
[0004] With the booming development of emerging technologies such as augmented reality, virtual reality, autonomous driving, and robotics, point cloud data has become one of the main data forms due to its concise representation of three-dimensional space. However, point cloud data is massive, and directly storing point cloud data consumes a lot of memory and is not conducive to transmission. Therefore, high-performance point cloud compression technology is essential.
[0005] In recent years, neural network and deep learning technologies have been widely applied in the field of point cloud geometry compression. However, in the Artificial Intelligence Point Cloud Coding (AI-PCC) standard, the current entropy encoding and decoding processes are not fully considered, which greatly reduces the entropy encoding and decoding performance of point clouds. Summary of the Invention
[0006] This application provides an encoding / decoding method, a method for transmitting code streams, an encoding / decoding method, and a storage medium, which can improve the entropy encoding / decoding performance of point clouds.
[0007] The technical solution of this application embodiment can be implemented as follows:
[0008] In a first aspect, embodiments of this application provide a decoding method applied to a decoder, the method comprising:
[0009] Decode the bitstream to determine the indication information of the current point cloud; wherein the indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on the geometric information of the learned AI-PCC; second syntax element identification information based on the attribute information of the learned AI-PCC;
[0010] The geometric reconstruction value of the current point cloud is determined based on the identifier information of the first syntax element;
[0011] The attribute reconstruction values of the current point cloud are determined based on the identification information of the second syntax element;
[0012] Reconstruct the current point cloud based on its geometric reconstruction values and attribute reconstruction values.
[0013] Secondly, embodiments of this application provide an encoding method applied to an encoder, the method comprising:
[0014] Determine the indication information of the current point cloud and write the indication information of the current point cloud into the bit stream;
[0015] in,
[0016] The current point cloud's indication information includes at least one or more of the following: first syntax element identification information based on the geometric information of the learned AI-PCC; and second syntax element identification information based on the attribute information of the learned AI-PCC.
[0017] The first syntax element identifier is used to determine the geometric reconstruction value of the current point cloud;
[0018] The second syntax element identifies the information used to determine the attribute reconstruction values of the current point cloud.
[0019] Thirdly, embodiments of this application provide a method for transmitting a bitstream, wherein the bitstream is generated based on the encoding method described in any one of the second aspects.
[0020] Fourthly, embodiments of this application provide an encoder, the encoder including a first determining unit, wherein:
[0021] The first determining unit is configured to determine the indication information of the current point cloud and write the indication information of the current point cloud into the code stream; wherein the indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on the geometric information of the learned AI-PCC; second syntax element identification information based on the attribute information of the learned AI-PCC; the first syntax element identification information is used to determine the geometric reconstruction value of the current point cloud; and the second syntax element identification information is used to determine the attribute reconstruction value of the current point cloud.
[0022] Fifthly, embodiments of this application provide an encoder, the encoder comprising:
[0023] A first memory for storing computer programs that can run on a first processor;
[0024] A first processor is configured to execute the encoding method as described in the second aspect when running a computer program.
[0025] Sixthly, embodiments of this application provide a decoder, the decoder comprising:
[0026] The second determining unit is configured to decode the bitstream and determine the indication information of the current point cloud; wherein the indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on the geometric information of the learned AI-PCC; second syntax element identification information based on the attribute information of the learned AI-PCC; determining the geometric reconstruction value of the current point cloud based on the first syntax element identification information; determining the attribute reconstruction value of the current point cloud based on the second syntax element identification information; and reconstructing the current point cloud based on the geometric reconstruction value and the attribute reconstruction value of the current point cloud.
[0027] In a seventh aspect, embodiments of this application provide a decoder, which includes a second memory and a second processor, wherein:
[0028] The second memory is used to store computer programs that can run on the second processor;
[0029] The second processor is used to execute the decoding method as described in the first aspect when running a computer program.
[0030] Eighthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the decoding method as described in the first aspect, or the encoding method as described in the second aspect.
[0031] In a ninth aspect, embodiments of this application provide a computer-readable storage medium having a bitstream stored thereon, the bitstream being generated by performing the steps of the encoding method as described in the second aspect.
[0032] In a tenth aspect, embodiments of this application provide a method for storing a bitstream, comprising generating a bitstream by performing the encoding method as described in the second aspect; and storing the bitstream.
[0033] In the eleventh aspect, embodiments of this application provide a method for reading a bitstream, including reading the bitstream and performing the decoding method as described in the first aspect to decode the bitstream and generate a video or image.
[0034] In a twelfth aspect, embodiments of this application provide a method for receiving a bitstream, including receiving the bitstream and performing the decoding method as described in the first aspect to decode the bitstream and generate a video or image.
[0035] In a thirteenth aspect, embodiments of this application provide a computer-readable storage medium having a computer program / instructions and a bitstream stored thereon, wherein the computer program / instructions, when executed by a processor, implement the steps of the encoding method as described in the second aspect to generate the bitstream.
[0036] In a fourteenth aspect, embodiments of this application provide a computer-readable storage medium having a computer program / instructions and a bitstream stored thereon, wherein the computer program / instructions, when executed by a processor, implement the steps of the decoding method as described in the first aspect to decode the bitstream and generate a video or image.
[0037] In a fifteenth aspect, embodiments of this application provide a bitstream generated by bit encoding based on information to be encoded; wherein the information to be encoded includes at least one of the following: indication information of the current point cloud; wherein the indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on geometric information of the learned AI-PCC; and second syntax element identification information based on attribute information of the learned AI-PCC.
[0038] This application provides an encoding / decoding method, a method for transmitting a bitstream, an encoder / decoder, and a storage medium. The decoder decodes the bitstream to determine indication information of the current point cloud. The indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on learned AI-PCC geometric information; second syntax element identification information based on learned AI-PCC attribute information; geometric reconstruction values of the current point cloud determined based on the first syntax element identification information; attribute reconstruction values of the current point cloud determined based on the second syntax element identification information; and the current point cloud is reconstructed based on the geometric reconstruction values and attribute reconstruction values. The encoder determines the indication information of the current point cloud and writes it into the bitstream. The indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on learned AI-PCC geometric information; second syntax element identification information based on learned AI-PCC attribute information; the first syntax element identification information is used to determine the geometric reconstruction values of the current point cloud; and the second syntax element identification information is used to determine the attribute reconstruction values of the current point cloud. Therefore, in the embodiments of this application, the indication information transmitted in the bitstream may include first syntax element identification information of geometric information based on learned AI-PCC and / or second syntax element identification information of attribute information based on learned AI-PCC. That is, the technical solution of this application provides a syntax structure or encoding sequence format that can be used to carry compressed data, realizes multiple support for multi-attribute datasets of learning-based artificial intelligence point cloud coding (AI-PCC), and thus can improve the entropy encoding and decoding performance of point clouds. Attached Figure Description
[0039] Figure 1A This is a schematic diagram of a 3D point cloud image;
[0040] Figure 1B This is a magnified view of a portion of a 3D point cloud image;
[0041] Figure 2AA schematic diagram of the six viewing dimensions of a point cloud image;
[0042] Figure 2B This is a schematic diagram illustrating the structure of the header information section and the data section of a file;
[0043] Figure 3 A schematic diagram of PCG voxelization;
[0044] Figure 4 This is a schematic diagram of the MP-POV processing procedure;
[0045] Figure 5 This is a schematic diagram of a sparse geometric coding framework;
[0046] Figure 6 A schematic diagram of the scaling up and down of the voxel sampling layer;
[0047] Figure 7 This is a schematic diagram of the classic ResNet network architecture;
[0048] Figure 8 This is a schematic diagram of the SOPA upscaling process;
[0049] Figure 9 A schematic diagram of the multi-level SOPA scaling process;
[0050] Figure 10 A schematic diagram of probability thresholding for point clouds of dense objects;
[0051] Figure 11 A schematic diagram illustrating the position offset adjustment of a sparse LiDAR point cloud;
[0052] Figure 12 A schematic diagram of 1 / 3 / 8 levels for grouping 8 labeled MP-POVs;
[0053] Figure 13 Schematic diagram of SLNE-enhanced Level 1 SOPA;
[0054] Figure 14 A schematic diagram of a general architecture with m scales encoded in lossless mode and (Nm) scales encoded in lossy mode;
[0055] Figure 15 This is a schematic diagram illustrating the application scenarios of a geometric encoder;
[0056] Figure 16 A flowchart illustrating the implementation details of OPU;
[0057] Figure 17 This is a schematic diagram of the structure of IRN and NPFormer used to form DNN blocks;
[0058] Figure 18A flowchart illustrating a decoding method provided in an embodiment of this application;
[0059] Figure 19 A flowchart illustrating an encoding method provided in an embodiment of this application;
[0060] Figure 20 A schematic diagram of the composition structure of an encoder provided in an embodiment of this application;
[0061] Figure 21 This is a schematic diagram of the specific hardware structure of an encoder provided in an embodiment of this application;
[0062] Figure 22 A schematic diagram of the composition structure of a decoder provided in an embodiment of this application;
[0063] Figure 23 This is a schematic diagram of the specific hardware structure of a decoder provided in an embodiment of this application;
[0064] Figure 24 This is a schematic diagram of the composition structure of an encoding / decoding system provided in an embodiment of this application. Detailed Implementation
[0065] In order to gain a more detailed understanding of the features and technical content of the embodiments of this application, the implementation of the embodiments of this application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for reference and illustration only and are not intended to limit the embodiments of this application.
[0066] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0067] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0068] It should also be noted that the terms "first, second, and third" used in the embodiments of this application are only used to distinguish similar objects and do not represent a specific order of objects. It is understood that "first, second, and third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0069] It should be understood that point cloud is a three-dimensional representation of an object's surface. Point cloud (data) of an object's surface can be collected using acquisition devices such as photoelectric radar, lidar, laser scanner, and multi-view camera.
[0070] A point cloud is a set of randomly distributed discrete points in three-dimensional space that represent the spatial structure and surface properties of a three-dimensional object or scene. These points contain geometric information for representing spatial location and attribute information for representing the appearance and texture of the point cloud. Figure 1A Showing 3D point cloud images and Figure 1B The image shows a magnified view of a 3D point cloud, revealing that the point cloud surface is composed of densely distributed points.
[0071] Two-dimensional images contain information at each pixel, and their distribution is regular, so there's no need to record their position information separately. However, the distribution of points in a point cloud in three-dimensional space is random and irregular, so it's necessary to record the position of each point in space to fully represent objects in three-dimensional space. Similar to two-dimensional images, each location during the acquisition process has corresponding attribute information, typically including color and reflectance information. Color information reflects the object's color and is usually represented by RGB; reflectance information reflects the object's surface material and is usually represented by reflectionance. Point cloud data typically consists of geometric information (x, y, z) representing three-dimensional spatial position information, and attribute information such as color information (r, g, b) and reflectance information. For example, reflectance information can be one-dimensional reflectance information (r); color information can be information in any color space, or it can be three-dimensional color information, such as RGB information. Here, R represents red (Red, R), G represents green (Green, G), and B represents blue (Blue, B). For example, color information can be luminance and chromaticity (YCbCr, YUV) information. Here, Y represents luminance (Luma), Cb(U) represents blue color difference, and Cr(V) represents red color difference.
[0072] For example, a point cloud obtained based on laser measurement principles may contain points whose three-dimensional coordinates and reflectance information are included. Similarly, a point cloud obtained based on photogrammetry principles may contain points whose three-dimensional coordinates and three-dimensional color information are included. Furthermore, a point cloud obtained by combining laser measurement and photogrammetry principles may contain points whose three-dimensional coordinates, reflectance information, and three-dimensional color information are included.
[0073] like Figure 2A and Figure 2B The image shown is a point cloud image and its corresponding data storage format. Figure 2AIt provides six viewing dimensions for point cloud images. Figure 2B It consists of a header information section and a data section. The header information includes the data format, data representation type, total number of points in the point cloud, and the content represented by the point cloud. For example, the point cloud file format is ".ply", represented by ASCII code, with a total of 207242 points, and each point has three-dimensional coordinate information (x, y, z) and three-dimensional color information (r, g, b).
[0074] Before providing a further detailed description of the embodiments of this application, the nouns and terms that may be involved in the embodiments of this application will be explained. The nouns and terms involved in the embodiments of this application shall be interpreted as follows:
[0075] Geometric information: A set of three-dimensional (x, y, z) coordinates of a vertex describing the position associated with that of a mesh vertex. The (x, y, z) coordinates representing the position should have defined precision and dynamic range.
[0076] Attribute information: A set of attribute primary colors R, G, B, transformation primary colors Y, U, V, reflectivity, transparency, and / or normals associated with the vertex's three-dimensional (x, y, z) coordinates describing attribute values. Among these, the values of (R, G, B, reflectivity) representing attribute information should have definite precision and dynamic range.
[0077] Point cloud: A collection of discrete data points in space. Points can represent three-dimensional shapes or objects. Each point's location contains its geometric information and, optionally, at least one attribute.
[0078] Quantization is the process of mapping continuous infinite values to a smaller set of discrete finite values.
[0079] Rate-distortion optimization quantization: It is encountered in the source coding of lossy data compression algorithms, and its purpose is to manage distortion within the bit rate range supported by the communication channel or storage medium.
[0080] Arithmetic Encoding (AE);
[0081] Arithmetic Decoding (AD);
[0082] Positively-Occupied Voxels (POV);
[0083] Unoccupied (empty) voxels (NOV);
[0084] Most-Probable Positively-Occupied Voxels (MP-POV);
[0085] Point cloud geometry (PCG);
[0086] Convolutional Neural Networks (CNN);
[0087] Sparse CNN-based Occupancy (SOPA);
[0088] Sparse CNN based Local Neighborhood Embedding (SLNE);
[0089] Deep neural network (DNN);
[0090] Voxel Sampling Layer (VSL);
[0091] Deep Feature Aggregation (DFA);
[0092] Occupancy Output Layer (OOL);
[0093] Occupancy Processing Unit (OPU): A module used for geometric compression in the form of binary octree occupancy codes;
[0094] Attribute Processing Unit (APU): A module used for attribute compression;
[0095] Geometry Coding Unit (GCU): A module used for geometric compression;
[0096] Dyadic Downscaling (DDS);
[0097] Multilayer Perception (MLP);
[0098] Conditional Probability Approximation (CPA);
[0099] Classical residual neural network (Inception ResNet, IRN);
[0100] Neighborhood Point Attention (NPA);
[0101] k-Nearest Neighbor (kNN);
[0102] A voxel is an image of a three-dimensional spatial region constrained by a given size. A voxel has its own nodal coordinates (geometry) in a recognized coordinate system, its own form, and the characteristics (attributes) of the modeling region.
[0103] It should also be understood that a point cloud is a collection of non-uniformly and sparsely distributed points that can be characterized (if applicable) using their three-dimensional coordinates (e.g., (x, y, z)) and attributes (e.g., RGB color, reflectivity, etc.). Unlike the well-constructed pixel grids of two-dimensional image planes or video frames, point clouds rely on the unconstrained displacement of points to flexibly represent three-dimensional objects of arbitrary shapes. However, the difficulty in representing and utilizing the cross-correlation between irregularly scattered points in free three-dimensional space presents problems for the efficient encoding of geometric occupancy.
[0104] Typically, the compression efficiency of points in a point cloud is closely related to the probability approximation of that point conditioned on its available neighbors. The more accurate the context modeling, the closer the probability distribution is to the real data, and the fewer bits are consumed.
[0105] The voxelized representation of PCG is as follows: Figure 3 As shown, (a) is a uniform voxel representation, and (b) is an octree representation. In (a), a densely sampled uniform voxel grid (which can be called a "uniform voxel") is converted into a binary occupancy state (or occupancy state) representation; (b) describes whether the geometric position of the current voxel is occupied (POV, occupancy state = 1) or unoccupied (NOV, occupancy state = 0).
[0106] Non-uniformly distributed point-of-view (POV) inputs to a PCG can be approximated using sparse tensors by caching the geometric coordinates and attribute information of MP-POVs (if applicable), which are generated from previous lower-scale POVs through voxel scaling, such as... Figure 4As shown, (a) represents the sparse tensor, (b) represents binary voxel downsampling, and (c) represents binary voxel upsampling. Scale-sparse tensors are formed within a multi-scale representation framework by progressively scaling down the original PCG to a multi-resolution PCG and then scaling it up accordingly for hierarchical reconstruction, as detailed below. Figure 5 As shown.
[0107] exist Figure 5 In the encoding process, the PCG tensor is regressively reduced, and the occupied voxels at each scale are encoded into a binary bitstream (or "code stream") based on their occupancy probabilities. During decoding, the occupied voxels are reconstructed by decoding the binary bitstream according to their occupancy probabilities.
[0108] Occupancy probability estimation can leverage prior information across and within the same scale and embed it into the encoder and decoder to achieve precise bit-by-bit matching. This occupancy probability approximation based on sparse CNNs is known as "SOPA". Depending on the trade-off between performance and complexity, SOPA can be implemented in a single-stage or multi-stage manner.
[0109] Cross-scale context modeling.
[0110] The occupancy of MP-POVs at each scale is encoded into a compressed bitstream and decoded accordingly to reconstruct POVs at the same scale. In the multi-scale representation model, cross-scale context modeling is performed by approximating the occupancy probability of each MP-POV at the current scale using decoded POVs from previous lower scales, and cross-scale context modeling is limited to between two consecutive scales.
[0111] Sparse convolution.
[0112] 3D sparse convolution can be effectively applied to sparse tensors. 3D sparse convolution is similar to commonly used 3D dense convolution, but it only uses effective MP-POV for convolution, thus making full use of the sparsity properties of point clouds.
[0113] A set of coordinates can be used and associated features This is used to construct sparse tensors. Therefore, sparse convolution can be expressed by the formula:
[0114]
[0115] in, and These are the input and output coordinates. If resolution is maintained, then... and same. and These are the input feature vector and the output feature vector at coordinates u(xu, yu, zu), respectively. A value centered at u is defined in... The method employs a 3D convolutional kernel with an offset of k. By setting different values for k, it obtains a predefined 3D receptive field (e.g., k×k×k or k). 3 Adjacent POVs or MP-POVs within the kernel are used for information aggregation. Wi represents the weight of the kernel.
[0116] Sparse convolution (SConv) or transposed sparse convolution (TSConv) can be formatted using "K, C, S", where K = k × k × k (k 3 S is the receptive field of the 3D convolution; C represents the number of channels; for downscaling, S can be s×s×s↓(s 3 ↓); or for upscaling, S can be s×s×s↑(s 3 ↑). Accordingly, the downscaling (upscaling) operator is associated with SConv (TSConv).
[0117] Voxel sampling layers (VSLs) can be used by, for example Figure 6 The "SConv 2" mentioned 3 C, 2 3 ↓ or TSConv 2 3 C, 2 3 "↑" represents a binary approach to upscaling or downscaling voxels. Deep Feature Aggregation (DFA) is used to characterize and embed information from spatial neighbors within the receptive domain. Occupied Output Layer (OOL) is used for probability or offset derivation.
[0118] Figure 5 The sparse geometric coding framework shown comprises a series of binary voxel samples, for example, S = 2×2×2↑ during upscaling, or S = 2×2×2↓ during downscaling. Binary voxel sampling can be integrated with sparse CNN blocks for better information embedding during resampling. This progressive scaling mechanism allows multi-scale representations to easily support resolution scalability. Since the processing is identical between two adjacent scales, it can be achieved through methods such as... Figure 5The two-scale example shown, from the (i-1)th scale to the ith scale, comprehensively explains the entire sparse PCGC. Each POV from the previous scale is subdivided into 8 MP-POVs, where some MP-POVs can be POVs, and the rest are NOVs. The ground truth is known in the encoder but unknown in the decoder. At each scale, the occupancy state of each MP-POV is compressed in the encoder and reconstructed in the decoder from the upscaled MP-POVs by parsing the compressed bitstream syntax. Thus, the same cross-scale context model is shared between the encoder and decoder, which is generated by SOPA using the decoded POVs from the previous scale to produce the occupancy probabilities pMP-POVs of the MP-POVs in the upscaled tensor for arithmetic coding.
[0119] Network architecture.
[0120] Convolutional layers and (non-linear) activation layers are typically stacked to form sparse CNN blocks, including Figure 6 The voxel sampling layer (VSL) shown and Figure 7 The Deep Feature Aggregation (DFA) and Occupied Output Layer (OOL) are shown. All computations are performed using POV or MP-POV with a sparse distribution in a sparse tensor. Figure 7 For example, (a) represents deep feature aggregation, and (b) represents occupying the output layer. These will be described in detail below.
[0121] (1) Voxel sampling layer (VSL).
[0122] Voxel upscaling is embedded in SOPA, and voxel downscaling is embedded in SLNE to perform cross-scale messaging. For downscaling in SLNE, "SConv 2" is applied. 3 C, 2 3 ↓” to merge 8 spatially connected voxels (e.g., 2×2×2) into one voxel ( Figure 6 This can be simply referred to as "VSL 2". 3 ↓".
[0123] For example, Figure 8 The upscaling process of SOPA is described, in which the application Figure 8 The upscaling "TSConv 2" in SOPA is shown. 3 C, 2 3 The up arrow indicates that the voxel can be subdivided into 8 sub-voxels (or sub-nodes), i.e., VSL 2. 3 ↑. C is an adjustable model parameter representing the number of channels. The upscaling process of multi-level SOPA is as follows: Figure 9 As shown, this is to raise the (i-1)th scale to the i-th scale.
[0124] (2) Deep Feature Aggregation (DFA).
[0125] For most convolutional layers that do not have resolution scaling, apply "SConv k" 3 C” is a feature that aggregates neighboring voxels at the same scale. To characterize the spatial dependencies between voxel neighbors, it can be applied to a deep IRN block consisting of 3 basic IRN units.
[0126] (3) Occupy output layer (OOL).
[0127] The OPA model is then described in detail. For the last Occupied Output Layer (OOL), the OPA model can consist of 3 convolutional layers and 1 activation layer (Sigmoid layer) to derive the occupancy probability p in the range [0, 1] for dense objects using PCG, such as... Figure 7 As shown in (b) of the diagram.
[0128] To compress sparse LiDAR point clouds, the OOL embedded in the SOPA (location) model has its Sigmoid layer removed, and the number of output channels of the last convolutional layer is reduced from 1. Figure 10 ) Extended to 3 ( Figure 11 This allows for the direct derivation of the three coordinate offsets. Among them, Figure 10 An example of probability thresholding for dense object point clouds. Figure 11 Example of Position Offset Adjustment (POA) for sparse LiDAR point clouds.
[0129] Lossless SOPA.
[0130] The compression performance of a learning-based encoder depends on the efficiency of the underlying SOPA engine used for context modeling. Multi-level SOPA iteratively estimates the pMP-POV by leveraging the correlations between previous lower-scale POVs and causal neighbors at the same scale. Causal neighbors have already been processed and used as prior knowledge for processing the current element. Subsequently, the multi-level SOPA architecture primarily relies on permuting MP-POV groups (e.g., Figure 12 The method (from G1 to G8) is used to explore inter-group dependencies. Figure 12 Examples of 1 / 3 / 8 levels are shown, grouping 8 labeled MP-POVs.
[0131] Here, different groups can be arranged to achieve multi-level computation and process groups of elements in the same stage simultaneously. The POV determined in the previous stage is used as a priori for processing MP-POV in subsequent stages, thus effectively revealing neighborhood correlations.
[0132] Level 8 SOPA.
[0133] Considering any set of “8 MP-POVs” sampled from the corresponding POV at the previous scale, an intuitive approach is to classify each element into a separate group, for example, such as Figure 12 The labels G1, G2, ..., G8 are shown. The same labels are applied to all sets of "8 MP-POVs". Therefore, the MP-POV labeled "1" (i.e., G1 MP-POV) will be processed simultaneously in the first level (i.e., "S.1"), followed by the MP-POV labeled "2" (i.e., G2 MP-POV in the second level "S.2"), and so on, until all 8 groups have been traversed.
[0134] To achieve Figure 9 The steps for the 8-level SOPA are as follows:
[0135] Step 1: Upscale all MP-POVs from their corresponding POVs at a lower previous scale using stacked DFA and VSL blocks, and divide all MP-POVs into 8 groups from G1 to G8. Multi-level SOPA is applied to process grouped elements from the first group G1 to the last group G8, and elements within the same group are processed in parallel.
[0136] Step 2: In Level 1, stacked DFA and OOL blocks are used to process the G1 MP-POV to determine the pMP-POV of the G1 MP-POV, so that the real value voxel occupancy information can be compressed into the bitstream in the encoder, or the bitstream can be parsed in the decoder to identify the POV and NOV. It should be noted that the decoded POV and associated features (i.e., ...) are preserved. Figure 9 The dark gray filled "1" in level 2 is used to process MP-POV in subsequent levels and immediately trim NOV, i.e., the top left "1" in level 1 is... Figure 9 NOV to be removed in Level 2.
[0137] Step 3: For Level 2 and the remaining levels, the calculations are the same as for Level 1, but the input data is slightly different. As mentioned above, the POV and its features from previous levels are used as priors for processing MP-POV in subsequent levels to better estimate the occupancy probability.
[0138] SLNE-enhanced SOPA.
[0139] Figure 13 The SLNE between the i-th and (i-1)-th scales is shown. Similar to the SOPA model, the same SLNE is applied to any two adjacent scales. As can be seen, except that the sparse tensor POV is... i Geometric downscaling to POV i-1 In addition, SLNE aggregates the local neighborhood changes of each POV as the feature attribute F of the POV. i-1 Therefore, POVi-1 The occupancy status and characteristic attributes of each POV in the bitstream are compressed into the bitstream.
[0140] The two-scale SLNE model uses 3 DFA blocks and 2 "VSL 2" blocks. 3 ↓ blocks are interleaved for feature downscaling and embedding; correspondingly, a pair of “VSL 2” blocks are interleaved. 3 The "↑" and DFA blocks are used for feature upscaling. Quantization (Q), widely used in learning-based image / video coding, is applied, where uniform noise injection is used during training, and rounding is used during inference. A factorization entropy model is used to compress features. At each scale, the occupancy state of the sparse tensor and feature attributes are encoded and multiplexed separately in the bitstream.
[0141] The decoder receives the reconstructed occupancy state and features of each POV at the (i-1)th scale, where the reconstructed occupancy state and features of each POV at the (i-1)th scale are fed into the SOPA engine to derive the occupancy probability of the MP-POV at the ith scale. Occupancy reconstruction is achieved by encoding and decoding the bitstream associated with the occupancy state using the estimated MP-POV probabilities; correspondingly, feature reconstruction stacks DFA blocks and VSL blocks to increase the decoded feature scale by a factor of 2 across the three axes. Figure 13 As shown, SLNE-related feature attribute processing (e.g., downscaling, encoding, decoding, and upscaling) is contained within two consecutive scales. Based on extensive simulations, staggering the stacking of 3 DFA blocks and 2 VSL blocks before quantization provides a good balance between the compactness and efficiency of the quantized features, optimally estimating the occupancy probability of the MP-POV in the tensor at the i-th scale. The stacking of 1 DFA block and 1 VSL block to upscale the decoded features is intended to match previous lower scales (e.g., Figure 13 The geometric resolution of the sparse tensor at the (i-1)th scale.
[0142] It is detrimental to SOPA.
[0143] (1) Probability threshold of dense point cloud.
[0144] By using Figure 8 The SOPA shown uses OOL blocks to derive the occupancy probability of each MP-POV in lossless mode, while lossy mode sequentially applies probability thresholds t to classify and determine binary occupancy states. For example, as... Figure 10 As shown, if the probability pMP-POV > t, then MP-POV is a POV; otherwise, MP-POV is a NOV. t can be adaptively set according to the number of POVs at each scale.
[0145] (2) Used for position offset adjustment of sparse point clouds.
[0146] Lossy SOPA based on probability thresholds performs well in compressing dense point clouds. To handle sparse point clouds as LiDAR, occupancy position adjustment replaces the occupancy probability approximation in the native SOPA model. To clarify, in Figure 11 In this method, the position adjustment is referred to as "SOPA (Position)". This approach includes a DFA block and a modified OOL block to directly estimate the coordinate offset. In the OOL block, the Sigmoid layer is removed, and the output layer is configured with three channels to generate position offsets for adjustment.
[0147] like Figure 11 As shown, each sparse tensor at the m-th scale is scaled up to the N-th scale in one step (N>m). For tensors at the m-th scale located at (x... m ,y m ,z m Given a POV at position (), its corresponding POV at the Nth scale is located at:
[0148]
[0149] in, As the position offset estimated based on the suggested SOPA (position).
[0150] For example, in Figure 14 This paper describes a unified learning framework that is sequentially applied to lossy and lossless stages. A general architecture is provided here with m scales encoded in lossless mode and (Nm) scales encoded in lossy mode. Figure 15 This is a schematic diagram illustrating an application scenario for a geometric encoder. For example... Figure 15 As shown, (a) is lossless static coding, (b) is lossy static coding, and (c) is lossy dynamic coding.
[0151] For example, Figure 16 A flowchart illustrating the implementation details of OPU. (For example...) Figure 16 As shown, (a) is the lossless OPU, (b) is the lossy OPU, and (c) is the dynamically encoded component. The lossless OPU uses multi-level CPA (MsCPA) to estimate the value from the first group of Os. l g1 to group 8 Os l Os at the subscale level of g8 l The occupancy probability of elements; lossy OPUs use fractal dimension to guide CPR-D to refine shifted POVs, or to guide CPR-V to recover vanished POVs. As an example, the lossless and lossy OPUs mentioned above are statically encoded only by utilizing low-scale spatial priors. By distorting the temporal prior information of lossless and lossy modes, dynamic encoding can be flexibly implemented.
[0152] Lossless OPU.
[0153] For a given geometric tensor Os l 7. Lossless OPU input with lower-scale prior Os l-1 To perform Os for arithmetic coding l The conditional probability approximation (CPA) for each relevant element in the superset, for example,
[0154] P(Os l |Of s l-1 )=p(Os l |Os l-1 ) = CPA(Os l-1 ).
[0155] The rigorous definition of lower-scale priors should be Of s l-1 , because Of s l-1 It can be achieved through decoding and reconstruction using potential compressed noise, for example... Alternatively, it can enhance high-dimensional features learned from spatial neighbors or temporal references, such as Figure 16 (a) and (b) in the example. In lossless mode, Of s l-1 =Os l-1 .
[0156] Os l-1 Each POV in the image is magnified twice to be contained within a local 2×2×2 of 8 MP-POVs (e.g. Figure 16 (a)). Such an MP-POV patch is used with Os l The corresponding 2×2×2 voxel cubes composed of interconnected non-POVs and POVs are geometrically aligned. Each MP-POV represents a probability that can then be used to losslessly compress the corresponding voxel element (e.g., POV or non-POV) or to perform probability determination to classify the MP-POV as POV or non-POV, thereby providing a probability for predictive reconstruction in a lossy mode.
[0157] like Figure 17 As shown, a DNN block can be implemented by simply stacking three classic ResNet (IRN) blocks or three NPFormer (Neighborhood Point Former) blocks. IRN involves sparse convolutions, while NPFormer relies on Neighborhood Point Attention (NPA).
[0158] Such DNN blocks can be used to estimate shift residuals, such as shift residuals in the PR cells of CPR-D, for coordinate refinement;
[0159] This DNN block can also be applied to PM units to generate conditional contexts for compressed quantized latent features Fe^ in FeCPA.
[0160] Binary downsampling (DS↓) is often integrated with DNN blocks to build AT (analysis transform) or extractor units for neighborhood relevance representation and embedding.
[0161] Binary upsampling (US↑) is used to scale s l-1 POV extended to scales below l The corresponding 2×2×2MP-POV patch is shown below (as shown in the CPA model).
[0162] Neighborhood point attention.
[0163] NPA architecture as follows Figure 17 As shown on the right. Assume the input to the NPA layer is composed of coordinates. and characteristics The sparse tensor O is formed by performing a kNN search on each element of O to form a tensor. Then, the tensor is expanded using relative positions, i.e.:
[0164]
[0165] This is called positional embedding.
[0166] Assume Q, K kNN and V kNN These are the query vector, key vector, and value vector, respectively. Quantization representation. From F in Derived from a linear transformation; and By F e The other two separate linear transformations are used for calculation. The weights of these three linear transformations are respectively... and NPA is:
[0167]
[0168] Output {C out ,F out}. C out -C in The same resolution is involved in NPA. The kNN neighborhood and attention mechanism in NPA help the network adaptively utilize local relevance regardless of how the density of the underlying content changes. In contrast, the fixed receptive field setting in sparse convolutions may not include enough (and effective) neighbors, especially for sparse content.
[0169] Currently, related technologies compress individual point clouds or a series of individual point clouds (dynamic point cloud sequences). However, current common methods cannot provide a syntax structure or encoding sequence format that can be used to carry compressed data.
[0170] Based on this, embodiments of this application provide an encoding / decoding method, a method for transmitting a bitstream, an encoder / decoder, and a storage medium. The decoder decodes the bitstream to determine indication information of the current point cloud. The indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on learned AI-PCC geometric information; second syntax element identification information based on learned AI-PCC attribute information; geometric reconstruction values of the current point cloud determined based on the first syntax element identification information; attribute reconstruction values of the current point cloud determined based on the second syntax element identification information; and the current point cloud is reconstructed based on the geometric reconstruction values and attribute reconstruction values. The encoder determines the indication information of the current point cloud and writes it into the bitstream. The indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on learned AI-PCC geometric information; second syntax element identification information based on learned AI-PCC attribute information; the first syntax element identification information is used to determine the geometric reconstruction values of the current point cloud; and the second syntax element identification information is used to determine the attribute reconstruction values of the current point cloud. Therefore, in the embodiments of this application, the indication information transmitted in the bitstream may include first syntax element identification information of geometric information based on learned AI-PCC and / or second syntax element identification information of attribute information based on learned AI-PCC. That is, the technical solution of this application provides a syntax structure or encoding sequence format that can be used to carry compressed data, realizes multiple support for multi-attribute datasets of learning-based artificial intelligence point cloud coding (AI-PCC), and thus can improve the entropy encoding and decoding performance of point clouds.
[0171] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0172] It should be noted that the method in the embodiments of this application can be applied to an encoder, a decoder, or even to both encoders and decoders, but no specific limitation is made here.
[0173] In one embodiment of this application, Figure 18 This is a flowchart illustrating a decoding method provided in an embodiment of this application; as shown below. Figure 18 As shown, the method may include:
[0174] Step 1801: Decode the bitstream to determine the indication information of the current point cloud; wherein, the indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on the geometric information of the learned AI-PCC; second syntax element identification information based on the attribute information of the learned AI-PCC.
[0175] In the embodiments of this application, at the decoding end, the decoder can determine the indication information of the current point cloud by parsing the bit stream; wherein, the indication information of the current point cloud can be used to reconstruct the geometric information and / or attribute information of the current point cloud.
[0176] In the embodiments of this application, the indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on the geometric information of the learned AI-PCC; and second syntax element identification information based on the attribute information of the learned AI-PCC.
[0177] In some embodiments, the first syntax element identification information can be understood as including relevant information and parameters of geometric information used to reconstruct the current point cloud, which includes learning-based artificial intelligence point cloud coding (AI-PCC).
[0178] In some embodiments, the second syntax element identification information can be understood as including relevant information and parameters of the attribute information used to reconstruct the current point cloud based on learning-based AI-PCC.
[0179] In other words, in the embodiments of this application, the determination and / or transmission of multi-attribute datasets based on learning-based AI-PCC are supported.
[0180] In some embodiments, the current point cloud may allow for the reconstruction of geometric and / or attribute information using a learning-based AI-PCC encoding and decoding method.
[0181] For example, in some embodiments, the first syntax element identification information can be determined by parsing the code stream, wherein the first syntax element identification information can be used for the reconstruction of geometric information based on the learned AI-PCC, that is, for the determination of the geometric reconstruction value based on the learned AI-PCC.
[0182] For example, in some embodiments, the second syntax element identification information can be determined by parsing the code stream, wherein the second syntax element identification information can be used for the reconstruction of attribute information based on the learned AI-PCC, that is, for the determination of attribute reconstruction values based on the learned AI-PCC.
[0183] In the embodiments of this application, the order in which the first syntax element identification information and / or the second syntax element identification information are obtained is not specifically limited.
[0184] In embodiments of this application, the first syntax element identification information may be one or more of the following: information in the sequence parameter set (sps); information in the picture parameter set (pps); information in the video parameter set (vps); information in the adaptive parameter set (aps); information in the picture header (ph); information in the slice header (sh); and information in the CTU layer.
[0185] In other words, in the embodiments of this application, the first syntax element identification information can be one or more syntax element information of SPS level, PPS level, VPS level, layer, APS level, PH level, SH level, and CTU level. Accordingly, parsing the corresponding first syntax element identification information can determine the relevant decoding information of the geometric information of the current point cloud based on learning-based AI-PCC.
[0186] In the embodiments of this application, the second syntax element identification information can be one or more of the following: information in SPS; information in PPS; information in VPS; information in APS; information in PH; information in SH; information in the CTU layer.
[0187] In other words, in the embodiments of this application, the second syntax element identification information can be one or more syntax element information of SPS level, PPS level, VPS level, layer, APS level, PH level, SH level, and CTU level. Accordingly, parsing the corresponding second syntax element identification information can determine the relevant decoding information of the learning-based AI-PCC attribute information of the current point cloud.
[0188] In the embodiments of this application, the geometric information of the current point cloud may include, but is not limited to, lossless geometric components and / or lossy geometric components. Lossless geometric components can be understood as lossless components of geometric information, and lossy geometric components can be understood as lossy components of geometric information.
[0189] Step 1802: Determine the geometric reconstruction value of the current point cloud based on the identification information of the first syntax element.
[0190] In the embodiments of this application, after decoding the bitstream and determining the indication information of the current point cloud, the geometric reconstruction value of the current point cloud can be further determined based on the first syntax element identification information included in the indication information.
[0191] In embodiments of this application, the first grammatical element identification information used for reconstructing geometric information based on learned AI-PCC may include at least one or more of the following: first information, second information, third information, fourth information, and fifth information.
[0192] In the embodiments of this application, the first syntax element identification information includes first information. When the value of the first information is a first value, it is determined that the geometric information of the current point cloud has a lossy component; when the value of the first information is a second value, it is determined that the geometric information of the current point cloud does not have a lossy component.
[0193] In other words, in the embodiments of this application, the first information is used to indicate whether there are lossy components in the geometric information of the current point cloud. Or, the first information can be used to indicate whether there are lossy geometric components in the current point cloud.
[0194] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by parsing the values of the syntax elements in the bitstream, the relevant decoding information of the current point cloud can be determined, such as whether there are lossy components in the geometric information of the current point cloud, that is, whether the lossy components of the geometric information of the current point cloud exist in the bitstream.
[0195] In some embodiments, the first syntax element identification information in the bitstream is parsed; wherein the first information included in the first syntax element identification information is used to indicate whether there are lossy components in the geometric information of the current point cloud, that is, to indicate whether there are lossy components in the geometric information of the current point cloud in the bitstream.
[0196] In this embodiment of the application, if the value of the first information is a first value, then the geometric information of the current point cloud has a lossy component; if the value of the first information is a second value, then the geometric information of the current point cloud does not have a lossy component.
[0197] For example, the first value can be set to 1 and the second value can be set to 0; or, the first value can be set to true and the second value can be set to false. This application does not impose any specific limitations.
[0198] For example, in some embodiments, the first information may be represented as oi_lossy_component_present_flag, which is used to indicate whether the lossy component of the geometric information exists in the bitstream, that is, to indicate whether the geometric information has a lossy component (whether there is a lossy geometric component).
[0199] In the embodiments of this application, the first syntax element identification information includes second information. When determining the geometric reconstruction value of the current point cloud based on the first syntax element identification information, if it is determined that the geometric information of the current point cloud has a lossy component, the second information is determined; the length of the lossy component is determined based on the second information; and the geometric reconstruction value of the current point cloud is determined based on the length of the lossy component.
[0200] In other words, in the embodiments of this application, the second information can be used to determine the length of the lossy component of the geometric information of the current point cloud. Here, the lossy component of the geometric information can be understood as the length of the lossy geometric component, and the length of the lossy component of the geometric information can be understood as the length of the lossy geometric component.
[0201] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by parsing the values of the syntax elements in the bitstream, the relevant decoding information of the current point cloud can be determined, for example, the length of the lossy components of the geometric information of the current point cloud.
[0202] In some embodiments, the first syntax element identification information in the parsed bitstream is parsed; wherein the second information included in the first syntax element identification information can be used to determine the length of the lossy component of the geometric information of the current point cloud.
[0203] For example, in some embodiments, the second information can be represented as oi_lossy_component_length_in_bytes, which indicates the length of the lossy component of the geometric information of the current point cloud. The length unit determined based on the value of oi_lossy_component_length_in_bytes can be bytes.
[0204] Of course, the length unit of the lossy component of geometric information is not limited to bytes, but can be other length units, and this application does not make specific limitations.
[0205] In some embodiments, after determining the length of the lossy component based on the second information, the geometric information of the current point cloud can be further reconstructed based on the length of the lossy component. For example, the lossy component of the geometric information of the current point cloud can be reconstructed to finally determine the geometric reconstruction value of the current point cloud.
[0206] In some embodiments, the determination of the second information may depend on the first information. For example, if the first information indicates that the geometric information of the current point cloud contains lossy components, then the second information can be further determined by parsing the bitstream, as shown in Table 1:
[0207] Table 1
[0208]
[0209]
[0210] If the oi_lossy_component_present_flag indicates the presence of a lossy geometric component, then oi_lossy_component_length_in_bytes can be further parsed to obtain it.
[0211] In this embodiment of the application, the first syntax element identification information includes third information. When determining the geometric reconstruction value of the current point cloud based on the first syntax element identification information, the component length of the geometric information of the current point cloud can be determined based on the third information; the geometric reconstruction value of the current point cloud is determined based on the component length of the geometric information.
[0212] In other words, in the embodiments of this application, the third information can be used to determine the component lengths of the geometric information of the current point cloud. Here, the geometric information may include lossy geometric components and / or lossless geometric components.
[0213] For example, in some embodiments, the component lengths of geometric information can also be used to determine the lengths of lossy geometric components and / or lossless geometric components.
[0214] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by parsing the values of the syntax elements in the bitstream, the relevant decoding information of the current point cloud can be determined, for example, the component lengths of the geometric information of the current point cloud.
[0215] In some embodiments, the first syntax element identification information in the parsed bitstream is parsed; wherein the third information included in the first syntax element identification information can be used to determine the component length of the geometric information of the current point cloud.
[0216] For example, in some embodiments, the third information can be represented as oi_length_in_bytes, which indicates the component length of the geometric information of the current point cloud. The length unit determined based on the value of oi_length_in_bytes can be bytes.
[0217] Of course, the length unit of the geometric information component is not limited to bytes, and can be other length units, which are not specifically limited in this application.
[0218] In some embodiments, after determining the component length of the geometric information based on the third information, the geometric information of the current point cloud can be further reconstructed based on the component length of the geometric information to finally determine the geometric reconstruction value of the current point cloud.
[0219] In the embodiments of this application, the first syntax element identification information includes fourth information. When determining the geometric reconstruction value of the current point cloud based on the first syntax element identification information, the number of layers of double-value downsampling of the lossless component of the geometric information of the current point cloud can be determined based on the fourth information; then the geometric reconstruction value of the current point cloud is determined based on the number of layers of double-value downsampling of the lossless component.
[0220] In other words, in the embodiments of this application, the fourth information can be used to determine the number of layers of the lossless geometric components of the current point cloud, that is, to determine the number of layers of the lossless geometric components of the current point cloud.
[0221] For example, in some embodiments, the number of layers of diadic downsampling of lossless geometric components can be understood as the number of layers of diadic downsampling of lossless geometric components.
[0222] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by parsing the values of the syntax elements in the bitstream, the relevant decoding information of the current point cloud can be determined. For example, by determining the number of layers of double-value downsampling of the lossless components of the geometric information of the current point cloud, the number of layers of double-value downsampling of the lossless geometric components of the current point cloud can be determined.
[0223] In some embodiments, the first syntax element identification information in the parsed bitstream is parsed; wherein the fourth information included in the first syntax element identification information can be used to determine the number of layers of double-value downsampling of the lossless components of the geometric information of the current point cloud.
[0224] For example, in some embodiments, the fourth information can be represented as oi_num_lossless_dds_levels_minus1, which indicates the number of layers of double-value downsampling of the lossless components of the geometric information of the current point cloud. The number of layers of double-value downsampling of the lossless geometric components can be determined based on the value of oi_num_lossless_dds_levels_minus1.
[0225] In some embodiments, after determining the number of layers of the lossless components of the geometric information based on the fourth information, the geometric information of the current point cloud can be further reconstructed based on the number of layers of the lossless components of the geometric information to finally determine the geometric reconstruction value of the current point cloud.
[0226] In the embodiments of this application, the first syntax element identification information includes fifth information. When determining the geometric reconstruction value of the current point cloud based on the number of layers of lossless component double-value downsampling, the component length of the i-th layer can be determined first based on the fifth information of the i-th layer of lossless component double-value downsampling; wherein, the value of i is less than or equal to the number of layers of lossless component double-value downsampling; i is an integer greater than 0; then, the geometric reconstruction value of the i-th layer of lossless component double-value downsampling is determined based on the component length of the i-th layer; finally, the geometric reconstruction value of the current point cloud is determined based on the geometric reconstruction value of the i-th layer of lossless component double-value downsampling.
[0227] In other words, in the embodiments of this application, the fifth information can be used to determine the component length of the i-th layer of the lossless downsampling of the geometric information of the current point cloud.
[0228] In the embodiments of this application, i is an integer greater than 0. The value of i is less than or equal to the number of layers in the lossless component's double-value downsampling.
[0229] In some embodiments, after determining the number of layers of lossless downsampling of the geometric information of the current point cloud based on the fourth information, for any layer, such as the i-th layer, the fifth information corresponding to the i-th layer can be determined, thereby determining the component length of the i-th layer based on the fifth information corresponding to the i-th layer.
[0230] For example, in some embodiments, the i-th layer of lossless component double downsampling can be understood as the component of the i-th layer double downsampling, and correspondingly, the component length of the i-th layer of lossless component double downsampling can be understood as the length of the component of the i-th layer double downsampling.
[0231] In this embodiment, some syntax element-based indication information or flags can be written into the bitstream. By parsing the values of these syntax elements in the bitstream, the relevant decoding information of the current point cloud can be determined. For example, determining the length of the i-th layer component of the lossless geometric components of the current point cloud through double-value downsampling can determine the length of the i-th layer double-value downsampled component of the lossless geometric components of the current point cloud.
[0232] In some embodiments, the first syntax element identification information in the parsed bitstream is parsed; wherein the fifth information included in the first syntax element identification information can be used to determine the component length of the i-th layer of the lossless downsampling of the geometric information of the current point cloud.
[0233] For example, in some embodiments, the fifth information can be represented as oi_lossless_dds_level_length_in_bytes, which indicates the component length of the i-th layer of the lossless downsampled binary downsampled geometric components of the current point cloud. The length unit determined based on the value of oi_lossless_dds_level_length_in_bytes can be bytes.
[0234] Of course, the length unit of the i-th layer component length of the lossless component of geometric information is not limited to bytes, but can be other length units, and this application does not make specific limitations.
[0235] In some embodiments, after determining the component length of the i-th layer of the lossless downsampled geometric information based on the fifth information, the i-th layer of the lossless downsampled geometric information can be further reconstructed based on the component length of the i-th layer of the lossless downsampled geometric information, and the geometric reconstruction value of the i-th layer of the lossless downsampled geometric information can be finally determined. Finally, the geometric reconstruction value of the current point cloud can be determined based on the geometric reconstruction value of the i-th layer of the lossless downsampled geometric information.
[0236] In some embodiments, the number of fifth pieces of information can depend on the fourth pieces of information. For example, if the fourth pieces of information indicate that the number of layers of the lossless downsampling of the geometric information of the current point cloud is N1, then the number of fifth pieces of information determined by parsing the bitstream will not exceed N1. That is, for the lossless downsampling of the N1 layers of lossless components, a maximum of N1 pieces of fifth pieces of information can be determined, as shown in Table 2.
[0237] Table 2
[0238] oi_num_lossless_dds_levels_minus1 u(8) for(i = 0; i < oi_num_lossless_dds_levels_minus_1) oi_lossless_dds_level_length_in_bytes[i] ue(v)
[0239] Here, oi_num_lossless_dds_levels_minus1 defines the number of layers in which the lossless geometrically coded components are diadic downsampled. oi_lossless_dds_level_length_in_bytes defines the length of the ith-th layer diadic downsampled component in the geometric bitstream.
[0240] For example, in some embodiments, the parsing of the first syntax element identification information can refer to Table 3:
[0241] Table 3
[0242]
[0243] oi_length_in_bytes defines the length (in bytes) of the geometric components of an AI-PCC encoded frame.
[0244] oi_num_lossless_dds_levels_minus1 defines the number of layers for diadic downsampling of the lossless geometrically encoded components.
[0245] oi_lossless_dds_level_length_in_bytes defines the length (in bytes) of each double-valued downsampled component of the geometric bitstream.
[0246] The oi_lossy_component_present_flag indicates whether the lossy component of the geometric information exists in the encoded bitstream.
[0247] oi_lossy_component_length_in_bytes indicates the length (in bytes) of the lossy geometric component.
[0248] In embodiments of this application, the bitstreams of the geometric information components can be arranged in ascending order, and each of these bitstreams can be decoded independently. The decoding process can terminate at any dynamically downsampled component level.
[0249] In the embodiments of this application, the indication information of the current point cloud can be determined by the encoder's rate-distortion optimization algorithm.
[0250] Step 1803: Determine the attribute reconstruction values of the current point cloud based on the identification information of the second syntax element.
[0251] In the embodiments of this application, after decoding the bitstream and determining the indication information of the current point cloud, the attribute reconstruction value of the current point cloud can be further determined based on the second syntax element identification information included in the indication information.
[0252] In embodiments of this application, the second syntax element identification information used for reconstructing attribute information based on learned AI-PCC may include at least one or more of the following: sixth information, seventh information, eighth information, ninth information, tenth information, eleventh information, twelfth information, thirteenth information, and fourteenth information.
[0253] In the embodiments of this application, the second syntax element identification information includes a sixth piece of information. When the value of the sixth piece of information is the third value, it is determined that the current point cloud has attribute information; when the value of the sixth piece of information is the fourth value, it is determined that the current point cloud does not have attribute information.
[0254] In other words, in the embodiments of this application, the sixth piece of information is used to indicate whether the current point cloud has attribute information. Or, the sixth piece of information can be used to indicate whether the current point cloud has attribute components.
[0255] In this embodiment, some syntax element-based indication information or flag bits can be written into the bitstream. In this way, by parsing the values of the syntax elements in the bitstream, the relevant decoding information of the current point cloud can be determined, such as whether the current point cloud has attribute information, that is, whether the attribute information of the current point cloud exists in the bitstream.
[0256] In some embodiments, the second syntax element identification information in the bitstream is parsed; wherein the sixth information included in the second syntax element identification information is used to indicate whether the current point cloud has attribute information, that is, to indicate whether the attribute information of the current point cloud exists in the bitstream.
[0257] In this embodiment of the application, if the value of the sixth information is the third value, then the current point cloud has attribute information; if the value of the sixth information is the fourth value, then the current point cloud does not have attribute information.
[0258] For example, the third value can be set to 1 and the fourth value can be set to 0; or, the third value can be set to true and the fourth value can be set to false. This application does not impose any specific limitations.
[0259] For example, in some embodiments, the sixth information can be represented as ai_attributes_present_flag, which is used to indicate whether the attribute information exists in the bitstream, that is, to indicate whether the attribute information exists (whether the encoded sequence contains attribute information).
[0260] In the embodiments of this application, the second syntax element identification information includes the seventh information. When determining the attribute reconstruction value of the current point cloud based on the second syntax element identification information, if it is determined that the current point cloud has attribute information, the seventh information is determined; the number of attribute information is determined based on the seventh information; and the attribute reconstruction value of the current point cloud is determined based on the number of attribute information.
[0261] In other words, in the embodiments of this application, the seventh information can be used to determine the number of attribute information of the current point cloud.
[0262] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by parsing the values of the syntax elements in the bitstream, the relevant decoding information of the current point cloud can be determined, such as the number of attribute information of the current point cloud.
[0263] In some embodiments, the second syntax element identification information in the parsed code stream is parsed; wherein the seventh information included in the second syntax element identification information can be used to determine the number of attribute information of the current point cloud.
[0264] For example, in some embodiments, the seventh piece of information can be represented as ai_num_attributes, which indicates the number of attribute information in the current point cloud. The number of attribute information in the current point cloud can be further determined based on the value of ai_num_attributes.
[0265] In some embodiments, after determining the number of attribute information based on the seventh information, the attribute information of the current point cloud can be further reconstructed based on the number of attribute information to finally determine the attribute reconstruction value of the current point cloud.
[0266] In some embodiments, the determination of the seventh information may depend on the sixth information. For example, if the sixth information indicates the existence of attribute information for the current point cloud, then the seventh information can be further determined by parsing the bitstream, as shown in Table 4.
[0267] Table 4
[0268] ai_attributes_present_flag u(1) if(ai_attributes_present_flag){ ai_num_attributes
[0269] If ai_attributes_present_flag indicates the existence of attribute information, then ai_num_attributes can be further parsed to obtain it.
[0270] In the embodiments of this application, the second syntax element identification information may include at least one or more of the following: eighth information, ninth information, tenth information, eleventh information, and twelfth information. When determining the attribute reconstruction value of the current point cloud based on the number of attribute information, the component length of the j-th attribute information is determined based on the eighth information of the j-th attribute information; where the value of j is less than or equal to the number of attribute information; j is an integer greater than 0; and / or, the number of channels of the j-th attribute information is determined based on the ninth information of the j-th attribute information; and / or, the type of the j-th attribute information is determined based on the tenth information of the j-th attribute information; and / or, the model information of the j-th attribute information is determined based on the eleventh information of the j-th attribute information; and / or, the number of layers for double-value downsampling of the j-th attribute information is determined based on the twelfth information of the j-th attribute information; the attribute reconstruction value of the j-th attribute information is determined based on one or more of the component length, the number of channels, the type, the model information, and the number of layers for double-value downsampling of the j-th attribute information; and the attribute reconstruction value of the current point cloud is determined based on the attribute reconstruction value of the j-th attribute information.
[0271] In other words, in the embodiments of this application, one or more of the following information can be determined based on one or more of the eighth, ninth, tenth, eleventh, and twelfth information included in the second syntax element identification information: the component length of the j-th attribute information, the number of channels of the j-th attribute information, the type of the j-th attribute information, the model information of the j-th attribute information, and the number of layers of the dual-value downsampling of the j-th attribute information. Then, the attribute reconstruction value of the j-th attribute information can be further determined based on one or more of the above information. Finally, the attribute reconstruction value of the current point cloud is determined based on the attribute reconstruction value of the j-th attribute information.
[0272] In the embodiments of this application, j is an integer greater than 0. The value of j is less than or equal to the number of determined attribute information items.
[0273] In some embodiments, the eighth information can be used to determine the component length of the j-th attribute information of the current point cloud.
[0274] In some embodiments, after determining the number of attribute information of the current point cloud based on the seventh information, for any attribute information, such as the j-th attribute information, the eighth information corresponding to the j-th attribute information can be determined respectively, thereby determining the component length of the j-th attribute information based on the eighth information corresponding to the j-th attribute information.
[0275] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by parsing the values of the syntax elements in the bitstream, the relevant decoding information of the current point cloud can be determined. For example, by determining the length of the j-th attribute information component of the current point cloud, the length of the attribute component of the current point cloud can be determined.
[0276] In some embodiments, the second syntax element identification information in the parsed bitstream is parsed; wherein the eighth information included in the second syntax element identification information can be used to determine the component length of the j-th attribute information of the current point cloud.
[0277] For example, in some embodiments, the eighth information can be represented as ai_length_in_bytes[j], which indicates the component length of the j-th attribute information of the current point cloud. The length unit determined based on the value of ai_length_in_bytes[j] can be bytes.
[0278] Of course, the length unit of the component length of the j-th attribute information is not limited to bytes, and can also be other length units. This application does not make specific limitations.
[0279] In some embodiments, after determining the component length of the j-th attribute information based on the eighth information, the j-th attribute information of the current point cloud can be reconstructed based on the component length of the j-th attribute information to finally determine the attribute reconstruction value of the j-th attribute information. Finally, the attribute reconstruction value of the current point cloud can be determined based on the attribute reconstruction value of the j-th attribute information.
[0280] In some embodiments, the number of eighth pieces of information may depend on the seventh pieces of information. For example, if the seventh pieces of information indicates that the number of attribute information in the current point cloud is N², then the number of eighth pieces of information determined by parsing the bitstream will not exceed N², that is, corresponding to N² attribute information, a maximum of N² eighth pieces of information can be determined, as shown in Table 5.
[0281] Table 5
[0282] ai_num_attributes for (j = 0; j < ai_num_attributes; j++) { ai_length_in_bytes[j] ue(v)
[0283] Here, ai_num_attributes defines the number of attribute information in the current point cloud. ai_length_in_bytes[j] defines the length (in bytes) of the j-th attribute information.
[0284] In some embodiments, the ninth information can be used to determine the number of channels for the j-th attribute information of the current point cloud.
[0285] In some embodiments, after determining the number of attribute information of the current point cloud based on the seventh information, for any attribute information, such as the j-th attribute information, the ninth information corresponding to the j-th attribute information can be determined respectively, thereby determining the number of channels of the j-th attribute information based on the ninth information corresponding to the j-th attribute information.
[0286] In some embodiments, the attribute information may include, but is not limited to, the following types of attribute information: color, reflectivity, intensity, normal, etc. Each type of attribute information may correspond to a different number of channels, and the number of channels for the attribute information can be determined by the total number of channels for all types of attribute information.
[0287] For example, in some embodiments, it is assumed that the attribute information of the reflectivity type has 1 channel and the attribute information of the color type has 4 channels (including 1 channel for the red component, 1 channel for the green component, 1 channel for the blue component, and 1 channel for the transparency component). In this case, the number of channels of the attribute information can be 5.
[0288] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by parsing the values of the syntax elements in the bitstream, the relevant decoding information of the current point cloud can be determined. For example, by determining the number of channels of the j-th attribute information of the current point cloud, the number of channels of the attribute components of the current point cloud can be determined.
[0289] In some embodiments, the second syntax element identification information in the parsed bitstream is used to determine the number of channels for the j-th attribute information of the current point cloud.
[0290] For example, in some embodiments, the ninth information can be represented as ai_attribute_num_channels[j], which indicates the number of channels for the j-th attribute information of the current point cloud.
[0291] In some embodiments, after determining the number of channels for the j-th attribute information based on the ninth information, the j-th attribute information of the current point cloud can be reconstructed based on the number of channels for the j-th attribute information, and the attribute reconstruction value of the j-th attribute information can be determined. Finally, the attribute reconstruction value of the current point cloud can be determined based on the attribute reconstruction value of the j-th attribute information.
[0292] In some embodiments, the number of ninth pieces of information may depend on the seventh pieces of information. For example, if the seventh pieces of information indicates that the number of attribute information in the current point cloud is N², then the number of ninth pieces of information determined by parsing the bitstream will not exceed N², that is, corresponding to N² attribute information, a maximum of N₁₂ ninth pieces of information can be determined, as shown in Table 6.
[0293] Table 6
[0294] ai_num_attributes for (j = 0; j < ai_num_attributes; j++) { ai_attribute_num_channels[j] ue(v)
[0295] Here, ai_num_attributes defines the number of attribute information for the current point cloud. ai_attribute_num_channels[j] defines the number of channels used for the j-th attribute.
[0296] In some embodiments, the tenth information can be used to determine the type of the j-th attribute information of the current point cloud.
[0297] In some embodiments, after determining the number of attribute information of the current point cloud based on the seventh information, for any attribute information, such as the j-th attribute information, the tenth information corresponding to the j-th attribute information can be determined, thereby determining the type of the j-th attribute information based on the tenth information corresponding to the j-th attribute information.
[0298] In some embodiments, the attribute information may include, but is not limited to, the following types of attribute information: color, reflectivity, intensity, normal, material_id, etc.
[0299] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by parsing the values of the syntax elements in the bitstream, the relevant decoding information of the current point cloud can be determined. For example, by determining the type of the j-th attribute information of the current point cloud, the type of the attribute component of the current point cloud can be determined.
[0300] In some embodiments, the second syntax element identification information in the parsed code stream is parsed; wherein the tenth information included in the second syntax element identification information can be used to determine the type of the j-th attribute information of the current point cloud.
[0301] For example, in some embodiments, the tenth information can be represented as ai_attribute type[j], which indicates the type of the j-th attribute information of the current point cloud.
[0302] In some embodiments, after determining the type of the j-th attribute information based on the tenth information, the j-th attribute information of the current point cloud can be further reconstructed based on the type of the j-th attribute information to finally determine the attribute reconstruction value of the j-th attribute information. Finally, the attribute reconstruction value of the current point cloud can be determined based on the attribute reconstruction value of the j-th attribute information.
[0303] In some embodiments, the number of tenth pieces of information may depend on the seventh pieces of information. For example, if the seventh pieces of information indicates that the number of attribute information in the current point cloud is N², then the number of tenth pieces of information determined by parsing the bitstream will not exceed N², that is, for N² attribute information, a maximum of N² tenth pieces of information can be determined, as shown in Table 7.
[0304] Table 7
[0305] ai_num_attributes for(j=0;j<ai_num_attributes;j++){ ai_attribute type[j] ue(v)
[0306] Here, `ai_num_attributes` defines the number of attribute information for the current point cloud. `ai_attribute type[j]` defines the type of the j-th attribute. The type of attribute information can include color, reflectivity, intensity, normal, material_id, etc.
[0307] In some embodiments, the eleventh information can be used to determine the model information of the j-th attribute information of the current point cloud.
[0308] In some embodiments, after determining the number of attribute information of the current point cloud based on the seventh information, for any attribute information, such as the j-th attribute information, the eleventh information corresponding to the j-th attribute information can be determined, thereby determining the model information of the j-th attribute information based on the eleventh information corresponding to the j-th attribute information.
[0309] In some embodiments, the model information of the attribute information can be used to determine the model used by the attribute information. The model used by the attribute information includes, but is not limited to, a preset model and other models besides the preset model.
[0310] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by parsing the values of the syntax elements in the bitstream, the relevant decoding information of the current point cloud can be determined. For example, by determining the model information of the j-th attribute of the current point cloud, the model information of the attribute components of the current point cloud can be determined.
[0311] In some embodiments, the second syntax element identification information in the parsed code stream is used to parse the model information of the j-th attribute of the current point cloud.
[0312] In the embodiments of this application, the second syntax element identification information includes an eleventh piece of information. When the value of the eleventh piece of information is the fifth value, it is determined that the j-th attribute information uses a preset model; when the value of the eleventh piece of information is the sixth value, it is determined that the j-th attribute information does not use a preset model.
[0313] In other words, in the embodiments of this application, the eleventh information is used to indicate whether the j-th attribute information of the current point cloud uses a preset model.
[0314] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by parsing the values of the syntax elements in the bitstream, the relevant decoding information of the current point cloud can be determined, for example, whether the j-th attribute information of the current point cloud uses a preset model.
[0315] In some embodiments, the second syntax element identification information in the parsed code stream is used to indicate whether the j-th attribute information of the current point cloud uses a preset model.
[0316] In this embodiment, if the eleventh information is the fifth value, then the j-th attribute information of the current point cloud uses the preset model; if the eleventh information is the sixth value, then the j-th attribute information of the current point cloud does not use the preset model.
[0317] For example, the fifth value can be set to 1 and the sixth value can be set to 0; or, the fifth value can be set to true and the sixth value can be set to false. This application does not impose any specific limitations.
[0318] For example, in some embodiments, the eleventh information can be represented as ai_model_override_flag[j], which is used to indicate whether the j-th attribute information of the current point cloud uses a preset model.
[0319] In the embodiments of this application, the type and structure of the preset model are not specifically limited. For example, the preset model can be a model of mst.
[0320] In some embodiments, after determining the model information of the j-th attribute information based on the eleventh information, the j-th attribute information of the current point cloud can be reconstructed based on the model information of the j-th attribute information, and the attribute reconstruction value of the j-th attribute information can be determined. Finally, the attribute reconstruction value of the current point cloud can be determined based on the attribute reconstruction value of the j-th attribute information.
[0321] In some embodiments, the number of eleventh pieces of information can depend on the seventh pieces of information. For example, if the seventh pieces of information indicates that the number of attribute information in the current point cloud is N², then the number of eleventh pieces of information determined by parsing the bitstream will not exceed N², that is, corresponding to N² attribute information, a maximum of eleventh pieces of information can be determined, as shown in Table 8.
[0322] Table 8
[0323] ai_num_attributes for(j=0;j<ai_num_attributes;j++){ ai_model_override_flag[j] ue(v)
[0324] Here, ai_num_attributes defines the number of attribute information for the current point cloud. ai_model_override_flag[j] defines the model information for the j-th attribute.
[0325] In the embodiments of this application, the second syntax element identification information includes thirteenth information. When determining the model information of the j-th attribute information based on the eleventh information of the j-th attribute information, the thirteenth information is determined when it is determined based on the eleventh information that the j-th attribute information does not use a preset model. The model index of the j-th attribute information is determined based on the thirteenth information, and the model information of the j-th attribute information is determined based on the model index.
[0326] In other words, in the embodiments of this application, the thirteenth information can be used to determine the model index of the j-th attribute information of the current point cloud, and then the model indicated by the model index of the j-th attribute information can be determined as the model used by the j-th attribute information, that is, the model information of the j-th attribute information is determined.
[0327] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by parsing the values of the syntax elements in the bitstream, the relevant decoding information of the current point cloud can be determined, for example, the model index of the j-th attribute information of the current point cloud.
[0328] In some embodiments, the second syntax element identification information in the parsed bitstream is used to parse the second syntax element identification information; wherein the thirteenth information included in the second syntax element identification information can be used to determine the model index of the j-th attribute information of the current point cloud.
[0329] For example, in some embodiments, the thirteenth piece of information can be represented as ai_model_index[j], which indicates the index of the model used for the j-th attribute information of the current point cloud. The model index of the j-th attribute information can be determined based on the value of ai_model_index[j].
[0330] In some embodiments, the determination of the thirteenth information may depend on the eleventh information. For example, if the eleventh information indicates that the j-th attribute information of the current point cloud does not use a preset model, then the thirteenth information can be further determined by parsing the bitstream, as shown in Table 9.
[0331] Table 9
[0332] ai_model_override_flag[j] u(1) if(ai_model_override_flag[j]) ai_model_index[j] u(3)
[0333] If ai_model_override_flag[j] indicates that the j-th attribute information does not use the preset model, then ai_model_index[j] can be further parsed to obtain it.
[0334] In some embodiments, the twelfth information can be used to determine the number of layers of double-value downsampling of the j-th attribute information of the current point cloud.
[0335] In some embodiments, after determining the number of attribute information of the current point cloud based on the seventh information, for any attribute information, such as the j-th attribute information, the twelfth information corresponding to the j-th attribute information can be determined, thereby determining the number of layers for double-value downsampling of the j-th attribute information based on the twelfth information corresponding to the j-th attribute information.
[0336] For example, in some embodiments, the number of layers of double-value downsampling of the j-th attribute information can be understood as the number of layers of double-value downsampling of the attribute information.
[0337] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by parsing the values of the syntax elements in the bitstream, the relevant decoding information of the current point cloud can be determined. For example, by determining the number of double-value downsampling layers of the j-th attribute information of the current point cloud, the number of double-value downsampling layers of the attribute components of the current point cloud can be determined.
[0338] In some embodiments, the second syntax element identification information in the parsed bitstream is used to parse the second syntax element identification information; wherein the twelfth information included in the second syntax element identification information can be used to determine the number of layers of double-value downsampling of the j-th attribute information of the current point cloud.
[0339] For example, in some embodiments, the twelfth information may be represented as ai_num_lossless_dds_levels_minus1[j], which indicates the number of layers of double-value downsampling of the j-th attribute information of the current point cloud.
[0340] In some embodiments, after determining the number of double-value downsampling layers of the j-th attribute information based on the twelfth information, the j-th attribute information of the current point cloud can be reconstructed based on the number of double-value downsampling layers of the j-th attribute information, and the attribute reconstruction value of the j-th attribute information can be determined. Finally, the attribute reconstruction value of the current point cloud can be determined based on the attribute reconstruction value of the j-th attribute information.
[0341] In some embodiments, the number of twelfth pieces of information can depend on the seventh piece of information. For example, if the seventh piece of information indicates that the number of attribute information of the current point cloud is N2, then the number of twelfth pieces of information determined by parsing the bitstream will not exceed N2. That is, for N2 attribute information, at most 12th pieces of information can be determined, as shown in Table 10.
[0342] Table 10
[0343] ai_num_attributes for(j=0;j<ai_num_attributes;j++){ ai_num_lossless_dds_levels_minus1[j] ue(v)
[0344] Here, ai_num_attributes defines the number of attribute information in the current point cloud. ai_num_lossless_dds_levels_minus1[j] indicates the number of layers used for double-value downsampling of the j-th attribute information.
[0345] In the embodiments of this application, the second syntax element identification information includes the fourteenth information. When reconstructing the j-th attribute information of the current point cloud based on the number of layers of double-value downsampling based on the j-th attribute information, the component length of the k-th layer is first determined based on the fourteenth information of the k-th layer of double-value downsampling based on the j-th attribute information; then, the attribute reconstruction value of the k-th layer of double-value downsampling based on the component length of the k-th layer is determined; finally, the attribute reconstruction value of the j-th attribute information is determined based on the attribute reconstruction value of the k-th layer of double-value downsampling based on the j-th attribute information.
[0346] In other words, in the embodiments of this application, the fourteenth information can be used to determine the component length of the k-th layer of the double-value downsampling of the j-th attribute information of the current point cloud.
[0347] In the embodiments of this application, k is an integer greater than 0. The value of k is less than or equal to the number of layers in the double-value downsampling of the j-th attribute information.
[0348] In some embodiments, after determining the number of double-valued downsampling layers of the j-th attribute information of the current point cloud based on the twelfth information, for any layer, such as the k-th layer, the fourteenth information corresponding to the k-th layer can be determined, thereby determining the component length of the k-th layer based on the fourteenth information corresponding to the k-th layer.
[0349] For example, in some embodiments, the k-th layer of the double-valued downsampling of the j-th attribute information can be understood as the component of the k-th layer double-valued downsampling, and correspondingly, the component length of the k-th layer of the double-valued downsampling of the j-th attribute information can be understood as the length of the component of the k-th layer double-valued downsampling.
[0350] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by parsing the values of the syntax elements in the bitstream, the relevant decoding information of the current point cloud can be determined. For example, by determining the length of the k-th layer component of the double-value downsampled j-th attribute information of the current point cloud, the length of the k-th layer double-value downsampled component of the double-value downsampled j-th attribute information of the current point cloud can be determined.
[0351] In some embodiments, the second syntax element identification information in the parsed bitstream is used to parse the fourteenth information included in the second syntax element identification information, which can be used to determine the component length of the k-th layer of the double-value downsampling of the j-th attribute information of the current point cloud.
[0352] For example, in some embodiments, the fourteenth information can be represented as ai_lossless_dds_level_length_in_bytes[j][k], which indicates the length of the k-th layer component of the j-th attribute information of the current point cloud after double-value downsampling. The length unit determined based on the value of ai_lossless_dds_level_length_in_bytes[j][k] can be bytes.
[0353] Of course, the length unit of the component length of the k-th layer in the double-value downsampling of the j-th attribute information is not limited to bytes, and can also be other length units. This application does not make specific limitations.
[0354] In some embodiments, after determining the component length of the k-th layer of the double-valued downsampled j-th attribute information based on the fourteenth information, the k-th layer of the double-valued downsampled j-th attribute information of the current point cloud can be reconstructed based on the component length of the k-th layer of the double-valued downsampled j-th attribute information, and finally the attribute reconstruction value of the k-th layer of the double-valued downsampled j-th attribute information can be determined. Finally, the attribute reconstruction value of the j-th attribute information can be determined based on the attribute reconstruction value of the k-th layer of the double-valued downsampled j-th attribute information.
[0355] In some embodiments, the number of fourteenth pieces of information may depend on the twelfth pieces of information. For example, if the twelfth pieces of information indicates that the number of layers of double-value downsampling of the j-th attribute information of the current point cloud is N3, then the number of fourteenth pieces of information determined by parsing the bitstream will not exceed N3. That is, for double-value downsampling of the j-th attribute information corresponding to layer N3, a maximum of N3 fourteenth pieces of information can be determined, as shown in Table 11.
[0356] Table 11
[0357] ai_num_lossless_dds_levels_minus1[j] u(8) for(k=0;k<ai_num_lossless_dds_levels_minus1;k++){ ai_lossless_dds_level_length_in_bytes[j][k]
[0358] Where ai_num_lossless_dds_levels_minus1[j] indicates the number of layers used for double-value downsampling of the j-th encoded information. ai_lossless_dds_level_length_in_bytes[j][k] indicates the length (in bytes) of the k-th layer for double-value downsampling of the j-th encoded information.
[0359] For example, in some embodiments, the parsing of the second syntax element identification information can refer to Table 12:
[0360] Table 12
[0361]
[0362] ai_num_attributes defines the number of the current attributes.
[0363] ai_length_in_bytes[j] defines the length (in bytes) of the j-th encoded attribute.
[0364] ai_attribute_num_channels[j] defines the number of channels used for the j-th attribute (e.g., reflectivity can be 1 channel, color can be 4 channels consisting of red, green, blue and alpha components, etc.).
[0365] ai_attribute type[j] defines the type of the j-th attribute (the type of the attribute can be color, reflectivity, intensity, normal, material_id, etc.).
[0366] ai_model_override_flag[j] indicates a model that is different from the default model mst for frame decoding.
[0367] ai_model_index[j] indicates the index of the NN model used for decoding the j-th attribute component.
[0368] ai_num_lossless_dds_levels_minus1[j] indicates the number of attribute components that are downsampled in the encoded bitstream for the j-th encoded attribute.
[0369] ai_lossless_dds_level_length_in_bytes[j][k] indicates the length (in bytes) of the k-th level component of the double-valued downsampled j-th encoded attribute.
[0370] Step 1804: Reconstruct the current point cloud based on the geometric reconstruction values and / or the attribute reconstruction values of the current point cloud.
[0371] In the embodiments of this application, after the reconstruction of geometric information based on the learned AI-PCC is completed by the first syntax element identification information obtained by parsing the code stream, and / or the reconstruction of attribute information based on the learned AI-PCC is completed by the second syntax element identification information obtained by parsing the code stream, the current point cloud can be further reconstructed based on the geometric reconstruction value and / or the attribute reconstruction value of the current point cloud.
[0372] In summary, the embodiments of this application propose various supports for multi-attribute datasets of learning-based artificial intelligence point cloud coding (AI-PCC). The encoding and decoding methods proposed in this application can be used in AI-PCC coding standards such as AVS and MPEG. When implementing the encoding and decoding methods proposed in this application, modifications can be made to the bitstream structure, syntax, constraints, and mappings used to generate decoded point clouds for standardization.
[0373] This application provides a decoding method in which a decoder decodes the bitstream to determine the indication information of the current point cloud. The indication information of the current point cloud includes at least one or more of the following: a first syntax element identifier based on the geometric information of the learned AI-PCC; a second syntax element identifier based on the attribute information of the learned AI-PCC; a geometric reconstruction value of the current point cloud determined based on the first syntax element identifier; an attribute reconstruction value of the current point cloud determined based on the second syntax element identifier; and the current point cloud reconstructed based on the geometric reconstruction value and the attribute reconstruction value. Therefore, in the embodiments of this application, the indication information transmitted in the bitstream may include the first syntax element identifier based on the geometric information of the learned AI-PCC and / or the second syntax element identifier based on the attribute information of the learned AI-PCC. That is, the technical solution of this application provides a syntax structure or encoding sequence format that can be used to carry compressed data, achieving multiple supports for multi-attribute datasets of AI-PCC, thereby improving the entropy encoding and decoding performance of point clouds.
[0374] In one embodiment of this application, Figure 19 This is a flowchart illustrating an encoding method provided in an embodiment of this application; as shown below. Figure 19 As shown, the method may include:
[0375] Step 1901: Determine the indication information of the current point cloud and write the indication information of the current point cloud into the code stream; wherein, the indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on the geometric information of the learned AI-PCC; second syntax element identification information based on the attribute information of the learned AI-PCC.
[0376] In the embodiments of this application, at the encoding end, the encoder can write the determined indication information of the current point cloud into the bit stream; wherein, the indication information of the current point cloud can be determined by the encoder through a rate-distortion optimization algorithm, and the indication information of the current point cloud can be used for the reconstruction of the geometric information and / or attribute information of the current point cloud.
[0377] In the embodiments of this application, the indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on the geometric information of the learned AI-PCC; and second syntax element identification information based on the attribute information of the learned AI-PCC.
[0378] In some embodiments, the first syntax element identification information can be understood as including relevant information and parameters of geometric information used to reconstruct the current point cloud, which includes learning-based artificial intelligence point cloud coding (AI-PCC).
[0379] In some embodiments, the second syntax element identification information can be understood as including relevant information and parameters of the attribute information used to reconstruct the current point cloud based on learning-based AI-PCC.
[0380] In other words, in the embodiments of this application, the determination and / or transmission of multi-attribute datasets based on learning-based AI-PCC are supported.
[0381] In some embodiments, the current point cloud may allow for the reconstruction of geometric and / or attribute information using a learning-based AI-PCC encoding and decoding method.
[0382] For example, in some embodiments, the first syntax element identification information can be used for the reconstruction of geometric information based on the learned AI-PCC, i.e. for the determination of geometric reconstruction values based on the learned AI-PCC.
[0383] For example, in some embodiments, the second syntax element identification information can be used for the reconstruction of attribute information based on the learned AI-PCC, that is, for the determination of attribute reconstruction values based on the learned AI-PCC.
[0384] In the embodiments of this application, the encoding order of the first syntax element identification information and / or the second syntax element identification information is not specifically limited.
[0385] In the embodiments of this application, the first syntax element identification information can be one or more of the following: information in SPS; information in PPS; information in VPS; information in APS; information in PH; information in SH; information in the CTU layer.
[0386] In other words, in the embodiments of this application, the first syntax element identification information can be one or more syntax element information of SPS level, PPS level, VPS level, layer, APS level, PH level, SH level, and CTU level. Accordingly, the encoded first syntax element identification information can indicate the relevant encoded information of the geometric information of the current point cloud based on the learning-based AI-PCC.
[0387] In the embodiments of this application, the second syntax element identification information can be one or more of the following: information in SPS; information in PPS; information in VPS; information in APS; information in PH; information in SH; information in the CTU layer.
[0388] In other words, in the embodiments of this application, the second syntax element identification information can be one or more syntax element information of SPS level, PPS level, VPS level, layer, APS level, PH level, SH level, and CTU level. Accordingly, the encoded second syntax element identification information can indicate the relevant encoded information of the learning-based AI-PCC attribute information of the current point cloud.
[0389] In the embodiments of this application, the geometric information of the current point cloud may include, but is not limited to, lossless geometric components and / or lossy geometric components. Lossless geometric components can be understood as lossless components of geometric information, and lossy geometric components can be understood as lossy components of geometric information.
[0390] In embodiments of this application, the first grammatical element identification information used for reconstructing geometric information based on learned AI-PCC may include at least one or more of the following: first information, second information, third information, fourth information, and fifth information.
[0391] In the embodiments of this application, the first syntax element identification information includes first information. When the geometric information of the current point cloud has lossy components, the value of the first information is a first value; when the geometric information of the current point cloud does not have lossy components, the value of the first information is a second value.
[0392] In other words, in the embodiments of this application, the first information is used to indicate whether there are lossy components in the geometric information of the current point cloud. Or, the first information can be used to indicate whether there are lossy geometric components in the current point cloud.
[0393] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by taking the values of the syntax elements in the bitstream, the relevant encoding information of the current point cloud can be indicated. For example, it can be determined whether there are lossy components in the geometric information of the current point cloud, that is, whether the lossy components of the geometric information of the current point cloud exist in the bitstream.
[0394] In some embodiments, the first information included in the first syntax element identification information is used to indicate whether there are lossy components in the geometric information of the current point cloud, that is, to indicate whether there are lossy components in the geometric information of the current point cloud in the bitstream.
[0395] In this embodiment of the application, if the value of the first information is a first value, then the geometric information of the current point cloud has a lossy component; if the value of the first information is a second value, then the geometric information of the current point cloud does not have a lossy component.
[0396] For example, the first value can be set to 1 and the second value can be set to 0; or, the first value can be set to true and the second value can be set to false. This application does not impose any specific limitations.
[0397] For example, in some embodiments, the first information may be represented as oi_lossy_component_present_flag, which is used to indicate whether the lossy component of the geometric information exists in the bitstream, that is, to indicate whether the geometric information has a lossy component (whether there is a lossy geometric component).
[0398] In the embodiments of this application, the first syntax element identification information includes second information. When it is determined that the geometric information of the current point cloud has a lossy component, the second information is determined based on the length of the lossy component.
[0399] In other words, in the embodiments of this application, the second information can be used to indicate the length of the lossy component of the geometric information of the current point cloud. Here, the lossy component of the geometric information can be understood as the length of the lossy geometric component, and the length of the lossy component of the geometric information can be understood as the length of the lossy geometric component.
[0400] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by taking the values of the syntax elements in the bitstream, the relevant encoding information of the current point cloud can be indicated, for example, the length of the lossy components of the geometric information of the current point cloud can be determined.
[0401] In some embodiments, the second information included in the first syntax element identification information can be used to indicate the length of the lossy component of the geometric information of the current point cloud.
[0402] For example, in some embodiments, the second information can be represented as oi_lossy_component_length_in_bytes, which indicates the length of the lossy component of the geometric information of the current point cloud. The length unit determined based on the value of oi_lossy_component_length_in_bytes can be bytes.
[0403] Of course, the length unit of the lossy component of geometric information is not limited to bytes, but can be other length units, and this application does not make specific limitations.
[0404] In some embodiments, after determining the length of the lossy component, the geometric information of the current point cloud can be further reconstructed based on the length of the lossy component. For example, the lossy component of the geometric information of the current point cloud can be reconstructed to finally determine the geometric reconstruction value of the current point cloud.
[0405] In some embodiments, the determination of the second information may depend on the first information. For example, if the first information indicates that the geometric information of the current point cloud contains lossy components, then the second information can be further written into the bitstream, such as in Table 1. In the embodiments of this application, the first syntax element identification information includes third information, which is determined based on the component lengths of the geometric information of the current point cloud.
[0406] In other words, in the embodiments of this application, the third information can be used to indicate the component lengths of the geometric information of the current point cloud. Here, the geometric information may include lossy geometric components and / or lossless geometric components.
[0407] For example, in some embodiments, the component lengths of geometric information can also be used to determine the lengths of lossy geometric components and / or lossless geometric components.
[0408] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by taking the values of the syntax elements in the bitstream, the relevant encoding information of the current point cloud can be indicated, for example, the component lengths of the geometric information of the current point cloud can be determined.
[0409] In some embodiments, the third information included in the first syntax element identification information can be used to indicate the component length of the geometric information of the current point cloud.
[0410] For example, in some embodiments, the third information can be represented as oi_length_in_bytes, which indicates the component length of the geometric information of the current point cloud. The length unit determined based on the value of oi_length_in_bytes can be bytes.
[0411] Of course, the length unit of the geometric information component is not limited to bytes, and can be other length units, which are not specifically limited in this application.
[0412] In some embodiments, after determining the component lengths of the geometric information, the geometric information of the current point cloud can be further reconstructed based on the component lengths of the geometric information to finally determine the geometric reconstruction value of the current point cloud.
[0413] In the embodiments of this application, the first syntax element identification information includes fourth information, which is determined by the number of layers of double-value downsampling of the lossless components of the geometric information of the current point cloud.
[0414] In other words, in the embodiments of this application, the fourth information can be used to indicate the number of layers of the lossless geometric components of the current point cloud, that is, it can indicate the number of layers of the lossless geometric components of the current point cloud.
[0415] For example, in some embodiments, the number of layers of diadic downsampling of lossless geometric components can be understood as the number of layers of diadic downsampling of lossless geometric components.
[0416] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, the values of the syntax elements in the bitstream can indicate the relevant encoding information of the current point cloud. For example, determining the number of layers of double downsampling of the lossless geometric components of the current point cloud can determine the number of layers of double downsampling of the lossless geometric components of the current point cloud.
[0417] In some embodiments, the fourth information included in the first syntax element identification information can be used to indicate the number of layers of the lossless downsampling of the geometric information of the current point cloud.
[0418] For example, in some embodiments, the fourth information can be represented as oi_num_lossless_dds_levels_minus1, which indicates the number of layers of double-value downsampling of the lossless components of the geometric information of the current point cloud. The number of layers of double-value downsampling of the lossless geometric components can be determined based on the value of oi_num_lossless_dds_levels_minus1.
[0419] In some embodiments, after determining the number of layers of double-value downsampling of the lossless components of the geometric information, the geometric information of the current point cloud can be further reconstructed based on the number of layers of double-value downsampling of the lossless components of the geometric information, and the geometric reconstruction value of the current point cloud can be finally determined.
[0420] In the embodiments of this application, the first syntax element identification information includes fifth information, which is determined based on the component length of the i-th layer of the lossless component's double-value downsampling; wherein, the value of i is less than or equal to the number of layers of the lossless component's double-value downsampling; and i is an integer greater than 0.
[0421] In other words, in the embodiments of this application, the fifth information can be used to indicate the component length of the i-th layer of the lossless downsampling of the geometric information of the current point cloud.
[0422] In the embodiments of this application, i is an integer greater than 0. The value of i is less than or equal to the number of layers in the lossless component's double-value downsampling.
[0423] In some embodiments, after determining the number of layers of lossless downsampling of the geometric information of the current point cloud, for any one of the layers, such as the i-th layer, the fifth information corresponding to the i-th layer can be determined based on the component length of the i-th layer.
[0424] For example, in some embodiments, the i-th layer of lossless component double downsampling can be understood as the component of the i-th layer double downsampling, and correspondingly, the component length of the i-th layer of lossless component double downsampling can be understood as the length of the component of the i-th layer double downsampling.
[0425] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, the relevant encoding information of the current point cloud can be indicated by the value of the syntax elements in the bitstream. For example, by determining the length of the i-th layer of the lossless geometric components of the current point cloud, the length of the i-th layer of the lossless geometric components of the current point cloud can be determined.
[0426] In some embodiments, the fifth information included in the first syntax element identification information can be used to indicate the component length of the i-th layer of the lossless downsampling of the geometric information of the current point cloud.
[0427] For example, in some embodiments, the fifth information can be represented as oi_lossless_dds_level_length_in_bytes, which indicates the component length of the i-th layer of the lossless downsampled binary downsampled geometric components of the current point cloud. The length unit determined based on the value of oi_lossless_dds_level_length_in_bytes can be bytes.
[0428] Of course, the length unit of the i-th layer component length of the lossless component of geometric information is not limited to bytes, but can be other length units, and this application does not make specific limitations.
[0429] In some embodiments, after determining the component length of the i-th layer of the lossless downsampling of the geometric information, the i-th layer of the lossless downsampling of the geometric information can be further reconstructed based on the component length of the i-th layer of the lossless downsampling of the geometric information, and finally the geometric reconstruction value of the i-th layer of the lossless downsampling of the geometric information can be determined. Finally, the geometric reconstruction value of the current point cloud can be determined based on the geometric reconstruction value of the i-th layer of the lossless downsampling of the geometric information.
[0430] In some embodiments, the number of fifth pieces of information may depend on the fourth pieces of information. For example, if the fourth pieces of information indicates that the number of layers of the lossless downsampling of the geometric information of the current point cloud is N1, then the number of fifth pieces of information will not exceed N1. That is, for the lossless downsampling of the N1 layers, a maximum of N1 pieces of fifth pieces of information can be determined, as shown in Table 2.
[0431] For example, in some embodiments, the encoding of the first syntax element identification information can be referred to Table 3.
[0432] In embodiments of this application, the bitstreams of the geometric information components can be arranged in ascending order, and each of these bitstreams can be encoded independently. The encoding process can terminate at any dynamically downsampled component level.
[0433] In the embodiments of this application, the indication information of the current point cloud can be determined by the encoder's rate-distortion optimization algorithm.
[0434] In embodiments of this application, the second syntax element identification information used for reconstructing attribute information based on learned AI-PCC may include at least one or more of the following: sixth information, seventh information, eighth information, ninth information, tenth information, eleventh information, twelfth information, thirteenth information, and fourteenth information.
[0435] In the embodiments of this application, the second syntax element identification information includes a sixth information. When the current point cloud has attribute information, the value of the sixth information is a third value; when the current point cloud does not have attribute information, the value of the sixth information is a fourth value.
[0436] In other words, in the embodiments of this application, the sixth piece of information is used to indicate whether the current point cloud has attribute information. Or, the sixth piece of information can be used to indicate whether the current point cloud has attribute components.
[0437] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by taking the values of the syntax elements in the bitstream, the relevant encoding information of the current point cloud can be indicated, for example, determining whether the current point cloud has attribute information, that is, determining whether the attribute information of the current point cloud exists in the bitstream.
[0438] In some embodiments, the sixth information included in the second syntax element identification information is used to indicate whether the current point cloud has attribute information, that is, to indicate whether the attribute information of the current point cloud exists in the bitstream.
[0439] In this embodiment of the application, if the value of the sixth information is the third value, then the current point cloud has attribute information; if the value of the sixth information is the fourth value, then the current point cloud does not have attribute information.
[0440] For example, the third value can be set to 1 and the fourth value can be set to 0; or, the third value can be set to true and the fourth value can be set to false. This application does not impose any specific limitations.
[0441] For example, in some embodiments, the sixth information can be represented as ai_attributes_present_flag, which is used to indicate whether the attribute information exists in the bitstream, that is, to indicate whether the attribute information exists (whether the encoded sequence contains attribute information).
[0442] In the embodiments of this application, the second syntax element identification information includes the seventh information, which is determined based on the number of attribute information when it is determined that the current point cloud has attribute information.
[0443] In other words, in the embodiments of this application, the seventh information can be used to indicate the number of attribute information of the current point cloud.
[0444] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by taking the values of the syntax elements in the bitstream, the relevant encoding information of the current point cloud can be indicated, for example, the number of attribute information of the current point cloud can be determined.
[0445] In some embodiments, the seventh information included in the second syntax element identification information can be used to indicate the number of attribute information of the current point cloud.
[0446] For example, in some embodiments, the seventh piece of information can be represented as ai_num_attributes, which indicates the number of attribute information in the current point cloud. The number of attribute information in the current point cloud can be further determined based on the value of ai_num_attributes.
[0447] In some embodiments, after determining the number of attribute information, the attribute information of the current point cloud can be further reconstructed based on the number of attribute information to finally determine the attribute reconstruction value of the current point cloud.
[0448] In some embodiments, the determination of the seventh information may depend on the sixth information. For example, if the sixth information indicates that the attribute information of the current point cloud exists, then the seventh information can be further written into the bitstream, as shown in Table 4.
[0449] In embodiments of this application, the second syntax element identification information may include at least one or more of the following: eighth information, ninth information, tenth information, eleventh information, and twelfth information. The eighth information of the j-th attribute information is determined based on the component length of the j-th attribute information; wherein the value of j is less than or equal to the number of attribute information; j is an integer greater than 0; and / or,
[0450] Based on the number of channels of the j-th attribute information, determine the ninth information of the j-th attribute information; and / or, based on the type of the j-th attribute information, determine the tenth information of the j-th attribute information; and / or, based on the model information of the j-th attribute information, determine the eleventh information of the j-th attribute information; and / or, based on the number of layers of double-value downsampling of the j-th attribute information, determine the twelfth information of the j-th attribute information.
[0451] In the embodiments of this application, the attribute reconstruction value of the j-th attribute information can be determined based on one or more of the following information: the component length of the j-th attribute information, the number of channels of the j-th attribute information, the type of the j-th attribute information, the model information of the j-th attribute information, and the number of layers of the dual-value downsampling of the j-th attribute information; finally, the attribute reconstruction value of the current point cloud is determined based on the attribute reconstruction value of the j-th attribute information.
[0452] In the embodiments of this application, j is an integer greater than 0. The value of j is less than or equal to the number of determined attribute information items.
[0453] In some embodiments, the eighth information can be used to indicate the component length of the j-th attribute information of the current point cloud.
[0454] In some embodiments, after determining the number of attribute information in the current point cloud, for any attribute information, such as the j-th attribute information, the eighth information corresponding to the j-th attribute information can be determined based on the component length of the j-th attribute information.
[0455] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, the relevant encoding information of the current point cloud can be indicated by the value of the syntax elements in the bitstream. For example, by determining the length of the j-th attribute information component of the current point cloud, the length of the attribute component of the current point cloud can be determined.
[0456] In some embodiments, the eighth information included in the second syntax element identification information can be used to indicate the component length of the j-th attribute information of the current point cloud.
[0457] For example, in some embodiments, the eighth information can be represented as ai_length_in_bytes[j], which indicates the component length of the j-th attribute information of the current point cloud. The length unit determined based on the value of ai_length_in_bytes[j] can be bytes.
[0458] Of course, the length unit of the component length of the j-th attribute information is not limited to bytes, and can also be other length units. This application does not make specific limitations.
[0459] In some embodiments, after determining the component length of the j-th attribute information, the j-th attribute information of the current point cloud can be reconstructed based on the component length of the j-th attribute information, and the attribute reconstruction value of the j-th attribute information can be determined. Finally, the attribute reconstruction value of the current point cloud can be determined based on the attribute reconstruction value of the j-th attribute information.
[0460] In some embodiments, the number of eighth pieces of information may depend on the seventh pieces of information. For example, if the seventh pieces of information indicates that the number of attribute information of the current point cloud is N2, then the number of eighth pieces of information will not exceed N2. That is, for N2 attribute information, at most 8 pieces of eighth information can be determined, as shown in Table 5.
[0461] In some embodiments, the ninth information can be used to indicate the number of channels for the j-th attribute information of the current point cloud.
[0462] In some embodiments, after determining the number of attribute information in the current point cloud, for any attribute information, such as the j-th attribute information, the ninth information corresponding to the j-th attribute information can be determined based on the number of channels of the j-th attribute information.
[0463] In some embodiments, the attribute information may include, but is not limited to, the following types of attribute information: color, reflectivity, intensity, normal, etc. Each type of attribute information may correspond to a different number of channels, and the number of channels for the attribute information can be determined by the total number of channels for all types of attribute information.
[0464] For example, in some embodiments, it is assumed that the attribute information of the reflectivity type has 1 channel and the attribute information of the color type has 4 channels (including 1 channel for the red component, 1 channel for the green component, 1 channel for the blue component, and 1 channel for the transparency component). In this case, the number of channels of the attribute information can be 5.
[0465] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, the relevant encoding information of the current point cloud can be indicated by the value of the syntax elements in the bitstream. For example, by determining the number of channels of the j-th attribute information of the current point cloud, the number of channels of the attribute component of the current point cloud can be determined.
[0466] In some embodiments, the ninth information included in the second syntax element identification information can be used to indicate the number of channels of the j-th attribute information of the current point cloud.
[0467] For example, in some embodiments, the ninth information can be represented as ai_attribute_num_channels[j], which indicates the number of channels for the j-th attribute information of the current point cloud.
[0468] In some embodiments, after determining the number of channels for the j-th attribute information, the j-th attribute information of the current point cloud can be reconstructed based on the number of channels for the j-th attribute information, and the attribute reconstruction value of the j-th attribute information can be determined. Finally, the attribute reconstruction value of the current point cloud can be determined based on the attribute reconstruction value of the j-th attribute information.
[0469] In some embodiments, the number of ninth pieces of information may depend on the seventh pieces of information. For example, if the seventh piece of information indicates that the number of attribute information of the current point cloud is N2, then the number of ninth pieces of information shall not exceed N2. That is, for N2 attribute information, at most a number of ninth pieces of information can be determined, as shown in Table 6.
[0470] In some embodiments, the tenth information can be used to indicate the type of the j-th attribute information of the current point cloud.
[0471] In some embodiments, after determining the number of attribute information in the current point cloud, for any attribute information, such as the j-th attribute information, the tenth information corresponding to the j-th attribute information can be determined based on the type of the j-th attribute information.
[0472] In some embodiments, the attribute information may include, but is not limited to, the following types of attribute information: color, reflectivity, intensity, normal, material_id, etc.
[0473] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, the values of the syntax elements in the bitstream can indicate the relevant encoding information of the current point cloud. For example, by determining the type of the j-th attribute information of the current point cloud, the type of the attribute component of the current point cloud can be determined.
[0474] In some embodiments, the tenth information included in the second syntax element identification information can be used to indicate the type of the j-th attribute information of the current point cloud.
[0475] For example, in some embodiments, the tenth information can be represented as ai_attribute type[j], which indicates the type of the j-th attribute information of the current point cloud.
[0476] In some embodiments, after determining the type of the j-th attribute information based on the tenth information, the j-th attribute information of the current point cloud can be further reconstructed based on the type of the j-th attribute information to finally determine the attribute reconstruction value of the j-th attribute information. Finally, the attribute reconstruction value of the current point cloud can be determined based on the attribute reconstruction value of the j-th attribute information.
[0477] In some embodiments, the number of tenth pieces of information may depend on the seventh pieces of information. For example, if the seventh piece of information indicates that the number of attribute information of the current point cloud is N2, then the number of tenth pieces of information will not exceed N2. That is, for N2 attribute information, at most 10th pieces of information can be determined, as shown in Table 7.
[0478] In some embodiments, the eleventh information can be used to indicate model information for the j-th attribute of the current point cloud.
[0479] In some embodiments, after determining the number of attribute information of the current point cloud, for any attribute information, such as the j-th attribute information, the eleventh information corresponding to the j-th attribute information can be determined, thereby determining the model information of the j-th attribute information based on the eleventh information corresponding to the j-th attribute information.
[0480] In some embodiments, the model information of the attribute information can be used to determine the model used by the attribute information. The model used by the attribute information includes, but is not limited to, a preset model and other models besides the preset model.
[0481] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, the relevant encoding information of the current point cloud can be indicated by the value of the syntax elements in the bitstream. For example, determining the model information of the j-th attribute of the current point cloud can determine the model information of the attribute components of the current point cloud.
[0482] In some embodiments, the eleventh information included in the second syntax element identification information can be used to indicate model information of the j-th attribute information of the current point cloud.
[0483] In the embodiments of this application, the second syntax element identification information includes an eleventh information. When the j-th attribute information uses a preset model, the value of the eleventh information is the fifth value; when the j-th attribute information does not use a preset model, the value of the eleventh information is the sixth value.
[0484] In other words, in the embodiments of this application, the eleventh information is used to indicate whether the j-th attribute information of the current point cloud uses a preset model.
[0485] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by taking the values of the syntax elements in the bitstream, the relevant encoding information of the current point cloud can be indicated, for example, determining whether the j-th attribute information of the current point cloud uses a preset model.
[0486] In some embodiments, the eleventh information included in the second syntax element identification information is used to indicate whether the j-th attribute information of the current point cloud uses a preset model.
[0487] In this embodiment, if the eleventh information is the fifth value, then the j-th attribute information of the current point cloud uses the preset model; if the eleventh information is the sixth value, then the j-th attribute information of the current point cloud does not use the preset model.
[0488] For example, the fifth value can be set to 1 and the sixth value can be set to 0; or, the fifth value can be set to true and the sixth value can be set to false. This application does not impose any specific limitations.
[0489] For example, in some embodiments, the eleventh information can be represented as ai_model_override_flag[j], which is used to indicate whether the j-th attribute information of the current point cloud uses a preset model.
[0490] In the embodiments of this application, the type and structure of the preset model are not specifically limited. For example, the preset model can be a model of mst.
[0491] In some embodiments, after determining the model information of the j-th attribute information, the j-th attribute information of the current point cloud can be further reconstructed based on the model information of the j-th attribute information, and the attribute reconstruction value of the j-th attribute information can be determined. Finally, the attribute reconstruction value of the current point cloud can be determined based on the attribute reconstruction value of the j-th attribute information.
[0492] In some embodiments, the number of eleventh information may depend on the seventh information. For example, if the seventh information indicates that the number of attribute information of the current point cloud is N2, then the number of eleventh information shall not exceed N2, that is, for N2 attribute information, at most eleventh information can be determined, as shown in Table 8.
[0493] In the embodiments of this application, the second syntax element identification information includes thirteenth information. When it is determined that the j-th attribute information does not use a preset model, the model index of the j-th attribute information is determined based on the model information of the j-th attribute information; the thirteenth information is determined based on the model index.
[0494] In other words, in the embodiments of this application, the thirteenth information can be used to indicate the model index of the j-th attribute information of the current point cloud, wherein the model index of the j-th attribute information can be determined according to the model used by the j-th attribute information.
[0495] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, by taking the values of the syntax elements in the bitstream, the relevant encoding information of the current point cloud can be indicated, for example, determining the model index of the j-th attribute information of the current point cloud.
[0496] In some embodiments, the thirteenth information included in the second syntax element identification information can be used to indicate the model index of the j-th attribute information of the current point cloud.
[0497] For example, in some embodiments, the thirteenth piece of information can be represented as ai_model_index[j], which indicates the index of the model used for the j-th attribute information of the current point cloud. The model index of the j-th attribute information can be determined based on the value of ai_model_index[j].
[0498] In some embodiments, the determination of the thirteenth information may depend on the eleventh information. For example, if the eleventh information indicates that the j-th attribute information of the current point cloud does not use a preset model, then the thirteenth information can be further written into the bitstream, as shown in Table 9.
[0499] In some embodiments, the twelfth information may be used to indicate the number of layers of double-value downsampling of the j-th attribute information of the current point cloud.
[0500] In some embodiments, after determining the number of attribute information in the current point cloud, for any attribute information, such as the j-th attribute information, the twelfth information corresponding to the j-th attribute information can be determined based on the number of layers of double-value downsampling of the j-th attribute information.
[0501] For example, in some embodiments, the number of layers of double-value downsampling of the j-th attribute information can be understood as the number of layers of double-value downsampling of the attribute information.
[0502] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, the relevant encoding information of the current point cloud can be indicated by the value of the syntax elements in the bitstream. For example, by determining the number of double-value downsampling layers of the j-th attribute information of the current point cloud, the number of double-value downsampling layers of the attribute components of the current point cloud can be determined.
[0503] In some embodiments, the twelfth information included in the second syntax element identification information can be used to indicate the number of layers of double-value downsampling of the j-th attribute information of the current point cloud.
[0504] For example, in some embodiments, the twelfth information may be represented as ai_num_lossless_dds_levels_minus1[j], which indicates the number of layers of double-value downsampling of the j-th attribute information of the current point cloud.
[0505] In some embodiments, after determining the number of layers of double-value downsampling of the j-th attribute information, the j-th attribute information of the current point cloud can be reconstructed based on the number of layers of double-value downsampling of the j-th attribute information, and the attribute reconstruction value of the j-th attribute information can be determined. Finally, the attribute reconstruction value of the current point cloud can be determined based on the attribute reconstruction value of the j-th attribute information.
[0506] In some embodiments, the number of twelfth information may depend on the seventh information. For example, if the seventh information indicates that the number of attribute information of the current point cloud is N2, then the number of twelfth information shall not exceed N2, that is, for N2 attribute information, at most one twelfth information can be determined, as shown in Table 10.
[0507] In the embodiments of this application, the second syntax element identification information includes the fourteenth information. Based on the component length of the double-value downsampling of the j-th attribute information, the fourteenth information of the k-th layer of the double-value downsampling of the j-th attribute information is determined; wherein, the value of k is less than or equal to the number of layers of the double-value downsampling of the j-th attribute information; k is an integer greater than 0.
[0508] In other words, in the embodiments of this application, the fourteenth information can be used to indicate the component length of the k-th layer of the double-value downsampling of the j-th attribute information of the current point cloud.
[0509] In the embodiments of this application, k is an integer greater than 0. The value of k is less than or equal to the number of layers in the double-value downsampling of the j-th attribute information.
[0510] In some embodiments, after determining the number of layers of the double-value downsampling of the j-th attribute information of the current point cloud, for any layer, such as the k-th layer, the fourteenth information corresponding to the k-th layer can be determined based on the component length of the k-th layer.
[0511] For example, in some embodiments, the k-th layer of the double-valued downsampling of the j-th attribute information can be understood as the component of the k-th layer double-valued downsampling, and correspondingly, the component length of the k-th layer of the double-valued downsampling of the j-th attribute information can be understood as the length of the component of the k-th layer double-valued downsampling.
[0512] In this embodiment, some instruction information in the form of syntax elements or flags can be written into the bitstream. In this way, the relevant encoding information of the current point cloud can be indicated by the value of the syntax elements in the bitstream. For example, by determining the length of the k-th layer component of the double-value downsampled j-th attribute information of the current point cloud, the length of the k-th layer double-value downsampled component of the double-value downsampled j-th attribute information of the current point cloud can be determined.
[0513] In some embodiments, the fourteenth information included in the second syntax element identification information can be used to indicate the component length of the k-th layer of the double-value downsampling of the j-th attribute information of the current point cloud.
[0514] For example, in some embodiments, the fourteenth information can be represented as ai_lossless_dds_level_length_in_bytes[j][k], which indicates the length of the k-th layer component of the j-th attribute information of the current point cloud after double-value downsampling. The length unit determined based on the value of ai_lossless_dds_level_length_in_bytes[j][k] can be bytes.
[0515] Of course, the length unit of the component length of the k-th layer in the double-value downsampling of the j-th attribute information is not limited to bytes, and can also be other length units. This application does not make specific limitations.
[0516] In some embodiments, after determining the component length of the k-th layer of the double-valued downsampling of the j-th attribute information, the k-th layer of the double-valued downsampling of the j-th attribute information can be further reconstructed based on the component length of the k-th layer of the double-valued downsampling of the j-th attribute information, and finally the attribute reconstruction value of the k-th layer of the double-valued downsampling of the j-th attribute information can be determined.
[0517] In some embodiments, the number of fourteenth pieces of information may depend on the twelfth pieces of information. For example, if the twelfth pieces of information indicates that the number of layers of the double-value downsampling of the j-th attribute information of the current point cloud is N3, then the number of fourteenth pieces of information will not exceed N3. That is, the double-value downsampling of the j-th attribute information corresponding to the N3 layers can determine at most N3 pieces of fourteenth pieces of information, as shown in Table 11.
[0518] For example, in some embodiments, the encoding of the second syntax element identification information may refer to Table 12.
[0519] In summary, the embodiments of this application propose various supports for multi-attribute datasets of learning-based artificial intelligence point cloud coding (AI-PCC). The encoding and decoding methods proposed in this application can be used in AI-PCC coding standards such as AVS and MPEG. When implementing the encoding and decoding methods proposed in this application, modifications can be made to the bitstream structure, syntax, constraints, and mappings used to generate decoded point clouds for standardization.
[0520] This application provides an encoding method in which an encoder determines the indication information of the current point cloud and writes the indication information of the current point cloud into the bitstream. The indication information of the current point cloud includes at least one or more of the following: a first syntax element identifier based on the geometric information of the learned AI-PCC; a second syntax element identifier based on the attribute information of the learned AI-PCC; the first syntax element identifier is used to determine the geometric reconstruction value of the current point cloud; and the second syntax element identifier is used to determine the attribute reconstruction value of the current point cloud. Therefore, in the embodiments of this application, the indication information transmitted in the bitstream may include the first syntax element identifier based on the geometric information of the learned AI-PCC and / or the second syntax element identifier based on the attribute information of the learned AI-PCC. That is, the technical solution of this application provides a syntax structure or encoding sequence format that can be used to carry compressed data, realizing multiple supports for multi-attribute datasets of learning-based artificial intelligence point cloud encoding (AI-PCC), thereby improving the entropy encoding and decoding performance of point clouds.
[0521] Based on the above embodiments, in one embodiment of this application, in the implementation of compressing a single point cloud or a series of single point clouds (dynamic point cloud sequence), in view of the problem that related technologies cannot provide a syntax structure or encoding sequence format that can be used to carry compressed data, this application embodiment provides a method for effectively representing sparse tensors as an intermediate representation of a learning-based model to achieve compressed point cloud carrying.
[0522] In embodiments of this application, rate-distortion-based quantization allows for the creation of an efficient data storage mechanism while maintaining the fidelity of point cloud reconstruction.
[0523] In the embodiments of this application, the main components of the point cloud to be encoded or decoded may include one or more of the following: lossless geometric components; lossy geometric components; attribute components.
[0524] In embodiments of this application, the indication information, including but not limited to: information on component coding parameters, the selected NN model for compression, and decoding information, is represented by signals in the header of the encoded file.
[0525] In embodiments of this application, the concept of an encoded sequence header is introduced as a first modification to the encoded bitstream used for a learning-based solution.
[0526] In the embodiments of this application, the V3C parameter set AI-PCC extended syntax is shown in Table 13:
[0527] Table 13
[0528]
[0529]
[0530] Here, `oi_length_in_bytes` defines the length (in bytes) of the geometric components of the AI-PCC encoded frame. `oi_num_lossless_dds_levels_minus1` defines the number of layers of the lossless geometric encoded components with dual-valued downsampling. `oi_lossless_dds_level_length_in_bytes` defines the length (in bytes) of each dual-valued downsampled component of the geometric bitstream. The geometric bitstream components are arranged in ascending order, and each of these bitstreams can be decoded independently. The decoding process can terminate at any dynamically downsampled component level. `oi_lossy_component_present_flag` indicates whether lossy components of the geometric information exist in the encoded bitstream. `oi_lossy_component_length_in_bytes` indicates the length (in bytes) of the lossy geometric components. `ai_attributes_present_flag` indicates whether the encoded sequence contains attribute information. `ai_num_attributes` defines the number of current attributes. `ai_length_in_bytes[j]` defines the length (in bytes) of the j-th encoded attribute. `ai_attribute_num_channels[j]` defines the number of channels used for the j-th attribute (e.g., reflectivity can be 1 channel, color can be 4 channels consisting of red, green, blue, and alpha components, etc.). `ai_attribute type[j]` defines the type of the j-th attribute (type can be color, reflectivity, intensity, normal, material_id, etc.). `ai_model_override_flag[j]` indicates a model different from the default model `mst` used for frame decoding. `ai_model_index[j]` indicates the index of the NN model used for decoding the j-th attribute component. `ai_num_lossless_dds_levels_minus1[j]` indicates the number of double-downsampled attribute components in the encoded bitstream used for the j-th encoded attribute. `ai_lossless_dds_level_length_in_bytes[j][k]` indicates the length (in bytes) of the k-th level component of the double-downsampled j-th encoded attribute.
[0531] In summary, the embodiments of this application propose various supports for multi-attribute datasets of learning-based artificial intelligence point cloud coding (AI-PCC). The encoding and decoding methods proposed in this application can be used in AI-PCC coding standards such as AVS and MPEG. When implementing the encoding and decoding methods proposed in this application, modifications can be made to the bitstream structure, syntax, constraints, and mappings used to generate decoded point clouds for standardization.
[0532] This application provides an encoding and decoding method. The indication information transmitted in the bitstream may include first syntax element identification information based on the geometric information of AI-PCC learned data and / or second syntax element identification information based on the attribute information of AI-PCC learned data. That is, the technical solution of this application provides a syntax structure or encoding sequence format that can be used to carry compressed data, realizes multiple supports for multi-attribute datasets of AI-PCC learned data, and thus improves the entropy encoding and decoding performance of point clouds.
[0533] Based on the above embodiments, in another embodiment of this application, based on the same inventive concept as the foregoing embodiments, Figure 20 This is a schematic diagram of the composition structure of an encoder provided in an embodiment of this application, as shown below. Figure 20 As shown, the encoder 100 may include a first determining portion 1101, wherein:
[0534] The first determining part 1101 is configured to determine the indication information of the current point cloud and write the indication information of the current point cloud into the code stream; wherein, the indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on the geometric information of the learned AI-PCC; second syntax element identification information based on the attribute information of the learned AI-PCC; the first syntax element identification information is used to determine the geometric reconstruction value of the current point cloud; the second syntax element identification information is used to determine the attribute reconstruction value of the current point cloud.
[0535] Understandably, in the embodiments of this application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular component. Furthermore, the components in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit described above can be implemented in hardware or as a software functional module.
[0536] If the integrated unit is implemented as a software functional module and not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0537] Therefore, this application provides a computer-readable storage medium applied to an encoder 100, the computer-readable storage medium storing a computer program that, when executed by a first processor, implements the method described in any of the foregoing embodiments.
[0538] Based on the composition of the encoder 100 and the computer-readable storage medium described above, Figure 21 This is a schematic diagram of the specific hardware structure of an encoder provided in an embodiment of this application, as shown below. Figure 21 As shown, the encoder 100 may include a first communication interface 1201, a first memory 1202, and a first processor 1203; the various components are coupled together via a first bus system 1204. It is understood that the first bus system 1204 is used to implement communication between these components. In addition to a data bus, the first bus system 1204 also includes a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as the first bus system 1204 in the figure.
[0539] The first communication interface 1201 is used for receiving and sending signals during the process of sending and receiving information with other external network elements;
[0540] The first memory 1202 is used to store computer programs that can run on the first processor 1203;
[0541] The first processor 1203 is configured to, when running the computer program, perform the following: determine indication information of the current point cloud and write the indication information of the current point cloud into the code stream; wherein the indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on the geometric information of the learned AI-PCC; second syntax element identification information based on the attribute information of the learned AI-PCC; the first syntax element identification information is used to determine the geometric reconstruction value of the current point cloud; and the second syntax element identification information is used to determine the attribute reconstruction value of the current point cloud.
[0542] It is understood that the first memory 1202 in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The first memory 1202 of the system and method described in this application is intended to include, but is not limited to, these and any other suitable types of memory.
[0543] The first processor 1203 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the first processor 1203 or by instructions in software form. The first processor 1203 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the first memory 1202. The first processor 1203 reads the information in the first memory 1202 and completes the steps of the above method in conjunction with its hardware.
[0544] It is understood that the embodiments described in this application can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or combinations thereof. For software implementation, the technology described in this application can be implemented through modules (e.g., procedures, functions, etc.) that perform the functions described in this application. Software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.
[0545] Alternatively, as another embodiment, the first processor 1203 is further configured to execute the method described in any of the foregoing embodiments when running the computer program.
[0546] In yet another embodiment of this application, based on the same inventive concept as the foregoing embodiments, Figure 22 This is a schematic diagram of the composition structure of a decoder provided in an embodiment of this application, as shown below. Figure 22 As shown, the decoder 200 may include a second determining portion 2101, wherein:
[0547] The second determining part 2101 is configured to decode the bitstream and determine the indication information of the current point cloud; wherein, the indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on the geometric information of the learned AI-PCC; second syntax element identification information based on the attribute information of the learned AI-PCC; determining the geometric reconstruction value of the current point cloud based on the first syntax element identification information; determining the attribute reconstruction value of the current point cloud based on the second syntax element identification information; and reconstructing the current point cloud based on the geometric reconstruction value and the attribute reconstruction value of the current point cloud.
[0548] Understandably, in this embodiment, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular component. Furthermore, the components in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional module.
[0549] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, this embodiment provides a computer-readable storage medium applied to the decoder 200. This computer-readable storage medium stores a computer program, which, when executed by a second processor, implements the method described in any of the foregoing embodiments.
[0550] Based on the composition of the decoder 200 and the computer-readable storage medium described above, Figure 23 This is a schematic diagram of the specific hardware structure of a decoder provided in an embodiment of this application, such as... Figure 23As shown, the decoder 200 may include: a second communication interface 2201, a second memory 2202, and a second processor 2203; the various components are coupled together via a second bus system 2204. It is understood that the second bus system 2204 is used to implement communication between these components. In addition to a data bus, the second bus system 2204 also includes a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as the second bus system 2204 in the figure.
[0551] The second communication interface 2201 is used for receiving and sending signals during the process of sending and receiving information with other external network elements;
[0552] The second memory 2202 is used to store computer programs that can run on the second processor 2203;
[0553] The second processor 2203 is configured to, when running the computer program, perform the following: determining a first candidate list; performing preset processing on the first candidate list to obtain a second candidate list; wherein the second candidate list includes one or more pairwise averaged candidates, and the one or more pairwise averaged candidates are generated based on the first candidate list; decoding the bitstream to determine indication information of the current point cloud; wherein the indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on the geometric information of the learned AI-PCC; second syntax element identification information based on the attribute information of the learned AI-PCC; determining the geometric reconstruction value of the current point cloud based on the first syntax element identification information; determining the attribute reconstruction value of the current point cloud based on the second syntax element identification information; and reconstructing the current point cloud based on the geometric reconstruction value and the attribute reconstruction value of the current point cloud.
[0554] It is understood that the second memory 2202 has similar hardware functions to the first memory 1202, and the second processor 2203 has similar hardware functions to the first processor 1203; details will not be elaborated here.
[0555] In yet another embodiment of this application, Figure 24 This is a schematic diagram of the composition structure of an encoding / decoding system provided in an embodiment of this application, such as... Figure 24 As shown, the encoding / decoding system 300 may include an encoder 100 and a decoder 200.
[0556] In the embodiments of this application, the encoder 100 may be any one of the encoders described in the foregoing embodiments, and the decoder 200 may be any one of the decoders described in the foregoing embodiments.
[0557] In some embodiments, this application also provides a method for transmitting a bitstream, wherein the bitstream is generated based on the encoding method described in any of the foregoing embodiments.
[0558] In some embodiments, this application also provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the method as described in any of the foregoing embodiments. Specifically, when executed by a first processor, the computer program implements the encoding method as described in any of the foregoing embodiments, or when executed by a second processor, it implements the decoding method as described in any of the foregoing embodiments.
[0559] In some embodiments, this application also provides a computer program product, including a computer program or instructions. When executed by a processor, the computer program or instructions implement the method as described in any of the foregoing embodiments. Specifically, when executed by a first processor, the computer program or instructions implement the encoding method as described in any of the foregoing embodiments, or when executed by a second processor, they implement the decoding method as described in any of the foregoing embodiments.
[0560] In some embodiments, this application also provides a computer program that, when executed by a processor, implements the method as described in any of the foregoing embodiments. Specifically, when executed by a first processor, the computer program or instructions implement the encoding method as described in any of the foregoing embodiments, or when executed by a second processor, implement the decoding method as described in any of the foregoing embodiments.
[0561] In some embodiments, this application also provides a computer-readable storage medium storing a bitstream thereon. The bitstream is generated by performing the steps of the encoding method as described in any of the foregoing embodiments.
[0562] In this embodiment of the application, the information to be encoded in the encoding method includes at least one of the following: indication information of the current point cloud; wherein, the indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on the geometric information of the learned AI-PCC; and second syntax element identification information based on the attribute information of the learned AI-PCC.
[0563] In some embodiments, this application provides a method for storing a bitstream, which involves generating a bitstream by performing an encoding method as described in any of the foregoing embodiments; and storing the bitstream.
[0564] In some embodiments, this application provides a method for reading a bitstream, including reading the bitstream and performing a decoding method as described in any of the foregoing embodiments to decode the bitstream and generate a video or image.
[0565] In some embodiments, this application provides a method for receiving a bitstream, including receiving the bitstream and performing a decoding method as described in any of the foregoing embodiments to decode the bitstream and generate a video or image.
[0566] In some embodiments, this application provides a computer-readable storage medium storing a computer program / instructions and a bitstream thereon, wherein the computer program / instructions, when executed by a processor, implement the steps of the encoding method as described in any of the foregoing embodiments to generate a bitstream.
[0567] In some embodiments, this application provides a computer-readable storage medium storing a computer program / instructions and a bitstream thereon. When the computer program / instructions are executed by a processor, they implement the steps of the decoding method as described in any of the foregoing embodiments to decode the bitstream and generate a video or image.
[0568] In some embodiments, this application provides a bitstream generated by bit encoding based on information to be encoded; wherein the information to be encoded includes at least one of the following: indication information of the current point cloud; wherein the indication information of the current point cloud includes at least one or more of the following: first syntax element identification information based on the geometric information of the learned AI-PCC; and second syntax element identification information based on the attribute information of the learned AI-PCC.
[0569] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0570] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0571] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0572] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0573] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0574] It should be noted that, in this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0575] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0576] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined to obtain new method embodiments without conflict. The features disclosed in the several product embodiments provided in this application can be arbitrarily combined to obtain new product embodiments without conflict. The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined to obtain new method embodiments or device embodiments without conflict.
[0577] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A decoding method applied to a decoder, the method comprising: decoding a bitstream to determine indication information of a current point cloud; wherein the indication information of the current point cloud comprises at least one or more of: first syntax element identification information of geometry information of an artificial intelligence point cloud coding (AI-PCC) based on learning; and second syntax element identification information of attribute information of the AI-PCC based on learning; determining a geometry reconstruction value of the current point cloud based on the first syntax element identification information; determining an attribute reconstruction value of the current point cloud according to the second syntax element identification information; reconstructing the current point cloud based on the geometry reconstruction value of the current point cloud and / or the attribute reconstruction value of the current point cloud.
2. The method of claim 1, wherein, The first syntax element identification information comprises first information, and the method further comprises: in a case where a value of the first information is a first value, determining that the geometry information of the current point cloud has a lossy component; in a case where the value of the first information is a second value, determining that the geometry information of the current point cloud does not have a lossy component.
3. The method of claim 2, wherein, The first syntax element identification information comprises second information, and the determining of the geometry reconstruction value of the current point cloud based on the first syntax element identification information comprises: in a case where it is determined that the geometry information of the current point cloud has a lossy component, determining the second information; determining a length of the lossy component based on the second information; determining the geometry reconstruction value of the current point cloud based on the length of the lossy component.
4. The method of any one of claims 1 to 3, wherein, The first syntax element identification information comprises third information, and the determining of the geometry reconstruction value of the current point cloud based on the first syntax element identification information comprises: determining a component length of the geometry information of the current point cloud based on the third information; determining the geometry reconstruction value of the current point cloud based on the component length of the geometry information.
5. The method of any one of claims 1 to 4, wherein, The first syntax element identification information comprises fourth information, and the determining of the geometry reconstruction value of the current point cloud based on the first syntax element identification information comprises: determining a number of layers of bi-value down-sampling of a lossless component of the geometry information of the current point cloud based on the fourth information; determining the geometry reconstruction value of the current point cloud based on the number of layers of bi-value down-sampling of the lossless component.
6. The method of claim 5, wherein, The first syntax element identification information comprises fifth information, and the determining of the geometry reconstruction value of the current point cloud based on the number of layers of bi-value down-sampling of the lossless component comprises: determining a component length of an i-th layer based on the fifth information of the i-th layer of bi-value down-sampling of the lossless component; wherein i has a value less than or equal to the number of layers of bi-value down-sampling of the lossless component; i is an integer greater than 0; determining a geometry reconstruction value of the i-th layer of bi-value down-sampling of the lossless component based on the component length of the i-th layer; determining the geometry reconstruction value of the current point cloud based on the geometry reconstruction value of the i-th layer of bi-value down-sampling of the lossless component.
7. The method of any one of claims 1 to 6, wherein, The second syntax element identification information comprises sixth information, and the method further comprises: in a case where a value of the sixth information is a third value, determining that the attribute information of the current point cloud exists; In a case where the sixth information has a fourth value, it is determined that the current point cloud does not have the attribute information.
8. The method of claim 7, wherein, The second syntax element identification information includes seventh information, and the determining of the attribute reconstruction value of the current point cloud based on the second syntax element identification information includes: In a case where it is determined that the current point cloud has the attribute information, the seventh information is determined; The number of the attribute information is determined based on the seventh information; The attribute reconstruction value of the current point cloud is determined based on the number of the attribute information.
9. The method of claim 8, wherein, The second syntax element identification information includes one or more of the following: eighth information, ninth information, tenth information, eleventh information, twelfth information, and the determining of the attribute reconstruction value of the current point cloud based on the number of the attribute information includes: The component length of the jth attribute information is determined based on the eighth information of the jth attribute information, where j has a value less than or equal to the number of the attribute information, and j is an integer greater than 0; and / or, The channel number of the jth attribute information is determined based on the ninth information of the jth attribute information; and / or, The type of the jth attribute information is determined based on the tenth information of the jth attribute information; and / or, The model information of the jth attribute information is determined based on the eleventh information of the jth attribute information; and / or, The number of layers of the binary value downsampling of the jth attribute information is determined based on the twelfth information of the jth attribute information; The attribute reconstruction value of the jth attribute information is determined based on one or more of the component length of the jth attribute information, the channel number of the jth attribute information, the type of the jth attribute information, the model information of the jth attribute information, and the number of layers of the binary value downsampling of the jth attribute information; The attribute reconstruction value of the current point cloud is determined based on the attribute reconstruction value of the jth attribute information.
10. The method of claim 9, wherein, The method further includes: In a case where the eleventh information has a fifth value, it is determined that the jth attribute information uses a preset model; In a case where the eleventh information has a sixth value, it is determined that the jth attribute information does not use a preset model.
11. The method of claim 9 or 10, wherein, The second syntax element identification information includes thirteenth information, and the determining of the model information of the jth attribute information based on the eleventh information of the jth attribute information includes: In a case where it is determined based on the eleventh information that the jth attribute information does not use a preset model, the thirteenth information is determined; The model index of the jth attribute information is determined based on the thirteenth information, and the model information of the jth attribute information is determined based on the model index.
12. The method of claim 11, wherein, The second syntax element identification information includes fourteenth information, and the method further includes: The component length of the kth layer of the binary value downsampling of the jth attribute information is determined based on the fourteenth information of the kth layer of the binary value downsampling of the jth attribute information, where k has a value less than or equal to the number of layers of the binary value downsampling of the jth attribute information, and k is an integer greater than 0; determine a k-th layer attribute reconstruction value of the bi-value down-sampling of the j-th attribute information based on the k-th layer attribute reconstruction value of the bi-value down-sampling of the j-th attribute information; determine the attribute reconstruction value of the j-th attribute information based on the k-th layer attribute reconstruction value of the bi-value down-sampling of the j-th attribute information.
13. An encoding method applied to an encoder, the method comprising: determining indication information of a current point cloud, and writing the indication information of the current point cloud into a bitstream; wherein, the indication information of the current point cloud comprises one or more of the following: first syntax element identification information of geometry information of a learning-based AI-PCC; second syntax element identification information of attribute information of the learning-based AI-PCC; the first syntax element identification information is used to determine a geometry reconstruction value of the current point cloud; the second syntax element identification information is used to determine an attribute reconstruction value of the current point cloud.
14. The method of claim 12, wherein, the first syntax element identification information comprises first information, and the method further comprises: in a case where there is a lossy component in the geometry information of the current point cloud, the first information has a first value; in a case where there is no lossy component in the geometry information of the current point cloud, the first information has a second value.
15. The method of claim 14, wherein, the first syntax element identification information comprises second information, and the method further comprises: in a case where it is determined that there is a lossy component in the geometry information of the current point cloud, the second information is determined based on a length of the lossy component.
16. The method of any one of claims 12 to 15, wherein, the first syntax element identification information comprises third information, and the method further comprises: the third information is determined based on a component length of the geometry information of the current point cloud.
17. The method of any one of claims 12 to 16, wherein, the first syntax element identification information comprises fourth information, and the method further comprises: the fourth information is determined based on a number of layers of bi-value down-sampling of lossless components of the geometry information of the current point cloud.
18. The method of claim 17, wherein, the fifth information of the i-th layer of bi-value down-sampling of the lossless component is determined based on a component length of the i-th layer of bi-value down-sampling of the lossless component; wherein i has a value less than or equal to the number of layers of bi-value down-sampling of the lossless component; i is an integer greater than 0.
19. The method of any one of claims 12 to 18, wherein, the second syntax element identification information comprises sixth information, in a case where the attribute information exists in the current point cloud, the sixth information has a third value; in a case where the attribute information does not exist in the current point cloud, the sixth information has a fourth value.
20. The method of claim 19, wherein, the second syntax element identification information comprises seventh information, and the method further comprises: in a case where it is determined that the attribute information exists in the current point cloud, the seventh information is determined based on a number of the attribute information.
21. The method of claim 20, wherein, the second syntax element identification information comprises one or more of the following: eighth information, ninth information, tenth information, eleventh information, twelfth information, and the method further comprises: the eighth information of the j-th attribute information is determined based on a component length of the j-th attribute information; wherein j has a value less than or equal to the number of the attribute information; j is an integer greater than 0; and / or, determine the ninth information of the jth attribute information based on a channel number of the jth attribute information; and / or, determine the tenth information of the jth attribute information based on a type of the jth attribute information; and / or, determine the eleventh information of the jth attribute information based on model information of the jth attribute information; and / or, determine the twelfth information of the jth attribute information based on a number of layers of binary value downsampling of the jth attribute information.
22. The method of claim 21, wherein, The method further comprises: in a case where the jth attribute information uses a preset model, the eleventh information has a fifth value; in a case where the jth attribute information does not use a preset model, the eleventh information has a sixth value.
23. The method of claim 9 or 10, wherein, The second syntax element identification information comprises thirteenth information, and the method further comprises: in a case where it is determined that the jth attribute information does not use a preset model, determining a model index of the jth attribute information based on model information of the jth attribute information; determining the thirteenth information based on the model index.
24. The method of claim 23, wherein, The second syntax element identification information comprises fourteenth information, and the method further comprises: determining the fourteenth information of the kth layer of binary value downsampling of the jth attribute information based on a component length of binary value downsampling of the jth attribute information; wherein, the value of k is less than or equal to the number of layers of binary value downsampling of the jth attribute information; k is an integer greater than 0.
25. A method of transmitting a bitstream, for transmitting a bitstream, wherein, The code stream is generated based on the encoding method in any one of claims 13 to 24.
26. An encoder, comprising a first determining unit, wherein: The first determining unit is configured to determine indication information of a current point cloud, and write the indication information of the current point cloud into a code stream; wherein, the indication information of the current point cloud comprises one or more of the following: first syntax element identification information of geometry information of a learning-based AI-PCC; second syntax element identification information of attribute information of the learning-based AI-PCC; the first syntax element identification information is used to determine a geometry reconstruction value of the current point cloud; and the second syntax element identification information is used to determine an attribute reconstruction value of the current point cloud.
27. An encoder, comprising a first memory and a first processor, wherein: The first memory is used to store a computer program capable of running on the first processor; The first processor is used to execute the method in any one of claims 13 to 24 when the computer program is running.
28. A decoder, comprising a second determining unit, wherein: The second determining unit is configured to decode the code stream and determine indication information of a current point cloud; wherein the indication information of the current point cloud at least includes one or more of the following: first syntax element identification information of geometry information of a learning-based AI-PCC; second syntax element identification information of attribute information of the learning-based AI-PCC; determining a geometry reconstruction value of the current point cloud based on the first syntax element identification information; determining an attribute reconstruction value of the current point cloud based on the second syntax element identification information; and reconstructing the current point cloud based on the geometry reconstruction value of the current point cloud and the attribute reconstruction value of the current point cloud.
29. A decoder comprising a second memory and a second processor, wherein: the second memory is for storing a computer program capable of running on the second processor; the second processor is configured to execute the method according to any one of claims 1 to 12 when running the computer program.
30. A computer readable storage medium having stored thereon a computer program, wherein, The computer program, when executed by the processor, implements the method according to any one of claims 1 to 12, or the method according to any one of claims 13 to 24.
31. A computer-readable storage medium having a code stream stored thereon, wherein, The code stream is generated by executing the steps of the encoding method according to any one of claims 13 to 24. The computer program, when executed by the processor, implements the method according to any one of claims 1 to 12, or the method according to any one of claims 13 to 24. The code stream is generated by executing the steps of the encoding method according to any one of claims 13 to 24.