Coding method, decoding method, coders, decoders, code stream and storage medium
Multi-scale deep encoding of point clouds is solved through multi-channel encoding method, which solves the problem of insufficient encoding performance of point cloud attribute information, realizes more efficient storage and transmission, and improves the compression efficiency of point cloud data.
Patent Information
- Application Number
- PCT/CN2024/070905
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-10
AI Technical Summary
In the prior art, the attribute information encoding performance of point clouds needs to be improved, especially when processing large-scale point cloud data, the storage and transmission efficiency are low.
The multi-channel encoding method is used to perform multi-scale deep encoding of the attribute information of the point cloud. By processing the first point cloud, the second point cloud is obtained, the number of point clouds is reduced, and the attribute features of the first and second point clouds are extracted respectively for encoding, and the feature reduction and reconstruction are used for neural networks.
It improves the encoding performance and coding accuracy of point cloud attribute features, reduces resource consumption for storage and transmission, and improves the compression efficiency of point cloud data.
Smart Images

Figure CN2024070905_10072025_PF_FP_ABST
Abstract
Description
Coding and decoding method, codec, code stream and storage medium Technical Field
[0001] The present application relates to the field of point cloud coding and decoding technology, and in particular to a coding and decoding method, a codec, a bit stream, and a storage medium. Background Art
[0002] Related technologies encode attribute information of point clouds based on region adaptive hierarchical transform (RAHT), but the encoding performance of point cloud attribute information needs to be further improved.
[0003] Summary of the Invention
[0004] The embodiments of the present application provide a coding and decoding method, a codec, a bit stream, and a storage medium. The following introduces various aspects of the present application.
[0005] In a first aspect, a decoding method is provided, which is applied to a decoder, including: parsing a code stream to obtain a first attribute feature and a second attribute feature; performing feature restoration on the first attribute feature to obtain first attribute information; performing feature restoration on the second attribute feature to obtain second attribute information; and processing the first attribute information and the second attribute information to obtain attribute reconstruction information.
[0006] In a second aspect, an encoding method is provided, which is applied to an encoder, including: processing a first point cloud based on attribute information to obtain a second point cloud, where the second point cloud and the first point cloud occupy the same space size, and the number of points in the second point cloud is less than the number of points in the first point cloud; performing feature extraction on the first point cloud and the second point cloud respectively to obtain a first attribute feature and a second attribute feature, wherein the first attribute feature is the attribute feature of the first point cloud, and the second attribute feature is the attribute feature of the second point cloud; encoding the first attribute feature and the second attribute feature respectively, and writing the obtained encoding bits into a first code stream segment and a second code stream segment respectively.
[0007] According to a third aspect, a decoder is provided, comprising: a parsing unit for parsing a code stream to obtain a first attribute feature and a second attribute feature; a first feature restoration unit for performing feature restoration on the first attribute feature to obtain first attribute information; a second feature restoration unit for performing feature restoration on the second attribute feature to obtain second attribute information; and a processing unit for processing the first attribute information and the second attribute information to obtain attribute reconstruction information.
[0008] According to a fourth aspect, a decoder is provided, comprising: a memory for storing a computer program; and a processor for executing the method according to the first aspect when running the computer program.
[0009] In a fifth aspect, an encoder is provided, including: a processing unit, used to process a first point cloud based on attribute information to obtain a second point cloud, the second point cloud and the first point cloud occupy the same space size, and the number of points in the second point cloud is less than the number of points in the first point cloud; a feature extraction unit, used to extract features from the first point cloud and the second point cloud respectively, to obtain a first attribute feature and a second attribute feature, wherein the first attribute feature is the attribute feature of the first point cloud, and the second attribute feature is the attribute feature of the second point cloud; an encoding unit, used to encode the first attribute feature and the second attribute feature respectively, and write the obtained encoding bits into the first code stream segment and the second code stream segment respectively.
[0010] In a sixth aspect, an encoder is provided, comprising: a memory for storing a computer program; and a processor for executing the method described in the second aspect when running the computer program.
[0011] In a seventh aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed, the method as described in the first aspect or the second aspect is implemented.
[0012] In an eighth aspect, a non-volatile computer-readable storage medium for storing a bit stream is provided, wherein the bit stream is generated by an encoding method using an encoder, or the bit stream is decoded by a decoding method using a decoder, wherein the decoding method is the method described in the first aspect, and the encoding method is the method described in the second aspect.
[0013] According to a ninth aspect, a code stream is provided, comprising a code stream generated according to the method of the second aspect.
[0014] The embodiment of the present application extracts features from the attribute information of the point cloud through multiple channels, and encodes the different attribute features obtained, thereby realizing multi-scale deep encoding of the attribute information of the point cloud. Among them, the point clouds input by the multiple channels are different. Taking the input of the multiple channels including the first point cloud and the second point cloud as an example, the second point cloud can be obtained by processing the first point cloud, and the number of points in the second point cloud is less than the number of points in the first point cloud. By inputting different point clouds, the point cloud can be targeted for feature extraction, which helps to improve the accuracy and comprehensiveness of the feature extraction, thereby helping to improve the encoding performance of the point cloud attribute features. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] FIG1A is a schematic diagram of a three-dimensional point cloud image.
[0016] FIG1B is a partially enlarged view of a three-dimensional point cloud image.
[0017] FIG2A is a schematic diagram of six viewing angles of a point cloud image.
[0018] FIG2B is a schematic diagram of a data storage format corresponding to a point cloud image.
[0019] FIG3 is a schematic diagram of a network architecture for point cloud encoding and decoding.
[0020] FIG4A is a schematic diagram of a composition framework of a G-PCC encoder.
[0021] FIG4B is a schematic diagram of a composition framework of a G-PCC decoder.
[0022] FIG5 is a schematic diagram showing the use of binary symbols to indicate whether a voxel is occupied.
[0023] FIG6 is an example diagram of the RAHT transformation process.
[0024] FIG7 is a schematic diagram of an encoding process based on 3DAC.
[0025] FIG8 is a flowchart of a decoding method provided in an embodiment of the present application.
[0026] FIG9 is a flow chart of the encoding method provided in an embodiment of the present application.
[0027] Figure 10 is a structural diagram of an offset attention module provided in an embodiment of the present application.
[0028] FIG11 is a structural example diagram of a feature extraction module provided in an embodiment of the present application.
[0029] FIG12 is a schematic diagram of the structure of a weight module provided in an embodiment of the present application.
[0030] FIG13 is a schematic diagram of the composition of a first processing module provided in an embodiment of the present application.
[0031] Figure 14 is a schematic diagram of the composition of a self-attention module provided in an embodiment of the present application.
[0032] Figure 15A is a schematic diagram of a point cloud coding framework provided in an embodiment of the present application.
[0033] Figure 15B is a schematic diagram of a point cloud decoding framework provided in an embodiment of the present application.
[0034] FIG. 16A is a structural diagram illustrating an example of the attribute encoding module in FIG. 15A .
[0035] FIG. 16B is a structural diagram illustrating an example of the attribute decoding module in FIG. 15B .
[0036] FIG17 is an example diagram of a point cloud encoding and decoding process provided in an embodiment of the present application.
[0037] Figure 18 is a schematic diagram of geometric assisted coding provided in an embodiment of the present application.
[0038] FIG19 is a schematic diagram of the structure of a decoder provided in an embodiment of the present application.
[0039] FIG20 is a schematic diagram of the structure of a decoder provided in another embodiment of the present application.
[0040] FIG21 is a schematic diagram of the structure of an encoder provided in one embodiment of the present application.
[0041] FIG22 is a schematic diagram of the structure of an encoder provided in another embodiment of the present application. DETAILED DESCRIPTION
[0042] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below with reference to the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present application.
[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0044] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0045] It should also be pointed out that the terms "first\second\third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0046] A point cloud is a three-dimensional representation of an object's surface. Point clouds (data) on an object's surface can be collected using acquisition equipment such as photoelectric radars, lidars, laser scanners, and multi-view cameras.
[0047] A point cloud is a set of irregularly distributed discrete points in space that express the spatial structure and surface properties of a three-dimensional object or scene. Figure 1A shows a three-dimensional point cloud image and Figure 1B shows a partially enlarged view of the three-dimensional point cloud image. It can be seen that the point cloud surface is composed of densely distributed points.
[0048] In a two-dimensional image, each pixel contains information and is distributed regularly, so there's no need to record its location. However, the distribution of points in a point cloud in three-dimensional space is random and irregular, so recording the location of each point in space is necessary to fully represent the point cloud. Similar to a two-dimensional image, each location in the acquisition process has corresponding attribute information, typically an RGB color value, which reflects the object's color. For a point cloud, in addition to color information, each point's corresponding attribute information often includes reflectance values, which reflect the surface texture of the object. Therefore, point cloud data typically includes both point location information and point attribute information. Point location information can also be referred to as point geometric information. For example, point geometric information can be the point's three-dimensional coordinates (x, y, z). Point attribute information can include color information and / or reflectance. For example, reflectance can be one-dimensional reflectance information (r). Color information can be information in any color space, or it can be three-dimensional color information, such as RGB. Here, R represents red (red), G represents green (green), and B represents blue (blue). For another example, the color information may be luminance and chrominance (YCbCr, YUV) information, where Y represents brightness (luma), Cb (U) represents blue color difference, and Cr (V) represents red color difference.
[0049] For example, a point cloud generated using laser measurement principles can include both its 3D coordinate information and its reflectivity. For another example, a point cloud generated using photogrammetry principles can include both its 3D coordinate information and its 3D color information. For another example, a point cloud generated using a combination of laser measurement and photogrammetry principles can include both its 3D coordinate information, its reflectivity value, and its 3D color information.
[0050] Figures 2A and 2B show a point cloud image and its corresponding data storage format. Figure 2A provides six viewing angles of the point cloud image, while Figure 2B consists of a file header and data. The header includes the data format, data representation type, the total number of points in the point cloud, and the content represented by the point cloud. For example, the point cloud is in ".ply" format, represented by ASCII code, with a total of 207,242 points. Each point has 3D coordinate information (x, y, z) and 3D color information (r, g, b).
[0051] Point clouds can be divided into the following categories according to the acquisition method:
[0052] Static point cloud: the object is stationary and the device that acquires the point cloud is also stationary;
[0053] Dynamic point cloud: The object is moving, but the device that obtains the point cloud is stationary;
[0054] Dynamic point cloud acquisition: The device used to acquire the point cloud is in motion.
[0055] For example, point clouds can be divided into two categories according to their usage:
[0056] Category 1: Machine perception point cloud, which can be used in scenarios such as autonomous navigation systems, real-time inspection systems, geographic information systems, visual sorting robots, and disaster relief robots;
[0057] Category 2: Human eye perception point cloud, which can be used in point cloud application scenarios such as digital cultural heritage, free viewpoint broadcasting, 3D immersive communication, and 3D immersive interaction.
[0058] Point clouds can flexibly and conveniently express the spatial structure and surface properties of three-dimensional objects or scenes. Moreover, since point clouds are obtained by directly sampling real objects, they can provide a strong sense of reality while ensuring accuracy. Therefore, they are widely used, including virtual reality games, computer-aided design, geographic information systems, automatic navigation systems, digital cultural heritage, free viewpoint broadcasting, three-dimensional immersive remote presentation, and three-dimensional reconstruction of biological tissues and organs.
[0059] Point clouds are primarily collected through computer generation, 3D laser scanning, and 3D photogrammetry. Computers can generate point clouds of virtual 3D objects and scenes; 3D laser scanning can obtain point clouds of static real-world 3D objects or scenes, generating millions of point clouds per second; and 3D photogrammetry can obtain point clouds of dynamic real-world 3D objects or scenes, generating tens of millions of point clouds per second. These technologies reduce the cost and time required to acquire point cloud data while improving data accuracy. While changes in point cloud data acquisition methods have made it possible to acquire large amounts of point cloud data, the processing of this massive amount of 3D point cloud data is facing bottlenecks due to storage space and transmission bandwidth constraints, as application demands grow.
[0060] For example, taking a point cloud video with a frame rate of 30 frames per second (fps), each frame contains 700,000 points, and each point has coordinate information (xyz, float) and color information (RGB, uchar). Therefore, the data volume of a 10-second point cloud video is approximately 0.7 million × (4 bytes × 3 + 1 byte × 3) × 30 fps × 10 seconds = 3.15 GB, where 1 byte is 10 bits. For a 1280 × 720 2D video with a YUV sampling format of 4:2:0 and a frame rate of 24 fps, the data volume for 10 seconds is approximately 1280 × 720 × 12 bits × 24 fps × 10 seconds ≈ 0.33 GB. The data volume of a 10-second two-view 3D video is approximately 0.33 × 2 = 0.66 GB. This shows that the data volume of a point cloud video far exceeds that of 2D and 3D videos of the same length. Therefore, in order to better realize data management, save server storage space, and reduce the transmission traffic and transmission time between the server and the client, point cloud compression has become a key issue in promoting the development of the point cloud industry.
[0061] That is to say, since the point cloud is a collection of massive points, storing the point cloud not only consumes a lot of memory, but is also not conducive to transmission. There is also not enough bandwidth to support direct transmission of the point cloud at the network layer without compression. Therefore, the point cloud needs to be compressed.
[0062] Currently, point cloud coding frameworks that can compress point clouds can be the geometry-based point cloud compression (G-PCC) codec framework or the video-based point cloud compression (V-PCC) codec framework provided by the Moving Picture Experts Group (MPEG), or the AVS-PCC codec framework provided by AVS. The G-PCC codec framework can be used to compress the first type of static point clouds and the third type of dynamically acquired point clouds, and it can be based on the point cloud compression test platform (test model compression 13, TMC13). The V-PCC codec framework can be used to compress the second type of dynamic point clouds, and it can be based on the point cloud compression test platform (test model compression 2, TMC2). Therefore, the G-PCC codec framework is also called the point cloud codec TMC13, and the V-PCC codec framework is also called the point cloud codec TMC2.
[0063] An embodiment of the present application provides a network architecture of a point cloud encoding and decoding system including a decoding method and an encoding method. FIG3 is a schematic diagram of a network architecture of a point cloud encoding and decoding system provided by an embodiment of the present application. As shown in FIG3 , the network architecture includes one or more electronic devices 13 to 1N and a communication network 01, wherein the electronic devices 13 to 1N can perform video interaction through the communication network 01. During the implementation process, the electronic device can be various types of devices with point cloud encoding and decoding functions. For example, the electronic device can include a mobile phone, a tablet computer, a personal computer, a personal digital assistant, a navigator, a digital phone, a video phone, a television, a sensor device, a server, etc., which is not limited by the embodiment of the present application. Among them, the decoder or encoder in the embodiment of the present application can be the above-mentioned electronic device.
[0064] Among them, the electronic device in the embodiment of the present application has a point cloud encoding and decoding function, generally including a point cloud encoder (ie, encoder) and a point cloud decoder (ie, decoder).
[0065] The following describes the related technologies using the G-PCC codec framework and the AVS codec framework as examples.
[0066] As you can understand, in the point cloud G-PCC codec framework, the point cloud data to be encoded is first divided into multiple slices through slice partitioning. In each slice, the geometric information of the point cloud and the attribute information corresponding to each point are encoded separately.
[0067] Figure 4A shows a schematic diagram of the G-PCC encoder's architecture. As shown in Figure 4A, during the geometry encoding process, the geometric information is transformed so that the entire point cloud is contained within a bounding box. Quantization is then performed. This quantization step primarily serves a scaling purpose. Due to quantization rounding, the geometric information of some point clouds becomes identical. Parameters are then used to determine whether to remove duplicate points. This process of quantization and removing duplicate points is also known as voxelization. The bounding box is then partitioned into an octree or constructed as a prediction tree. During this process, arithmetic coding is performed on the points within the leaf nodes of the partition to generate a binary geometry bitstream. Alternatively, arithmetic coding is performed on the intersections (vertex) generated by the partition (surface fitting is performed based on the intersections) to generate a binary geometry bitstream. During the attribute encoding process, after the geometry encoding is completed and the geometric information is reconstructed, color conversion is performed to convert the color information (i.e., attribute information) from the RGB color space to the YUV color space. The reconstructed geometry information is then used to recolor the point cloud so that the unencoded attribute information corresponds to the reconstructed geometry information. Attribute coding is mainly performed on color information. In the process of color information coding, there are two main transformation methods. One is the distance-based lifting transformation that relies on the level of detail (LOD) division, and the other is direct RAHT. Both methods convert color information from the spatial domain to the frequency domain, and obtain high-frequency coefficients and low-frequency coefficients through transformation. Finally, the coefficients are quantized and then arithmetic coding is performed on the quantized coefficients to generate a binary attribute bit stream.
[0068] Figure 4B shows a schematic diagram of the composition framework of a G-PCC decoder. As shown in Figure 4B, for the acquired binary bit stream, the geometric bit stream and attribute bit stream in the binary bit stream are first decoded independently. When decoding the geometric bit stream, the geometric information of the point cloud is obtained through arithmetic decoding-reconstruction of the octree / reconstruction of the prediction tree-reconstruction of the geometry-coordinate inverse conversion; when decoding the attribute bit stream, the attribute information of the point cloud is obtained through arithmetic decoding-inverse quantization-LOD partitioning / RAHT-color inverse conversion, and the point cloud data to be encoded (i.e., the output point cloud) is restored based on the geometric information and attribute information.
[0069] It should be noted that, as shown in FIG4A or FIG4B , the current geometric coding and decoding of G-PCC can be divided into octree-based geometric coding and decoding (marked by a dotted box) and prediction tree-based geometric coding and decoding (marked by a dotted box).
[0070] As mentioned earlier, there are two main transformation methods for encoding attribute information: lifting transformation and RAHT transformation. Lifting transformation predicts and transforms the point cloud based on the LOD generation order, while RAHT transformation adaptively transforms attribute information from the bottom up based on the octree construction hierarchy. The following describes RAHT transformation in conjunction with a point cloud encoding process (including steps 101-104).
[0071] In step 101, an octree is used for geometric coding.
[0072] An octree is a three-dimensional extension of a two-dimensional quadtree. Assuming the volume to be represented is a cube with a side length of W meters, in the first level of the octree, the point cloud is divided into eight smaller cubes with side lengths of W / 2 × W / 2 × W / 2. In the second level, each cube is further divided in the same manner, resulting in eight cubes with side lengths of W / 4 × W / 4 × W / 4. This process can be repeated for L levels.
[0073] During the geometry encoding process, a binary symbol can be used to indicate whether a voxel is occupied. Typically, occupied voxels account for less than 1% of the total volume. For each divided cube, check whether it is occupied: if so, mark the cube with a binary symbol 1, otherwise mark it with 0, as shown in Figure 5. In this way, the geometric information (the position of each occupied voxel) can be encoded separately and passed to the decoder.
[0074] In step 102, the attribute information is encoded using RAHT transformation.
[0075] The RAHT transform can transform the attribute information of the point cloud from the spatial domain to the frequency domain, thereby further reducing the correlation between the attribute information of the point cloud. Figure 6 is an example diagram of the RAHT transform process. As shown in Figure 6, the attribute information of the occupied nodes in the same parent node is recursively transformed in a bottom-up manner, and the nodes in each layer are transformed from the three dimensions of x, y, and z until they are transformed to the root node of the octree. In the process of hierarchical transformation, the direct current (DC) coefficients (or low-pass coefficients) obtained after the transformation of the nodes in the same layer are passed to the nodes in the upper layer for further transformation, and all alternating current (AC) coefficients (or high-pass coefficients) will be quantized and encoded.
[0076] RAHT transformation is performed on every two voxels. Assume that the voxel at level l is g l,x,y,z , where x, y, z are integers. g l,2x,y,z and gl,2x+1,y,z are grouped together, and their respective weights w l,2x,y,zand wl,2x+1,y,z, the corresponding DC coefficient g can be obtained by applying the following transformation matrix l-1,x,y,z and AC coefficient h l-1,x,y,z .
[0077] Where w1 = w l,2x,y,z , w2=wl,2x+1,y,z, and T w1w2 The following formula is satisfied.
[0078] It should be noted that the transformation matrix can be adjusted according to the weights to adapt to the number of leaf voxels that each element actually represents.
[0079] During the transformation at the last stage (i.e. the root node), the remaining two voxels are transformed into the final two coefficients:
[0080] It can be seen that in the case of uniform weights, the transformation process can be reduced to Haar transform. The DC coefficient is considered to form its own subband. Assume that there are N v occupied voxels, there are S sub-bands after transformation, each sub-band has N i transformation coefficients, we can get the transformation process of each channel: {f Y (m, n)}=H({Y i}) {f U (m, n)}=H({U i}) {f V (m, n)}=H({V i})
[0081] Where m is the subband index and n is the transform coefficient index.
[0082] In step 103, the transform coefficients obtained in step 102 are quantized using uniform quantization.
[0083] In step 104 , an arithmetic encoder is used to perform entropy coding on the quantized result obtained in step 103 .
[0084] The arithmetic encoder described above can be driven by a Laplace distribution, whose parameters are unique for each subband. These parameters can be passed from the encoder to the decoder.
[0085] Due to the linear characteristics of the Haar transform, the RAHT transform cannot perform nonlinear fitting on complex point cloud attribute information.
[0086] Figure 7 illustrates a schematic diagram of the encoding process for learning attribute compression for 3D point clouds (3DAC). Referring to Figure 7, 3DAC encoding employs RAHT for initial encoding, followed by an attribute-guided deep entropy model to estimate the probability distribution of transform coefficients. This method also explores channel-wise correlations between different attributes. The following describes the 3DAC encoding process in conjunction with steps 201 to 204.
[0087] In step 201, initial encoding is performed.
[0088] 3DAC uses RAHT for initial encoding to capture the spatial redundancy of point cloud attribute information by converting the attribute information into transform coefficients.
[0089] In step 202, feature extraction is performed using the RAHT tree.
[0090] RAHT is modeled using a tree structure for contextual feature extraction. If the low-frequency coefficient of a RAHT tree node has no neighbors in the transformation direction, it is passed to its parent node. Otherwise, the two nodes are merged into their parent node to generate both low-frequency and high-frequency coefficients.
[0091] Quantize all high-frequency coefficients and serialize the quantization results into a symbol stream by breadth-first traversal of the RAHT tree. For the DC coefficient, it can be transmitted directly.
[0092] In step 203, an estimated probability distribution of the transform coefficients is obtained.
[0093] First, a deep entropy model, such as the context feature extraction model used in the initial encoding, is used to obtain an estimated probability distribution p(R) of the transform coefficients. This estimated probability distribution p(R) is then used to approximate the unknown probability distribution q(R). A deep neural network, such as the inter-channel cross-entropy model, is then used to approximate the estimated probability distribution using a cross-entropy loss. The estimated probability distribution of the transform coefficients can be used to compress the transform coefficients.
[0094] In some embodiments, the context feature extraction model of the initial encoding can aggregate the context information from the high-frequency nodes and the context information from the low-frequency nodes to obtain an estimated probability distribution of the transform coefficients.
[0095] In some embodiments, the inter-channel cross entropy model may obtain the cross entropy loss by aggregating coefficient correlations between channels and spatial correlations between channels.
[0096] In step 204, entropy coding is performed.
[0097] The estimated probability distribution of the transform coefficients and the transform coefficients are passed to the arithmetic encoder to generate the final attribute bitstream for further compression. The transform coefficients can be serialized into a symbol stream before being passed to the arithmetic encoder. Accordingly, during the decoding stage, the arithmetic encoder can recover the transform coefficients using the same probability distribution generated by the deep entropy model.
[0098] In order to further improve the encoding performance of point cloud attribute information, an embodiment of the present application proposes an encoding method, including: processing a first point cloud based on attribute information to obtain a second point cloud, the second point cloud and the first point cloud occupying the same space size, and the number of points in the second point cloud is less than the number of points in the first point cloud; performing feature extraction on the first point cloud and the second point cloud respectively to obtain first attribute features and second attribute features, wherein the first attribute features are attribute features of the first point cloud, and the second attribute features are attribute features of the second point cloud; encoding the first attribute features and the second attribute features respectively, and writing the obtained encoding bits into the first code stream segment and the second code stream segment respectively.
[0099] An embodiment of the present application also provides a decoding method, including: parsing a code stream to obtain a first attribute feature and a second attribute feature; performing feature restoration on the first attribute feature to obtain first attribute information; performing feature restoration on the second attribute feature to obtain second attribute information; and processing the first attribute information and the second attribute information to obtain attribute reconstruction information.
[0100] The embodiment of the present application extracts features from the attribute information of the point cloud through multiple channels, and encodes the different attribute features obtained, thereby realizing multi-scale deep encoding of the attribute information of the point cloud. Among them, the point clouds input by the multiple channels are different. Taking the input of multiple channels including a first point cloud and a second point cloud as an example, the second point cloud can be obtained by processing the first point cloud, and the number of points in the second point cloud is less than the number of points in the first point cloud. By inputting different point clouds, feature extraction of the point cloud can be carried out in a targeted manner, which helps to improve the accuracy and comprehensiveness of feature extraction, thereby helping to improve the encoding performance of the point cloud attribute features.
[0101] The decoding method provided in the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0102] FIG8 is a flow chart illustrating a decoding method according to an embodiment of the present application. The decoding method shown in FIG8 can be applied to a decoder. The method shown in FIG8 can be used to decode attribute information of a point cloud, such as color information and / or reflectivity information. The method shown in FIG8 can include steps S810 to S840.
[0103] In step S810, the code stream is parsed to obtain the first attribute feature and the second attribute feature.
[0104] The above-mentioned attribute features may refer to attribute feature values and / or attribute feature matrices, etc. In the decoding method provided in the embodiments of the present application, the bitstream may carry multiple attribute features. The above-mentioned first attribute feature and second attribute feature may be any two different attribute features among the multiple attribute features carried in the bitstream. The number of attribute features carried in the bitstream may be the same as the number of point clouds input during encoding, or in other words, the number of attribute features is the same as the number of encoding channels.
[0105] In some embodiments, different attribute features may be carried in different code stream segments. By parsing different code stream segments, multiple attribute features may be obtained. For example, parsing a first code stream segment may obtain a first attribute feature, while parsing a second code stream segment may obtain a second attribute feature.
[0106] In some embodiments, a codestream segment may include a codestream segment identifier. On one hand, the codestream segment identifier can be used to distinguish codestream segments carrying different attribute features, facilitating parsing of the attribute features. During the decoding process, attribute features associated with different point clouds utilize different decoding methods. On the other hand, the codestream segment identifier can be used to indicate the point cloud associated with the codestream segment. By indicating the point cloud associated with the codestream segment, the relationship between the attribute features carried by the codestream segment and the point cloud can be determined, or in other words, the relationship between the attribute features carried by the codestream segment and the coding channel can be determined.
[0107] In some embodiments, the code stream may be parsed to obtain a code stream length identifier. Based on the code stream segment identifier and the code stream length identifier, the code stream segment may be parsed to obtain multiple attribute features. If the code stream segment identifier indicates that a first code stream segment is associated with a first point cloud and a second code stream segment is associated with a second point cloud, the code stream is parsed to obtain a code stream length identifier. Based on the code stream length identifier, the first code stream segment and the second code stream segment are parsed separately to obtain a first attribute feature and a second attribute feature.
[0108] For example, the code stream length identifier may include a code stream segment start point and a code stream segment length. For another example, the code stream length identifier may include a code stream segment start point and a code stream segment end point. For another example, the code stream length identifier may include a code stream segment start point, a code stream segment end point, and a code stream segment length.
[0109] It should be understood that in the decoding method provided in the embodiment of the present application, the code stream may include more code stream segments, and each code stream segment may be associated with a different point cloud, which is not limited in the present application.
[0110] In step S820, the first attribute feature is restored to obtain first attribute information.
[0111] In some embodiments, a first attribute feature is restored by performing a first feature restoration to obtain initial attribute information; a second feature restoration is performed on the initial attribute information to obtain initial attribute information; and the initial attribute information is updated based on the initial feature information to obtain first attribute information. The second feature restoration may also be referred to as attribute reconstruction, and the initial attribute information may also be referred to as initial attribute reconstruction information.
[0112] For example, the first attribute information can be obtained using the following formula. ij =A ij +(F ij ×W ij )
[0113] Among them, C ij is the first attribute information (matrix), A ij is the initial attribute information (matrix), F ij is the initial feature information (matrix), W ij is the weight matrix, i is the row index, and j is the column index. The weight matrix can be used to adjust the influence of the feature matrix. In practice, the weight matrix can be preconfigured or generated according to predefined rules.
[0114] In step S830, the second attribute feature is restored to obtain second attribute information.
[0115] In step S840, the first attribute information and the second attribute information are processed to obtain attribute reconstruction information.
[0116] In some embodiments, the first attribute information and the second attribute information are processed to obtain first residual information; the first residual information is subjected to third feature restoration (also known as attribute reconstruction) to obtain attribute reconstruction information. For example, the first attribute information is subjected to fourth feature restoration to obtain third attribute information; the third attribute information and the second attribute information are processed to obtain first residual information.
[0117] By performing attribute reconstruction layer by layer from coarse to fine, such as first obtaining the initial attribute reconstruction information, and then reconstructing the attribute again in combination with the second attribute feature to obtain the final attribute reconstruction information, it helps to reduce the error between the reconstructed attribute information and the original attribute information.
[0118] In some embodiments, the bitstream is parsed to obtain geometric information; the geometric information and attribute reconstruction information are processed to obtain a reconstructed first point cloud. For example, the attribute reconstruction information is associated with the geometric information to obtain the reconstructed first point cloud. As an example, the first point cloud can be an original point cloud.
[0119] In an embodiment of the present application, feature restoration can be implemented through a neural network. For example, the first feature restoration is implemented through a first neural network, the second feature restoration is implemented through a second neural network, the second neural network includes one or more convolutional layers; and / or, the first neural network includes one or more of the following: one or more convolutional layers; one or more activation layers; and one or more feature restoration modules. For another example, the third feature restoration is implemented through a third neural network, the fourth feature restoration is implemented through a fourth neural network, and the second attribute information is obtained through a fifth neural network, the third neural network includes one or more convolutional layers, and / or, the fourth neural network and / or the fifth neural network include one or more of the following: one or more convolutional layers; one or more activation layers; and one or more feature restoration modules.
[0120] In some embodiments, the first feature restoration and the fourth feature restoration can be implemented by the same neural network, that is, the first neural network and the fourth neural network can be the same. For example, the first neural network and the fourth neural network can include two convolutional layers, two activation layers, and two feature restoration modules.
[0121] In some embodiments, the first feature restoration and the fourth feature restoration may be implemented by different neural networks. For example, the first neural network and the fourth neural network may have different network depths. For example, the number of convolutional layers included in the fourth neural network is smaller than the number of convolutional layers included in the first neural network, and / or the number of feature restoration modules in the fourth neural network is smaller than the number of feature restoration modules in the first neural network.
[0122] In some embodiments, the second feature restoration and the third feature restoration can be implemented by the same neural network. For example, the second neural network and the third neural network can include a convolutional layer.
[0123] In some embodiments, the fifth neural network may include a feature restoration module, a convolutional layer, and an activation layer.
[0124] It should be noted that the convolution layer involved in the decoding method provided in the embodiment of the present application may refer to a transposed convolution layer.
[0125] The following introduces the convolutional layer, feature restoration module and activation layer that may be included in the above neural network.
[0126] The first neural network may include two transposed convolutional layers (i.e., sparse transposed matrices), two rectified linear unit (ReLU) activation layers, and two feature restoration modules. The order of the modules from input to output may be transposed convolutional layer, rectified linear unit, transposed convolutional layer, rectified linear unit, feature restoration module, and feature restoration module.
[0127] To improve the accuracy of feature restoration results, the feature restoration module includes multiple feature restoration channels. In some embodiments, different feature restoration channels have different network depths to effectively restore different features, such as high-frequency features and low-frequency features.
[0128] In some embodiments, different feature restoration channels may have different corresponding weights to effectively restore local features. For example, when restoring high-frequency attribute features, the feature restoration channel for high-frequency attribute features may have a larger weight. Conversely, when restoring low-frequency features, the feature restoration channel for low-frequency features may have a larger weight.
[0129] In some embodiments, the weights of the feature restoration channels can be the same as the weights of the feature extraction channels during the encoding process. For example, a first parameter is obtained by parsing the bitstream; the first parameter indicates the weights of multiple feature extraction channels. Based on the first parameter, the weights of the multiple feature restoration channels can be determined.
[0130] In some embodiments, the weights of different feature restoration channels can be controlled using different weight modules. For example, the weight module can consist of an average pooling layer and a fully convolutional layer, using a sigmoid function to output weights. Experiments have shown that the weight module can effectively control the weight distribution between inputs, effectively helping to restore local features.
[0131] The decoding method provided by the embodiment of the present application is described in detail above in conjunction with Figure 8. The encoding method provided by the embodiment of the present application is described in detail below in conjunction with Figure 9.
[0132] FIG9 is a flow chart of an encoding method according to an embodiment of the present application. The encoding method of FIG9 can be applied to an encoder. The encoding method of FIG9 can be used to encode attribute information of a point cloud. The encoding method shown in FIG9 can include steps S910 to S930.
[0133] In step S910, the first point cloud is processed based on the attribute information to obtain a second point cloud, wherein the second point cloud occupies the same space as the first point cloud and the number of points in the second point cloud is less than the number of points in the first point cloud.
[0134] As mentioned above, the encoding method provided in the embodiments of the present application achieves multi-scale deep encoding of point cloud attribute information by performing feature extraction on multiple groups of point clouds. In some embodiments, a first point cloud and a second point cloud can be obtained based on the original point cloud to be encoded. For example, the first point cloud can be the original point cloud, and the second point cloud can occupy the same spatial size as the original point cloud, with the number of points in the second point cloud being less than the number of points in the original point cloud.
[0135] In some embodiments, the second point cloud may include some of the points in the first point cloud, or in other words, the points in at least a portion of the second point cloud are sparser than the points in the same portion of the first point cloud. In other words, the resolution of at least a portion of the second point cloud is lower than the resolution of the same portion of the first point cloud.
[0136] In some embodiments, a second point cloud can be determined based on attribute features of points in the first point cloud. For example, points in the second point cloud may correspond to high-frequency features in the first point cloud. High-frequency features, for example, may have the same meaning as high-frequency coefficients of attribute information in related art. For another example, points corresponding to high-frequency features may refer to points with significant changes in attribute information or a high rate of change in attribute information.
[0137] The encoding method provided in the embodiment of the present application can perform feature extraction on multiple groups of different point clouds respectively, and the above-mentioned first point cloud and second point cloud can be any two groups of associated point clouds in the multiple groups of point clouds. For example, the input of the encoding method provided in the embodiment of the present application may include four groups of point clouds (the first group of point clouds to the fourth group of point clouds), wherein the first point cloud and the second point cloud can be any two groups of associated point clouds in the four groups of point clouds. The associated point clouds mentioned here refer to those that can be obtained based on one group of point clouds based on another group of point clouds. As an example, the first group of point clouds can be the original point cloud, and the second group of point clouds can include some points in the original point cloud; the third group of point clouds can be obtained through the second group of point clouds, such as the third group of point clouds can include some points in the second group of point clouds; the fourth group of point clouds can be obtained through the third group of point clouds, such as the fourth group of point clouds can include some points in the third group of point clouds.
[0138] The more point cloud groups are input, the higher the accuracy of the encoding result, but the system complexity also increases. Therefore, the applicant found through experiments that if the input point cloud is four groups (including the original point cloud), the accuracy of the encoding result and the complexity of the system can be balanced, and the encoding performance is better.
[0139] In step S920, feature extraction is performed on the first point cloud and the second point cloud to obtain a first attribute feature and a second attribute feature, wherein the first attribute feature is the attribute feature of the first point cloud and the second attribute feature is the attribute feature of the second point cloud.
[0140] A complete point cloud may contain millions of points. Compressing all the attribute information of the complete point cloud together consumes a huge amount of network computing and storage. Therefore, in some embodiments, based on geometric information, the first point cloud and the second point cloud are spatially divided to obtain multiple first sub-blocks of the first point cloud and multiple second sub-blocks of the second point cloud; feature extraction is performed on the multiple first sub-blocks and the multiple second sub-blocks to obtain first attribute features and second attribute features. For example, feature extraction can be performed on the multiple first sub-blocks and the multiple second sub-blocks respectively, and then the feature extraction results can be spliced, merged, etc. to obtain the first attribute features and the second attribute features.
[0141] For example, a point cloud with arbitrary points N is divided into blocks using the same method as the octree division in G-PCC. It should be noted that the stopping condition for the division and blocking of the first point cloud and the second point cloud in the embodiment of the present application is different from the stopping condition of the octree division in G-PCC (stop when the division reaches the size of the child node is 1×1×1). As an example, the stopping condition for the division and blocking of the first point cloud and the second point cloud may be until the division reaches the size of the child node is 8×8×8. Each node of the octree represents a spatial area, so blocks can be defined according to a specific level of the octree, and each block contains point cloud data from a specific spatial area.
[0142] In order to avoid the impact of sub-block division on the accuracy of attribute features, additional processing can be performed on the attribute information at the edges of sub-blocks, such as extracting features again on the attribute information at the edges of multiple sub-blocks, or using different methods to extract features, or extracting features on the attribute information at the edges of multiple sub-blocks based on the parent block of the first sub-block and / or the parent block of the second sub-block.
[0143] In some embodiments, the spatial division level of the first point cloud is different from the spatial division level of the second point cloud. For example, the lower the resolution of the point cloud, the shallower the division level, such as the spatial division level of the second point cloud is lower than the spatial division level of the first point cloud.
[0144] Through research and experiments, it was found that the geometric information of a point cloud, i.e., the spatial coordinate position, and attribute information have a certain correlation. Utilizing this correlation to assist in the encoding of attribute information can help further compress the attribute information, thereby saving transmission and storage resources. For example, in a local area of a point cloud, such as a tiny slice, the correlation between the attribute information and coordinates of the point cloud can be represented by a functional relationship. The smaller the slice size, the higher the accuracy of this correlation. For another example, in the process of feature extraction of attribute information, the weights occupied by the change features between the attribute information of points at different distances from the current point and the attribute information of the current point are different. The distance from the current point can be determined based on geometric information. For example, the weight occupied by the change features between the attribute information of points farther away from the current point and the attribute information of the current point is smaller, while the weight occupied by the change features between the attribute information of points closer to the current point and the attribute information of the current point is larger.
[0145] As an example, based on the correlation between attribute information and geometric information, feature extraction is performed on the first point cloud and the second point cloud to obtain first attribute features and second attribute features. The correlation between attribute information and geometric information assists in the feature extraction process of the first and second point clouds, helping to further compress the attribute features of the first and second point clouds.
[0146] Since the decoding end is usually unable to obtain the original geometric information of the point cloud, the correlation between the attribute information and the geometric information can be obtained based on the geometric reconstruction information, which helps to improve the encoding accuracy. For example, the correlation between the attribute information and the geometric information can be determined based on the association between the geometric reconstruction information of the first point cloud (or the original point cloud) and the attribute information of the first point cloud and the second point cloud, respectively, so as to facilitate implementation. For another example, the geometric reconstruction information of the second point cloud can be obtained based on the geometric reconstruction information of the first point cloud. Determining the correlation between the attribute information and the geometric information based on the association between the geometric reconstruction information of the first point cloud and the attribute information of the first point cloud can assist in the encoding of the attribute information of the first point cloud; and determining the correlation between the attribute information and the geometric information based on the geometric reconstruction information of the second point cloud and the association between the attribute information of the second point cloud can assist in the encoding of the attribute information of the second point cloud.
[0147] By learning the geometric structure of point clouds and the correlation between geometric structure and attribute information, we can assist in encoding the attribute information of point clouds, thereby helping to improve the encoding accuracy of attribute information. For example, geometric information is used as the input of an auxiliary channel and extracted together with the attribute information of the point cloud to generate a compact representation. The feature extraction results can include detailed local geometric and texture features of the point cloud, reflecting subtle structural changes and detailed surface characteristics in the point cloud.
[0148] In some embodiments, before feature extraction is performed on the attribute information of the first point cloud, the first point cloud can be processed to remove attribute features in the first point cloud that are repeated with the second point cloud, thereby making the feature extraction of the first point cloud more targeted, or in other words, facilitating the full capture of the features at the resolution level of the point cloud itself. For example, feature extraction is first performed on the second point cloud to obtain the second attribute feature; secondly, the attribute information of the first point cloud is processed based on the second attribute feature to obtain residual information; finally, feature extraction is performed on the first point cloud based on the residual information to obtain the first attribute feature. Processing the attribute information of the first point cloud based on the second attribute feature may include dimensional processing of the attribute information of the first point cloud, as well as residual operations, etc.
[0149] In some embodiments, before performing feature extraction on the first point cloud and the second point cloud, the initial attribute information of the first point cloud and the second point cloud can be preprocessed to obtain attribute information of the first point cloud and the second point cloud, wherein the attribute information value of the same point is less than the initial attribute information value. The preprocessing mentioned here can, for example, be scaling the attribute information, such as dividing the attribute information value of the point cloud by a maximum value so that the attribute information is within the range of 0 to 1 before feature extraction. As an example, if the attribute information of the point cloud is 8 bits, then the maximum value of the attribute information is 255. By preprocessing the initial attribute information of the point cloud, the encoding performance of the attribute information can be improved.
[0150] In step S930, the first attribute feature and the second attribute feature are encoded respectively, and the obtained encoding bits are written into the first code stream segment and the second code stream segment respectively.
[0151] In some embodiments, the encoding method of the first attribute feature and the second attribute feature may adopt the entropy coding mentioned above.
[0152] In some embodiments, a codestream segment may have a codestream segment identifier, such as a codestream segment identifier that may be included in a syntax element table of a codestream. On the one hand, the codestream segment identifier can be used to distinguish different codestream segments; on the other hand, the codestream segment identifier can be used to indicate the point cloud associated with the codestream segment, or in other words, the codestream segment identifier is used to indicate which set of point cloud attribute information encoding results the current codestream segment carries. For example, a first codestream segment identifier can be used to indicate that the current codestream segment is a first codestream segment, associated with a first point cloud, or carries the attribute information encoding results of the first point cloud; a second codestream segment identifier can be used to indicate that the current codestream segment is a second codestream segment, associated with a second point cloud, or carries the attribute information encoding results of the second point cloud.
[0153] In some embodiments, the length of a codestream segment can be written into the codestream. For example, the start point of a codestream segment and its length can be written into the codestream. In another example, the start point and end point of a codestream segment can be written into the codestream. In another example, the start point, end point, and length of a codestream segment can be written into the codestream.
[0154] As mentioned above, the encoding method provided in the embodiments of this application can encode a variety of attribute information. Therefore, in some embodiments, an attribute type can be written into the bitstream. The attribute type can be used to indicate the type of attribute information carried by the current bitstream, such as color information or reflectivity information.
[0155] There are multiple ways to encode the attribute information of a point cloud. Therefore, in some embodiments, an encoding method parameter can be written into the code stream. The encoding method parameter can be used to indicate the encoding method of the attribute information carried by the current code stream. For example, if the encoding method provided in the embodiment of the present application is used to encode the attribute information of the point cloud, the value of the encoding method parameter can be the first value. If the encoding method provided in the embodiment of the present application is not used to encode the attribute information of the point cloud, the value of the encoding method parameter is not the first value, and can be the second value.
[0156] In the embodiments of the present application, by processing the point cloud to be encoded into multiple groups of different point clouds, different features of attribute information, such as high-frequency features and low-frequency features, can be effectively extracted. Furthermore, feature extraction can be performed specifically for the characteristics of different point clouds. By performing feature extraction on different point clouds through multiple channels, the characteristics of each group of point clouds can be extracted to improve the comprehensiveness and accuracy of the extracted attribute features, thereby helping to improve the encoding accuracy of the point cloud attribute information.
[0157] In some embodiments, the processing involved in the above-described point cloud encoding method can be implemented using a neural network. For example, the first attribute feature and the second attribute feature are obtained using the sixth neural network and the seventh neural network, respectively. In another example, the second point cloud is obtained using the first processing module, which can be implemented using a neural network.
[0158] In some embodiments, the sixth neural network and / or the seventh neural network may include one or more of the following: one or more convolutional layers; one or more activation layers; one or more offset attention (OA) modules; and one or more feature extraction modules.
[0159] When the network depth is deep, the self-attention module may not be able to handle the problem of information loss. Therefore, the sixth neural network and / or the seventh neural network may include a bias attention module to correct the features. The input-output relationship of the bias attention module can satisfy the following formula: Fout =OA(F in )=γ(F in -F sa )+F in (Q,K,V)=F in ·(W q ,W k ,W v )
[0160] Among them, F in is the input of the offset attention module, F out is the output of the offset attention module, W q , W k and W v is the input feature F in The query matrix, key matrix and value matrix generated by linear transformation, d k is the dimension of the key vector. Softmax is a mathematical function that is often used to convert a set of arbitrary real numbers into real numbers representing a probability distribution. γ can be an LBR network, i.e., a linear network layer + a batch norm layer + a ReLU layer.
[0161] Figure 10 shows a schematic diagram of the structure of a shifted attention module. The relationship between the input and output of the network structure of the shifted attention module shown in Figure 10 satisfies the above formula. Here, MLP stands for multi-layer perception network (MLP), and T stands for transposition module.
[0162] In some embodiments, the network depth of the sixth neural network is different from that of the seventh neural network. For example, for a point cloud with a larger number of points (also referred to as a high-resolution point cloud), a deeper network layer can be used to effectively extract high-frequency information therein; for a point cloud with a smaller number of points (also referred to as a lower-resolution point cloud), a shallower network depth can be used to avoid the extracted features being too concentrated in a local area. As an example, the number of one or more of the feature extraction modules, offset attention modules, activation layers, and convolutional layers included in the seventh neural network can be less than the number of corresponding modules included in the sixth neural network.
[0163] For example, the sixth neural network may include two convolutional layers, two ReLU activation layers, a feature extraction module, and two offset attention modules; the seventh neural network may include two convolutional layers, two ReLU activation layers, a feature extraction module, and a offset attention module. For another example, the sixth neural network may include two convolutional layers, two ReLU activation layers, a feature extraction module, and a offset attention module; the seventh neural network may include one convolutional layer, one ReLU activation layer, a feature extraction module, and one offset attention module.
[0164] In some embodiments, the feature extraction module may include multiple feature extraction channels. For example, the network depths of the multiple feature extraction channels may be different to effectively extract features of the point cloud. For another example, the weights of the multiple feature extraction channels may be different to extract different local features.
[0165] In some embodiments, the weights of different feature extraction channels can be controlled using different weight modules. For example, the weight module can consist of an average pooling layer and a fully convolutional layer, using a sigmoid function to output weights. Experiments have shown that the weight module can effectively control the weight distribution between inputs and effectively extract local features.
[0166] Since the decoding end needs to use the same feature extraction channel weights to correctly decode the attribute information, the first parameter can be written into the bitstream, and the first parameter can be used to indicate the weights of multiple feature extraction channels.
[0167] Figure 11 shows an example structure of a feature extraction module. As shown in Figure 11 , the feature extraction module includes three feature extraction channels (channels 1101 to 1103). Different feature extraction channels have different network depths. For example, channel 1101 does not include any neural network layers, channel 1102 includes eight neural network layers, and channel 1103 includes four neural network layers. The neural network layers include a ReLU activation layer and an SConv C×3*3 convolutional layer (with a 3*3 convolution kernel size).
[0168] Referring again to Figure 11, each channel of the feature extraction module also includes a weight module CAM. Figure 12 shows a schematic diagram of the structure of a weight module.
[0169] In some embodiments, the second point cloud is obtained by a first processing module, which includes one or more of the following: one or more perception layers; one or more global pooling layers; one or more self-attention (SA) layers; and one or more convolutional layers.
[0170] Figure 13 shows a schematic diagram of the components of the first processing module. This module effectively removes low-frequency smooth regions from the first point cloud, retaining points corresponding to high-frequency features. This facilitates subsequent effective feature extraction and encoding of high-frequency information. The convolution kernel size of the convolution layer in the first processing module is 1*1.
[0171] The attention mechanism is widely used to distinguish the importance of features, so the first processing module can use the self-attention mechanism to identify and focus on the areas that the network wants to learn during training. The input-output relationship of the self-attention module can satisfy the following formula: (Q, K, V) = F in ·(W q ,W k ,W v ), F out =γ(F sa )+F in ,
[0172] Among them, F in is the input of the self-attention module, F out is the output of the self-attention module, W q , W k and W v is the input feature F in The query matrix, key matrix and value matrix generated by linear transformation, d k is the dimension of the key vector. Softmax is a mathematical function that is often used to convert a set of arbitrary real numbers into real numbers representing a probability distribution. γ can be an LBR network, i.e., a linear network layer + a batch norm layer + a ReLU layer.
[0173] Figure 14 shows a schematic diagram of the structure of a self-attention module. The relationship between the input and output of the self-attention module shown in Figure 14 satisfies the above formula.
[0174] In some embodiments, the parameters of the neural network involved in the encoding method provided in the embodiments of this application can be adjusted based on the encoding loss of the first point cloud and the second point cloud to improve encoding accuracy. For example, the encoding loss of the first point cloud and the second point cloud can be first obtained; then, the parameters of the neural network used in the encoding method can be adjusted based on the encoding loss.
[0175] In some embodiments, the encoding losses of the first point cloud and the second point cloud may be associated with one or more of: an attribute encoding loss; a structure encoding loss; and an entropy encoding loss.
[0176] Attribute encoding loss can also be called feature fidelity loss. Feature fidelity loss is used to ensure that the decoded attribute information maintains similarity with the original point cloud attributes in the feature space. The feature fidelity loss L FFL The following formula can be satisfied:
[0177] Where f(x) represents the original point cloud features, Represents the features of the decoded point cloud, N is the number of points, and i represents the i-th point.
[0178] As mentioned above, embodiments of the present application can use entropy coding to encode attribute features. In some embodiments, the first attribute feature and the second attribute feature can be entropy coded separately, that is, two entropy coding operations can be performed. Taking the aforementioned input point cloud comprising four groups of point clouds as an example, four entropy coding operations can be used to encode the attribute features of different point clouds respectively.
[0179] If entropy coding is used to encode attribute features, the encoding loss of the first point cloud and the second point cloud can be associated with the entropy coding loss. For example, the binary cross-entropy (BCE) loss can be used to measure the difference between the actual probability distribution p(R) and the model probability distribution q i (R) the difference between.
[0180] Since the first point cloud and the second point cloud are encoded separately, the entropy coding loss of the first point cloud and the entropy coding loss of the second point cloud can be calculated separately. i (R) represents, where i = 1, 2, 3, 4. Therefore, the entropy coding loss is L EL Satisfies the following formula: L EL =w1l1+w2l2+w3l3+w4l4
[0181] Among them, l i =-∑p(R)*log[q i (R)],i=1,2,3,4.
[0182] In some embodiments, encoding accuracy can be improved by controlling the contribution of each entropy model to the final loss. For example, the entropy model corresponding to the first point cloud may have a greater weight than the entropy model corresponding to the second point cloud. In another example, the entropy model corresponding to the original point cloud may have the largest weight.
[0183] If geometric information is used to assist in encoding the attribute information of the point cloud, the encoding loss of the first point cloud and the encoding loss of the second point cloud can be associated with the structural encoding loss. For example, the structural loss L SL The following formula can be satisfied:
[0184] Where N represents the number of points, i and j represent the i-th point and the j-th point.
[0185] If the coding loss is related to attribute coding loss, structure coding loss and entropy coding loss, then the total coding loss can be the weighted sum of the above three coding losses, that is, the total coding loss L total Satisfies the following formula: L total =αL EL +βL FFL +γL SL
[0186] Among them, α, β, and γ can be hyperparameters, that is, exponential decay weights are used, and the weight of each loss term will gradually change as the training progresses.
[0187] In some embodiments, the termination condition for neural network training may include one or more of the following: a total loss function is less than a preset threshold, an attribute coding loss is less than a preset threshold, a structure coding loss is less than a preset threshold, and an entropy coding loss is less than a preset threshold. For example, the termination condition for training may include the total loss function is less than a preset threshold. For another example, the termination condition for training may include the total loss function is less than a preset threshold, and the attribute coding loss is less than a preset threshold, the structure coding loss is less than a preset threshold, and the entropy coding loss is less than a preset threshold.
[0188] In some embodiments, the datasets for training the neural network can use the 8i dataset and the MVUB dataset specified by MPEG. For example, in the training process of the neural network, the Queen, Dancer, Boxer and All sequences of are used as training datasets, and all sequences of Longdress, Soldier, Loot, RedandBlack, David, Ricardo, Sarah and Andrew are used as test sets.
[0189] It should be noted that the embodiment of the present application does not limit the encoding method of geometric information.
[0190] The embodiment of the present application proposes a new end-to-end encoding method for point cloud attribute information, which helps to improve the compression accuracy of point cloud attribute information by performing targeted feature extraction on different point clouds.
[0191] Based on this, the present invention proposes a point cloud encoding and decoding framework. FIG15A is a schematic diagram of a point cloud encoding framework provided by the present invention, and FIG15B is a schematic diagram of a point cloud decoding framework provided by the present invention.
[0192] As shown in FIG15A , the input original point cloud is subjected to geometric encoding and attribute encoding respectively to generate a geometric information bit stream and an attribute information bit stream respectively.
[0193] The geometric coding method can adopt the G-PCC point cloud coding framework in related technologies. The geometric coding process can include preprocessing, octree-based coding, and arithmetic coding.
[0194] Preprocessing includes coordinate conversion and voxelization. Coordinate conversion transforms geometric coordinates in the "world coordinate system" into the "internal coordinate system" required by the encoder, allowing the point cloud to be contained within a bounding box for easier processing. Voxelization converts the converted geometric information from floating-point numbers to integers, facilitating subsequent octree partitioning and encoding.
[0195] After preprocessing, the coordinates of all points are in the range {0, 2 d}, the entire point cloud is in a circle with a side length of 2 d . At the same octree depth, a node is divided into 8 child nodes. If 1 indicates an occupied child node and 0 indicates an unoccupied child node, the occupancy status of the 8 child nodes can be used to form an 8-bit placeholder code. Only nodes that are non-empty and larger than 1×1×1 will be recursively octree-partitioned until the size of the child nodes is 1×1×1.
[0196] Taking the 8-bit placeholder code (b0b1b2b3b4b5b6b7) of the current coding node (i.e., the geometric information of the current coding node) as an example, conditional entropy is used to further improve compression efficiency in order to exploit local geometric correlation. Specifically, the first bit b0 can be directly encoded by the binary arithmetic encoder, the second bit b1 is encoded based on the value of b0, and bit b2 is encoded based on the value of b0b1, and so on. There are a total of 128 possible encoding options for bit b7.
[0197] Before attribute encoding of the original point cloud, the attribute information can be scaled to reduce storage and computational loads. Secondly, before attribute encoding of the original point cloud, the attribute information of the original point cloud (such as the scaled attribute information) can be spatially partitioned, such as by octree partitioning, to further reduce storage and computational loads. During attribute encoding, geometric reconstruction information can be used to assist in the encoding of attribute information. For example, attribute information can be encoded based on the correlation between geometric reconstruction information and attribute information.
[0198] It should be noted that the point cloud coding framework in FIG15A may include multiple attribute coding channels (not shown in the figure) to compress attribute information of different point clouds.
[0199] As shown in FIG15B , the input geometry bitstream and attribute bitstream may be parsed separately to obtain geometry reconstruction information and attribute reconstruction information.
[0200] The geometric decoding method corresponds to geometric coding and can use G-PCC's octree decoding. Based on the binary code stream obtained by octree coding, the geometric coordinates are restored layer by layer from the first layer to the last layer. For example, the first layer decodes from the header information to obtain the edge length of the current layer node as (256,256,256), and the edge length of the sub-node layer node is (128,128,128). The first layer code stream is 10000001. It can be calculated that the child nodes 0 and 7 of the first root node are occupied. The coordinates of child node 0 are (0+0,0+0,0+0), and the coordinates of child node 7 are (0+128,0+128,0+128). The decoding of geometric information is achieved by analogy.
[0201] During the decoding of the attribute bitstream, geometric reconstruction information can be used to assist in decoding the attribute information. For example, attribute information can be decoded based on the correlation between the geometric reconstruction information and the attribute information. Because the point cloud is spatially partitioned before attribute encoding, block merging is required after the attribute information is decoded to stitch together the attribute information from different parts. Furthermore, after completing the attribute information block merging, the attribute information needs to be inversely scaled to obtain the reconstructed attribute information.
[0202] Fig. 16A is a structural example diagram of the attribute encoding module in Fig. 15A. The input of the attribute encoding module may include the original point cloud and the first point cloud, that is, the attribute encoding module may include two attribute encoding channels.
[0203] The first processing unit processes the first point cloud to obtain a second point cloud, wherein the number of points in the second point cloud is smaller than the number of points in the first point cloud. The first processing unit may process the first point cloud based on the characteristics of the attribute information of the first point cloud. For example, the second point cloud may include points corresponding to high-frequency attribute features in the first point cloud.
[0204] The preprocessing unit may include a first preprocessing unit and a second preprocessing unit. In some embodiments, the preprocessing unit may be used to scale the attribute information of the point cloud to improve the training speed of the neural network involved in the encoding method provided in this application and to improve the performance of the neural network. In some embodiments, the preprocessing unit may also be used to perform spatial division on the input point cloud, such as octree division, to obtain a first sub-block and a second sub-block. Each node of the octree represents a spatial area, so the first sub-block and the second sub-block may be determined according to a specific level of the octree, and each sub-block may contain point cloud data from a specific spatial area. In other words, the output of the preprocessing unit may include the spatial coordinate range and size of the area, as well as point cloud data with additional attributes such as color and intensity.
[0205] The feature extraction unit may include a first feature extraction unit and a second feature extraction unit to extract attribute features from the input point cloud to obtain first attribute features and second attribute features, respectively. Because the second point cloud may include points corresponding to high-frequency features in the first point cloud, performing feature extraction on the first point cloud and the second point cloud separately can extract different attribute features in a targeted manner.
[0206] In some embodiments, geometric information and attribute information may be input into the feature extraction unit simultaneously, wherein geometric information, such as geometric reconstruction information, may assist in feature extraction of attribute information.
[0207] Before performing feature extraction on the first point cloud, the second processing unit may process the first point cloud to obtain residual information. For example, the second processing unit may first convert the dimension of the second attribute feature so that the dimension of the second attribute feature is the same as the dimension of the first point cloud attribute information; and then perform a residual operation to obtain the residual information.
[0208] The encoding unit may encode the first attribute feature and the second attribute feature respectively, and write the encoding results into the first code stream segment and the second code stream segment respectively.
[0209] Figure 16B illustrates the structure of the attribute decoding module in Figure 15B. When decoding the attribute information of a point cloud, the attribute information of the first point cloud (i.e., the first bitstream segment) and the attribute information of the second point cloud (i.e., the second bitstream segment) are decoded separately. In other words, the attribute decoding module's input may include the first bitstream segment and the second bitstream segment, and the first bitstream segment and the second bitstream segment are decoded separately.
[0210] The first feature restoration unit processes the first code stream segment to obtain initial feature information. The second feature restoration unit processes the initial feature information to obtain initial attribute information. The initial attribute information is associated with the geometric reconstruction information to obtain an initial reconstructed point cloud.
[0211] At the same time, the feature restoration unit processes the second stream segment to obtain second attribute information. Since the first stream segment and the second stream segment carry different types of attribute features, the attribute information can be updated and supplemented by combining the second attribute information with the initial feature information and the initial attribute information to obtain attribute reconstruction information.
[0212] Optionally, the initial feature information and the initial attribute information may first be processed by a third processing unit to obtain first attribute information. The operation performed by the third processing unit may include, for example, the update operation mentioned above. Secondly, the first attribute information may be restored by a fourth feature restoration unit to obtain third attribute information. The third attribute information and the second attribute information may then be processed by the fourth processing unit, such as by performing a residual operation. Subsequently, the residual information output by the fourth processing unit may be restored by the third feature restoration unit to obtain attribute reconstruction information.
[0213] Finally, the attribute reconstruction information is associated with the geometric reconstruction information to obtain the first reconstructed point cloud.
[0214] It should be noted that the processing units involved in the attribute decoding process and the attribute encoding process can be composed of neural networks. The neural network involved in the attribute decoding process can be a mirror network of the neural network involved in the attribute encoding process (except for the offset-attention module), and the parameters of the neural network can be shared.
[0215] Figure 17 is an example diagram of a point cloud encoding and decoding process provided by an embodiment of the present application. Referring to Figure 17, the input point cloud on the encoding side may include four groups of point clouds (hereinafter referred to as four groups of point clouds of different resolutions), each of which contains a different number of points. The number of points in each group of point clouds gradually decreases from the original point cloud downwards.
[0216] The other three groups of point clouds except the original point cloud can be obtained by the first processing unit mentioned above.
[0217] First, the original point cloud can be preprocessed, that is, the input of the feature extraction unit 1701 is converted into geometric information and color information of sparse tensors Here we use F 0 Represents the geometric information tensor, F 1 -F 3 Represents the color information tensor. Geometric information can be input as an auxiliary channel to provide correlation between points, thereby assisting in the compression of attribute information.
[0218] The preprocessing process can also include octree partitioning of the original point cloud based on geometric information. The octree partitioning level for high-resolution point clouds is higher than that for low-resolution point clouds. While different levels of partitioning for different resolution point clouds may result in a certain degree of information loss, experiments have shown that this information loss is within an acceptable range and has a positive impact on subsequent encoding and decoding performance.
[0219] The four groups of point clouds after preprocessing are respectively sent to the feature extraction unit 1701 to the feature extraction unit 1704 to extract features of different frequencies. For example, the features extracted by the feature extraction unit 1701 to the feature extraction unit 1704 can be respectively recorded as f 0 -f 3 .
[0220] For high-resolution point clouds, deeper network layers can be used to effectively extract high-frequency information. For lower-resolution point clouds, increasingly shallower network layers can be designed in sequence to avoid the extracted features being too concentrated in local areas. Feature extraction unit 1701 may include 2 sparse convolution layers, 2 Relu activation layers, a feature extraction module (A3CR-IRN) and two OA modules. Feature extraction unit 1702 may include 2 sparse convolution layers, 2 Relu activation layers, a feature extraction module (A3CR-IRN) and an OA module. Feature extraction unit 1703 may include 2 sparse convolution layers, 2 Relu activation layers, a feature extraction module (A3CR-IRN) and an OA module. Feature extraction unit 1704 may include 1 sparse convolution layer, 1 Relu activation layer, a feature extraction module (A3CR-IRN) and an OA module.
[0221] Optionally, before feature extraction, each high-resolution point cloud can first perform a residual operation with the features extracted from the lower-resolution point cloud, so as to help it fully capture the features of its own resolution level.
[0222] Finally, f 0 -f 3 They are sent to the entropy model for entropy coding to obtain 4 attribute code stream segments
[0223] As mentioned above, geometric information can assist in encoding point cloud attribute information. FIG18 is a schematic diagram of the geometry-assisted encoding provided by an embodiment of the present application. Referring to FIG18, by utilizing the geometric structure information of the point cloud data, a better understanding of the intrinsic shape of the data is provided, thereby improving the accuracy of encoding and decoding. The geometric information is used as the input of the auxiliary channel together with the attribute information of the original point cloud (i.e., the F 0 -F 3) for feature extraction and serve together as the input of the encoder to generate a compact representation.
[0224] Referring back to Figure 17, the attribute decoding of the point cloud adopts a layer-by-layer recovery method from coarse to fine. Reconstruction restores the initial reconstructed attribute information, which is combined with the feature code stream extracted from the point cloud at the next resolution After the combination, the next attribute reconstruction is performed. This reconstruction is repeated three times until the best attribute information is reconstructed. The decoded attribute information is associated with the decoded geometric information to obtain the final reconstructed point cloud.
[0225] Referring to Figure 17, module 1705 can be the first feature restoration unit mentioned in Figure 16B, module 1706 can be the second feature restoration unit mentioned in Figure 16B, module 1707 can be the fourth feature restoration unit mentioned in Figure 16B, module 1708 can be the third feature restoration unit mentioned in Figure 16B, module 1709 can be the feature restoration unit mentioned in Figure 16B, module 1710 can be the third processing unit mentioned in Figure 16B, and module 1711 can be the fourth processing unit mentioned in Figure 16B.
[0226] The following introduces the input and output dimensions of the neural network involved in Figure 17.
[0227] The input dimension of the feature extraction unit 1701 is (N, 4), corresponding to F 0 -F 3 , where N represents the number of points in the point cloud. The first convolution layer consists of 16 3×3 convolution kernels with a stride of 1 and no padding, resulting in an output dimension of (N, 16). The second convolution layer consists of 32 3×3 convolution kernels with a stride of 1 and no padding, resulting in an output dimension of (N, 32). The output dimension of A3CR-INR is (N, 256). The output dimension of the OA module is (N, 256).
[0228] The input dimensions of the feature extraction unit 1702 and the feature extraction unit 1703 are (N, 4), corresponding to F 0 -F 3 , where N represents the number of points in the point cloud. The first convolution layer consists of 16 3×3 convolution kernels with a stride of 1 and no padding, resulting in an output dimension of (N, 16). The second convolution layer consists of 32 3×3 convolution kernels with a stride of 1 and no padding, resulting in an output dimension of (N, 32). The output dimension of A3CR-INR is (N, 256). The output dimension of the OA module is (N, 256).
[0229] The input dimension of the feature extraction unit 1704 is (N, 4), corresponding to F 0 -F 3, where N represents the number of points in the point cloud. The first convolution layer consists of 16 3×3 kernels with a stride of 1 and no padding, resulting in an output dimension of (N, 16). The output dimension of A3CR-INR is (N, 128). The output dimension of the OA module is (N, 128).
[0230] The input dimension of module 1705 is (N, 256). The sparse transposed matrix of the first layer, SConv C×3*3, is (k, k, 256, 128), where k is the convolution kernel size, which is 3 here. The output dimension of the ReLU activation layer is (N, 128). The sparse transposed matrix of the second layer, (k, k, 128, 64), where k is the convolution kernel size, which is 3 here, has an output dimension of (N, 64). The output dimension of the second feature restoration module, A3CR-INR, is (N, 16).
[0231] Module 1706 includes a sparse transposed matrix (k, k, 16, 3), where k is the convolution kernel size, which is 1 here, and the output dimension is (N, 3).
[0232] The input dimension of module 1707 is (N, 16). The first layer sparse transposed matrix is (k, k, 16, 16), where k is the convolution kernel size, which is 3 here, and the padding is 1. The second layer sparse transposed matrix is (k, k, 16, 16), where k is the convolution kernel size, which is 3 here, and the padding is 1. The output dimension of the second layer A3CR-INR is (N, 16).
[0233] The input dimension of module 1709 is (N, 256). The first layer sparse transposed matrix is (k, k, 256, 128), where k is the convolution kernel size, which is 3 here, and padded with 0. The output dimension of A3CR-INR is (N, 16).
[0234] Module 1708 includes a sparse transposed matrix (k, k, 16, 3), where k is the convolution kernel size, which is 1 here, padded with 0, and outputs (N, 3).
[0235] It should be noted that the decoding method of the third and fourth code stream segments in FIG. 17 is similar to the decoding method of the first and second code stream segments, and for the sake of brevity, they are not described here in detail.
[0236] The method embodiment of the present application is described in detail above in conjunction with Figures 1 to 18 , and the device embodiment of the present application is described in detail below in conjunction with Figures 19 to 22 . It should be understood that the description of the method embodiment corresponds to the description of the device embodiment, and therefore, for portions not described in detail, reference can be made to the above method embodiment.
[0237] FIG19 is a schematic diagram of the structure of a decoder provided by an embodiment of the present application. As shown in FIG19 , the decoder 1900 may include a parsing unit 1910 , a first feature restoration unit 1920 , a second feature restoration unit 1930 , and a processing unit 1940 .
[0238] The parsing unit 1910 is configured to parse the code stream to obtain the first attribute feature and the second attribute feature.
[0239] The first feature restoration unit 1920 is configured to perform feature restoration on the first attribute feature to obtain first attribute information.
[0240] The second feature restoration unit 1930 is configured to perform feature restoration on the second attribute feature to obtain second attribute information.
[0241] The processing unit 1940 is configured to process the first attribute information and the second attribute information to obtain attribute reconstruction information.
[0242] In some embodiments, the parsing unit is further configured to parse the code stream to obtain a code stream segment identifier; wherein the code stream segment identifier is used to indicate a point cloud associated with the code stream segment.
[0243] In some embodiments, if the code stream segment identifier indicates that the first code stream segment is associated with the first point cloud, and the second code stream segment is associated with the second point cloud, then parsing the code stream to obtain the first attribute feature and the second attribute feature includes: parsing the code stream to obtain a code stream length identifier; and parsing the first code stream segment and the second code stream segment respectively based on the code stream length identifier to obtain the first attribute feature and the second attribute feature.
[0244] In some embodiments, the feature restoration of the first attribute feature to obtain the first attribute information includes: performing a first feature restoration on the first attribute feature to obtain initial feature information; performing a second feature restoration on the initial feature information to obtain initial attribute information; and updating the initial attribute information based on the initial feature information to obtain the first attribute information.
[0245] In some embodiments, the processing of the first attribute information and the second attribute information to obtain attribute reconstruction information includes: processing the first attribute information and the second attribute information to obtain first residual information; and performing third feature restoration on the first residual information to obtain the attribute reconstruction information.
[0246] In some embodiments, the processing of the first attribute information and the second attribute information to obtain the first residual information includes: performing a fourth feature restoration on the first attribute information to obtain the third attribute information; and processing the third attribute information and the second attribute information to obtain the first residual information.
[0247] In some embodiments, the parsing unit is further configured to: parse the code stream to obtain a coding mode parameter; and if the coding mode parameter is a first value, decode the code stream using the decoding method.
[0248] In some embodiments, the parsing unit is further configured to: parse the code stream to obtain geometric information; and process the geometric information and the attribute reconstruction information to obtain a reconstructed first point cloud.
[0249] In some embodiments, the first feature restoration is implemented by a first neural network, the second feature restoration is implemented by a second neural network, the second neural network includes one or more convolutional layers; and / or, the first neural network includes one or more of the following: one or more convolutional layers; one or more activation layers; and one or more feature restoration modules.
[0250] In some embodiments, the third feature restoration is implemented through a third neural network, the fourth feature restoration is implemented through a fourth neural network, the second attribute information is obtained through a fifth neural network, the third neural network includes one or more convolutional layers, and / or the fourth neural network and / or the fifth neural network includes one or more of the following: one or more convolutional layers; one or more activation layers; and one or more feature restoration modules.
[0251] In some embodiments, the feature restoration module includes multiple feature restoration channels, and the network depths of the multiple feature restoration channels are different.
[0252] In some embodiments, the weights corresponding to the multiple feature restoration channels are different.
[0253] In some embodiments, the parsing unit is further used to: parse the code stream to obtain a first parameter; wherein the first parameter is used to indicate the weights of the multiple feature extraction channels.
[0254] It is understandable that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and of course it can also be a module, or it can be non-modular. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.
[0255] If the integrated unit is implemented in the form of a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0256] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the decoder 1900. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the decoding method described in any one of the aforementioned embodiments.
[0257] Based on the composition of the above-mentioned decoder 1900 and the computer-readable storage medium, refer to Figure 20, which shows a specific hardware structure diagram of the encoder 1900 provided in an embodiment of the present application. As shown in Figure 20, the encoder 2000 may include: a communication interface 2010, a memory 2020 and a processor 2030; each component is coupled together through a bus system 2040. It can be understood that the bus system 2040 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 2040 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as bus systems 2040 in Figure 19. Among them,
[0258] Communication interface 2010, used for sending and receiving signals when sending and receiving information with other external network elements;
[0259] Memory 2020, for storing computer programs;
[0260] The processor 2030 is configured to, when running the computer program, execute:
[0261] Parse the code stream to obtain a first attribute feature and a second attribute feature; perform feature restoration on the first attribute feature to obtain first attribute information; perform feature restoration on the second attribute feature to obtain second attribute information; and process the first attribute information and the second attribute information to obtain attribute reconstruction information.
[0262] It is understood that the memory 2020 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronized dynamic random access memory (SLDRAM), and direct rambus RAM (DRRAM). The memory 2020 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0263] Processor 2030 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 2030. The above-mentioned processor 2030 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 2020, and the processor 2030 reads the information in the memory 2020 and completes the steps of the above method in combination with its hardware.
[0264] It is understood that the embodiments described herein can be implemented with hardware, software, firmware, middleware, microcode or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein or a combination thereof. For software implementation, the technology described herein can be implemented by a module (such as a process, a function, etc.) that performs the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in a processor or outside a processor.
[0265] Optionally, as another embodiment, the processor 2030 is further configured to execute the decoding method described in any one of the aforementioned embodiments when running the computer program.
[0266] FIG21 is a schematic diagram of the structure of an encoder provided by an embodiment of the present application. As shown in FIG21 , the encoder 2100 includes a processing unit 2110 , a feature extraction unit 2120 , and an encoding unit 2130 .
[0267] The processing unit 2110 is used to process the first point cloud based on the attribute information to obtain a second point cloud, where the second point cloud occupies the same space as the first point cloud, and the number of points in the second point cloud is less than the number of points in the first point cloud.
[0268] The feature extraction unit 2120 is used to perform feature extraction on the first point cloud and the second point cloud respectively to obtain a first attribute feature and a second attribute feature, wherein the first attribute feature is the attribute feature of the first point cloud, and the second attribute feature is the attribute feature of the second point cloud.
[0269] The encoding unit 2130 is configured to encode the first attribute feature and the second attribute feature respectively, and write the obtained encoding bits into the first code stream segment and the second code stream segment respectively.
[0270] In some embodiments, the points in the second point cloud are points corresponding to high-frequency attribute features in the first point cloud.
[0271] In some embodiments, the feature extraction is performed on the first point cloud and the second point cloud respectively to obtain the first attribute feature and the second attribute feature, including: based on geometric information, spatially dividing the first point cloud and the second point cloud respectively to obtain multiple first sub-blocks of the first point cloud and multiple second sub-blocks of the second point cloud; and feature extraction is performed on the multiple first sub-blocks and the multiple second sub-blocks respectively to obtain the first attribute feature and the second attribute feature.
[0272] In some embodiments, the spatial division level of the first point cloud is different from the spatial division level of the second point cloud.
[0273] In some embodiments, the feature extraction of the first point cloud and the second point cloud respectively to obtain the first attribute feature and the second attribute feature includes: based on the correlation between attribute information and geometric information, the feature extraction of the first point cloud and the second point cloud respectively to obtain the first attribute feature and the second attribute feature.
[0274] In some embodiments, the correlation between the attribute information and the geometric information is determined based on associations between the geometric reconstruction information of the first point cloud and the attribute information of the first point cloud and the second point cloud, respectively.
[0275] In some embodiments, the feature extraction is performed on the first point cloud and the second point cloud respectively to obtain the first attribute feature and the second attribute feature, including: feature extraction on the second point cloud to obtain the second attribute feature; processing the attribute information of the first point cloud based on the second attribute feature to obtain residual information; and feature extraction on the first point cloud based on the residual information to obtain the first attribute feature.
[0276] In some embodiments, the encoder further includes a preprocessing unit for preprocessing the initial attribute information of the first point cloud and the second point cloud before extracting features from the first point cloud and the second point cloud, respectively, to obtain the attribute information of the first point cloud and the second point cloud, wherein the attribute information value of the same point is less than the initial attribute information value.
[0277] In some embodiments, the first attribute feature and the second attribute feature are obtained through a sixth neural network and a seventh neural network, respectively, and the network depths of the sixth neural network and the seventh neural network are different.
[0278] In some embodiments, the sixth neural network and / or the seventh neural network includes one or more of the following: one or more convolutional layers; one or more activation layers; one or more offset attention modules; and one or more feature extraction modules.
[0279] In some embodiments, the feature extraction module includes multiple feature extraction channels, and the network depths of the multiple feature extraction channels are different.
[0280] In some embodiments, the weights corresponding to the multiple feature extraction channels are different.
[0281] In some embodiments, the encoder further includes a writing unit for writing a first parameter into a bitstream, where the first parameter is used to indicate weights of the multiple feature extraction channels.
[0282] In some embodiments, the second point cloud is obtained by a first processing module, which includes one or more of the following: one or more perception layers; one or more global pooling layers; one or more self-attention layers; and one or more convolutional layers.
[0283] In some embodiments, the encoder further includes an acquisition unit for acquiring encoding losses of the first point cloud and the second point cloud; and an adjustment unit for adjusting parameters of the neural network used in the encoding method based on the encoding losses.
[0284] In some embodiments, the first attribute feature and the second attribute feature are encoded using an entropy coding model, and the coding loss is associated with one or more of the following: attribute coding loss; structure coding loss; and entropy coding loss.
[0285] In some embodiments, the writing unit is further used to write one or more of the following parameters into the bitstream: the length of the bitstream segment; a bitstream segment identifier, used to indicate the point cloud associated with the bitstream segment; an attribute type, used to indicate the type of attribute information carried by the current bitstream; and an encoding method parameter, used to indicate the encoding method of the attribute information carried by the current bitstream.
[0286] In some embodiments, if the encoding method is used to encode the attribute information of the point cloud, the value of the encoding method parameter is a first value.
[0287] In some embodiments, the first point cloud is an original point cloud.
[0288] It is understandable that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and of course it can also be a module, or it can be non-modular. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.
[0289] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, ROM, RAM, a magnetic disk, or an optical disk.
[0290] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the encoder 2000. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the decoding method described in any one of the aforementioned embodiments.
[0291] Based on the composition of the above-mentioned encoder 2100 and the computer-readable storage medium, refer to Figure 22, which shows a specific hardware structure diagram of the encoder 2100 provided in an embodiment of the present application. As shown in Figure 22, the encoder 2200 may include: a communication interface 2210, a memory 2220 and a processor 2230; each component is coupled together through a bus system 2240. It can be understood that the bus system 2240 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 2240 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as bus systems 2240 in Figure 22. Among them,
[0292] Communication interface 2210, used for sending and receiving signals during the process of sending and receiving information with other external network elements;
[0293] Memory 2220, used for storing computer programs;
[0294] The processor 2230 is configured to, when running the computer program, execute:
[0295] Determining a target coding mode corresponding to a current layer from a plurality of coding modes, the plurality of coding modes comprising RAHT transform coding and RAHT prediction combined with transform coding;
[0296] The attribute information of the nodes of the current layer is encoded according to the target coding mode.
[0297] It will be appreciated that the memory 2220 in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a ROM, PROM, EPROM, EEPROM, or flash memory. The volatile memory may be a RAM, which serves as an external cache. By way of example and not limitation, many forms of RAM are available, such as SRAM, DRAM, SDRAM, DDRSDRAM, ESDRAM, SLDRAM, and DRRAM. The memory 2220 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0298] Processor 2230 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method may be performed by hardware integrated logic circuits within processor 2230 or by software instructions. Processor 2230 may be a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. It may implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of this application may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules within the decoding processor. The software modules may be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 2220. Processor 2230 reads information from memory 2220 and, in conjunction with its hardware, completes the steps of the above method.
[0299] It is understood that the embodiments described herein can be implemented with hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more ASICs, DSPs, DSPDs, PLDs, FPGAs, general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or a combination thereof. For software implementation, the technology described herein can be implemented by modules (e.g., processes, functions, etc.) that perform the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0300] Optionally, as another embodiment, the processor 2230 is further configured to execute the encoding method described in any one of the aforementioned embodiments when running the computer program.
[0301] An embodiment of the present application also provides a computer-readable storage medium, which is a non-volatile computer-readable storage medium for storing a bit stream. The bit stream can be generated by an encoding method of an encoder, or the bit stream can be decoded by a decoding method of a decoder, wherein the decoding method can be the decoding method described in any of the foregoing embodiments, and the encoding method can be the encoding method described in any of the foregoing embodiments.
[0302] It should be noted that, in this application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0303] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0304] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0305] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0306] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0307] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A decoding method, applied to a decoder, comprising: Parsing a bitstream to obtain a first attribute feature and a second attribute feature; Performing feature restoration on the first attribute feature to obtain first attribute information; Performing feature restoration on the second attribute feature to obtain second attribute information; Processing the first attribute information and the second attribute information to obtain attribute reconstruction information.
2. The method according to claim 1, wherein The method further comprises: Parsing the bitstream to obtain a bitstream segment identifier; wherein the bitstream segment identifier is used to indicate the point cloud associated with the bitstream segment.
3. The method according to claim 2, wherein, If the bitstream segment identifier indicates that a first bitstream segment is associated with a first point cloud and a second bitstream segment is associated with a second point cloud, then the parsing the bitstream to obtain the first attribute feature and the second attribute feature comprises: Parsing the bitstream to obtain a bitstream length identifier; Based on the bitstream length identifier, respectively parsing the first bitstream segment and the second bitstream segment to obtain the first attribute feature and the second attribute feature.
4. The method according to any one of claims 1 to 3, wherein The performing feature restoration on the first attribute feature to obtain first attribute information comprises: Performing a first feature restoration on the first attribute feature to obtain initial feature information; Performing a second feature restoration on the initial feature information to obtain initial attribute information; Updating the initial attribute information based on the initial feature information to obtain the first attribute information.
5. The method according to any one of claims 1 - 4, wherein The processing the first attribute information and the second attribute information to obtain attribute reconstruction information comprises: Processing the first attribute information and the second attribute information to obtain first residual information; Performing a third feature restoration on the first residual information to obtain the attribute reconstruction information.
6. The method according to claim 5, wherein The processing the first attribute information and the second attribute information to obtain first residual information comprises: Performing a fourth feature restoration on the first attribute information to obtain third attribute information; Processing the third attribute information and the second attribute information to obtain the first residual information.
7. The method according to any one of claims 1-6, wherein The method further comprises: Parsing the bitstream to obtain an encoding mode parameter; If the value of the encoding mode parameter is a first value, then decoding the bitstream using the decoding method.
8. The method according to any one of claims 1-7, wherein, The method further comprises: Parsing the bitstream to obtain geometric information; Processing the geometric information and the attribute reconstruction information to obtain a reconstructed first point cloud.
9. The method according to claim 4, wherein The first feature restoration is implemented by a first neural network, the second feature restoration is implemented by a second neural network, and the second neural network includes one or more convolutional layers; and / or, the first neural network includes one or more of the following: One or more convolutional layers; One or more activation layers; And One or more feature restoration modules.
10. The method according to claim 6, wherein, The third feature restoration is implemented by a third neural network, the fourth feature restoration is implemented by a fourth neural network, the second attribute information is obtained by a fifth neural network, the third neural network includes one or more convolutional layers, and / or, the fourth neural network and / or the fifth neural network includes one or more of the following: One or more convolutional layers; One or more activation layers; and One or more feature restoration modules.
11. The method according to claim 9 or 10, wherein, The feature restoration module includes a plurality of feature restoration channels with different network depths.
12. The method according to claim 11, wherein The weights corresponding to the plurality of feature restoration channels are different.
13. The method according to claim 12, wherein, The method further includes: Analyzing the bitstream to obtain a first parameter; wherein the first parameter is used to indicate the weights of the plurality of feature extraction channels.
14. An encoding method applied to an encoder, including: Processing a first point cloud based on attribute information to obtain a second point cloud, where the second point cloud occupies the same spatial size as the first point cloud, and the number of points in the second point cloud is less than the number of points in the first point cloud; Performing feature extraction on the first point cloud and the second point cloud respectively to obtain a first attribute feature and a second attribute feature, where the first attribute feature is the attribute feature of the first point cloud, and the second attribute feature is the attribute feature of the second point cloud; Encoding the first attribute feature and the second attribute feature respectively, and writing the obtained encoded bits into a first bitstream segment and a second bitstream segment respectively.
15. The method according to claim 14, wherein, The points in the second point cloud are the points corresponding to the high-frequency attribute features in the first point cloud.
16. The method according to claim 14 or 15, wherein, The performing feature extraction on the first point cloud and the second point cloud respectively to obtain a first attribute feature and a second attribute feature includes: Based on geometric information, spatially partitioning the first point cloud and the second point cloud respectively to obtain a plurality of first sub-blocks of the first point cloud and a plurality of second sub-blocks of the second point cloud; Performing feature extraction on the plurality of first sub-blocks and the plurality of second sub-blocks respectively to obtain the first attribute feature and the second attribute feature.
17. The method according to claim 16, wherein, The spatial partitioning levels of the first point cloud and the second point cloud are different.
18. The method according to any one of claims 14-17, wherein, The performing feature extraction on the first point cloud and the second point cloud respectively to obtain a first attribute feature and a second attribute feature includes: Based on the correlation between attribute information and geometric information, performing feature extraction on the first point cloud and the second point cloud respectively to obtain the first attribute feature and the second attribute feature.
19. The method according to claim 18, wherein The correlation between the attribute information and the geometric information is determined based on the association relationship between the geometric reconstruction information of the first point cloud and the attribute information of the first point cloud and the second point cloud respectively.
20. The method according to any one of claims 14 - 19, wherein The performing feature extraction on the first point cloud and the second point cloud respectively to obtain a first attribute feature and a second attribute feature includes: Performing feature extraction on the second point cloud to obtain the second attribute feature; Processing the attribute information of the first point cloud based on the second attribute feature to obtain residual information; Performing feature extraction on the first point cloud based on the residual information to obtain the first attribute feature.
21. The method according to any one of claims 14-20, wherein, Before performing feature extraction on the first point cloud and the second point cloud respectively, the method further includes: Preprocessing the initial attribute information of the first point cloud and the second point cloud respectively to obtain the attribute information of the first point cloud and the second point cloud, where the attribute information value of the same point is less than the initial attribute information value.
22. The method according to any one of claims 14-21, wherein, The first attribute feature and the second attribute feature are obtained through a sixth neural network and a seventh neural network respectively, and the network depths of the sixth neural network and the seventh neural network are different.
23. The method according to claim 22, wherein, The sixth neural network and / or the seventh neural network includes one or more of the following: One or more convolutional layers; One or more activation layers; One or more offset attention modules; And One or more feature extraction modules.
24. The method according to claim 23, wherein The feature extraction module includes a plurality of feature extraction channels, and the network depths of the plurality of feature extraction channels are different.
25. The method according to claim 24, wherein The weights corresponding to the plurality of feature extraction channels are different.
26. The method according to claim 25, wherein, The method further includes: Writing a first parameter into the bitstream, where the first parameter is used to indicate the weights of the plurality of feature extraction channels.
27. The method according to any one of claims 14-26, wherein, The second point cloud is obtained through a first processing module, and the first processing module includes one or more of the following: One or more perception layers; One or more global pooling layers; One or more self-attention layers; And One or more convolutional layers.
28. The method according to any one of claims 14 - 27, wherein, The method further includes: Obtaining the encoding loss of the first point cloud and the second point cloud; Adjusting the parameters of the neural network used in the encoding method based on the encoding loss.
29. The method according to claim 28, wherein The first attribute feature and the second attribute feature are encoded using an entropy encoding model, and the encoding loss is associated with one or more of the following: Attribute encoding loss; Structure encoding loss; and Entropy encoding loss.
30. The method according to any one of claims 14 - 29, wherein The method further includes: Writing one or more of the following parameters into the bitstream: The length of the bitstream segment; A bitstream segment identifier for indicating the point cloud associated with the bitstream segment; An attribute type for indicating the type of attribute information carried by the current bitstream; and An encoding method parameter for indicating the encoding method of the attribute information carried by the current bitstream.
31. The method according to claim 30, wherein, If the encoding method is used to encode the attribute information of the point cloud, the value of the encoding method parameter is a first value.
32. The method according to any one of claims 14 - 31, wherein, The first point cloud is the original point cloud.
33. A decoder, comprising: A parsing unit for parsing the bitstream to obtain a first attribute feature and a second attribute feature; A first feature restoration unit for restoring the first attribute feature to obtain first attribute information; A second feature restoration unit for restoring the second attribute feature to obtain second attribute information; A processing unit for processing the first attribute information and the second attribute information to obtain attribute reconstruction information.
34. A decoder, the decoder comprising: A memory for storing a computer program; A processor for executing the method according to any one of claims 1 to 13 when running the computer program.
35. An encoder, comprising: A processing unit for processing a first point cloud based on attribute information to obtain a second point cloud, where the second point cloud occupies the same spatial size as the first point cloud, and the number of points in the second point cloud is less than the number of points in the first point cloud; A feature extraction unit, configured to perform feature extraction on the first point cloud and the second point cloud respectively, to obtain a first attribute feature and a second attribute feature, where the first attribute feature is an attribute feature of the first point cloud, and the second attribute feature is an attribute feature of the second point cloud; An encoding unit, configured to encode the first attribute feature and the second attribute feature respectively, and write the obtained encoded bits into a first bitstream segment and a second bitstream segment respectively.
36. An encoder, the encoder comprising: A memory, configured to store a computer program; A processor, configured to execute the method according to any one of claims 14-32 when running the computer program.
37. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 13 or the method according to any one of claims 14-32 is implemented.
38. A bitstream, the bitstream comprising the bitstream generated by the method according to any one of claims 14 to 32.
39. A non-volatile computer-readable storage medium storing a bitstream, the bitstream being generated by an encoding method using an encoder, or the bitstream being decoded by a decoding method using a decoder, wherein, The decoding method is the method according to any one of claims 1-13, and the encoding method is the method according to any one of claims 14-32.
Citation Information
Patent Citations
Cross-modal data compression method and device, equipment and medium
CN117014633A
Point cloud geometric compression method based on depth auto-encoder
US20210019918A1
Point cloud attribute decoding method and point cloud attribute encoding method
WO2022188164A1
Point cloud decoding and upsampling and model training methods and apparatus
WO2022246724A1