Point group encoding method, decoding method, device, and electronic device
By predicting occupancy information and determining contexts for subnodes in point cloud sequences, the encoding method improves encoding efficiency and geometric compression performance.
Patent Information
- Application Number
- JP2024531439
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-03
- Filing Date
- 2022-12-01
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2042-12-01
AI Technical Summary
Existing point cloud encoding methods result in low coding efficiency for sequences with multiple planar features due to unoccupied spaces, particularly in sparse and dense point cloud sequences.
The encoding method predicts occupancy information of nodes based on reference nodes and their positions, determining a context for subnodes using tree structures to improve encoding efficiency by entropy-coding these subnodes.
This approach enhances geometric compression performance and encoding efficiency by effectively utilizing occupancy information in point cloud sequences.
Smart Images

Figure 0007763956000002 
Figure 0007763956000003 
Figure 0007763956000004
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority from Chinese Patent Application No. 202111466682.1, filed in China on December 3, 2021, the entire contents of which are incorporated herein by reference.
[0002] The present application relates to the technical field of codecs, and more particularly to a point cloud encoding method, a decoding method, an apparatus, and an electronic device. [Background technology]
[0003] In the framework of the point cloud digital audio video coding standard (AVS) encoder, the geometric information of a point cloud and the attribute information corresponding to each point are coded separately. Currently, spatial occupation coding employs a context-based adaptive binary arithmetic encoder, and sparse point cloud sequences and dense point cloud sequences are coded using different context models. However, for point cloud sequences with multiple planar features, unoccupied spaces exist in such point cloud sequences, and coding using methods based on sparse point cloud sequences and dense point cloud sequences still results in low coding efficiency. Summary of the Invention
[0004] The embodiments of the present application provide a point cloud encoding method, a decoding method, an apparatus, and an electronic device that can solve the problem of low encoding efficiency in related art.
[0005] In the first aspect, An encoding side obtains a node to be encoded in a point cloud sequence and m encoded reference nodes in the point cloud sequence, where m is a positive integer; a step in which the encoding side determines a context of the subnode to be encoded based on occupancy information of the m reference nodes and a position of the subnode to be encoded in the node to be encoded, the subnode to be encoded being any of the subnodes obtained by dividing the node to be encoded based on a tree structure; the encoding side entropy-encoding the subnode to be encoded based on the context to generate a target code stream.
[0006] In the second aspect, A decoding side obtains a node to be decoded in a point cloud sequence and m decoded reference nodes in the point cloud sequence, where m is a positive integer; a step in which the decoding side determines a context of the subnode to be decoded based on occupancy information of the m reference nodes and a position of the subnode to be decoded in the node to be decoded, the subnode to be decoded being any of the subnodes obtained by dividing the node to be decoded based on a tree structure; the decoding side entropy-decoding the subnode to be decoded based on the context to generate a target code stream.
[0007] In the third aspect, A first acquisition module for acquiring a node to be coded in a point cloud sequence and m coded reference nodes in the point cloud sequence, where m is a positive integer; a first determination module for determining a context of the subnode to be coded based on occupancy information of the m reference nodes and a position of the subnode to be coded in the node to be coded, the subnode to be coded being any one of subnodes obtained by dividing the node to be coded based on a tree structure; a coding module for entropy coding the subnode to be coded based on the context to generate a target code stream.
[0008] In the fourth aspect, A second acquisition module for acquiring a node to be decoded in a point cloud sequence and m decoded reference nodes in the point cloud sequence, where m is a positive integer; a second determination module for determining a context of the subnode to be decoded based on occupancy information of the m reference nodes and a position of the subnode to be decoded in the node to be decoded, wherein the subnode to be decoded is any subnode obtained by dividing the node to be decoded based on a tree structure; a decoding module for entropy decoding the subnode to be decoded based on the context to generate a target code stream.
[0009] In a fifth aspect, there is provided an electronic device including a processor and a memory, wherein the memory stores a program or command executable by the processor, and when the program or command is executed by the processor, steps of the point cloud encoding method described in the first aspect are realized, or steps of the point cloud decoding method described in the second aspect are realized.
[0010] In a sixth aspect, there is provided a readable storage medium having stored thereon a program or command that, when executed by a processor, causes the steps of the point cloud encoding method described in the first aspect to be realized or the steps of the point cloud decoding method described in the second aspect to be realized.
[0011] In a seventh aspect, there is provided a chip including a processor and a communication interface, the communication interface and the processor being coupled together, the processor executing a program or command to implement the method described in the first aspect or to be used to implement the method described in the second aspect.
[0012] In an eighth aspect, there is provided a computer program / program product stored on a storage medium and executed by at least one processor to implement the method described in the first aspect or to implement the method described in the second aspect.
[0013] In a ninth aspect, there is provided a communications device configured to execute and implement a method according to the first aspect or to implement a method according to the second aspect.
[0014] In the embodiments of the present application, the encoding side can predict the occupancy information of the node to be encoded based on the occupancy information of the encoded reference node, and determine the context of the subnode to be encoded based on the prediction result of the node to be encoded and the position of the subnode to be encoded in the node to be encoded, thereby making better use of the occupancy information of the encoded nodes in the point cloud sequence to improve the geometric compression performance of the point cloud and increase the encoding efficiency of the encoding side. [Brief explanation of the drawings]
[0015] [Figure 1] FIG. 1 is a framework diagram of an AVS encoder. [Figure 2] 1 is a flowchart of a point cloud encoding method provided in an embodiment of the present application; [Figure 3a] FIG. 2 is a schematic diagram of a node to be coded and a reference node in a point cloud sequence. [Figure 3b] FIG. 1 is a schematic diagram of a node in a point cloud sequence. [Figure 4a] FIG. 10 is a schematic diagram of a subnode to be coded and adjacent subnodes in a point cloud sequence. [Figure 4b] FIG. 1 is a schematic diagram of a node to be coded and adjacent nodes in a point cloud sequence. [Figure 5a] 1 is a schematic diagram (part 1) of a node to be coded, a sub-node to be coded, and adjacent nodes in a point cloud sequence. [Figure 5b]10 is a schematic diagram (part 2) of a node to be coded, a subnode to be coded, and adjacent subnodes in a point cloud sequence. [Figure 6] 1 is a flowchart of a point cloud decoding method provided in an embodiment of the present application; [Figure 7] FIG. 1 is a configuration diagram of a point cloud encoding device provided in an embodiment of the present application. [Figure 8] FIG. 1 is a configuration diagram of a point cloud decoding device provided in an embodiment of the present application. [Figure 9] 1 is a block diagram of an electronic device provided in an embodiment of the present application. [Figure 10] FIG. 2 is a configuration diagram of a terminal provided in an embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION
[0016] The technical solutions in the embodiments of the present application will be clearly explained below with reference to the drawings in the embodiments of the present application. Of course, the described embodiments are only a part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments in the present application fall within the scope of protection of the present application.
[0017] The terms "first," "second," etc., used in the specification and claims of this application are not intended to describe a particular order or chronology, but rather to distinguish between similar objects. It should be understood that terms used in this manner may be interchanged where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein. It should also be understood that objects distinguished by "first" and "second" generally belong to a single category, and the number of objects is not limited; for example, the first object may be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the symbol " / " generally indicates that the related objects before and after are in an "or" relationship.
[0018] In order to better understand the technical solution of the present application, the following briefly describes the related art to the technical solution of the present application.
[0019] Figure 1 shows the AVS encoder framework. In the point cloud AVS encoder framework, the geometric information of a point cloud and the corresponding attribute information for each point are encoded separately. First, the point cloud is preprocessed. A bounding box is constructed, which is the smallest rectangular box that contains all points in the input point cloud. The origin coordinates of the bounding box are the minimum values of the x, y, and z dimensions of each point in the point cloud. Next, coordinate transformation (i.e., coordinate translation shown in Figure 1) is performed on the points in the point cloud. Using this coordinate origin as a reference, the original coordinates of the points are converted to coordinates relative to the coordinate origin. Then, coordinate quantization is performed on the geometric coordinates of the points, which mainly serves the purpose of scaling. After quantization and rounding, the geometric information of some points becomes the same, and a parameter determines whether or not to remove points with overlapping geometric information. Next, the preprocessed point cloud is partitioned into a structured tree (e.g., an octree, quadtree, or binary tree) according to breadth-first traversal order. Taking the octree partitioning shown in Figure 1 as an example, the preprocessed bounding box is used as the root node, which is then divided into eight equal parts to generate eight sub-cubes as its sub-nodes. The occupancy information (also called occupancy information) of each sub-node is represented by eight bits, called the space occupancy code. If a sub-cube contains a point, it indicates that the sub-node is occupied, and the corresponding occupancy bit is set to 1; otherwise, it is set to 0. The occupied sub-cube is then divided until the resulting leaf node is a 1x1x1 unit cube, completing the encoding of the geometric octree. During the octree encoding process, the generated space occupancy code and the number of points contained in the final leaf node are entropy coded to obtain the output code stream. In the geometric decoding process based on the octree, the decoding side follows the order of breadth-first traversal, continuously analyzes to obtain the occupancy code of each node, and continuously divides the node sequentially, stopping the division until a 1x1x1 unit cube is obtained, analyzes to obtain the number of points contained in each leaf node, and finally restores it to obtain the geometric reconstruction point cloud information.
[0020] After geometric encoding is complete, the geometric information is reconstructed. Currently, attribute encoding is primarily performed on attribute information such as color and reflectance. First, a color space conversion is determined. If color space conversion is performed, the color information is converted from the red, green, and blue (RGB) color space to the luminance-chromaticity (YUV) color space. Next, the reconstructed point cloud is recolored using the original point cloud, so that the unencoded attribute information corresponds to the reconstructed geometric information. Color information encoding is divided into two modules: attribute prediction and attribute conversion. The attribute prediction process first sorts the point cloud, and then performs differential prediction. Here, the sorting method is Hilbert sorting. Attribute prediction is performed on the sorted point cloud using differential methods. Finally, the prediction residual is quantized and entropy coded to generate a binary code stream. In the attribute transformation process, first, a wavelet transform is performed on the point cloud attributes, and the transform coefficients are quantized. Then, the attribute reconstruction values are obtained by inverse quantization and inverse wavelet transform. After that, the difference between the original attributes and the attribute reconstruction values is calculated to obtain the attribute residuals, which are then quantized. Finally, the quantized transform coefficients and attribute residuals are entropy coded to generate a binary code stream.
[0021] The point cloud encoding method provided in the embodiments of the present application will be described in detail below with reference to the drawings according to several embodiments and application scenarios.
[0022] 2 is a flowchart of the point cloud encoding method provided in the embodiment of the present application. As shown in FIG. 2, the method includes the following steps 201, 202 and 203.
[0023] In step 201, the encoding side obtains a node to be encoded in a point cloud sequence and m encoded reference nodes in the point cloud sequence.
[0024] Here, m is a positive integer. The encoding side may be, for example, an electronic device such as a mobile phone, a tablet, or a computer, and is not specifically limited thereto in this application.
[0025] It should be noted that the encoding side can sequentially encode nodes in the point cloud sequence according to a preset encoding order, and the node to be encoded may be the uncoded node arranged first in the preset encoding order among the uncoded nodes in the point cloud sequence. Here, the reference node is an encoded node in the point cloud sequence, for example, the reference node may be any one of the encoded nodes, or the reference node may be any one of the encoded nodes belonging to the same node division hierarchical level as the node to be encoded, or the reference node may be an encoded node adjacent to the node to be encoded.
[0026] Here, the node split hierarchical level refers to the node hierarchical level at which a structural tree (e.g., binary tree, quad tree, octree, etc.) is split for a node. For example, node 1 obtains nodes 11 and 12 based on binary tree split, node 11 obtains nodes 111 and 112 based on binary tree split, and node 12 obtains nodes 121 and 122 based on binary tree split, where node 11 and node 12 belong to the same node split hierarchical level, and node 111, node 112, node 121, and node 122 belong to the same node split hierarchical level.
[0027] In step 202, the encoding side determines the context of the subnode to be encoded based on the occupancy information of the m reference nodes and the position of the subnode to be encoded in the node to be encoded, and the subnode to be encoded is one of the subnodes obtained by dividing the node to be encoded based on the structural tree.
[0028] Here, the occupancy information of a node refers to the occupancy status of each subnode based on the structural tree division of the node, including whether it is occupied or unoccupied. The occupancy information of the reference node may also refer to the occupancy status of each subnode of the reference node. For example, if the reference node obtains subnode 1 and subnode 2 based on binary tree division, the occupancy information of the reference node may indicate that subnode 1 is occupied and subnode 2 is unoccupied. Alternatively, the occupancy information of the reference node may indicate the number of occupied subnodes and the number of unoccupied subnodes in the reference node. For example, if the reference node obtains eight subnodes based on octree division, the occupancy information of the reference node may indicate that the number of occupied subnodes is three and the number of unoccupied subnodes is five. Alternatively, the occupancy information of the reference node may refer to the occupancy status of each subnode of the reference node in the preset area. For example, if the reference node obtains eight subnodes based on octree division, the occupancy information of the reference node can be expressed as the occupancy status in the low plane area and the occupancy status in the high plane area. The low plane area and the high plane area may be two plane areas on the target direction of the reference node. As shown in Figure 3a, in the z-axis direction of the coordinate system, the subnodes in the low plane area of the reference node (the filled subnodes in Figure 3a) are occupied, while the subnodes in the high plane area (the unfilled subnodes in Figure 3a) are unoccupied. Optionally, the occupancy information of the reference node may be expressed in other ways, and detailed descriptions thereof will be omitted in this application.
[0029] It is understood that the reference nodes are encoded nodes in the point cloud sequence, and the occupancy information of the encoded reference nodes can be further obtained. In the embodiment of the present application, after obtaining m reference nodes, the encoding side predicts the occupancy information of the node to be encoded according to the occupancy information of the m reference nodes, and then determines the prediction result of the node to be encoded.
[0030] For example, in Figure 3a, the dashed frame indicates the node to be coded, and the solid frame indicates the reference node. In Figure 3a, there are three reference nodes, and the filled-in portion indicates occupied subnodes. Assuming that these nodes are divided into eight subnodes based on octree partitioning, the lower four subnodes in the z-axis direction are subnodes in the low planar region, and the upper four subnodes are subnodes in the high planar region. If the occupancy information of the reference nodes is the occupancy status of the low planar region and the occupancy status of the high planar region of the reference nodes, the occupied subnodes of the three reference nodes in Figure 3a are all in the low planar region, and the unoccupied subnodes are all in the high planar region. Based on the occupancy information of the three reference nodes, it can be predicted that the occupied subnodes of the node to be coded will also be in the low planar region, thereby obtaining a prediction result for the occupancy information of the node to be coded.
[0031] It should be noted that after obtaining a node to be coded, the coding side can perform tree division on the node to be coded using a tree division method such as an octree, a quadtree, or a binary tree. For example, if an encoded node belonging to the same node division hierarchical level as the node to be coded obtains eight subnodes based on the octree division, the node to be coded will also obtain eight subnodes based on the octree division. Furthermore, the coding side sequentially codes each subnode divided from the node to be coded based on a predetermined coding order, and the subnode to be coded is any uncoded subnode obtained by dividing the node to be coded based on the tree.
[0032] In an embodiment of the present application, the encoding side predicts the occupancy information of the node to be encoded based on m reference nodes, and after obtaining the prediction result of the node to be encoded, the encoding side determines the context of the subnode to be encoded based on the prediction result and the position of the subnode to be encoded in the node to be encoded.
[0033] In step 203, the encoding side entropy encodes the subnode to be encoded based on the context to generate a target codestream.
[0034] Alternatively, the encoding side may determine the context of the subnode to be encoded, and then assign an adaptive probability model to the subnode to be encoded, and then arithmetically encode, i.e., entropy encode, the occupied bit code of the subnode to be encoded based on the adaptive probability model to generate a target code stream, for example, a binary code stream.
[0035] In an embodiment of the present application, the encoding side predicts the occupancy information of the node to be encoded based on the occupancy information of the encoded reference node, and then determines the context of the subnode to be encoded based on the prediction result of the node to be encoded and the position of the subnode to be encoded in the node to be encoded, thereby making better use of the occupancy information of the encoded nodes in the point cloud sequence to improve the geometric compression performance of the point cloud and increase the encoding efficiency of the encoding side.
[0036] Optionally, step 201 includes: The encoding side acquires a node to be encoded in a point cloud sequence; The encoding side obtains k encoded nodes before the encoding target node based on a node encoding order, where k is a positive integer; If at least one coded node among the previous k coded nodes has a target plane feature, the coding side may obtain m coded reference nodes in the point cloud sequence.
[0037] In the embodiment of the present application, the encoding side obtains a node to be encoded in the point cloud sequence, and then obtains k encoded nodes before the node to be encoded according to the node encoding order. For example, assuming that eight nodes are arranged in the node encoding order as node 0, node 1, node 2, ... node 7, and the node to be encoded is node 7, that is, when the nodes before node 7 are encoded, assuming that the value of k is 3, the three encoded nodes before the node to be encoded are node 4, node 5, and node 6, respectively.
[0038] Furthermore, if at least one coded node among the previous k coded nodes has a target plane feature, the coding side obtains m coded reference nodes in the point cloud sequence. Optionally, the m reference nodes may be nodes among the k coded nodes, where m≦k, or the m reference nodes may not belong to the k coded nodes, or only some of them may belong to the k coded nodes. For example, the m reference nodes may be m coded nodes adjacent to the node to be coded, and in this case, these reference nodes are not related to the coding order, so some of the reference nodes are not among the previous k coded nodes.
[0039] Alternatively, the target plane characteristic may mean that a node is divided into n subnodes according to an n-ary tree and split into two planes along the target coordinate axis, with all occupied subnodes located in one plane and all subnodes in the other plane being unoccupied, and the node is therefore a node with the target plane characteristic. For example, the node in Figure 3b obtains eight subnodes through octree division, with subnodes numbered 0, 2, 4, and 6 constituting the first plane and subnodes numbered 1, 3, 5, and 7 constituting the second plane. If at least one of the four subnodes in the first plane is occupied but none of the four subnodes in the second plane are occupied, or if none of the four subnodes in the first plane are occupied but at least one of the four subnodes in the second plane is occupied, the node is a node with the target plane characteristic.
[0040] In an embodiment of the present application, if at least one of the k coded nodes preceding the node to be coded has a target plane feature, it is considered that the node to be coded may also have the target plane feature. For example, all occupied subnodes in the node to be coded are located on the first or second plane, and the coding side obtains m coded reference nodes in the point cloud sequence. Based on the occupancy information of the m reference nodes, the occupancy information of the node to be coded is predicted. By determining whether at least one of the k coded nodes has a target plane feature, the accuracy of prediction of the occupancy information of the node to be coded can be improved.
[0041] Alternatively, if the previous k coded nodes do not all have the target plane features, the coding side does not need to obtain the m coded reference nodes, and further does not predict the occupancy information of the node to be coded based on the occupancy information of the reference node. The coding side can use related conventional means to perform entropy coding of the node to be coded, and detailed description is omitted here.
[0042] Optionally, when at least one coded node among the previous k coded nodes has a target plane feature, the step of the coding side obtaining m coded reference nodes in the point cloud sequence includes: When the number of encoded nodes having target plane features among the previous k encoded nodes is equal to or greater than a first threshold, the encoding side obtains m encoded reference nodes in the point cloud sequence; Here, the first threshold is a positive integer less than k.
[0043] In an embodiment of the present application, the encoding side obtains k encoded nodes before the node to be encoded, and then obtains the number of encoded nodes that have the target plane feature among the k encoded nodes before. If the number of encoded nodes that have the target plane feature is equal to or greater than a first threshold, it is considered that the node to be encoded may also have the target plane feature. For example, if all occupied subnodes in the node to be encoded are located on the first plane or the second plane, the encoding side further obtains m encoded reference nodes in the point cloud sequence.
[0044] Alternatively, the first threshold value may be a value preset by a user, or may be an empirical value obtained by the encoding side based on a limited number of tests.
[0045] Optionally, step 201 further specifically includes: The encoding side acquires a node to be encoded in a point cloud sequence; a step of determining a target coordinate system based on the coordinate values of the node to be encoded by the encoding side; a step in which the encoding side determines, as a candidate reference node, a node in the point cloud sequence that belongs to the same node division hierarchical level as the encoding target node and has the same coordinate value on a target coordinate axis as the encoding target node, wherein the target coordinate axis is any coordinate axis in the target coordinate system and is perpendicular to a target plane; The encoding side may obtain m reference nodes from the candidate reference nodes.
[0046] The geometric information of a node in a point cloud sequence can be expressed by its coordinate values (e.g., x1, y1, z1) in a Cartesian coordinate system. In this case, it can be understood that each node in a point cloud sequence includes corresponding coordinate values. In an embodiment of the present application, after obtaining a node to be encoded in a point cloud sequence, the encoding side can also obtain the coordinate values corresponding to the node to be encoded accordingly, and determine a target coordinate system based on the coordinate values, where the target coordinate system is a coordinate system corresponding to the coordinate values of the node to be encoded. Based on the target coordinate system, the encoding side can determine the coordinate origin and each coordinate axis of the target coordinate system, and can also obtain the coordinate values of other nodes in the point cloud sequence, and determine the relative position of the node and the node to be encoded based on the coordinate values of the node.
[0047] The encoding side determines, as candidate reference nodes, nodes in the point cloud sequence that belong to the same node split hierarchical level as the encoding target node and have the same coordinate values on the target coordinate axis as the encoding target node. For example, assuming that the target coordinate axis is the z-axis and the coordinate values of the encoding target node are (x1, y1, z1), the candidate reference node is the encoded node with the z-axis coordinate value z1, and the coordinate values of the reference node can be expressed as (x1-a*xNodeSize, y1-b*yNodeSize, z1), where a and b are integers greater than or equal to 0 and cannot simultaneously take the value 0, xNodeSize is the node edge length of the encoding target node in the x-axis direction, and yNodeSize is the node edge length of the encoding target node in the y-axis direction. Furthermore, m reference nodes are selected from these reference nodes.
[0048] It should be noted that the target plane feature can be determined based on the target coordinate axis. For example, if the target coordinate axis is the z-axis, the target plane is a plane perpendicular to the z-axis. If a subnode divided based on the structure tree of a node is located on the target plane, the node is a node with a target plane feature.
[0049] In the embodiment of the present application, the encoding side predicts the occupancy information of the node to be encoded based on the occupancy information of m reference nodes, where the m reference nodes are selected from nodes in the point cloud sequence that belong to the same node division hierarchical level as the node to be encoded and have the same coordinate value on the target coordinate axis as the node to be encoded. The m reference nodes are nodes surrounding the node to be encoded, which can effectively utilize the spatial geometric relationship between nodes in the point cloud sequence and further improve the prediction accuracy of the occupancy status of the node to be encoded.
[0050] Optionally, the occupancy information of the encoded m reference nodes in the point cloud sequence includes the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, and the method includes: a step in which the encoding side determines a first plane and a second plane of a node, the first plane and the second plane are both parallel to the target plane, the first plane and the second plane are stacked in an extension direction of the target coordinate axis, the first plane is closer to a coordinate origin than the second plane, and the coordinate origin is the coordinate origin of the target coordinate system; The method further includes a step in which the encoding side obtains the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, where the first subnodes are subnodes located in the first plane and the second subnodes are subnodes located in the second plane.
[0051] Optionally, step 202 specifically includes: the encoding side predicting occupancy information of the encoding target node based on occupancy information of the m reference nodes, and determining a prediction result of the encoding target node; The encoding side determines a context of the subnode to be encoded based on the prediction result of the node to be encoded and the position of the subnode to be encoded in the node to be encoded.
[0052] Optionally, the step of the encoding side predicting the occupancy information of the encoding target node based on the occupancy information of the m reference nodes and determining the prediction result of the encoding target node includes: a step in which the encoding side determines a first plane and a second plane of a node, the first plane and the second plane are both parallel to the target plane, the first plane and the second plane are stacked in an extension direction of the target coordinate axis, the first plane is closer to a coordinate origin than the second plane, and the coordinate origin is the coordinate origin of the target coordinate system; The encoding side obtains the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, where the first subnodes are subnodes located in the first plane and the second subnodes are subnodes located in the second plane; The encoding side predicts occupancy information of the node to be encoded based on the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, and determines the prediction result of the node to be encoded.
[0053] It should be noted that the encoding side can divide any node in the point cloud sequence into a first plane and a second plane using the above method, including but not limited to a reference node and a node to be encoded. For example, in Figure 3a, the dashed frame indicates the node to be encoded, and the solid frame indicates the reference node. In Figure 3a, there are three reference nodes, and these nodes obtain eight subnodes based on octree division. Assuming that the target coordinate axis is the z-axis, taking the leftmost reference node as an example, the first plane of the reference node is the plane where the four lower nodes are located, and the second plane is the plane where the four upper nodes are located. Similarly, the first plane of the other reference nodes is the plane where the four lower nodes are located, and the second plane is the plane where the four upper nodes are located.
[0054] Alternatively, the target coordinate axis may be the x-axis or the y-axis, and accordingly, the first plane and the second plane are also planes perpendicular to the x-axis or the y-axis, and a detailed description thereof will be omitted here.
[0055] In an embodiment of the present application, the node is divided into a first plane and a second plane, and the number of occupied first subnodes in the first plane and the number of occupied second subnodes in the second plane of m reference nodes are obtained, and then the occupancy information of the node to be coded is predicted based on the number of occupied first subnodes and the number of occupied second subnodes, and the predicted result of the node to be coded can be determined.
[0056] Continuing with Figure 3a, the dashed frame indicates the node to be coded, and the solid frame indicates the reference node. In Figure 3a, there are three reference nodes, and the filled in subnodes are occupied. That is, in Figure 3a, the number of occupied first subnodes is 10, and the number of occupied second subnodes is 0. Since it can be seen that the number of occupied first subnodes of the reference node is greater than the number of occupied second subnodes, and all of the occupied subnodes of the reference node are located in the first plane, it is predicted that the occupancy information of the node to be coded may also be that the occupied subnodes are located in the first plane, and the subnodes in the second plane are unoccupied.
[0057] Optionally, the step of the encoding side predicting occupancy information of the node to be encoded based on the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, and determining the prediction result of the node to be encoded, includes: a step of the encoding side predicting that at least one of the first subnodes in the node to be encoded is occupied and the second subnode in the node to be encoded is unoccupied, and determining that the prediction result of the node to be encoded is a first prediction result, when the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes meet a first condition, wherein the first condition is that the number of occupied first subnodes in the m reference nodes is greater than a second threshold and the number of occupied second subnodes is less than a third threshold; a step of the encoding side predicting that at least one of the second subnodes in the node to be encoded is occupied and the first subnode in the node to be encoded is unoccupied, and determining that the prediction result of the node to be encoded is a second prediction result, when the number of the occupied first subnodes and the number of the occupied second subnodes in the m reference nodes meet a second condition, wherein the second condition is that the number of the occupied second subnodes in the m reference nodes is greater than the second threshold and the number of the occupied first subnodes is less than the third threshold; and if the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes do not meet either the first condition or the second condition, the encoding side predicts that at least one of the first subnodes and at least one of the second subnodes in the node to be encoded is occupied, and determines that the prediction result of the node to be encoded is a third prediction result.
[0058] In an embodiment of the present application, if the number of occupied subnodes, i.e., occupied first subnodes, in the first plane of m reference nodes is greater than a second threshold, and the number of occupied subnodes, i.e., occupied second subnodes, in the second plane is less than a third threshold, the predicted result of the node to be coded is a first predicted result that at least one first subnode in the first plane of the node to be coded is occupied, and all second subnodes in the second plane of the node to be coded are unoccupied.
[0059] If the number of occupied first subnodes in the first plane of the m reference nodes is less than the third threshold and the number of occupied second subnodes in the second plane is greater than the second threshold, the prediction result of the node to be coded is a second prediction result that all first subnodes in the first plane of the node to be coded are unoccupied and at least one second subnode in the second plane of the node to be coded is occupied.
[0060] If the number of occupied first subnodes in the first plane and the number of occupied second subnodes in the second plane of the m reference nodes do not meet either the first or second condition, for example, the number of occupied first subnodes in the m reference nodes is less than the second threshold and the number of occupied second subnodes is greater than the third threshold, or the number of occupied first subnodes in the m reference nodes is less than the second threshold and the number of occupied second subnodes is less than the third threshold, the predicted result of the node to be coded is a third predicted result that at least one first subnode in the first plane of the node to be coded is occupied and at least one second subnode in the second plane is occupied.
[0061] In the embodiment of the present application, the occupancy information of the node to be coded is predicted by comparing the occupied first subnode and the occupied second subnode with the magnitude of the second threshold and the third threshold, thereby effectively improving the prediction accuracy of the occupancy status of the node to be coded.
[0062] Optionally, if the first subnode and the second subnode in the reference node are divided based on an n-ary tree, the second threshold is a positive integer less than or equal to n / 2×m-1, and the third threshold is a positive integer less than or equal to n / 2×m, where n is a positive integer.
[0063] For example, the first and second subnodes in the reference node are divided based on an octree, i.e., when the value of n is 8, the number of reference nodes is 3, i.e., the value of m is 3, then the second threshold is a positive integer less than or equal to 11, and the third threshold is a positive integer less than or equal to 12.
[0064] In an embodiment of the present application, the encoding side determines a prediction result of the node to be encoded, and then determines a context of the subnode to be encoded based on the prediction result and the position of the subnode to be encoded in the node to be encoded.
[0065] Optionally, step 202 specifically includes: When the prediction result of the node to be coded is the first prediction result and the sub-node to be coded is located on a second plane of the node to be coded, the coding side determines a first preset model as a context of the sub-node to be coded; When the prediction result of the node to be coded is the second prediction result and the sub-node to be coded is located on a first plane of the node to be coded, the coding side determines a second preset model as a context of the sub-node to be coded; When the prediction result of the node to be coded is the first prediction result and the subnode to be coded is located on a first plane of the node to be coded, the coding side obtains adjacent nodes of the subnode to be coded, and determines a context of the subnode to be coded based on the occupancy status of the adjacent nodes; When the prediction result of the node to be coded is the second prediction result and the subnode to be coded is located on a second plane of the node to be coded, the coding side obtains adjacent nodes of the subnode to be coded, and determines a context of the subnode to be coded according to the occupancy status of the adjacent nodes; and if the prediction result of the node to be encoded is the third prediction result, the encoding side obtains adjacent nodes of the subnode to be encoded and determines the context of the subnode to be encoded based on the occupancy status of the adjacent nodes.
[0066] Specifically, when the encoding side determines that the prediction result of the node to be encoded is a first prediction result, i.e., at least one first subnode in the first plane of the node to be encoded is occupied and all second subnodes in the second plane of the node to be encoded are not occupied, for a subnode to be encoded located in the second plane of the node to be encoded, the encoding side determines a first preset model as the context of such subnode to be encoded, i.e., assigns the first preset model to the subnode to be encoded, and the first preset model is an adaptive probability model, and the encoding side arithmetically encodes the occupied bit code of the subnode to be encoded based on the adaptive probability model to generate a target code stream.
[0067] When the encoding side determines that the prediction result of the node to be encoded is the second prediction result, i.e., all first subnodes in the first plane of the node to be encoded are unoccupied and at least one second subnode in the second plane of the node to be encoded is occupied, the encoding side determines a second preset model as the context of the subnode to be encoded located in the first plane of the node to be encoded. That is, the encoding side assigns the second preset model to the subnode to be encoded, and the second preset model is also an adaptive probability model. The encoding side arithmetically codes the occupied bit code of the subnode to be encoded based on the adaptive probability model to generate a target code stream. Alternatively, the first preset model and the second preset model may be the same probability model, for example, both may be adaptive probability models.
[0068] If the encoding side determines that the prediction result of the node to be encoded is the first prediction result, for a subnode to be encoded located on the first plane in the node to be encoded, the encoding side obtains adjacent nodes of the subnode to be encoded, and determines the context of the subnode to be encoded based on the occupancy status of the adjacent nodes.
[0069] If the encoding side determines that the prediction result of the node to be encoded is the second prediction result, for a subnode to be encoded located on the second plane in the node to be encoded, the encoding side can similarly obtain adjacent nodes of the subnode to be encoded, and determine the context of the subnode to be encoded based on the occupancy status of the adjacent nodes.
[0070] When the encoding side determines that the prediction result of the node to be encoded is the third prediction result, that is, at least one first subnode in the first plane of the node to be encoded is occupied and at least one second subnode in the second plane of the node to be encoded, for the subnode to be encoded in the node to be encoded, the encoding side can similarly obtain adjacent nodes of the subnode to be encoded, and determine the context of the subnode to be encoded based on the occupation status of the adjacent nodes.
[0071] Alternatively, the encoding side can obtain the adjacent nodes of the subnode to be encoded and determine the context of the subnode to be encoded based on the occupancy status of the adjacent nodes in two different ways, which will be described in detail below.
[0072] Method 1 Taking octree partitioning as an example, in the partitioning method of octree breadth-first traversal, the adjacent information that the encoding side can obtain for the subnode to be encoded in the node to be encoded includes adjacent subnodes in three target directions. Illustratively, adjacent subnodes in three directions (left, front, bottom, left, front) of the subnode to be encoded are obtained, including three adjacent subnodes coplanar with the current subnode to be encoded, three collinear adjacent subnodes, and one co-point adjacent subnode.
[0073] Regarding the context design of the subnode hierarchy, for a subnode to be coded, the coding side searches for the occupancy status of three coplanar adjacent subnodes, three collinear adjacent subnodes, one copoint adjacent subnode in the left front-bottom direction at the same level as the subnode to be coded, and the adjacent subnode that is two subnode edge lengths away from the current subnode to be coded in the negative direction in the dimension with the shortest subnode edge length. Taking the subnode with the shortest edge length in the x-axis direction as an example, the reference node selected by each subnode is as shown in Figure 4a, where the dashed frame node is the current node to be coded, the filled node is the current subnode to be coded, and the solid frame node is the adjacent subnode selected by each subnode.
[0074] The coding side considers in detail the occupancy status of the three coplanar adjacent subnodes, the three collinear subnodes, and the subnode that is two subnode side lengths away from the current coding target subnode in the negative direction on the shortest dimension of the subnode side length. The occupancy status of these seven subnodes is a total of 2 7 = 128 types. If at least one is occupied, there are a total of 2 7 There are 127 possible situations (-1 = 127), and one context is assigned to each situation. If all seven subnodes are unoccupied, the occupancy status of the co-point adjacent subnode is further considered. The co-point adjacent subnode can be either occupied or unoccupied. One context is assigned individually to a situation where the co-point adjacent subnode is occupied. If the co-point adjacent subnode is also unoccupied, the occupancy status of the adjacent nodes in the hierarchy of the node to be coded is further considered. This allows a total of 127 + 2 - 1 = 128 contexts to be obtained based on the occupancy status of the adjacent subnodes in the hierarchy of the subnode to be coded.
[0075] If the eight adjacent subnodes of the target node are not occupied, the occupancy status of four adjacent nodes in the target node hierarchy is obtained as shown in Figure 4b. Among them, the node enclosed by the dashed line is the target node, and the node enclosed by the solid line is the adjacent node. The context of the target node hierarchy is determined by the following steps 1 and 2.
[0076] 1. First, obtain the coplanar adjacent nodes in the three preset directions of the node to be coded. For example, obtain the three coplanar adjacent nodes on the upper right and rear of the node to be coded. The occupancy status of the three coplanar adjacent nodes on the upper right and rear of the node to be coded is 2 3 There are 8 possible cases, and one context is assigned to a situation where at least one is occupied. When this is combined with the position of the subnode to be coded in the node to be coded, a total of (8-1) x 8 = 56 contexts are provided for the set of coplanar neighboring nodes. If the three coplanar neighboring nodes to the upper right and rear of the node to be coded are all unoccupied, the occupancy statuses of the other three sets of neighboring nodes in the hierarchy of the node to be coded (i.e., the coplanar neighbor to the lower left front, the colinear neighbor to the upper right and rear, and the colinear neighbor to the lower left front in Figure 4b) are further obtained.
[0077] 2. Obtain the distance between the nearest occupied node and the current node. The correspondence between the occupancy status of specific adjacent nodes and the distance is shown in Table 1.
[0078] [Table 1]
[0079] As can be seen from Table 1, there are a total of three distance values, and when one context is assigned to each of these three values and further combined with the positional status of the subnode to be coded in the node to be coded, there are a total of 3 x 8 = 24 contexts.
[0080] Up to now, the total number of contexts determined by the above method 1 is 128+56+24=208, and the encoding side assigns one adaptive probability model to each context.
[0081] Method 2 After the encoding side determines the node to be encoded, for each subnode to be encoded, the encoding side can obtain six adjacent nodes that are coplanar and collinear with the node to be encoded in the hierarchy of the node to be encoded, as shown in Figure 5a. In Figure 5a, the node framed in the dashed line is the node to be encoded, the filled nodes are each subnode to be encoded, and the nodes framed in the solid line are coplanar and collinear adjacent nodes of the node to be encoded. For the three coplanar adjacent nodes, a total of 2 nodes are obtained by considering the distribution status of each node. 3 For the remaining three collinear adjacent nodes, only the number of occupied nodes among the three adjacent nodes is obtained, i.e., 0, 1, 2, and 3, for a total of four situations. Together, these give a total of 4 x 8 = 32 situations, and if one context is set for each situation, a total of 32 contexts are obtained for the node hierarchy to be coded.
[0082] Furthermore, for each subnode to be coded, adjacent nodes in the target direction at the same level as the subnode are obtained. For example, as shown in Figure 5b, three coplanar adjacent subnodes are obtained to the left, front, and bottom (negative direction of each coordinate axis) of the subnode to be coded. In Figure 5b, the node framed in the dashed line is the node to be coded, the filled node is the subnode to be coded, and the node framed in the solid line is the coplanar adjacent node at the same level as the subnode to be coded. The occupancy status of these three coplanar adjacent nodes at the same level as the subnode to be coded is 2 3 = 8 types, and if one context is assigned to each situation, a total of 8 contexts are provided for the subnode to be coded.
[0083] Since there is no interference between the contexts of the node hierarchy to be coded and the subnode hierarchy to be coded, the total number of contexts that can be determined by Method 2 is 32 x 8 = 256, and the coding side assigns one adaptive probability model to each context.
[0084] Alternatively, if the prediction result of the node to be coded is the third prediction result, or if the prediction result of the node to be coded is the first prediction result and the subnode to be coded is located in the first plane of the node to be coded, or if the prediction result of the node to be coded is the second prediction result and the subnode to be coded is located in the second plane of the node to be coded, the coding side can determine the context of the subnode to be coded by the above method 1 or method 2.
[0085] The embodiment of the present application further provides a point cloud decoding method. Figure 6 is a flowchart of the point cloud decoding method provided in the embodiment of the present application. As shown in Figure 6, the method includes the following steps 601, 601, and 603:
[0086] In step 601, the decoding side obtains a node to be decoded in a point cloud sequence and m decoded reference nodes in the point cloud sequence, where m is a positive integer.
[0087] In step 602, the decoding side determines the context of the subnode to be decoded based on the occupancy information of the m reference nodes and the position of the subnode to be decoded in the node to be decoded, and the subnode to be decoded is one of the subnodes obtained by dividing the node to be decoded based on the structural tree.
[0088] In step 603, the decoding side entropy decodes the subnode to be decoded based on the context to generate a target code stream.
[0089] Optionally, step 601 specifically includes: The decoding side acquires a node to be decoded in a point cloud sequence; The decoding side obtains k decoded nodes before the node to be decoded based on a node decoding order, where k is a positive integer; and if at least one decoded node among the previous k decoded nodes has a target plane feature, the decoding side obtains m decoded reference nodes in the point cloud sequence.
[0090] Optionally, when at least one decoded node among the previous k decoded nodes has a target plane feature, the step of the decoding side obtaining m decoded reference nodes in the point cloud sequence includes: The method includes a step in which, when the number of decoded nodes having the target plane feature among the previous k decoded nodes is equal to or greater than a first threshold, the decoding side determines that the previous k decoded nodes match the target plane feature; Here, the first threshold is a positive integer less than k.
[0091] Optionally, step 601 further specifically includes: The decoding side acquires a node to be decoded in a point cloud sequence; a step of determining a target coordinate system based on the coordinate values of the decoding target node by the decoding side; a step in which the decoding side determines, as a candidate reference node, a node in the point cloud sequence that belongs to the same node split hierarchical level as the node to be decoded and has the same coordinate value on a target coordinate axis as the node to be decoded, wherein the target coordinate axis is any one of coordinate axes in the target coordinate system and is perpendicular to a target plane; The decoding side may obtain m reference nodes from the candidate reference nodes.
[0092] Optionally, the occupancy information of the encoded m reference nodes in the point cloud sequence includes the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, and the method includes: a step in which the decoding side determines a first plane and a second plane of a node, the first plane and the second plane being both parallel to the target plane, the first plane and the second plane being stacked along an extension direction of the target coordinate axis, the first plane being closer to a coordinate origin than the second plane, and the coordinate origin being the coordinate origin of the target coordinate system; The method further includes a step in which the decoding side obtains the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, wherein the first subnodes are subnodes located in the first plane and the second subnodes are subnodes located in the second plane.
[0093] Optionally, step 602 specifically includes: The decoding side predicts occupancy information of the node to be decoded based on occupancy information of the m reference nodes, and determines a prediction result of the node to be decoded; The decoding side determines a context of the subnode to be decoded based on the prediction result of the node to be decoded and the position of the subnode to be decoded in the node to be decoded.
[0094] Optionally, the step of determining the context of the subnode to be decoded by the decoding side based on the prediction result of the node to be decoded and the position of the subnode to be decoded in the node to be decoded includes: a step in which the decoding side determines a first plane and a second plane of a node, the first plane and the second plane being both parallel to the target plane, the first plane and the second plane being stacked along an extension direction of the target coordinate axis, the first plane being closer to a coordinate origin than the second plane, and the coordinate origin being the coordinate origin of the target coordinate system; the decoding side obtaining the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, wherein the first subnodes are subnodes located in the first plane and the second subnodes are subnodes located in the second plane; The decoding side predicts occupancy information of the node to be decoded based on the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, and determines the prediction result of the node to be decoded.
[0095] Optionally, the step of the decoding side predicting occupancy information of the node to be decoded based on the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, and determining the prediction result of the node to be decoded, a step of the decoding side predicting that at least one of the first subnodes in the node to be decoded is occupied and the second subnode in the node to be decoded is unoccupied, and determining that the prediction result of the node to be decoded is a first prediction result, when the number of the occupied first subnodes and the number of the occupied second subnodes in the m reference nodes meet a first condition, wherein the first condition is that the number of the occupied first subnodes in the m reference nodes is greater than a second threshold and the number of the occupied second subnodes is less than a third threshold; a step of the decoding side predicting that at least one of the second subnodes in the node to be decoded is occupied and the first subnode in the node to be decoded is unoccupied, and determining that the prediction result of the node to be decoded is a second prediction result, when the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes meet a second condition, wherein the second condition is that the number of occupied second subnodes in the m reference nodes is greater than the second threshold and the number of occupied first subnodes is less than the third threshold; and if the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes do not meet either the first condition or the second condition, the decoding side predicts that at least one of the first subnodes and at least one of the second subnodes in the node to be decoded is occupied, and determines that the prediction result of the node to be decoded is a third prediction result.
[0096] Optionally, step 602 specifically includes: When the prediction result of the node to be decoded is the first prediction result and the subnode to be decoded is located on a second plane of the node to be decoded, the decoding side determines a first preset model as a context of the subnode to be decoded; When the prediction result of the node to be decoded is the second prediction result and the subnode to be decoded is located on a first plane of the node to be decoded, the decoding side determines a second preset model as a context of the subnode to be decoded; When the prediction result of the node to be decoded is the first prediction result and the subnode to be decoded is located on a first plane of the node to be decoded, the decoding side obtains adjacent nodes of the subnode to be decoded, and determines a context of the subnode to be decoded based on the occupancy status of the adjacent nodes; When the prediction result of the node to be decoded is the second prediction result and the subnode to be decoded is located on a second plane of the node to be decoded, the decoding side obtains adjacent nodes of the subnode to be decoded, and determines a context of the subnode to be decoded according to the occupancy status of the adjacent nodes; If the prediction result of the node to be decoded is the third prediction result, the decoding side obtains adjacent nodes of the subnode to be decoded and determines the context of the subnode to be decoded based on the occupancy status of the adjacent nodes.
[0097] Optionally, if the first subnode and the second subnode in the reference node are divided based on an n-ary tree, the second threshold is a positive integer less than or equal to n / 2×m-1, and the third threshold is a positive integer less than or equal to n / 2×m, where n is a positive integer.
[0098] In the embodiments of the present application, the decoding side can predict the occupancy information of the node to be decoded based on the occupancy information of the decoded reference node, and determine the context of the subnode to be decoded based on the prediction result of the node to be decoded and the position of the subnode to be decoded in the node to be decoded, thereby making better use of the occupancy information of the decoded nodes in the point cloud sequence, improving the decompression performance of the point cloud, and increasing the decoding efficiency of the decoding side.
[0099] It should be noted that the point cloud decoding method provided in the embodiments of the present application differs from the above point cloud encoding method only in the execution body, and the specific execution steps and realization process can refer to the specific description in the above point cloud encoding method, and the description will be omitted here.
[0100] The point cloud encoding method provided in the embodiments of the present application may be executed by a point cloud encoding device. In the embodiments of the present application, the point cloud encoding device will be described taking the point cloud encoding method as an example.
[0101] FIG. 7 is a block diagram of a point cloud encoding device provided in an embodiment of the present application. As shown in FIG. 7, the point cloud encoding device 700 includes: A first obtaining module 701 for obtaining a node to be coded in a point cloud sequence and m coded reference nodes in the point cloud sequence, where m is a positive integer; a first determination module 702 for determining a context of the subnode to be coded based on occupancy information of the m reference nodes and the position of the subnode to be coded in the node to be coded, the first determination module 702 being one of the subnodes obtained by dividing the node to be coded based on a tree structure; and an encoding module 703 for entropy encoding the subnode to be encoded based on the context to generate a target codestream.
[0102] Optionally, the first acquisition module 701 further comprises: obtaining a node to be coded in a point cloud sequence; obtaining k coded nodes before the node to be coded based on a node coding order, where k is a positive integer; and obtaining m coded reference nodes in the point cloud sequence if at least one coded node among the previous k coded nodes has a target plane feature.
[0103] Optionally, the first acquisition module 701 further comprises: If the number of coded nodes having the target plane feature among the previous k coded nodes is equal to or greater than a first threshold, it is used to obtain coded m reference nodes in the point cloud sequence; Here, the first threshold is a positive integer less than k.
[0104] Optionally, the first acquisition module 701 further comprises: obtaining a node to be coded in a point cloud sequence; determining a target coordinate system based on the coordinate values of the encoding target node; determining, as a candidate reference node, a node in the point cloud sequence that belongs to the same node split hierarchical level as the encoding target node and has the same coordinate value on a target coordinate axis as the encoding target node, wherein the target coordinate axis is any coordinate axis in the target coordinate system and is perpendicular to a target plane; and obtaining m reference nodes from the candidate reference nodes.
[0105] Optionally, the occupancy information of the encoded m reference nodes in the point cloud sequence includes the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, and the point cloud encoding device 700: a third determination module for determining a first plane and a second plane of a node, wherein the first plane and the second plane are both parallel to the target plane, and the first plane and the second plane are stacked in an extension direction of the target coordinate axis, the first plane is closer to a coordinate origin than the second plane, and the coordinate origin is the coordinate origin of the target coordinate system; and a third acquisition module for acquiring the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, wherein the first subnodes are subnodes located in the first plane and the second subnodes are subnodes located in the second plane.
[0106] Optionally, the first determination module 702 further comprises: predicting occupancy information of the encoding target node based on occupancy information of the m reference nodes, and determining a prediction result of the encoding target node; determining a context of the subnode to be coded based on the prediction result of the node to be coded and the position of the subnode to be coded in the node to be coded.
[0107] Optionally, the first determination module 702 further comprises: a step in which the encoding side determines a first plane and a second plane of a node, the first plane and the second plane are both parallel to the target plane, the first plane and the second plane are stacked in an extension direction of the target coordinate axis, the first plane is closer to a coordinate origin than the second plane, and the coordinate origin is the coordinate origin of the target coordinate system; The encoding side obtains the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, where the first subnodes are subnodes located in the first plane and the second subnodes are subnodes located in the second plane; The encoding side predicts the occupancy information of the node to be encoded based on the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, and determines the prediction result of the node to be encoded.
[0108] Optionally, the first determination module 702 further comprises: predicting that at least one of the first subnodes in the node to be encoded is occupied and that the second subnode in the node to be encoded is unoccupied, and determining that the prediction result of the node to be encoded is a first prediction result, if the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes meet a first condition, wherein the first condition is that the number of occupied first subnodes in the m reference nodes is greater than a second threshold and the number of occupied second subnodes is less than a third threshold; predicting that at least one second subnode in the node to be encoded is occupied and that the first subnode in the node to be encoded is unoccupied, and determining that the prediction result of the node to be encoded is a second prediction result, if the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes meet a second condition, wherein the second condition is that the number of occupied second subnodes in the m reference nodes is greater than the second threshold and the number of occupied first subnodes is less than the third threshold; and if the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes do not meet either the first condition or the second condition, predicting that at least one of the first subnodes and at least one of the second subnodes in the node to be encoded is occupied, and determining that the prediction result of the node to be encoded is a third prediction result.
[0109] Optionally, the first determination module 702 further comprises: determining a first preset model as a context of the subnode to be encoded when the prediction result of the node to be encoded is the first prediction result and the subnode to be encoded is located on a second plane of the node to be encoded; determining a second preset model as a context of the subnode to be encoded when the prediction result of the node to be encoded is the second prediction result and the subnode to be encoded is located on a first plane of the node to be encoded; if the prediction result of the node to be coded is the first prediction result and the subnode to be coded is located on a first plane of the node to be coded, obtaining adjacent nodes of the subnode to be coded, and determining a context of the subnode to be coded based on the occupancy status of the adjacent nodes; if the prediction result of the node to be coded is the second prediction result and the subnode to be coded is located on a second plane of the node to be coded, obtaining adjacent nodes of the subnode to be coded, and determining a context of the subnode to be coded based on the occupancy status of the adjacent nodes; If the prediction result of the node to be coded is the third prediction result, obtaining adjacent nodes of the subnode to be coded and determining the context of the subnode to be coded based on the occupancy status of the adjacent nodes.
[0110] Optionally, if the first subnode and the second subnode in the reference node are divided based on an n-ary tree, the second threshold is a positive integer less than or equal to n / 2×m-1, and the third threshold is a positive integer less than or equal to n / 2×m, where n is a positive integer.
[0111] In an embodiment of the present application, the point cloud encoding device 700 can predict the occupancy information of the node to be encoded based on the occupancy information of the encoded reference node, and determine the context of the subnode to be encoded based on the prediction result of the node to be encoded and the position of the subnode to be encoded in the node to be encoded, thereby making better use of the occupancy information of the encoded nodes in the point cloud sequence to improve the performance of geometric compression of the point cloud and increase the encoding efficiency of the point cloud encoding device 700.
[0112] The point cloud encoding device 700 in the embodiments of the present application may be an electronic device, for example, an electronic device having an operating system, or a component of an electronic device, for example, an integrated circuit or a chip. The electronic device may be a terminal or other devices other than a terminal. Exemplarily, the terminal may include, but is not limited to, the types of terminal 11 listed above. The other devices may be, for example, a server, a network attached storage (NAS), etc., and are not specifically limited in the embodiments of the present application.
[0113] The point cloud encoding device 700 provided in the embodiment of the present application can implement each process implemented in the embodiment of the method described in Figure 1 and achieve the same technical effects, so that the description will be omitted here to avoid repetition.
[0114] The point cloud decoding method provided in the embodiments of the present application may be executed by a point cloud decoding device. In the embodiments of the present application, the point cloud decoding device provided in the embodiments of the present application will be described by taking the point cloud decoding method executed by the point cloud decoding device as an example.
[0115] FIG. 8 is a block diagram of a point cloud decoding device provided in an embodiment of the present application. As shown in FIG. 8, the point cloud decoding device 800 includes: A second acquisition module 801 for acquiring a node to be decoded in a point cloud sequence and m decoded reference nodes in the point cloud sequence, where m is a positive integer; a second determination module 802 for determining a context of the subnode to be decoded based on occupancy information of the m reference nodes and a position of the subnode to be decoded in the node to be decoded, the second determination module 802 being one of the subnodes obtained by dividing the node to be decoded based on a tree structure; a decoding module 803 for entropy decoding the subnode to be decoded based on the context to generate a target code stream.
[0116] Optionally, the second acquisition module 801 further comprises: obtaining a node to be decoded in a point cloud sequence; obtaining k decoded nodes before the node to be decoded based on a node decoding order, where k is a positive integer; If at least one decoded node among the previous k decoded nodes has a target plane feature, obtaining m decoded reference nodes in the point cloud sequence.
[0117] Optionally, the second acquisition module 801 further comprises: If the number of decoded nodes having target plane features among the previous k decoded nodes is equal to or greater than a first threshold, the number of decoded m reference nodes in the point cloud sequence is used to obtain the decoded m reference nodes; Here, the first threshold is a positive integer less than k.
[0118] Optionally, the second acquisition module 801 further comprises: obtaining a node to be decoded in a point cloud sequence; determining a target coordinate system based on the coordinate values of the decoding target node; determining, as a candidate reference node, a node in the point cloud sequence that belongs to the same node split hierarchical level as the node to be decoded and has the same coordinate value on a target coordinate axis as the node to be decoded, wherein the target coordinate axis is any one of coordinate axes in the target coordinate system and is perpendicular to a target plane; and obtaining m reference nodes from the candidate reference nodes.
[0119] Optionally, the occupancy information of the encoded m reference nodes in the point cloud sequence includes the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, and the point cloud decoding device 800: a fourth determination module for determining a first plane and a second plane of the node, wherein the first plane and the second plane are both parallel to the target plane, and the first plane and the second plane are stacked along an extension direction of the target coordinate axis, the first plane is closer to a coordinate origin than the second plane, and the coordinate origin is the coordinate origin of the target coordinate system; and a fourth acquisition module for acquiring the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, wherein the first subnodes are subnodes located in the first plane and the second subnodes are subnodes located in the second plane.
[0120] Optionally, the second determination module 802 further comprises: predicting occupancy information of the node to be decoded based on occupancy information of the m reference nodes, and determining a prediction result of the node to be decoded; and determining a context of the subnode to be decoded based on the prediction result of the node to be decoded and the position of the subnode to be decoded in the node to be decoded.
[0121] Optionally, the second determination module 802 further comprises: a step of determining a first plane and a second plane of the node, wherein the first plane and the second plane are both parallel to the target plane, and the first plane and the second plane are stacked along an extension direction of the target coordinate axis, the first plane is closer to a coordinate origin than the second plane, and the coordinate origin is the coordinate origin of the target coordinate system; obtaining the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, wherein the first subnodes are subnodes located in the first plane and the second subnodes are subnodes located in the second plane; and a step of predicting occupancy information of the node to be decoded based on the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, and determining the prediction result of the node to be decoded.
[0122] Optionally, the second determination module 802 further comprises: predicting that at least one of the first subnodes in the node to be decoded is occupied and the second subnode in the node to be decoded is unoccupied, and determining that the prediction result of the node to be decoded is a first prediction result, when the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes meet a first condition, wherein the first condition is that the number of occupied first subnodes in the m reference nodes is greater than a second threshold and the number of occupied second subnodes is less than a third threshold; a step of predicting that at least one second subnode in the node to be decoded is occupied and the first subnode in the node to be decoded is unoccupied, and determining that the prediction result of the node to be decoded is a second prediction result, when the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes meet a second condition, wherein the second condition is that the number of occupied second subnodes in the m reference nodes is greater than the second threshold and the number of occupied first subnodes is less than the third threshold; and if the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes do not meet either the first condition or the second condition, predicting that at least one of the first subnodes and at least one of the second subnodes in the node to be decoded is occupied, and determining that the prediction result of the node to be decoded is a third prediction result.
[0123] Optionally, the second determination module 802 further comprises: determining a first preset model as a context of the subnode to be decoded when the prediction result of the node to be decoded is the first prediction result and the subnode to be decoded is located on a second plane of the node to be decoded; When the prediction result of the node to be decoded is the second prediction result and the subnode to be decoded is located on a first plane of the node to be decoded, determining a second preset model as a context of the subnode to be decoded; When the prediction result of the node to be decoded is the first prediction result and the subnode to be decoded is located on a first plane of the node to be decoded, obtaining adjacent nodes of the subnode to be decoded, and determining a context of the subnode to be decoded based on the occupancy status of the adjacent nodes; if the prediction result of the node to be decoded is the second prediction result and the subnode to be decoded is located on a second plane of the node to be decoded, obtaining adjacent nodes of the subnode to be decoded, and determining a context of the subnode to be decoded based on the occupancy status of the adjacent nodes; If the prediction result of the node to be decoded is the third prediction result, obtaining adjacent nodes of the subnode to be decoded and determining the context of the subnode to be decoded based on the occupancy status of the adjacent nodes.
[0124] Optionally, if the first subnode and the second subnode in the reference node are divided based on an n-ary tree, the second threshold is a positive integer less than or equal to n / 2×m-1, and the third threshold is a positive integer less than or equal to n / 2×m, where n is a positive integer.
[0125] In an embodiment of the present application, the point cloud decoding device 800 can predict the occupancy information of the node to be decoded based on the occupancy information of the decoded reference node, and determine the context of the subnode to be decoded based on the prediction result of the node to be decoded and the position of the subnode to be decoded in the node to be decoded, thereby making better use of the occupancy information of the decoded nodes in the point cloud sequence to improve the decompression performance of the point cloud and increase the decoding efficiency of the point cloud decoding device 800.
[0126] The point cloud decoding device 800 in the embodiment of the present application may be an electronic device, for example, an electronic device having an operating system, or may be a component of an electronic device, for example, an integrated circuit or a chip. The electronic device may be a terminal or other devices other than a terminal. Illustratively, the terminal includes, but is not limited to, the types of terminal 11 listed above, and the other devices may be, for example, a server, a network attached storage (NAS), etc., which are not specifically limited in the embodiment of the present application.
[0127] The point cloud decoding device 800 provided in the embodiment of the present application can implement each process implemented in the embodiment of the method described in Figure 6, and achieve the same technical effect. In order to avoid repetition, the description will be omitted here.
[0128] Optionally, as shown in FIG. 9, an embodiment of the present application further provides a communication device 900. The communication device includes a processor 901 and a memory 902, and the memory 902 stores a program or command executable on the processor 901. For example, when the communication device 900 is an encoding side, the program or command is executed by the processor 901 to realize each process of the embodiment of the method described in FIG. 1 above, and the same technical effect can be achieved. When the communication device 900 is a decoding side, the program or command is executed by the processor 901 to realize each process of the embodiment of the method described in FIG. 6 above, and the same technical effect can be achieved. To avoid repetition, the description will be omitted here.
[0129] The embodiments of the present application further provide a terminal. The implementation processes and realization modes of the method embodiments shown in Figures 1 and 6 above can all be applied to the embodiments of the terminal, and the same technical effects can be achieved. Specifically, Figure 10 is a hardware configuration schematic diagram of a terminal implementing the embodiments of the present application.
[0130] The terminal 1000 includes at least some components such as, but not limited to, a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009, and a processor 1010.
[0131] Those skilled in the art will understand that the terminal 1000 may further include a power source (e.g., a battery) for supplying power to each component, and the power source may be logically connected to the processor 1010 via a power management system, which may further realize functions such as charge / discharge management and power consumption management. The structure of the terminal shown in Figure 10 does not limit the terminal, and the terminal may include more or fewer components than those shown, or a combination of some components, or a different component arrangement, and description thereof will be omitted here.
[0132] It should be understood that in the embodiment of the present application, the input unit 1004 may include a graphics processing unit (GPU) 10041 for processing image data of static or video images acquired by an image acquisition device (e.g., a camera) in a video acquisition mode or an image acquisition mode, and a microphone 10042. The display unit 1006 may include a display panel 10061, which may be arranged in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1007 includes at least one of a touch panel 10071 and other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include two parts: a touch detection device and a touch controller. The other input devices 10072 may include, but are not limited to, a physical keyboard, function buttons (e.g., volume control buttons, switch buttons, etc.), a trackball, a mouse, and a control lever, and description thereof will be omitted here.
[0133] In the embodiment of the present application, the radio frequency unit 1001 can receive downlink data from the network side device and then transmit it to the processor 1010 for processing, and can also transmit uplink data to the network side device. Typically, the radio frequency unit 1001 includes, but is not limited to, an antenna, an amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, etc.
[0134] The memory 1009 can be used to store software programs or commands and various data. The memory 1009 may mainly include a first storage area for storing programs or commands and a second storage area for storing data, where the first storage area can store an operating system, an application or command required for at least one function (e.g., audio playback function, image playback function, etc.), etc. The memory 1009 may include volatile memory or nonvolatile memory, or may include both volatile memory and nonvolatile memory. Among them, the nonvolatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synch link dynamic random access memory (SLDRAM), and direct Rambus random access memory (DRRAM). Memory 1009 in embodiments of the present application includes, but is not limited to, these and any other suitable memory.
[0135] The processor 1010 may include one or more processing units, and optionally, an application processor that mainly processes operations related to an operating system, a user interface, applications, etc., and a modem processor that mainly processes wireless communication signals, such as a baseband processor, are integrated into the processor 1010. It is understood that the modem processor need not be integrated into the processor 1010.
[0136] Here, when the terminal 1000 is the encoding side, the processor 1010 Obtaining a node to be coded in a point cloud sequence and m coded reference nodes in the point cloud sequence, where m is a positive integer; a step of determining a context of the subnode to be coded based on occupancy information of the m reference nodes and a position of the subnode to be coded in the node to be coded, wherein the subnode to be coded is any one of the subnodes obtained by dividing the node to be coded based on a tree structure; and entropy coding the subnode to be coded based on the context to generate a target codestream.
[0137] Optionally, the processor 1010 further obtaining a node to be coded in a point cloud sequence; obtaining k coded nodes before the node to be coded based on a node coding order, where k is a positive integer; and obtaining m coded reference nodes in the point cloud sequence if at least one coded node among the previous k coded nodes has a target plane feature.
[0138] Optionally, the processor 1010 further If the number of coded nodes having the target plane feature among the previous k coded nodes is equal to or greater than a first threshold, it is used to obtain coded m reference nodes in the point cloud sequence; Here, the first threshold is a positive integer less than k.
[0139] Optionally, the processor 1010 further obtaining a node to be coded in a point cloud sequence; determining a target coordinate system based on the coordinate values of the encoding target node; determining, as a candidate reference node, a node in the point cloud sequence that belongs to the same node split hierarchical level as the encoding target node and has the same coordinate value on a target coordinate axis as the encoding target node, wherein the target coordinate axis is any coordinate axis in the target coordinate system and is perpendicular to a target plane; and obtaining m reference nodes from the candidate reference nodes.
[0140] Optionally, the occupancy information of the encoded m reference nodes in the point cloud sequence includes the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, and the processor 1010 further: determining a first plane and a second plane of the node, wherein the first plane and the second plane are both parallel to the target plane, and the first plane and the second plane are stacked in an extension direction of the target coordinate axis, the first plane is closer to a coordinate origin than the second plane, and the coordinate origin is the coordinate origin of the target coordinate system; and obtaining the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, wherein the first subnodes are subnodes located in the first plane and the second subnodes are subnodes located in the second plane.
[0141] Optionally, the processor 1010 further is used to determine a context of the subnode to be coded based on the occupancy information of the m reference nodes and the position of the subnode to be coded in the node to be coded; predicting occupancy information of the encoding target node based on occupancy information of the m reference nodes, and determining a prediction result of the encoding target node; and determining a context of the subnode to be coded based on the prediction result of the node to be coded and the position of the subnode to be coded in the node to be coded.
[0142] Optionally, the processor 1010 further determining a first plane and a second plane of the node, wherein the first plane and the second plane are both parallel to the target plane, and the first plane and the second plane are stacked in an extension direction of the target coordinate axis, the first plane is closer to a coordinate origin than the second plane, and the coordinate origin is the coordinate origin of the target coordinate system; obtaining the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, wherein the first subnodes are subnodes located in the first plane and the second subnodes are subnodes located in the second plane; and a step of predicting occupancy information of the node to be coded based on the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, and determining a prediction result of the node to be coded.
[0143] Optionally, the processor 1010 further predicting that at least one of the first subnodes in the node to be encoded is occupied and that the second subnode in the node to be encoded is unoccupied, and determining that the prediction result of the node to be encoded is a first prediction result, if the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes meet a first condition, wherein the first condition is that the number of occupied first subnodes in the m reference nodes is greater than a second threshold and the number of occupied second subnodes is less than a third threshold; predicting that at least one second subnode in the node to be encoded is occupied and that the first subnode in the node to be encoded is unoccupied, and determining that the prediction result of the node to be encoded is a second prediction result, if the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes meet a second condition, wherein the second condition is that the number of occupied second subnodes in the m reference nodes is greater than the second threshold and the number of occupied first subnodes is less than the third threshold; and if the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes do not meet either the first condition or the second condition, predicting that at least one of the first subnodes and at least one of the second subnodes in the node to be encoded is occupied, and determining that the prediction result of the node to be encoded is a third prediction result.
[0144] Optionally, the processor 1010 further determining a first preset model as a context of the subnode to be encoded when the prediction result of the node to be encoded is the first prediction result and the subnode to be encoded is located on a second plane of the node to be encoded; determining a second preset model as a context of the subnode to be encoded when the prediction result of the node to be encoded is the second prediction result and the subnode to be encoded is located on a first plane of the node to be encoded; if the prediction result of the node to be coded is the first prediction result and the subnode to be coded is located on a first plane of the node to be coded, obtaining adjacent nodes of the subnode to be coded, and determining a context of the subnode to be coded based on the occupancy status of the adjacent nodes; if the prediction result of the node to be coded is the second prediction result and the subnode to be coded is located on a second plane of the node to be coded, obtaining adjacent nodes of the subnode to be coded, and determining a context of the subnode to be coded based on the occupancy status of the adjacent nodes; If the prediction result of the node to be coded is the third prediction result, the step of obtaining adjacent nodes of the subnode to be coded and determining the context of the subnode to be coded based on the occupancy status of the adjacent nodes is used.
[0145] Optionally, if the first subnode and the second subnode in the reference node are divided based on an n-ary tree, the second threshold is a positive integer less than or equal to n / 2×m-1, and the third threshold is a positive integer less than or equal to n / 2×m, where n is a positive integer.
[0146] Alternatively, when the terminal 1000 is the decoding side, the processor 1010 Obtaining a node to be decoded in a point cloud sequence and m decoded reference nodes in the point cloud sequence, where m is a positive integer; a step of determining a context of the subnode to be decoded based on occupancy information of the m reference nodes and a position of the subnode to be decoded in the node to be decoded, wherein the subnode to be decoded is any one of the subnodes obtained by dividing the node to be decoded based on a structure tree; and entropy decoding the subnode to be decoded based on the context to generate a target code stream.
[0147] Optionally, the processor 1010 further obtaining a node to be decoded in a point cloud sequence; obtaining k decoded nodes before the node to be decoded based on a node decoding order, where k is a positive integer; If at least one decoded node among the previous k decoded nodes has a target plane feature, obtaining m decoded reference nodes in the point cloud sequence.
[0148] Optionally, the processor 1010 further If the number of decoded nodes having target plane features among the previous k decoded nodes is equal to or greater than a first threshold, the number of decoded m reference nodes in the point cloud sequence is used to obtain the decoded m reference nodes; Here, the first threshold is a positive integer less than k.
[0149] Optionally, the processor 1010 further obtaining a node to be decoded in a point cloud sequence; determining a target coordinate system based on the coordinate values of the decoding target node; determining, as a candidate reference node, a node in the point cloud sequence that belongs to the same node split hierarchical level as the node to be decoded and has the same coordinate value on a target coordinate axis as the node to be decoded, wherein the target coordinate axis is any one of coordinate axes in the target coordinate system and is perpendicular to a target plane; and obtaining m reference nodes from the candidate reference nodes.
[0150] Optionally, the occupancy information of the encoded m reference nodes in the point cloud sequence includes the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, and the processor 1010 further: a step of determining a first plane and a second plane of the node, wherein the first plane and the second plane are both parallel to the target plane, and the first plane and the second plane are stacked along an extension direction of the target coordinate axis, the first plane is closer to a coordinate origin than the second plane, and the coordinate origin is the coordinate origin of the target coordinate system; and obtaining the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, wherein the first subnodes are subnodes located in the first plane and the second subnodes are subnodes located in the second plane.
[0151] Optionally, the processor 1010 further predicting occupancy information of the node to be decoded based on occupancy information of the m reference nodes, and determining a prediction result of the node to be decoded; and determining a context of the subnode to be decoded based on the prediction result of the node to be decoded and the position of the subnode to be decoded in the node to be decoded.
[0152] Optionally, the processor 1010 further a step of determining a first plane and a second plane of the node, wherein the first plane and the second plane are both parallel to the target plane, and the first plane and the second plane are stacked along an extension direction of the target coordinate axis, the first plane is closer to a coordinate origin than the second plane, and the coordinate origin is the coordinate origin of the target coordinate system; obtaining the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, wherein the first subnodes are subnodes located in the first plane and the second subnodes are subnodes located in the second plane; and a step of predicting occupancy information of the node to be decoded based on the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, and determining the prediction result of the node to be decoded.
[0153] Optionally, the processor 1010 further predicting that at least one of the first subnodes in the node to be decoded is occupied and the second subnode in the node to be decoded is unoccupied, and determining that the prediction result of the node to be decoded is a first prediction result, when the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes meet a first condition, wherein the first condition is that the number of occupied first subnodes in the m reference nodes is greater than a second threshold and the number of occupied second subnodes is less than a third threshold; a step of predicting that at least one second subnode in the node to be decoded is occupied and the first subnode in the node to be decoded is unoccupied, and determining that the prediction result of the node to be decoded is a second prediction result, when the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes meet a second condition, wherein the second condition is that the number of occupied second subnodes in the m reference nodes is greater than the second threshold and the number of occupied first subnodes is less than the third threshold; and if the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes do not meet either the first condition or the second condition, predicting that at least one of the first subnodes and at least one of the second subnodes in the node to be decoded is occupied, and determining that the prediction result of the node to be decoded is a third prediction result.
[0154] Optionally, the processor 1010 further determining a first preset model as a context of the subnode to be decoded when the prediction result of the node to be decoded is the first prediction result and the subnode to be decoded is located on a second plane of the node to be decoded; When the prediction result of the node to be decoded is the second prediction result and the subnode to be decoded is located on a first plane of the node to be decoded, determining a second preset model as a context of the subnode to be decoded; When the prediction result of the node to be decoded is the first prediction result and the subnode to be decoded is located on a first plane of the node to be decoded, obtaining adjacent nodes of the subnode to be decoded, and determining a context of the subnode to be decoded based on the occupancy status of the adjacent nodes; if the prediction result of the node to be decoded is the second prediction result and the subnode to be decoded is located on a second plane of the node to be decoded, obtaining adjacent nodes of the subnode to be decoded, and determining a context of the subnode to be decoded based on the occupancy status of the adjacent nodes; If the prediction result of the node to be decoded is the third prediction result, the step of obtaining adjacent nodes of the subnode to be decoded and determining the context of the subnode to be decoded based on the occupancy status of the adjacent nodes is used.
[0155] Optionally, if the first subnode and the second subnode in the reference node are divided based on an n-ary tree, the second threshold is a positive integer less than or equal to n / 2×m-1, and the third threshold is a positive integer less than or equal to n / 2×m, where n is a positive integer.
[0156] The terminal 1000 provided in the embodiments of the present application can better utilize the occupancy information of the coded points and decoded points in the point cloud sequence, improve the geometric compression performance of the point cloud, and increase the coding efficiency and decoding efficiency.
[0157] The embodiments of the present application further provide a readable storage medium, which may be non-volatile or volatile, and stores a program or command, which, when executed by a processor, realizes each process of the method embodiments shown in Figure 1 or Figure 6, and can achieve the same technical effect. To avoid repetition, the description will be omitted here.
[0158] Wherein, the processor is the processor in the terminal described in the above embodiment. The readable storage medium includes, for example, a computer readable storage medium such as a computer read only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0159] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, and the communication interface and the processor are coupled together, and the processor is used to execute programs or commands to realize each process of the method embodiments shown in Figure 1 or Figure 6, and can achieve the same technical effects. To avoid repetition, the description will be omitted here.
[0160] It should be understood that the chips referred to in the embodiments of this application may also be referred to as system level chips, system chips, chip systems, or system-on-chips, etc.
[0161] The embodiments of the present application further provide a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement each process of the method embodiments shown in Figure 1 or Figure 6, and can achieve the same technical effects. To avoid repetition, the description will be omitted here.
[0162] It should be noted that, as used herein, the terms "comprise," "consist," or any other variation thereof, are intended to include a non-exclusive inclusion, whereby a process, method, article, or apparatus comprising a set of elements includes not only those elements but also other elements not expressly specified or inherent in such process, method, article, or apparatus. Unless otherwise specified, elements qualified by the phrase "comprise..." do not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element. It should also be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may include performing functions substantially simultaneously or in the reverse order, depending on such functionality. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Furthermore, features described with reference to one example may be combined in other examples.
[0163] From the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be realized in the form of a combination of software and a necessary common hardware platform. Of course, hardware implementation is also possible, but in many cases the former is a more preferred embodiment. Based on this understanding, the technical solutions of the present application can be substantially embodied in the form of a computer software product, which is stored in a storage medium (e.g., ROM / RAM, magnetic disk, optical disk) and includes a plurality of commands that cause a terminal (which may be a mobile phone, computer, server, air conditioner, network device, etc.) to execute the methods described in each embodiment of the present application.
[0164] Although the examples of the present application have been described above with reference to the drawings, the present application is not limited to the above-mentioned specific embodiments, which are merely illustrative and not limiting. Based on the suggestions of the present application, many forms that a person skilled in the art can make without departing from the spirit of the present application and the scope of protection of the claims are all within the scope of protection of the present application.
Claims
1. An encoding side obtains a node to be encoded in a point cloud sequence and m encoded reference nodes in the point cloud sequence, where m is a positive integer; a step in which the encoding side determines a context of the subnode to be encoded based on occupancy information of the m reference nodes and a position of the subnode to be encoded in the node to be encoded, the subnode to be encoded being any of the subnodes obtained by dividing the node to be encoded based on a tree structure; and a step in which the encoding side entropy codes the subnode to be encoded based on the context to generate a target code stream.
2. The step of the encoding side obtaining a node to be encoded in a point cloud sequence and m encoded reference nodes in the point cloud sequence includes: The encoding side acquires a node to be encoded in a point cloud sequence; a step of obtaining k coded nodes before the coding target node based on a node coding order by the coding side, where k is a positive integer; and if at least one coded node among the previous k coded nodes has a target plane feature, the coding side obtains m coded reference nodes in the point cloud sequence.
3. When at least one coded node among the previous k coded nodes has a target plane feature, the step of the coding side obtaining m coded reference nodes in the point cloud sequence includes: When the number of encoded nodes having target plane features among the previous k encoded nodes is equal to or greater than a first threshold, the encoding side obtains m encoded reference nodes in the point cloud sequence; The point cloud encoding method according to claim 2 , wherein the first threshold is a positive integer less than k.
4. The step of the encoding side obtaining a node to be encoded in a point cloud sequence and m encoded reference nodes in the point cloud sequence includes: The encoding side acquires a node to be encoded in a point cloud sequence; a step of determining a target coordinate system based on the coordinate values of the node to be encoded by the encoding side; a step in which the encoding side determines, as a candidate reference node, a node in the point cloud sequence that belongs to the same node division hierarchical level as the encoding target node and has the same coordinate value on a target coordinate axis as the encoding target node, wherein the target coordinate axis is any coordinate axis in the target coordinate system and is perpendicular to a target plane; The point cloud encoding method according to claim 1 , further comprising: a step in which the encoding side obtains m reference nodes from the candidate reference nodes.
5. The occupancy information of the encoded m reference nodes in the point cloud sequence includes the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes; The point cloud encoding method includes: a step in which the encoding side determines a first plane and a second plane of a node, the first plane and the second plane are both parallel to the target plane, the first plane and the second plane are stacked in an extension direction of the target coordinate axis, the first plane is closer to a coordinate origin than the second plane, and the coordinate origin is the coordinate origin of the target coordinate system; 5. The point cloud encoding method according to claim 4, further comprising: a step in which the encoding side obtains the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, wherein the first subnodes are subnodes located in the first plane and the second subnodes are subnodes located in the second plane.
6. The step of determining the context of the subnode to be coded based on the occupancy information of the m reference nodes and the position of the subnode to be coded in the node to be coded by the coding side includes: the encoding side predicting occupancy information of the encoding target node based on occupancy information of the m reference nodes, and determining a prediction result of the encoding target node; The point cloud encoding method according to claim 1, further comprising a step in which the encoding side determines a context of the subnode to be encoded based on a prediction result of the node to be encoded and a position of the subnode to be encoded in the node to be encoded.
7. The step of the encoding side predicting occupancy information of the encoding target node based on occupancy information of the m reference nodes and determining a prediction result of the encoding target node includes: a step in which the encoding side determines a first plane and a second plane of the node, the first plane and the second plane are both parallel to a target plane, the first plane and the second plane are stacked in an extension direction of a target coordinate axis, the first plane is closer to a coordinate origin than the second plane, and the coordinate origin is the coordinate origin of a target coordinate system; The encoding side obtains the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, where the first subnodes are subnodes located in the first plane and the second subnodes are subnodes located in the second plane; The point cloud encoding method according to claim 6, further comprising a step in which the encoding side predicts occupancy information of the node to be encoded based on the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, and determines a prediction result of the node to be encoded.
8. The step of the encoding side predicting occupancy information of the encoding target node based on the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes and determining the prediction result of the encoding target node, a step of the encoding side predicting that at least one of the first subnodes in the node to be encoded is occupied and the second subnode in the node to be encoded is unoccupied, and determining that the prediction result of the node to be encoded is a first prediction result, when the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes meet a first condition, wherein the first condition is that the number of occupied first subnodes in the m reference nodes is greater than a second threshold and the number of occupied second subnodes is less than a third threshold; a step of the encoding side predicting that at least one of the second subnodes in the node to be encoded is occupied and the first subnode in the node to be encoded is unoccupied, and determining that the prediction result of the node to be encoded is a second prediction result, when the number of the occupied first subnodes and the number of the occupied second subnodes in the m reference nodes meet a second condition, wherein the second condition is that the number of the occupied second subnodes in the m reference nodes is greater than the second threshold and the number of the occupied first subnodes is less than the third threshold; and if the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes do not meet either the first condition or the second condition, the encoding side predicts that at least one of the first subnodes and at least one of the second subnodes in the node to be encoded is occupied, and determines that the prediction result of the node to be encoded is a third prediction result.
9. The step of determining a context of the subnode to be coded based on the prediction result and the position of the subnode to be coded in the node to be coded by the coding side includes: When the prediction result of the node to be coded is the first prediction result and the sub-node to be coded is located on a second plane of the node to be coded, the coding side determines a first preset model as a context of the sub-node to be coded; When the prediction result of the node to be coded is the second prediction result and the sub-node to be coded is located on a first plane of the node to be coded, the coding side determines a second preset model as a context of the sub-node to be coded; When the prediction result of the node to be coded is the first prediction result and the subnode to be coded is located on a first plane of the node to be coded, the coding side obtains adjacent nodes of the subnode to be coded, and determines a context of the subnode to be coded based on the occupancy status of the adjacent nodes; When the prediction result of the node to be coded is the second prediction result and the subnode to be coded is located on a second plane of the node to be coded, the coding side obtains adjacent nodes of the subnode to be coded, and determines a context of the subnode to be coded based on the occupancy status of the adjacent nodes; and if the prediction result of the node to be encoded is the third prediction result, the encoding side obtains adjacent nodes of the subnode to be encoded and determines a context of the subnode to be encoded based on the occupancy status of the adjacent nodes.
10. A decoding side obtains a node to be decoded in a point cloud sequence and m decoded reference nodes in the point cloud sequence, where m is a positive integer; a step in which the decoding side determines a context of the subnode to be decoded based on occupancy information of the m reference nodes and a position of the subnode to be decoded in the node to be decoded, the subnode to be decoded being any of the subnodes obtained by dividing the node to be decoded based on a tree structure; the decoding side entropy-decoding the subnode to be decoded based on the context to generate a target code stream.
11. The step of the decoding side obtaining a node to be decoded in a point cloud sequence and m decoded reference nodes in the point cloud sequence includes: The decoding side acquires a node to be decoded in a point cloud sequence; The decoding side obtains k decoded nodes before the node to be decoded based on a node decoding order, where k is a positive integer; and if at least one decoded node among the previous k decoded nodes has a target plane feature, the decoding side obtains m decoded reference nodes in the point cloud sequence.
12. When at least one decoded node among the previous k decoded nodes has a target plane feature, the step of the decoding side obtaining m decoded reference nodes in the point cloud sequence includes: When the number of decoded nodes having target plane features among the previous k decoded nodes is equal to or greater than a first threshold, the decoding side obtains m decoded reference nodes in the point cloud sequence; The point cloud decoding method according to claim 11 , wherein the first threshold is a positive integer less than k.
13. The step of the decoding side obtaining a node to be decoded in a point cloud sequence and m decoded reference nodes in the point cloud sequence includes: The decoding side acquires a node to be decoded in a point cloud sequence; a step of determining a target coordinate system based on the coordinate values of the decoding target node by the decoding side; a step in which the decoding side determines, as a candidate reference node, a node in the point cloud sequence that belongs to the same node split hierarchical level as the node to be decoded and has the same coordinate value on a target coordinate axis as the node to be decoded, wherein the target coordinate axis is any one of coordinate axes in the target coordinate system and is perpendicular to a target plane; The point cloud decoding method according to claim 10, further comprising: the decoding side obtaining m reference nodes from the candidate reference nodes.
14. The occupancy information of the encoded m reference nodes in the point cloud sequence includes the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, and the point cloud decoding method includes: a step in which the decoding side determines a first plane and a second plane of a node, the first plane and the second plane being both parallel to the target plane, the first plane and the second plane being stacked along an extension direction of the target coordinate axis, the first plane being closer to a coordinate origin than the second plane, and the coordinate origin being the coordinate origin of the target coordinate system; The point cloud decoding method of claim 13, further comprising: a step in which the decoding side obtains the number of occupied first subnodes and the number of occupied second subnodes in the m reference nodes, wherein the first subnodes are subnodes located in the first plane and the second subnodes are subnodes located in the second plane.
15. 10. An electronic device comprising a processor and a memory, wherein the memory stores a program or command executable by the processor, and when the program or command is executed by the processor, the steps of the point cloud encoding method described in any one of claims 1 to 9 are realized.
16. An electronic device comprising a processor and a memory, wherein the memory stores a program or command executable by the processor, and when the program or command is executed by the processor, the steps of the point cloud decoding method described in any one of claims 10 to 14 are realized.
Citation Information
Patent Citations
Point cloud geometrical information encoding and decoding method
CN112565795A
Context determination for planar mode in octree-based point cloud coding
US10693492B1
Methods and devices for neighbourhood-based occupancy prediction in point cloud compression
US20210272324A1
Three-dimensional data coding method, three-dimensional data decoding method, three-dimensional data coding device, and three-dimensional data decoding device
WO2019198636A1
Point cloud encoding / decoding method, encoder, decoder, and storage medium
WO2021232251A1