Encoding, decoding method, apparatus and device
Patent Information
- Application Number
- CN202510351831.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2026-09-25
AI Technical Summary
[0002]相关技术中的算法并未充分考虑点云进行八叉树划分后的节点内的局部点云的特征,且默认径向分量不具备平面模式资格,导致无法充分利用局部点云的特征,使得部分符合平面模式的节点未被识别,进而限制了编解码的压缩效率
[0053]第十八方面,提供了一种计算机程序/程序产品,所述计算机程序/程序产品被存储在存储介质中,所述计算机程序/程序产品被至少一个处理器执行以实现如第一方面、第三方面、第五方面或第七方面所述的方法的步骤。
Smart Images

Figure CN122824907A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of video processing technology, specifically relating to an encoding and decoding method, apparatus, and device. Background Technology
[0002] The algorithms in the related technologies do not fully consider the characteristics of the local point cloud within the nodes after the point cloud is divided into octrees, and the radial component is not qualified for planar mode by default. This makes it impossible to make full use of the characteristics of the local point cloud, resulting in some nodes that conform to the planar mode not being recognized, which in turn limits the compression efficiency of encoding and decoding. Summary of the Invention
[0003] This application provides an encoding and decoding method, apparatus, and device to improve encoding, decoding, and compression efficiency.
[0004] Firstly, a decoding method is provided, which includes:
[0005] The decoding device performs octree reconstruction on the point cloud to be decoded to obtain the target node;
[0006] Based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device, the planar mode decoding qualification of the first component of the target node is determined, where the first component is the S-axis component or the T-axis component of the target node.
[0007] Based on the determination result of the planar mode decoding qualification of the first component, the S-axis component, T-axis component and V-axis component of the target node are decoded.
[0008] Secondly, a decoding device is provided, comprising:
[0009] The first processing module is used to reconstruct the point cloud to be decoded into an octree to obtain the target node;
[0010] The second processing module is used to determine the planar mode decoding qualification of the first component of the target node based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device. The first component is the S-axis component or the T-axis component of the target node.
[0011] The third processing module is used to decode the S-axis component, T-axis component and V-axis component of the target node based on the determination result of the planar mode decoding qualification of the first component.
[0012] Thirdly, an encoding method is provided, which includes:
[0013] The encoding device performs octree partitioning on the point cloud to be encoded to obtain the target node;
[0014] Based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device, the planar pattern encoding qualification of the first component of the target node is determined, where the first component is the S-axis component or the T-axis component of the target node.
[0015] Based on the determination result of the planar pattern encoding qualification of the first component, the S-axis component, T-axis component and V-axis component of the target node are encoded.
[0016] Fourthly, an encoding device is provided, comprising:
[0017] The seventh processing module is used to perform octree partitioning on the point cloud to be encoded to obtain the target nodes;
[0018] The eighth processing module is used to determine the planar pattern encoding qualification of the first component of the target node based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device. The first component is the S-axis component or the T-axis component of the target node.
[0019] The ninth processing module is used to encode the S-axis component, T-axis component, and V-axis component of the target node based on the determination result of the planar pattern encoding qualification of the first component.
[0020] Fifthly, a decoding method is provided, which includes:
[0021] The decoding device performs octree reconstruction on the point cloud to be decoded to obtain the target node;
[0022] Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, determine the processing order of the planar mode of the S-axis component, T-axis component and V-axis component of the target node;
[0023] According to the processing order, the S-axis component, T-axis component, and V-axis component of the target node are decoded.
[0024] Sixthly, a decoding device is provided, comprising:
[0025] The fourth processing module is used to reconstruct the point cloud to be decoded into an octree to obtain the target node;
[0026] The fifth processing module is used to determine the processing order of the planar modes of the S-axis, T-axis and V-axis components of the target node based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device.
[0027] The sixth processing module is used to decode the S-axis component, T-axis component and V-axis component of the target node according to the processing order.
[0028] Seventhly, an encoding method is provided, the method comprising:
[0029] The encoding device performs octree partitioning on the point cloud to be encoded to obtain the target node;
[0030] Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, determine the processing order of the planar mode of the S-axis component, T-axis component and V-axis component of the target node;
[0031] According to the processing order, the S-axis component, T-axis component, and V-axis component of the target node are encoded.
[0032] Eighthly, an encoding device is provided, comprising:
[0033] The tenth processing module is used to perform octree partitioning on the point cloud to be encoded to obtain the target node;
[0034] The eleventh processing module is used to determine the processing order of the planar modes of the S-axis, T-axis and V-axis components of the target node based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device.
[0035] The twelfth processing module is used to encode the S-axis component, T-axis component and V-axis component of the target node according to the processing order.
[0036] A ninth aspect provides a decoding device including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first or fifth aspect.
[0037] In a tenth aspect, a decoding device is provided, including a processor and a communication interface, wherein the processor is used to perform octree reconstruction on the point cloud to be decoded to obtain the target node;
[0038] Based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device, the planar mode decoding qualification of the first component of the target node is determined, where the first component is the S-axis component or the T-axis component of the target node.
[0039] Based on the determination result of the planar mode decoding qualification of the first component, the S-axis component, T-axis component and V-axis component of the target node are decoded.
[0040] Eleventhly, a decoding device is provided, including a processor and a communication interface, wherein the processor is used to perform octree reconstruction on the point cloud to be decoded to obtain the target node;
[0041] Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, determine the processing order of the planar mode of the S-axis component, T-axis component and V-axis component of the target node;
[0042] According to the processing order, the S-axis component, T-axis component, and V-axis component of the target node are decoded.
[0043] In a twelfth aspect, an encoding device is provided, including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the third or seventh aspect.
[0044] In a thirteenth aspect, an encoding device is provided, including a processor and a communication interface, wherein the processor is used to perform octree partitioning on a point cloud to be encoded to obtain target nodes;
[0045] Based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device, the planar pattern encoding qualification of the first component of the target node is determined, where the first component is the S-axis component or the T-axis component of the target node.
[0046] Based on the determination result of the planar pattern encoding qualification of the first component, the S-axis component, T-axis component and V-axis component of the target node are encoded.
[0047] In a fourteenth aspect, an encoding device is provided, including a processor and a communication interface, wherein the processor is used to perform octree partitioning on a point cloud to be encoded to obtain target nodes;
[0048] Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, determine the processing order of the planar mode of the S-axis component, T-axis component and V-axis component of the target node;
[0049] According to the processing order, the S-axis component, T-axis component, and V-axis component of the target node are encoded.
[0050] In a fifteenth aspect, a coding / decoding system is provided, comprising: an encoding device and a decoding device, wherein the encoding device is configured to perform the steps of the method as described in the third or seventh aspect, and the decoding device is configured to perform the steps of the method as described in the first or fifth aspect.
[0051] In a sixteenth aspect, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the methods described in the first, third, fifth, or seventh aspects.
[0052] In a seventeenth aspect, a chip is provided, the chip including a processor and a communication interface coupled to the processor, the processor being configured to run programs or instructions to implement the steps of the methods described in the first, third, fifth, or seventh aspects.
[0053] Eighteenth aspect, a computer program / program product is provided, the computer program / program product being stored in a storage medium, the computer program / program product being executed by at least one processor to implement the steps of the method as described in the first aspect, third aspect, fifth aspect or seventh aspect.
[0054] In this embodiment, the point cloud to be decoded is first reconstructed using an octree. Then, based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device, the planar mode decoding qualification of the S-axis component or T-axis component of the target node is determined. Subsequently, based on the determination result of the first component's planar mode decoding qualification, the S-axis component, T-axis component, and V-axis component of the target node are decoded. This embodiment determines the planar mode decoding qualification of the S-axis component or T-axis component based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device. This ensures the accuracy of the determination of the planar mode decoding qualification of the S-axis component or T-axis component, effectively improving the applicability of the planar mode encoding and decoding algorithm, and thus improving the compression efficiency of point cloud geometric information encoding and decoding. Attached Figure Description
[0055] Figure 1 This is a schematic diagram of the encoding / decoding system provided in an embodiment of this application;
[0056] Figure 2 This is a flowchart of the encoding process performed by an encoder based on the AVS-PCC encoding framework;
[0057] Figure 3 This is a flowchart of the encoding process performed by an encoder based on the MPEG G-PCC encoding framework;
[0058] Figure 4 This is a flowchart of the decoding process performed by the decoder based on the AVS-PCC decoding framework;
[0059] Figure 5 This is a flowchart of the decoding process performed by the decoder based on the MPEG G-PCC decoding framework;
[0060] Figure 6 This is one of the flowcharts illustrating the decoding method according to an embodiment of this application;
[0061] Figure 7 This is a schematic diagram of the decoding process according to an embodiment of this application;
[0062] Figure 8 This is a schematic diagram of the encoding process of an embodiment of this application;
[0063] Figure 9 This is one of the flowcharts illustrating the encoding method of this application embodiment;
[0064] Figure 10 This is a second schematic flowchart of the decoding method according to an embodiment of this application;
[0065] Figure 11 This is a second schematic flowchart of the encoding method according to an embodiment of this application;
[0066] Figure 12 This is one of the module schematic diagrams of the decoding device according to an embodiment of this application;
[0067] Figure 13 This is a second schematic diagram of the decoding device according to an embodiment of this application;
[0068] Figure 14 This is one of the schematic diagrams of the encoding device according to an embodiment of this application;
[0069] Figure 15 This is a second schematic diagram of the encoding device according to an embodiment of this application;
[0070] Figure 16 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application;
[0071] Figure 17 This is a schematic diagram of the terminal structure according to an embodiment of this application. Detailed Implementation
[0072] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0073] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, without limiting the number of objects; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, "A or B" covers three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0074] Before introducing the technical solutions provided in the embodiments of this application, the meanings of some terms will be explained first.
[0075] Point cloud: A point cloud is a set of discrete points in space that are randomly distributed and represent the spatial structure and surface properties of a three-dimensional object or scene. Point clouds can be classified into different categories according to different classification criteria. For example, according to the method of acquiring the point cloud, it can be divided into dense point clouds and sparse point clouds; or according to the temporal type of the point cloud, it can be divided into static point clouds and dynamic point clouds.
[0076] Point cloud data: Point cloud data is composed of the geometric coordinates and attribute information of each point. Geometric coordinate information, also known as 3D position information, refers to the spatial coordinates (x, y, z) of a point in the point cloud. This can include the coordinate values of the point along each coordinate axis of a 3D coordinate system, such as the coordinate value x along the X-axis, the coordinate value y along the Y-axis, and the coordinate value z along the Z-axis. The attribute information of a point in the point cloud can include at least one of the following: color information, material information, and laser reflection intensity information (also known as reflectivity). Typically, each point in the point cloud has the same number of attribute information. For example, each point in the point cloud can have both color information and laser reflection intensity information, or it can have color information, material information, and laser reflection intensity information.
[0077] Point cloud encoding (PCC) refers to the process of encoding the geometric coordinates and attribute information of each point in a point cloud to obtain a compressed bitstream. Point cloud encoding can include two main processes: geometric coordinate information encoding and attribute information encoding. Currently, point cloud encoding frameworks that can compress point clouds include the Geometry-Point Cloud Compression (G-PCC) codec framework provided by the Moving Picture Experts Group (MPEG) or the Video Point Cloud Compression (V-PCC) codec framework, or the AVS-PCC codec framework provided by the Audio Video Standard (AVS).
[0078] Point cloud decoding: Point cloud decoding refers to decoding the compressed bitstream obtained from point cloud encoding to reconstruct the point cloud. More specifically, it refers to the process of reconstructing the geometric coordinates and attribute information of each point in the point cloud based on the geometric bitstream and attribute bitstream in the compressed bitstream. After obtaining the compressed bitstream at the decoding end, for the geometric bitstream, entropy decoding is first performed to obtain the quantized information of each point in the point cloud, and then inverse quantization is performed to reconstruct the geometric coordinates of each point in the point cloud. For the attribute bitstream, entropy decoding is first performed to obtain the quantized attribute residual information or quantized transform coefficients of each point in the point cloud; then, inverse quantization is performed on the quantized attribute residual information to obtain the reconstructed residual information, and inverse quantization is performed on the quantized transform coefficients to obtain the reconstructed transform coefficients. The reconstructed transform coefficients are then inversely transformed to obtain the reconstructed residual information. Based on the reconstructed residual information of each point in the point cloud, the attribute information of each point in the point cloud can be reconstructed. The reconstructed attribute information of each point in the point cloud is then matched one-to-one with the reconstructed geometric coordinate information in sequence to reconstruct the point cloud.
[0079] Figure 1 This is a schematic diagram of the encoding / decoding system provided in an embodiment of this application. The technical solution of this application embodiment relates to encoding / decoding (CODEC) point cloud data (including encoding or decoding).
[0080] like Figure 1As shown, the encoding / decoding system includes a source device 100, which provides encoded point cloud data to be decoded and displayed by a destination device 110. Specifically, the source device 100 provides the point cloud data to the destination device 110 via a communication medium 120. The source device 100 and the destination device 110 may include any one or more of the following: desktop computer, laptop computer, tablet computer, set-top box, mobile phone, wearable device (e.g., smartwatch or wearable camera), television, camera, display device, in-vehicle device, virtual reality (VR) device, augmented reality (AR) device, mixed reality (MR) device, digital media player, video game console, video conferencing equipment, video streaming equipment, broadcast receiver equipment, broadcast transmitter equipment, spacecraft, aircraft, robot, satellite, etc.
[0081] exist Figure 1 In this example, source device 100 includes a data source 101, a memory 102, an encoder 200, and an output interface 104. Destination device 110 includes an input interface 111, a decoder 300, a memory 113, and a display device 114. Source device 100 represents an example of an encoding device, while destination device 110 represents an example of a decoding device. In other examples, source device 100 and destination device 110 may not include... Figure 1 Some components, or may include Figure 1 Other components besides the source device 100. For example, the source device 100 can acquire point cloud data through an external capture device. Similarly, the destination device 110 can interface with an external display device, without including an integrated display device. Furthermore, the memory 102 and memory 113 can be external memories.
[0082] Although Figure 1 Source device 100 and destination device 110 are illustrated as separate devices, but in some examples, they may also be integrated into a single device. In such embodiments, the same hardware or software, or separate hardware or software, or any combination thereof, may be used to implement the functionality corresponding to source device 100 and destination device 110.
[0083] In some examples, source device 100 and destination device 110 can perform unidirectional or bidirectional data transmission. In the case of bidirectional data transmission, source device 100 and destination device 110 can operate in a substantially symmetrical manner, i.e., each of source device 100 and destination device 110 includes an encoder and a decoder.
[0084] Data source 101 represents the source of point cloud data (i.e., raw, unencoded point cloud data) and provides the point cloud data to encoder 200, which encodes the point cloud data. Source device 100 may include capture devices (e.g., camera devices, sensing devices, or scanning devices), archives containing previously captured point cloud data, or feed interfaces for receiving point cloud data from data content providers. Camera devices may include ordinary cameras, stereo cameras, and light field cameras; sensing devices may include laser devices, radar devices, etc.; and scanning devices may include 3D laser scanning devices, etc. Point cloud data can be obtained by capturing real-world visual scenes using capture devices. Alternatively, data source 101 may generate computer graphics-based data as source data, or combine real-time data, archived data, and computer-generated data. For example, the data source may generate point cloud data based on virtual objects (e.g., virtual 3D objects and virtual 3D scenes obtained through 3D modeling).
[0085] Encoder 200 encodes captured, pre-captured, or computer-generated data. Encoder 200 can rearrange point cloud data from the received order (sometimes referred to as the "display order") according to the encoded order. Encoder 200 can generate a bitstream including the encoded point cloud data. Source device 100 can then output the encoded point cloud data to communication medium 120 via output interface 104 for reception or retrieval, for example, by input interface 111 of destination device 110.
[0086] The memory 102 of the source device 100 and the memory 113 of the destination device 110 represent general-purpose memory. In some examples, memory 102 may store raw data from data source 101, and memory 113 may store decoded point cloud data from decoder 300. Additionally or alternatively, memories 102 and 113 may respectively store software instructions executable by, for example, encoder 200 and decoder 300. Although memories 102 and 113 are shown separately from encoder 200 and decoder 300 in this example, it should be understood that encoder 200 and decoder 300 may also include internal memory for functionally similar or equivalent purposes. If encoder 200 and decoder 300 are deployed on the same hardware device, memories 102 and 113 may be the same memory. Furthermore, memories 102 and 113 may store, for example, encoded point cloud data output from encoder 200 and input to decoder 300. In some examples, portions of memories 102 and 113 may be allocated as one or more point cloud buffers, for example, to store raw, decoded, or encoded point cloud data.
[0087] In some examples, source device 100 can output encoded data from output interface 104 to memory 113. Similarly, destination device 110 can access encoded data from memory 113 via input interface 111. Memory 113 or memory 102 can include any of a variety of distributed or locally accessed data storage media, such as hard drives, Blu-ray discs, digital versatile discs (DVDs), compact disc read-only memory (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded point cloud data.
[0088] Output interface 104 may include any type of medium or device capable of transmitting encoded point cloud data from source device 100 to destination device 110. For example, output interface 104 may include a transmitter or transceiver, such as an antenna, configured to transmit encoded point cloud data directly from source device 100 to destination device 110 in real time. The encoded point cloud data may be modulated according to the communication standards of a wireless communication protocol and transmitted to destination device 110.
[0089] Communication medium 120 may include transient media, such as wireless broadcasting or wired network transmission. For example, communication medium 120 may include radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). Communication medium 120 may form part of a packet-based network (such as a local area network, a wide area network, or a global network such as the Internet). Communication medium 120 may also take the form of a storage medium (e.g., a non-transitory storage medium), such as a hard disk, flash drive, compact disk, digital point cloud disk, Blu-ray disc, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded point cloud data.
[0090] In some implementations, the communication medium 120 may include a router, switch, base station, or any other device that can be used to facilitate communication from source device 100 to destination device 110. For example, a server (not shown) may receive encoded point cloud data from source device 100 and provide it to destination device 110, for example, via network transmission. The server may include (e.g., a web server for a website), a server configured to provide file transfer protocol services (such as File Transfer Protocol (FTP) or File DeliveryOver Unidirectional Transport (FLUTE) protocol), a content delivery network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Services (MBMS) or Evolved Multimedia Broadcast Multicast Service (eMBMS) server, or a Network-attached Storage (NAS) device, etc. The server can implement one or more HTTP streaming protocols, such as MPEG Media Transport (MMT), Dynamic Adaptive Streaming over HTTP (DASH), HTTP Live Streaming (HLS), or Real Time Streaming Protocol (RTSP).
[0091] Destination device 110 can access encoded point cloud data from a server, for example, via a wireless channel (e.g., Wi-Fi connection) or a wired connection (e.g., Digital subscriber line (DSL), cable modem, etc.) for accessing encoded point cloud data stored on the server.
[0092] Output interface 104 and input interface 111 can represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to the IEEE 802.11 or IEEE 802.15 standard (e.g., ZigBee™), Bluetooth standard, or other physical components. In an example where output interface 104 and input interface 111 include wireless components, output interface 104 and input interface 111 can be configured to transmit data, such as encoded point cloud data, via Wi-Fi, Ethernet, or cellular networks (such as 4G, LTE (Long Term Evolution), Advanced LTE, 5G, 6G, etc.).
[0093] The technology provided in this application can be applied to support one or more of the following application scenarios: machine-perceived point clouds, which can be used in autonomous navigation systems, real-time inspection systems, geographic information systems, visual sorting robots, disaster relief robots, and other scenarios; human-perceived point clouds, which can be used in point cloud application scenarios such as digital cultural heritage, free-viewpoint broadcasting, 3D immersive communication, and 3D immersive interaction.
[0094] The input interface 111 of the destination device 110 receives an encoded bitstream from the communication medium 120. The encoded bitstream may include high-level syntax elements and encoded data units (e.g., sequences, image groups, images, slices, blocks, etc.), where the high-level syntax elements are used to decode the encoded data units to obtain decoded point cloud data. The display device 114 displays the decoded point cloud data to the user. The display device 114 may include a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices. In some examples, the destination device 110 may not have a display device 114; for example, if the decoded point cloud data is used to determine the location of a physical object, the display device 114 may be replaced by a processor.
[0095] The encoder 200 and decoder 300 can be implemented as one or more of various processing circuits, which may include microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. When the technology is implemented wholly or partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and use one or more processors to execute the instructions in hardware to perform the technology provided in the embodiments of this application.
[0096] The basic principles of the encoder 200 and decoder 300 provided in this application embodiment are introduced below, taking the G-PCC and AVS-PCC codec frameworks as examples.
[0097] The encoding and decoding frameworks of G-PCC and AVS-PCC are largely the same. For example... Figure 2 The diagram illustrates the encoding flowchart executed by the encoder within the AVS-PCC-based encoding framework, as shown below. Figure 3 This diagram illustrates the encoding flowchart performed by an encoder based on the MPEG G-PCC encoding framework. The encoder described above can be... Figure 1 The encoder 200 shown above. The above encoding framework can be broadly divided into a geometric coordinate information encoding process and an attribute information encoding process. In the geometric information encoding process, the geometric coordinate information of each point in the point cloud is encoded to obtain a geometric bitstream; in the attribute information encoding process, the attribute information of each point in the point cloud is encoded to obtain an attribute bitstream; the geometric bitstream and the attribute bitstream together constitute the compressed bitstream of the point cloud.
[0098] For the geometric information encoding process, the encoding flow executed by encoder 200 is as follows:
[0099] 1. Pre-processing: This can include coordinate transformation and voxelization. Through scaling and translation operations, pre-processing converts the point cloud data in 3D space into integer form and moves its smallest geometric position to the origin. In some examples, encoder 200 may not perform pre-processing.
[0100] 2. Geometric Coding: For the AVS-PCC coding framework, geometric coding includes two modes: octree-based geometric coding and prediction tree-based geometric coding. For the G-PCC coding framework, geometric coding includes three modes: octree-based geometric coding, trisoup-based geometric coding, and prediction tree-based prediction coding. Among them:
[0101] Octree-based geometric encoding: An octree is a tree-like data structure that uniformly divides a predefined bounding box in three-dimensional space, with each node having eight child nodes. By using "1" and "0" to indicate whether each child node of the octree is occupied, occupancy code information is obtained as the bitstream of point cloud geometric information.
[0102] Geometric coding based on prediction trees: A prediction tree is generated using a prediction strategy. Starting from the root node of the prediction tree, each node is traversed, and the residual coordinate value corresponding to each traversed node is encoded.
[0103] Geometric encoding based on triangulation: The point cloud is divided into blocks of a certain size, and the intersection points (called vertices) of the point cloud surface with the edges of the blocks are located. Geometric information is compressed by encoding whether there are intersection points on the edges of the blocks and the positions of the intersection points.
[0104] 3. Geometric Entropy Encoding: This method uses statistical compression encoding on the occupancy code information of the octree, the prediction residual information of the prediction tree, and the vertex information of the triangular representation, finally outputting a binary (0 or 1) compressed bitstream. Statistical coding is a lossless coding method that can effectively reduce the bit rate required to represent the same signal. A commonly used statistical coding method is Content Adaptive Binary Arithmetic Coding (CABAC).
[0105] 4. Geometric Reconstruction: Decoding and reconstructing the geometric information after geometric encoding.
[0106] For the attribute information encoding process, the encoding flow executed by encoder 200 is as follows:
[0107] 1. Color Transformation: Apply transformations to change the color information of an attribute to a different domain. For example, color information can be transformed from the RGB color space to the YCbCr color space.
[0108] 2. Attribute Recoloring: In lossy encoding, after encoding the geometric coordinate information, the encoding end needs to decode and reconstruct the geometric information, that is, restore the geometric information of each point in the point cloud. Attribute information corresponding to one or more neighboring points in the original point cloud is used as the attribute information for the reconstructed point.
[0109] In some examples, encoder 200 may not perform color transformation or attribute recoloring.
[0110] 3. Attribute information processing: In AVS-PCC, attribute information processing can include three modes: prediction coding, transformation coding, and prediction & transformation coding. These three coding modes can be used under different conditions.
[0111] Predictive coding refers to determining the neighboring points of the point to be coded as prediction points among the already coded points based on information such as distance or spatial relationships. Based on set criteria, the predicted attribute information of the point to be coded is calculated according to the attribute information of the prediction points. The difference between the actual attribute information and the predicted attribute information of the point to be coded is calculated as attribute residual information. This attribute residual information is then quantized, transformed (optional), and entropy encoded.
[0112] Transform coding refers to using transformation methods such as Discrete Cosine Transform (DCT) and Haar Transform (Haar) to group and transform attribute information, quantize the transformation coefficients, obtain attribute reconstruction information through inverse quantization and inverse transformation, calculate the difference between the real attribute information and the attribute reconstruction information to obtain attribute residual information and quantize it, and then entropy-encode the quantized transformation coefficients and attribute residuals.
[0113] Predictive transform coding refers to using the attribute residual information obtained from prediction to perform transformation, and then quantizing and entropy coding the transform coefficients.
[0114] In MPEG G-PCC, attribute information processing can include three modes: Prediction Transform coding, Lifting Transform coding, and Region Adaptive Hierarchical Transform (RAHT) coding. These three coding modes can be used under different conditions.
[0115] Predictive transform coding refers to dividing the point cloud into multiple different levels of detail (LoD) based on distance-selected subsets of points, achieving a multi-quality, hierarchical point cloud representation from coarse to fine. Bottom-up prediction is possible between adjacent layers, where neighboring points in the coarse layer predict the attribute information of points introduced in the fine layer, obtaining the corresponding attribute residual information. The points at the lowest level are encoded as reference information.
[0116] Lift transform coding refers to introducing a weight update strategy for neighboring points on the basis of LoD neighboring layer prediction, and finally obtaining the predicted attribute information of each point and the corresponding attribute residual information.
[0117] Hierarchical region adaptive transform coding refers to the process of transforming attribute information into the transform domain, which is called the transform coefficient.
[0118] 4. Attribute Quantization: The fineness of quantization is usually determined by the quantization parameters. The transformation coefficients or attribute residuals obtained from attribute information processing are quantized, and the quantized results are entropy-coded. For example, in predictive transform coding and boost transform coding, entropy coding is performed on the quantized attribute residuals; in RAHT, entropy coding is performed on the quantized transform coefficients.
[0119] 5. Entropy Coding: The quantized attribute residual information and / or transform coefficients are generally compressed using run-length coding and arithmetic coding. The corresponding coding mode, quantization parameters, and other information are also encoded using an entropy encoder.
[0120] The encoder 200 encodes the geometric coordinate information of each point in the point cloud to obtain a geometric bitstream, and encodes the attribute information of each point in the point cloud to obtain an attribute bitstream. The encoder 200 can transmit the encoded geometric bitstream and attribute bitstream together to the decoder 300.
[0121] Figure 4 The following is a flowchart illustrating the decoding process performed by the decoder in the AVS-PCC-based decoding framework: Figure 5 This diagram illustrates the decoding flowchart performed by the decoder within the MPEG G-PCC-based decoding framework. The decoder can be... Figure 1The decoder 300 is shown. After receiving the compressed bitstream (i.e., attribute bitstream and geometric bitstream) transmitted by the encoder 200, the decoder 300 decodes the geometric bitstream to reconstruct the geometric coordinate information of each point in the point cloud, and decodes the attribute bitstream to reconstruct the attribute information of each point in the point cloud.
[0122] The decoding process performed by decoder 300 is as follows:
[0123] 1. Entropy Decoding: Perform entropy decoding on the geometric bitstream and attribute bitstream respectively to obtain geometric syntax elements and attribute syntax elements.
[0124] 2. Geometric Decoding: For the AVS-PCC coding framework, geometric decoding includes two modes: octree-based geometric decoding and prediction tree-based geometric decoding. For the G-PCC coding framework, geometric decoding includes three modes: octree-based geometric decoding, trisoup-based geometric decoding, and prediction tree-based prediction decoding.
[0125] Octree-based geometric decoding: reconstructing the octree based on the geometric syntax elements obtained from parsing the geometric bitstream.
[0126] Geometric Decoding Based on Prediction Trees: Reconstructing the prediction tree based on the geometric syntax elements obtained from parsing the geometric bitstream.
[0127] Geometric Decoding Based on Triangle Representation: Reconstructing the triangular model based on the geometric syntax elements obtained from parsing the geometric bitstream.
[0128] 3. Geometric Reconstruction: Perform reconstruction to obtain the geometric coordinate information of the points in the point cloud.
[0129] 4. Inverse coordinate transformation: Perform an inverse transformation on the reconstructed geometric coordinate information to convert the reconstructed coordinates (positions) of points in the point cloud from the transformation domain back to the initial domain.
[0130] 5. Dequantization: Dequantizes attribute syntax elements.
[0131] 6. Attribute Information Processing: In AVS-PCC, attribute information processing determines the color information of points in the point cloud by predicting or predicting the transformation coefficients of the inverse-quantized prediction residuals, or by transforming the transformation coefficients of the inverse-quantized points.
[0132] In MPEG G-PCC, attribute information processing determines the color information of points in the point cloud by using RAHT to invert the attribute information, or by using LOD and inverse boosting to determine the color information of points in the point cloud.
[0133] 7. Inverse Color Transformation: Transforms color information from the YCbCr color space to the RGB color space. In some examples, the inverse color transformation operation may not be necessary.
[0134] The technologies related to the embodiments of this application will be described below.
[0135] 1. Octree-based geometric encoding and decoding
[0136] Encoding: First, perform coordinate transformation on the geometric information so that the entire point cloud is contained within a coordinate system consisting of two extreme points (0, 0, 0) and (2... d ,2 d ,2 d Within the bounding box determined by the algorithm, voxelization is performed, including quantization, rounding, and removal of duplicate points (depending on the parameters). Next, following a breadth-first traversal, octree partitioning is continuously performed on the non-empty sub-cubes (containing points from the point cloud) within the bounding box. At the same octree depth, a node is divided into 8 child nodes, continuing until the resulting leaf nodes form 1x1x1 unit cubes. The 8-bit binary code generated by determining whether a point in the sub-cube is occupied (1 for occupied, 0 for unoccupied) is called the occupancy code. The occupancy code of each node is encoded to generate a binary code stream.
[0137] Decoding: Following the breadth-first traversal order, the placeholder code of each node is obtained by continuously parsing, and the nodes are divided in turn until a 1x1x1 unit cube is obtained. The division stops when the division is stopped. The number of points contained in each leaf node is obtained by parsing, and finally the geometric reconstruction point cloud information is recovered.
[0138] 2. Trisoup-based geometric encoding and decoding
[0139] Encoding: First, an octree is partitioned. Unlike geometric information encoding based on an octree structure, this method does not need to partition the point cloud level by level down to the bottom-level leaf nodes with side length 1·1·1, but instead partitions leaf nodes with a specified side length. Then, the surface information composed of voxels within the node is represented by a series of triangle meshes. In GPCC, the parameter trisoupnode size represents the size of the block containing the triangle. When the trisoup node size is greater than 0, a geometric patch represents the set of voxels within the node. The maximum of twelve intersection points generated by the geometric patch and the twelve edges of the block are called vertices. The vertex coordinates of each block are encoded sequentially to generate a binary bitstream.
[0140] Decoding: To decode the geometric coordinates of the point cloud from the node triangular facets, it is necessary to check whether each voxel in the node cube intersects with the triangular facet. This technique is called triangulation. It uses 6 unit vectors (0,0,1), (0,0,1), (0,0,1), (0,0,1), (0,0,1), (0,0,1), (0,0,1) to perform an intersection check. If each unit vector intersects with the triangular facet, the intersection point is calculated and the decoded cube is output. The number of points generated in the decoder is determined by the grid distance d.
[0141] 3. Geometric Encoding and Decoding Based on Prediction Trees
[0142] Encoding: First, the input point cloud is sorted. Currently, sorting methods include unordered, Morton order, azimuth order, and radial distance order. At the encoding end, a prediction tree structure is built using two different methods: KD-Tree (high latency, slow mode) and using LiDAR calibration information to assign each point to a different laser and build a prediction structure according to the different lasers (low latency, fast mode). Next, based on the prediction tree structure, each node in the prediction tree is traversed. By selecting different prediction modes, the geometric position information of the node is predicted to obtain the prediction residual, and the geometric prediction residual is quantized using quantization parameters. Finally, through continuous iteration, the prediction residual of the prediction tree node position information, the prediction tree structure, and the quantization parameters are encoded to generate a binary code stream.
[0143] Decoding: The decoding end continuously parses the bitstream and reconstructs the prediction tree structure. Then, it obtains the geometric position prediction residual information and quantization parameters of each prediction node through parsing. It also performs inverse quantization on the prediction residual to recover the reconstructed geometric position information of each node, and finally completes the geometric reconstruction of the decoding end.
[0144] 4. Planar encoding mode
[0145] A common method for compressing geometric information is octree coding, which introduces a planar coding pattern during octree partitioning, a process that occurs within the octree geometric coding process. This method efficiently encodes nodes that meet the planar qualification criteria.
[0146] Planar occupancy coding (i.e., planar coding mode) decomposes the node occupancy bitmap into axis-aligned planes, as shown in Table 1. Each coded axis has two vertical planes that child nodes can occupy. For each coded axis that meets the plane condition, planar occupancy coding specifies whether one of the two planes is unoccupied. However, only certain axes are eligible for planar occupancy coding; the eligibility of the axis at index k is specified by the expression planarEligible[k]. In related algorithms, eligibility is determined as shown in Table 2. When angle mode is enabled, it depends on whether the node's angle context is valid; when angle mode is not enabled, it depends on the density condition at the current node's tree depth (the method proposed in this paper does not discuss the case where angle mode is not enabled, i.e., it follows the original method).
[0147] Table 1. Correspondence between plane, axis, and axis index k
[0148] k axis flat 0 S TV 1 T SV 2 V ST
[0149] Table 2. Axial Plane Coding Eligibility Comparison Table
[0150]
[0151] The method of determining eligibility using azimuth context applies to either the S-axis or the T-axis. The index of the contextualized axis is specified by the expression AzimuthAxis. It is determined by the position of the node relative to the angular origin, while the other axis will lose its planar encoding eligibility.
[0152] For ease of explanation in the following text, it is hereby noted that: non-azimuth axes are determined based on azimuth axes:
[0153] If the azimuth axis is S (AzimuthAxisIsT=1), then the non-azimuth axis is T.
[0154] If the azimuth axis is T (AzimuthAxisIsS=1), then the non-azimuth axis is S.
[0155] The specific encoding process begins by introducing a flag, the planar flag (is_planar_flag), which indicates whether the occupied child nodes of the current node belong to the same planar plane. If is_planar_flag is 1, then an additional flag, the plane position flag (plane_position), is needed. This flag uses one bit to indicate whether the plane is located in the low plane or the high plane; for example, 0 indicates the low plane, and 1 indicates the high plane.
[0156] When the planar coding mode was first introduced, it was only for a plane in one direction (i.e., the V direction). Later, considering that different point clouds have different distributions, it was expanded to a planar coding mode in three directions (i.e., the S, T, or V directions).
[0157] Specifically, firstly, a flag is introduced:
[0158] 1. Multi-planar flag: indicates whether the current node satisfies the condition that it forms a plane in all three directions.
[0159] 2. Planar flag (is_planar_flag[axisIdx] (axisIdx is 0, 1, 2, corresponding to the S, T, V directions respectively)): indicates whether the child nodes of the current node in a certain direction form a plane.
[0160] 3. Planar position flag (plane_position[axisIdx]): When is_planar_flag[axisIdx] is 1 in a certain direction, an extra bit is needed to encode plane_position to indicate whether the plane is in the lower plane (0) or the higher plane (1).
[0161] The encoding process is as follows:
[0162] If the current node meets the planar qualification criteria in all three directions, then the multi_planar_flag is encoded.
[0163] If multi_planar_flag is 1, it means that the child nodes of the current node form a plane in all three directions. Multi_planar_flag is directly encoded, eliminating the need for additional encoding of is_planar_flag[0], is_planar_flag[1], and is_planar_flag[2], thus saving encoding bits. The plane position information plane_position[axisIdx] is directly calculated and encoded.
[0164] If multi_planar_flag is 0, is_planar_flag[0] and is_planar_flag[1] need to be encoded sequentially, and is_planar_flag[2] is inferred based on the values of these two flag bits. If is_planar_flag[0] and is_planar_flag[1] are both 1, is_planar_flag[2] is directly inferred to be 0, no encoding is required, reducing encoding overhead; otherwise, is_planar_flag[2] is encoded normally. And when is_planar_flag[axisIdx] is 1, the planar position information plane_position[axisIdx] is calculated and encoded.
[0165] The decoding process is as follows:
[0166] If the current node meets the planar qualification criteria in all three directions, then the multi_planar_flag is decoded.
[0167] If multi_planar_flag is 1, it means that the current node is a three-plane (that is, the positions of the child nodes occupied by the current node form a plane in three directions). is_planar_flag[0], is_planar_flag[1] and is_planar_flag[2] can be directly inferred to be 1, and the plane position information plane_position[0], plane_position[1] and plane_position[2] can be decoded from the bitstream.
[0168] If multi_planar_flag is 0, is_planar_flag[0] and is_planar_flag[1] are decoded normally from the bitstream. If is_planar_flag[0] and is_planar_flag[1] are both 1, is_planar_flag[2] is directly inferred to be 0, no decoding is required, and the plane position information plane_position[0] and plane_position[1] are decoded from the bitstream; if is_planar_flag[0] and is_planar_flag[1] are not both 1, is_planar_flag[2] is decoded normally from the bitstream, and the plane position information plane_position[axisIdx] is decoded accordingly from the bitstream based on whether is_planar_flag[axisIdx] is 1.
[0169] When using an octree to partition and encode the point cloud of a rotating LiDAR, as the breadth-first traversal of the octree progresses, the size of the resulting octree nodes decreases continuously, and the spatial range contained in each node also decreases. At this point, the surface of irregularly shaped objects will be continuously subdivided by the nodes, and the shape of the object surface within each node will become closer and closer to a plane. It can be considered that the local point cloud contained in the smaller octree nodes will have stronger planar features, and therefore is more suitable for planar pattern coding algorithms.
[0170] A node is called a three-plane node when all three of its components are planar components. However, if a node is a three-plane node, but the planar pattern qualification inference algorithm does not assign planar pattern qualification to all three components (for example, omitting one component), then two of the node's components will be encoded using planar patterns, while the missing component will be encoded using placeholder codes. In this case, the representation overhead of the placeholder information for this octree node will increase, which will reduce the encoding efficiency of the octree.
[0171] When a rotating lidar acquires point clouds, the distribution characteristics of the point cloud differ in three directions—axial (elevation component V), tangential (azimuth component S or T), and radial (non-azimuth component T or S)—due to the rotation of the laser beam and multi-line scanning. However, existing planar qualification inference algorithms fail to fully consider these differences. They only roughly estimate whether a node is likely to be traversed by multi-line lasers with different elevation angles and determine whether the axial and tangential components of the node are qualified for planar mode based on this. For the radial direction, qualification is not assigned at all, thus significantly reducing the accuracy of planar mode qualification judgment for three-plane nodes.
[0172] Furthermore, in the relevant algorithm, the planar mode is encoded in a fixed order of S, T, and V components. For a three-planar-qualified node, a 1-bit planar mode joint flag, `multi_planar_flag`, needs to be encoded to indicate whether all three components are planar components. If `multi_planar_flag` is true, the planar positions (`plane_position`) of the three components are encoded sequentially. If none of the three components are planar components, then `multi_planar_flag` is false. In this case, the encoded information is insufficient to indicate whether each of the three components is a planar component. With only this information, the decoder cannot determine whether each component should be decoded using the planar mode or the placeholder code algorithm. Therefore, a 1-bit planar mode flag, `is_planar_flag`, needs to be encoded for each component that is qualified for the planar mode. In addition, if the planar mode flags of the S and Q components are both true, then the planar mode flag of the V component must be false. The decoder can rely on this logic to directly determine the value of the planar mode flag of the V component, so the encoder does not need to encode this bit. Therefore, it can be seen that encoding the three components in a plane mode according to the descending order of the plane component probabilities can reduce the representation overhead of the plane mode joint identifier and the plane mode identifier.
[0173] Based on the above analysis, the relevant algorithms have the following shortcomings:
[0174] 1. The relevant algorithms do not fully consider the planar features of the local point cloud within the node, and the radial component is not qualified for planar mode by default. This results in the inability to fully utilize the planar features of the local point cloud, causing some nodes that conform to the planar mode to be unrecognized, thus limiting the compression efficiency.
[0175] 2. Existing methods use a fixed encoding order, which fails to optimize the encoding order based on the actual distribution of point cloud data, thus increasing bit overhead.
[0176] The encoding and decoding methods, apparatuses, and devices provided in this application will be described in detail below with reference to the accompanying drawings and through some embodiments and application scenarios.
[0177] like Figure 6 As shown, this application provides a decoding method, including:
[0178] Step 601: The decoding device performs octree reconstruction on the point cloud to be decoded to obtain the target node;
[0179] Step 602: Determine the planar mode decoding qualification of the first component of the target node based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device. The first component is the S-axis component or T-axis component of the target node.
[0180] Step 603: Based on the determination result of the planar mode decoding qualification of the first component, decode the S-axis component, T-axis component and V-axis component of the target node.
[0181] It should be noted that, in this embodiment, the point cloud to be decoded is first reconstructed using an octree. Then, based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device, the planar mode decoding qualification of the S-axis component or T-axis component of the target node is determined. Subsequently, based on the determination result of the first component's planar mode decoding qualification, the S-axis component, T-axis component, and V-axis component of the target node are decoded. This embodiment determines the planar mode decoding qualification of the S-axis component or T-axis component based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device. This ensures the accuracy of the determination of the planar mode decoding qualification of the S-axis component or T-axis component, effectively improving the applicability of the planar mode encoding and decoding algorithm, and thus improving the compression efficiency of point cloud geometric information encoding and decoding.
[0182] Optionally, the target node mentioned in the embodiments of this application refers to at least one node obtained after the octree reconstruction.
[0183] It should be noted that the first component mentioned in the embodiments of this application can also be understood as a radial component or a non-azimuth component, and the V-axis component can be understood as an axial component or a pitch component.
[0184] Optionally, in one implementation, the specific implementation of determining the planar mode decoding qualification of the first component of the target node based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device includes:
[0185] If at least one of the planar mode and angle mode of the target node is enabled, the planar mode decoding qualification of the first component of the target node is determined according to the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device.
[0186] Optionally, the state of the target node's planar mode and angle mode can be pre-set in the configuration information of the point cloud data by means of a geometric planar mode enable flag and a geometric angle mode enable flag. Planar mode enable can be understood as the geometric planar mode enable flag being set to a first value (in this embodiment, the first value is 1 or true, which can be understood as the geometric planar mode enable flag being true). Angle mode enable can be understood as the geometric angle mode enable flag being set to a first value (for example, the first value is 1 or true, which can be understood as the geometric angle mode enable flag being true). In typical implementations, the planar mode decoding qualification of the target node's first component is determined only when both the geometric plane mode enable flag and the geometric angle mode enable flag are set to the first value. This determination is based on the slope range occupied by the target node's V-axis component and the slope resolution of the point cloud acquisition device. If the values of the geometric plane mode enable flag and the geometric angle mode enable flag are not all the first value or are all the second value (in this embodiment, the second value is 0 or false, i.e., false), the step of determining the planar mode decoding qualification of the target node's first component based on the slope range occupied by the target node's V-axis component and the slope resolution of the point cloud acquisition device is not executed. This ensures the effectiveness of the step of determining the planar mode decoding qualification of the target node's first component.
[0187] Optionally, the values of the geometric plane mode enable flag and the geometric angle mode enable flag are both the first value, which can be understood as a prerequisite for enabling the plane mode decoding qualification judgment algorithm of the first component (i.e., executing the algorithm to determine the plane mode decoding qualification of the first component of the target node based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device).
[0188] Optionally, the decoding end can determine the slope resolution of the point cloud acquisition device based on the slope of two adjacent laser lines.
[0189] For example, the slope resolution of a point cloud acquisition device can be determined according to Formula 1:
[0190] Formula 1, BeamDeltaGrad=|BeamElev i+1 -BeamElev i |;
[0191] Where BeamDeltaGrad is the slope resolution of the point cloud acquisition device, and BeamElev... i+1 BeamElev represents the slope of the (i+1)th laser beam. i Let be the slope of the i-th laser.
[0192] Optionally, the decoding end can determine the slope range occupied by the V-axis component of the target node based on the distance from the target point projected onto the target plane to the origin of the coordinate system of the point cloud acquisition device and the side length of the octree node.
[0193] Optionally, the target plane can be understood as the ST plane in the coordinate system of the point cloud acquisition device. The coordinate system of the point cloud acquisition device is a Cartesian coordinate system, which is formed by translating the origin from the corner to a certain position in the XYZ coordinate system, and mapping XYZ to STV according to the configuration information.
[0194] For example, the slope range occupied by the V-axis component of the target node can be determined according to Formula 2:
[0195]
[0196] Where elvIntvlMidGrad is the slope range occupied by the V-axis component, L is the side length of the octree node, and R... node The center point of the target node ((X) nodeCenter ,Y nodeCenter Z nodeCenter The target point (S) projected onto the target plane nodeCenter ,T nodeCenter The distance from the origin of the coordinate system of the point cloud acquisition device, where,
[0197]
[0198] Optionally, in one implementation, the specific implementation of determining the planar mode decoding qualification of the first component of the target node based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device includes one of A11 and A12:
[0199] A11. If the slope range occupied by the V-axis component of the target node is smaller than the slope resolution of the point cloud acquisition device, determine that the first component can be decoded in planar mode.
[0200] It should be noted that if the slope range occupied by the V-axis component of the target node is smaller than the slope resolution of the point cloud acquisition device, it indicates that the target node has a high probability of being a three-plane node (i.e., the positions of the child nodes occupied by the target node form a plane in three directions).
[0201] A12. If the slope range occupied by the V-axis component of the target node is greater than or equal to the slope resolution of the point cloud acquisition device, it is determined that the first component cannot be decoded in planar mode.
[0202] It should be noted that if the first component can be decoded in planar mode, the planar eligibility flag corresponding to the first component of the target node is set to true; otherwise, it is set to false.
[0203] Optionally, in one implementation, the specific implementation of decoding the S-axis, T-axis, and V-axis components of the target node based on the determination result of the planar mode decoding qualification of the first component includes S11-S13:
[0204] S11. Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, determine the first indication information. The first indication information is used to indicate the processing order of the planar mode of the S-axis component, T-axis component and V-axis component.
[0205] Optionally, the point cloud acquisition device mentioned in the embodiments of this application can be a lidar, such as a rotating lidar.
[0206] Optionally, the parameter settings of the point cloud acquisition device may include at least one of the following: the number of lines of the rotating lidar; the slope of each line of the lidar; and the number of sampling points per line of the lidar in one rotation.
[0207] Optionally, the parameter settings of the acquisition device may include at least one of the following: the angular spacing between adjacent laser beams, the radar rotation speed, and the data sampling frequency.
[0208] Rotating lidar acquires 3D point cloud data of the surrounding environment through a high-speed rotating laser scanning system. The radar's rotation axis is typically vertical, rotating at a constant angular velocity to scan the entire horizontal plane. Simultaneously, multiple laser emitters are arranged at fixed elevation angles, forming multiple scanning layers to cover different heights in the vertical direction. In point cloud acquisition, the elevation angle (slope = tan elevation angle) resolution is determined by the angular interval between adjacent laser beams, affecting the point cloud density in the vertical direction. Smaller elevation angle intervals provide finer vertical resolution, while the azimuth resolution is determined by the radar's rotation speed and data sampling frequency, determining the level of detail in the point cloud in the horizontal direction. Higher azimuth resolution allows for denser point cloud data acquisition, thereby improving the accuracy of target detection and environmental modeling.
[0209] Based on the above-described principle of rotating lidar point cloud acquisition, the sparsity of the rotating lidar point cloud in different directions depends on the elevation resolution θresolution of the multi-line laser emitter of the rotating lidar and the sampling frequency sampleFrequency of the rotating lidar during rotation. The sampling frequency during rotation can be converted into the azimuth resolution φresolution of the lidar, calculated as follows:
[0210]
[0211] The sparsity of point clouds acquired by rotating lidar in the axial and radial directions is affected by the elevation resolution, while the sparsity in the tangential direction is affected by the sampling frequency during rotation. When the elevation resolution of a rotating lidar is lower than its azimuth resolution (i.e., when the condition θresolution < φresolution is met), under the same real-world scenario, the planar feature intensity of the point cloud acquired by the rotating lidar in the axial direction is higher than that in the tangential direction.
[0212] When the planar feature intensity of a point cloud is higher in a certain component, the probability that component is a planar component is also higher. Therefore, the probability that the component of a point cloud satisfying this condition is a planar component in the axial direction is higher than the probability that it is a planar component in the tangential direction. The axial direction of the point cloud corresponds to the V component, the tangential direction corresponds to the azimuth component S or T, and the radial direction corresponds to the non-azimuth component T or S.
[0213] In other words, as described above, the axial and tangential resolutions can be calculated by setting the parameters of the acquisition device. The magnitude of the calculated resolution determines the magnitude of the axial and tangential planar feature intensity (a larger resolution results in a weaker planar feature intensity). The magnitude of the planar feature intensity determines the probability of the axial and tangential planar components (a larger planar feature intensity results in a higher probability of the planar components). Then, the processing order is determined according to the principle of processing higher probabilities.
[0214] Since the sparsity of the point cloud in the radial direction cannot be estimated using the equipment parameters of the rotating lidar, and considering the diversity of rotating lidar sequences in real-world scenarios, and in order to improve the generalization of the algorithm, this algorithm assumes that the density of the point cloud in the radial direction is the highest among the three directions, that is, the probability of the planar component of the point cloud in the radial direction is the lowest.
[0215] Based on the above conclusions, the representation cost of the three-plane qualification nodes is expected to be lowest when the planar component probabilities of the three components are encoded in descending order. Given that the axial direction of the rotating lidar point cloud has the highest single-plane probability, followed by the tangential direction, and the radial direction is assumed to be the lowest, an adaptive planar pattern encoding algorithm based on lidar resolution differences is used to encode the three-plane qualification nodes of the rotating lidar point cloud in the order of axial direction, tangential direction, and radial direction.
[0216] For example, based on the above method, and with the tangential (azimuth component) direction determined as the S-axis, the descending order of the processing sequence of the determined planar mode is: V-axis component, S-axis component, T-axis component.
[0217] S12. Based on the determination result and the first indication information, determine the processing order of the planar mode of the target node;
[0218] Optionally, the specific implementation of determining the processing order of the planar mode of the target node based on the determination result and the first indication information includes:
[0219] If the result indicates that the first component of the target node can be decoded in planar mode, the processing order of the planar mode of the S-axis component, T-axis component and V-axis component of the target node is determined according to the first indication information.
[0220] It should be noted that if the first component can be decoded in planar mode, it means that the other two components can also be decoded in planar mode. Therefore, the order indicated by the first indication information can be used as the processing order of the planar mode of the S-axis component, T-axis component, and V-axis component.
[0221] S13. According to the processing order, decode the S-axis component, T-axis component and V-axis component of the target node;
[0222] Optionally, the specific implementation of decoding the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following:
[0223] C11. Decode the multi-plane flag bit. If the multi-plane flag bit takes the first value, decode the planar position information of the S-axis component, T-axis component and V-axis component of the target node according to the processing order.
[0224] It should be noted that if the multi-plane flag is the first value, it means that the target node is a three-plane node. The plane flags (is_planar_flag) corresponding to the three components can be directly inferred to be 1 (i.e., is_planar_flag[0], is_planar_flag[1] and is_planar_flag[2] can be directly inferred to be 1), without the need for decoding. The decoding device can directly decode the plane position information from the bitstream according to the processing order, i.e., decode plane_position[0], plane_position[1] and plane_position[2]. The order of decoding plane_position[0], plane_position[1] and plane_position[2] is determined based on the previously determined processing order.
[0225] C12. Decode the multi-plane flag bit. If the multi-plane flag bit takes the second value, decode the plane flag bit of the S-axis component, T-axis component and V-axis component of the target node according to the processing order.
[0226] Optionally, the decoding of the planar flag bits of the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of C121 and C122:
[0227] C121. If, according to the processing order, the values of the plane flag bits corresponding to the first two components are not all the first value, then the plane flag bits corresponding to the third component are decoded.
[0228] It should be noted that if the plane flag bits corresponding to the first two components are all set to the first value according to the processing order, then the plane flag bits corresponding to the third component are not decoded, thereby reducing the decoding overhead.
[0229] C122. After decoding the plane flag of the target component, if the plane flag corresponding to the target component takes the first value, then the plane position information corresponding to the target component is decoded, wherein the target component is any one of the S-axis component, T-axis component and V-axis component.
[0230] It should be noted that if the multi-planar flag is the second value, it means that only two directions constitute a plane. If the corresponding is_planar_flag is decoded from the bitstream sequentially according to the processing order, if the first two components' is_planar_flags are both 1, it can be inferred that the third component's is_planar_flag is 0, and no decoding is needed; otherwise, the third component's is_planar_flag is decoded normally. Furthermore, when is_planar_flag[axisIdx] is 1, the plane position information plane_position[axisIdx] needs to be decoded from the bitstream. It should be noted that with adaptive order adjustments, this inference of 0 will increase, thus saving decoding overhead.
[0231] The corresponding implementation method on the encoding device side is as follows:
[0232] If the multi-plane flag bit takes the first value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are respectively calculated according to the processing order, and the plane position information is encoded.
[0233] If the multi-plane flag bit takes the second value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are encoded with plane flag bits according to the processing order.
[0234] Furthermore, the encoding of the multi-plane flag bits, which involves encoding the plane flag bits of the S-axis, T-axis, and V-axis components of the target node according to the processing order, includes at least one of the following:
[0235] If, according to the processing order, the values of the plane flag bits corresponding to the first two components of the encoding are not all the first value, then the plane flag bits corresponding to the third component are encoded.
[0236] After encoding the planar flag bit of the target component, if the value of the planar flag bit corresponding to the target component is the first value, then the planar position information of the target component is calculated and the planar position information is encoded. The target component is any one of the S-axis component, T-axis component and V-axis component.
[0237] It should be noted that, from the encoding perspective, if `multi_planar_flag` is 1, it means that the target node's child nodes form a plane in all three directions. In this case, `multi_planar_flag` can be directly encoded without additional encoding of `is_planar_flag[axisIdx]`, saving encoding bits. Specifically, the planar position information `plane_position[axisIdx]` of each of the target node's S-axis, T-axis, and V-axis components is calculated according to the processing order, and the corresponding `plane_position` is encoded. If `multi_planar_flag` is 0, it means that only two directions form a plane. `multi_planar_flag` is encoded first. Then, according to the processing order, the corresponding `is_planar_flag` is encoded sequentially. If the first two components' `is_planar_flag`s are both 1, it can be inferred that the third component's `is_planar_flag` is 0, and there is no need to encode the third component's `is_planar_flag`; otherwise, the third component's `is_planar_flag` is encoded normally. Furthermore, when is_planar_flag[axisIdx] is 1, the planar position information plane_position[axisIdx] is calculated and encoded. With adaptive order adjustments, the number of cases where this inference is 0 increases, thus saving bitstream.
[0238] It should be noted that, in this embodiment of the application, to address the shortcomings of related technologies, the following design is made for encoding and decoding rotating lidar sequences:
[0239] 1. Planar pattern qualification inference algorithm based on local planar features: By analyzing the relationship between the slope range occupied by the node and the slope resolution of the lidar, a new planar pattern qualification (for the decoding device, this planar pattern qualification is the planar pattern decoding qualification; for the encoding device, this planar pattern qualification is the planar pattern encoding qualification) inference algorithm is designed. This algorithm can more accurately select three-plane nodes, effectively improve the applicability of the planar pattern encoding algorithm, and thus improve the compression efficiency of point cloud geometric information.
[0240] 2. Adaptive Planar Pattern Processing Algorithm Based on Laser Resolution Differences: By analyzing the planar feature differences of octree nodes in the rotating lidar point cloud along the S-axis, T-axis, and V-axis, the influence of these planar feature differences on the coding efficiency of the joint planar pattern identifier and the planar pattern identifier is further analyzed. Based on this analysis, a new planar pattern processing order is designed, which can effectively reduce the average representation overhead of the joint planar pattern identifier and the planar pattern identifier of the octree nodes, thereby improving the compression efficiency of the planar pattern coding algorithm.
[0241] Specifically, let's first introduce a parameter:
[0242] The first instruction (planarOrder) is used to control the processing order of the planar pattern, determining which component to process first based on different point cloud distributions.
[0243] For example, the decoding process in this application embodiment is as follows: Figure 7 As shown, the main implementation process includes:
[0244] Following the G-PCC decoding process, the point cloud data is reconstructed into an octree.
[0245] Whether the current node is suitable for non-azimuth plane mode decoding is determined by the following two flag bits:
[0246] Geometric planar mode enable flag (geom_planar_mode_enabled_flag), geometric angle mode enable flag (geom_angular_mode_enabled_flag);
[0247] If both flags are set to the first value, then the step of determining the planar mode decoding qualification of the first component of the current node is required based on the slope range occupied by the V-axis component of the current node and the slope resolution of the point cloud acquisition device.
[0248] Contextual calculation:
[0249] Check if geometry angle mode is enabled (geom_angular_mode_enabled_flag). If enabled, calculate the context angles (contextAngle, contextAnglePhiX, contextAnglePhiY).
[0250] Determine whether to enable planar mode:
[0251] If a valid value can be calculated from the context angle (contextAngle, contextAnglePhiX, contextAnglePhiY), the corresponding planar mode's qualification flag (planarEligible) is set to true; otherwise, it is set to false. Generally, the context angles returned by the non-azimuth components of the radial direction (S or T, i.e., the first component) are invalid and will therefore be directly set to false.
[0252] Determine if a non-azimuth component planar mode is eligible:
[0253] Adjusting the planarEligible value of the first component based on the relationship between the slope range occupied by the V component and the slope resolution of the point cloud acquisition device can further improve the accuracy of filtering three-plane nodes. The specific judgment method is described above and will not be repeated here.
[0254] Decoding process:
[0255] If the current node is determined to be a three-plane node, the three components are decoded in plane mode according to the processing order indicated by the first instruction information, which can reduce the expected representation overhead.
[0256] Reconstructing point cloud geometric information
[0257] Based on the decoded planar pattern information, the geometric information of the nodes is reconstructed, all nodes of the octree are recursively processed, and finally the complete point cloud data is restored.
[0258] It should be noted that the encoding implementation is similar to the decoding implementation, for example, as... Figure 8 As shown, the main implementation process of the encoding flow includes:
[0259] Following the G-PCC encoding process, the point cloud data is recursively partitioned into an octree.
[0260] Whether the current node is suitable for non-azimuth plane mode coding is determined by the following two flag bits:
[0261] Geometric planar mode enable flag (geom_planar_mode_enabled_flag), geometric angle mode enable flag (geom_angular_mode_enabled_flag);
[0262] If both flags are set to the first value, then the step of determining the planar pattern encoding qualification of the first component of the current node is required based on the slope range occupied by the V-axis component of the current node and the slope resolution of the point cloud acquisition device.
[0263] Contextual calculation:
[0264] Check if geometry angle mode is enabled (geom_angular_mode_enabled_flag). If enabled, calculate the context angles (contextAngle, contextAnglePhiX, contextAnglePhiY).
[0265] Determine whether to enable planar mode:
[0266] If a valid value can be calculated from the context angle (contextAngle, contextAnglePhiX, contextAnglePhiY), the corresponding planar mode's qualification flag (planarEligible) is set to true; otherwise, it is set to false. Generally, the context angles returned by the non-azimuth components of the radial direction (S or T, i.e., the first component) are invalid and will therefore be directly set to false.
[0267] Determine if a non-azimuth component planar mode is eligible:
[0268] Adjusting the planarEligible value of the first component based on the relationship between the slope range occupied by the V component and the slope resolution of the point cloud acquisition device can further improve the accuracy of filtering three-plane nodes. The specific judgment method is described above and will not be repeated here.
[0269] Encoding processing:
[0270] If the current node is determined to be a three-plane node, the three components are encoded in a plane mode according to the processing order indicated by the first instruction information, which can reduce the expected representation overhead.
[0271] In summary, compared to traditional azimuth plane pattern encoding and decoding methods, this application proposes a plane pattern encoding and decoding algorithm for rotating lidar point clouds by improving plane pattern qualification determination and encoding order optimization. Experimental tests show that this proposed plane pattern encoding and decoding algorithm can effectively improve the compression efficiency of plane pattern encoding algorithms for rotating lidar point clouds and significantly reduce the encoding and decoding time complexity.
[0272] like Figure 9 As shown, the encoding method of this application embodiment includes:
[0273] Step 901: The encoding device performs octree partitioning on the point cloud to be encoded to obtain the target node;
[0274] Step 902: Based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device, determine the planar pattern encoding qualification of the first component of the target node, wherein the first component is the S-axis component or the T-axis component of the target node.
[0275] Step 903: Based on the determination result of the planar pattern encoding qualification of the first component, the S-axis component, T-axis component and V-axis component of the target node are encoded.
[0276] Optionally, determining the planar pattern encoding qualification of the first component of the target node based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device includes:
[0277] If at least one of the planar mode and angle mode of the target node is enabled, the planar mode encoding qualification of the first component of the target node is determined according to the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device.
[0278] Optionally, determining the planar pattern encoding qualification of the first component of the target node based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device includes one of the following:
[0279] If the slope range occupied by the V-axis component of the target node is smaller than the slope resolution of the point cloud acquisition device, it is determined that the first component can be encoded in a planar pattern.
[0280] If the slope range occupied by the V-axis component of the target node is greater than or equal to the slope resolution of the point cloud acquisition device, it is determined that the first component cannot be encoded in planar mode.
[0281] Optionally, the encoding processing of the S-axis, T-axis, and V-axis components of the target node based on the determination result of the planar pattern encoding qualification of the first component includes:
[0282] Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, the first indication information is determined. The first indication information is used to indicate the processing order of the planar mode of the S-axis component, T-axis component and V-axis component.
[0283] Based on the determination result and the first indication information, the processing order of the planar mode of the target node is determined;
[0284] According to the processing order, the S-axis component, T-axis component, and V-axis component of the target node are encoded.
[0285] Optionally, determining the processing order of the planar pattern of the target node based on the determination result and the first indication information includes:
[0286] If the result indicates that the first component of the target node can be encoded in a planar mode, the processing order of the planar modes of the S-axis component, T-axis component, and V-axis component of the target node is determined according to the first indication information.
[0287] Optionally, the encoding process for the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following:
[0288] If the multi-plane flag bit takes the first value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are respectively calculated according to the processing order, and the plane position information is encoded.
[0289] If the multi-plane flag bit takes the second value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are encoded with plane flag bits according to the processing order.
[0290] Optionally, the encoding of multi-plane flags, which involves encoding the S-axis, T-axis, and V-axis components of the target node according to the processing order, includes at least one of the following:
[0291] If, according to the processing order, the values of the plane flag bits corresponding to the first two components of the encoding are not all the first value, then the plane flag bits corresponding to the third component are encoded.
[0292] After encoding the planar flag bit of the target component, if the value of the planar flag bit corresponding to the target component is the first value, then the planar position information of the target component is calculated and the planar position information is encoded. The target component is any one of the S-axis component, T-axis component and V-axis component.
[0293] It should be noted that all descriptions of the encoding device side in the above embodiments are applicable to the embodiments of the encoding method applied to the encoding device side, and can achieve the same technical effect, so they will not be repeated here.
[0294] like Figure 10 As shown, the encoding method of this application embodiment includes:
[0295] Step 1001: The decoding device performs octree reconstruction on the point cloud to be decoded to obtain the target node;
[0296] Step 1002: Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, determine the processing order of the planar mode of the S-axis component, T-axis component and V-axis component of the target node.
[0297] Step 1003: According to the processing order, decode the S-axis component, T-axis component and V-axis component of the target node.
[0298] Optionally, the decoding process for the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following:
[0299] The multi-plane flag is decoded. If the multi-plane flag value is the first value, the S-axis component, T-axis component, and V-axis component of the target node are decoded according to the processing order.
[0300] The multi-plane flag is decoded. If the multi-plane flag value is the second value, the S-axis component, T-axis component and V-axis component of the target node are decoded according to the processing order.
[0301] Optionally, the decoding of the planar flag bits of the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following:
[0302] If, according to the processing order, the values of the plane flag bits corresponding to the first two components are not all the first value, then the plane flag bits corresponding to the third component are decoded.
[0303] After decoding the plane flag bit of the target component, if the plane flag bit corresponding to the target component takes the first value, then the plane position information corresponding to the target component is decoded. The target component is any one of the S-axis component, T-axis component, and V-axis component.
[0304] It should be noted that all descriptions of the processing order and decoding process in the above embodiments are applicable to this embodiment, and will not be repeated here to avoid repetition.
[0305] It should be noted that in this embodiment, the processing order of the planar modes of the S-axis, T-axis and V-axis components of the target node is determined by calculating the laser resolution based on the parameter settings of the point cloud acquisition device. Then, the S-axis, T-axis and V-axis components of the target node are decoded according to the processing order. This can optimize the processing order based on the actual distribution of the point cloud data and reduce encoding and decoding overhead.
[0306] like Figure 11 As shown, the encoding method of this application embodiment includes:
[0307] Step 1101: The encoding device performs octree partitioning on the point cloud to be encoded to obtain the target node;
[0308] Step 1102: Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, determine the processing order of the planar mode of the S-axis component, T-axis component and V-axis component of the target node.
[0309] Step 1103: According to the processing order, the S-axis component, T-axis component and V-axis component of the target node are encoded.
[0310] Optionally, the encoding process for the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following:
[0311] If the multi-plane flag bit takes the first value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are respectively calculated according to the processing order, and the plane position information is encoded.
[0312] If the multi-plane flag bit takes the second value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are encoded with plane flag bits according to the processing order.
[0313] Optionally, the encoding of multi-plane flags, which involves encoding the S-axis, T-axis, and V-axis components of the target node according to the processing order, includes at least one of the following:
[0314] If, according to the processing order, the values of the plane flag bits corresponding to the first two components of the encoding are not all the first value, then the plane flag bits corresponding to the third component are encoded.
[0315] After encoding the planar flag bit of the target component, if the value of the planar flag bit corresponding to the target component is the first value, then the planar position information of the target component is calculated and the planar position information is encoded. The target component is any one of the S-axis component, T-axis component and V-axis component.
[0316] It should be noted that all descriptions of the encoding device side in the above embodiments are applicable to the embodiments of the encoding method applied to the encoding device side, and can achieve the same technical effect, so they will not be repeated here.
[0317] The decoding method provided in this application can be executed by a decoding device. As an example, the device can be an electronic device or a component within an electronic device, such as a chip or circuit. This application uses the example of a decoding device executing the decoding method to illustrate the decoding device provided in this application.
[0318] like Figure 12 As shown, the decoding device 1200 of this application embodiment includes:
[0319] The first processing module 1201 is used to reconstruct the point cloud to be decoded into an octree to obtain the target node;
[0320] The second processing module 1202 is used to determine the planar mode decoding qualification of the first component of the target node based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device. The first component is the S-axis component or the T-axis component of the target node.
[0321] The third processing module 1203 is used to perform decoding processing on the S-axis component, T-axis component and V-axis component of the target node based on the determination result of the planar mode decoding qualification of the first component.
[0322] Optionally, the second processing module 1202 is used for:
[0323] If at least one of the planar mode and angle mode of the target node is enabled, the planar mode decoding qualification of the first component of the target node is determined according to the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device.
[0324] Optionally, the second processing module 1202 is configured to implement at least one of the following:
[0325] If the slope range occupied by the V-axis component of the target node is smaller than the slope resolution of the point cloud acquisition device, it is determined that the first component can be decoded in planar mode.
[0326] If the slope range occupied by the V-axis component of the target node is greater than or equal to the slope resolution of the point cloud acquisition device, it is determined that the first component cannot be decoded in planar mode.
[0327] Optionally, the third processing module 1203 is used for:
[0328] Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, the first indication information is determined. The first indication information is used to indicate the processing order of the planar mode of the S-axis component, T-axis component and V-axis component.
[0329] Based on the determination result and the first indication information, the processing order of the planar mode of the target node is determined;
[0330] According to the processing order, the S-axis component, T-axis component, and V-axis component of the target node are decoded.
[0331] Optionally, the implementation of determining the processing order of the planar mode of the target node based on the determination result and the first indication information includes:
[0332] If the result indicates that the first component of the target node can be decoded in planar mode, the processing order of the planar mode of the S-axis component, T-axis component and V-axis component of the target node is determined according to the first indication information.
[0333] Optionally, the implementation of decoding the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following:
[0334] The multi-plane flag is decoded. If the multi-plane flag value is the first value, the S-axis component, T-axis component, and V-axis component of the target node are decoded according to the processing order.
[0335] The multi-plane flag is decoded. If the multi-plane flag value is the second value, the S-axis component, T-axis component and V-axis component of the target node are decoded according to the processing order.
[0336] Optionally, the implementation of decoding the planar flag bits of the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following:
[0337] If, according to the processing order, the values of the plane flag bits corresponding to the first two components are not all the first value, then the plane flag bits corresponding to the third component are decoded.
[0338] After decoding the plane flag bit of the target component, if the plane flag bit corresponding to the target component takes the first value, then the plane position information corresponding to the target component is decoded. The target component is any one of the S-axis component, T-axis component, and V-axis component.
[0339] like Figure 13 As shown, the decoding device 1300 of this application embodiment includes:
[0340] The fourth processing module 1301 is used to reconstruct the point cloud to be decoded into an octree to obtain the target node;
[0341] The fifth processing module 1302 is used to determine the processing order of the planar modes of the S-axis component, T-axis component and V-axis component of the target node based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device.
[0342] The sixth processing module 1303 is used to decode the S-axis component, T-axis component and V-axis component of the target node according to the processing order.
[0343] Optionally, the sixth processing module 1303 is configured to implement at least one of the following:
[0344] The multi-plane flag is decoded. If the multi-plane flag value is the first value, the S-axis component, T-axis component, and V-axis component of the target node are decoded according to the processing order.
[0345] The multi-plane flag is decoded. If the multi-plane flag value is the second value, the S-axis component, T-axis component and V-axis component of the target node are decoded according to the processing order.
[0346] Optionally, the implementation of decoding the planar flag bits of the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following:
[0347] If, according to the processing order, the values of the plane flag bits corresponding to the first two components are not all the first value, then the plane flag bits corresponding to the third component are decoded.
[0348] After decoding the plane flag bit of the target component, if the plane flag bit corresponding to the target component takes the second value, then the plane position information corresponding to the target component is decoded. The target component is any one of the S-axis component, T-axis component, and V-axis component.
[0349] It should be noted that the above device embodiments correspond to the above methods, and all implementation methods in the above method embodiments are applicable to this device embodiment and can achieve the same technical effect.
[0350] The decoding device in this application embodiment can be an electronic device, such as an electronic device with an operating system, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or other devices besides a terminal. For example, a terminal can include, but is not limited to, mobile phones, tablet computers, laptop computers, notebook computers, personal digital assistants (PDAs), handheld computers, netbooks, ultra-mobile personal computers (UMPCs), mobile internet devices (MIDs), augmented reality (AR) devices, virtual reality (VR) devices, robots, wearable devices, flight vehicles, vehicle user equipment (VUEs), shipboard equipment, pedestrian user equipment (PUEs), smart home devices (home devices with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), game consoles, personal computers (PCs), ATMs, or self-service machines, etc. Wearable devices include: smartwatches, smart bracelets, smart earphones, smart glasses, smart jewelry (smart bracelets, smart necklaces, smart anklets, smart ankle chains, etc.), smart wristbands, smart clothing, etc. Among these, in-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units, etc.; other devices can be servers, network attached storage (NAS), etc., and this application does not specifically limit the scope of these devices.
[0351] like Figure 14 As shown, the encoding device 1400 of this application embodiment includes:
[0352] The seventh processing module 1401 is used to perform octree partitioning on the point cloud to be encoded to obtain the target node;
[0353] The eighth processing module 1402 is used to determine the planar pattern encoding qualification of the first component of the target node based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device. The first component is the S-axis component or the T-axis component of the target node.
[0354] The ninth processing module 1403 is used to encode the S-axis component, T-axis component and V-axis component of the target node based on the determination result of the planar pattern encoding qualification of the first component.
[0355] Optionally, the eighth processing module 1402 is used for:
[0356] If at least one of the planar mode and angle mode of the target node is enabled, the planar mode encoding qualification of the first component of the target node is determined according to the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device.
[0357] Optionally, the eighth processing module 1402 is configured to implement one of the following:
[0358] If the slope range occupied by the V-axis component of the target node is smaller than the slope resolution of the point cloud acquisition device, it is determined that the first component can be encoded in a planar pattern.
[0359] If the slope range occupied by the V-axis component of the target node is greater than or equal to the slope resolution of the point cloud acquisition device, it is determined that the first component cannot be encoded in planar mode.
[0360] Optionally, the ninth processing module 1403 is used for:
[0361] Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, the first indication information is determined. The first indication information is used to indicate the processing order of the planar mode of the S-axis component, T-axis component and V-axis component.
[0362] Based on the determination result and the first indication information, the processing order of the planar mode of the target node is determined;
[0363] According to the processing order, the S-axis component, T-axis component, and V-axis component of the target node are encoded.
[0364] Optionally, the implementation of determining the processing order of the planar mode of the target node based on the determination result and the first indication information includes:
[0365] If the result indicates that the first component of the target node can be encoded in a planar mode, the processing order of the planar modes of the S-axis component, T-axis component, and V-axis component of the target node is determined according to the first indication information.
[0366] Optionally, the implementation of encoding the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following:
[0367] If the multi-plane flag bit takes the first value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are respectively calculated according to the processing order, and the plane position information is encoded.
[0368] If the multi-plane flag bit takes the second value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are encoded with plane flag bits according to the processing order.
[0369] Optionally, the implementation of encoding the multi-plane flag bits, and encoding the S-axis, T-axis, and V-axis components of the target node according to the processing order, includes at least one of the following:
[0370] If, according to the processing order, the values of the plane flag bits corresponding to the first two components of the encoding are not all the first value, then the plane flag bits corresponding to the third component are encoded.
[0371] After encoding the planar flag bit of the target component, if the value of the planar flag bit corresponding to the target component is the first value, then the planar position information of the target component is calculated and the planar position information is encoded. The target component is any one of the S-axis component, T-axis component and V-axis component.
[0372] like Figure 15 As shown, the encoding device 1500 of this application embodiment includes:
[0373] The tenth processing module 1501 is used to perform octree partitioning on the point cloud to be encoded to obtain the target node;
[0374] The eleventh processing module 1502 is used to determine the processing order of the planar modes of the S-axis component, T-axis component and V-axis component of the target node based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device.
[0375] The twelfth processing module 1503 is used to encode the S-axis component, T-axis component and V-axis component of the target node according to the processing order.
[0376] Optionally, the twelfth processing module 1503 is configured to implement at least one of the following:
[0377] If the multi-plane flag bit takes the first value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are respectively calculated according to the processing order, and the plane position information is encoded.
[0378] If the multi-plane flag bit takes the second value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are encoded with plane flag bits according to the processing order.
[0379] Optionally, the implementation of encoding the multi-plane flag bits, and encoding the S-axis, T-axis, and V-axis components of the target node according to the processing order, includes at least one of the following:
[0380] If, according to the processing order, the values of the plane flag bits corresponding to the first two components of the encoding are not all the first value, then the plane flag bits corresponding to the third component are encoded.
[0381] After encoding the planar flag bit of the target component, if the value of the planar flag bit corresponding to the target component is the first value, then the planar position information of the target component is calculated and the planar position information is encoded. The target component is any one of the S-axis component, T-axis component and V-axis component.
[0382] It should be noted that this device embodiment corresponds to the above method, and all implementation methods in the above method embodiment are applicable to this device embodiment and can achieve the same technical effect.
[0383] The encoding device in this application embodiment can be an electronic device, such as an electronic device with an operating system, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or other devices besides a terminal. For example, a terminal can include, but is not limited to, mobile phones, tablet computers, laptop computers, notebook computers, personal digital assistants (PDAs), handheld computers, netbooks, ultra-mobile personal computers (UMPCs), mobile internet devices (MIDs), augmented reality (AR) devices, virtual reality (VR) devices, robots, wearable devices, flight vehicles, vehicle user equipment (VUEs), shipboard equipment, pedestrian user equipment (PUEs), smart home devices (home appliances with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), game consoles, personal computers (PCs), ATMs, or self-service machines, etc. Wearable devices include: smartwatches, smart bracelets, smart earphones, smart glasses, smart jewelry (smart bracelets, smart necklaces, smart anklets, smart ankle chains, etc.), smart wristbands, smart clothing, etc. Among these, in-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units, etc.; other devices can be servers, network attached storage (NAS), etc., and this application does not specifically limit the scope of these devices.
[0384] like Figure 16 As shown, this application embodiment also provides an electronic device 1600, including a processor 1601 and a memory 1602. The memory 1602 stores a program or instructions that can run on the processor 1601. For example, when the electronic device 1600 is an encoding device, the program or instructions executed by the processor 1601 implement the various steps of the above-described encoding method embodiment and achieve the same technical effect. When the electronic device 1600 is a decoding device, the program or instructions executed by the processor 1601 implement the various steps of the above-described decoding method embodiment and achieve the same technical effect. To avoid repetition, this will not be described again here. Optionally, the memory 1602 may be... Figure 1 The processor 1601 can implement the memory 102 or memory 113 in the illustrated embodiment. Figure 1 The function of the encoder or decoder in the illustrated embodiment.
[0385] This application also provides an electronic device, including: a memory configured to store video data; and a processing circuit configured to implement the steps of the encoding or decoding method embodiments described above. Optionally, the memory may be... Figure 1 The processing circuitry of memory 102 or memory 113 in the illustrated embodiment can implement... Figure 1 The function of the encoder or decoder in the illustrated embodiment.
[0386] This application embodiment also provides an electronic device, including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement, for example... Figure 6 , Figure 9 , Figure 10 or Figure 11 The steps in the method embodiment shown are illustrated. This device embodiment corresponds to the above method embodiment, and all implementation processes and methods of the above method embodiments can be applied to this terminal embodiment and achieve the same technical effect.
[0387] The processor or processing circuit in this application embodiment may include general-purpose processors, special-purpose processors, etc., such as central processing units (CPUs), microprocessors, digital signal processors (DSPs), artificial intelligence (AI) processors, graphics processing units (GPUs), application-specific integrated circuits (ASICs), network processors (NPs), field-programmable gate arrays (FPGAs), or other programmable logic devices, gate circuits, transistors, discrete hardware components, etc. The communication interface in this application embodiment may include transceivers, pins, circuits, buses, etc.
[0388] The aforementioned electronic devices can be terminals or other devices besides terminals, such as servers, network attached storage (NAS), etc.
[0389] The terminal can be a mobile phone, tablet computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile internet device (MID), augmented reality (AR), virtual reality (VR) device, mixed reality (MR) device, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipborne equipment, pedestrian user equipment (PUE), smart home (home devices with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), game console, personal computer (PC), ATM or self-service machine, etc. Wearable devices include: smartwatches, smart bracelets, smart earphones, smart glasses, smart jewelry (smart bracelets, smart chains, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. Among these, in-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units, etc. It should be noted that the embodiments in this application do not limit the specific type of terminal.
[0390] A server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server. A cloud server can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), or cloud computing services based on big data and artificial intelligence platforms.
[0391] For example, the aforementioned electronic devices may include, but are not limited to, those described above. Figure 1 The type of source device 100 or destination device 110 shown.
[0392] Taking electronic devices as terminals as an example, Figure 17 A schematic diagram of the hardware structure of a terminal to implement an embodiment of this application.
[0393] The terminal 1700 includes, but is not limited to, at least some of the following components: radio frequency unit 1701, network module 1702, audio output unit 1703, input unit 1704, sensor 1705, display unit 1706, user input unit 1707, interface unit 1708, memory 1709, and processor 1710.
[0394] Those skilled in the art will understand that the terminal 1700 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1710 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 17 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0395] It should be understood that, in this embodiment, the input unit 1704 may include a graphics processor 17041 and a microphone 17042. The graphics processor 17041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1706 may include a display panel 17061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1707 includes at least one of a touch panel 17071 and other input devices 17072. The touch panel 17071 is also called a touch screen. The touch panel 17071 may include a touch detection device and a touch controller. Other input devices 17072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0396] In this embodiment, after receiving downlink data from the access network device, the radio frequency unit 1701 can transmit it to the processor 1710 for processing; in addition, the radio frequency unit 1701 can send uplink data to the network-side device. Typically, the radio frequency unit 1701 includes, but is not limited to, antennas, amplifiers, transceivers, couplers, low-noise amplifiers, duplexers, etc.
[0397] The memory 1709 can be used to store software programs or instructions, as well as various data. The memory 1709 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1709 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1709 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0398] Processor 1710 may include one or more processing units; optionally, processor 1710 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1710.
[0399] In one case, the processor 1710 is used for:
[0400] The target node is obtained by reconstructing the point cloud to be decoded using an octree.
[0401] Based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device, the planar mode decoding qualification of the first component of the target node is determined, where the first component is the S-axis component or the T-axis component of the target node.
[0402] Based on the determination result of the planar mode decoding qualification of the first component, the S-axis component, T-axis component and V-axis component of the target node are decoded.
[0403] Optionally, the processor 1710 is configured to:
[0404] If at least one of the planar mode and angle mode of the target node is enabled, the planar mode decoding qualification of the first component of the target node is determined according to the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device.
[0405] Optionally, the processor 1710 is configured to implement one of the following:
[0406] If the slope range occupied by the V-axis component of the target node is smaller than the slope resolution of the point cloud acquisition device, it is determined that the first component can be decoded in planar mode.
[0407] If the slope range occupied by the V-axis component of the target node is greater than or equal to the slope resolution of the point cloud acquisition device, it is determined that the first component cannot be decoded in planar mode.
[0408] Optionally, the processor 1710 is configured to:
[0409] Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, the first indication information is determined. The first indication information is used to indicate the processing order of the planar mode of the S-axis component, T-axis component and V-axis component.
[0410] Based on the determination result and the first indication information, the processing order of the planar mode of the target node is determined;
[0411] According to the processing order, the S-axis component, T-axis component, and V-axis component of the target node are decoded.
[0412] Optionally, the processor 1710 is configured to:
[0413] If the result indicates that the first component of the target node can be decoded in planar mode, the processing order of the planar mode of the S-axis component, T-axis component and V-axis component of the target node is determined according to the first indication information.
[0414] Optionally, the processor 1710 is configured to implement at least one of the following:
[0415] The multi-plane flag is decoded. If the multi-plane flag value is the first value, the S-axis component, T-axis component, and V-axis component of the target node are decoded according to the processing order.
[0416] The multi-plane flag is decoded. If the multi-plane flag value is the second value, the S-axis component, T-axis component and V-axis component of the target node are decoded according to the processing order.
[0417] Optionally, the processor 1710 is configured to implement at least one of the following:
[0418] If, according to the processing order, the values of the plane flag bits corresponding to the first two components are not all the first value, then the plane flag bits corresponding to the third component are decoded.
[0419] After decoding the plane flag bit of the target component, if the plane flag bit corresponding to the target component takes the first value, then the plane position information corresponding to the target component is decoded. The target component is any one of the S-axis component, T-axis component, and V-axis component.
[0420] In another case, the processor 1710 is used for:
[0421] The point cloud to be encoded is partitioned into an octree to obtain the target nodes;
[0422] Based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device, the planar pattern encoding qualification of the first component of the target node is determined, where the first component is the S-axis component or the T-axis component of the target node.
[0423] Based on the determination result of the planar pattern encoding qualification of the first component, the S-axis component, T-axis component and V-axis component of the target node are encoded.
[0424] Optionally, the processor 1710 is configured to:
[0425] If at least one of the planar mode and angle mode of the target node is enabled, the planar mode encoding qualification of the first component of the target node is determined according to the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device.
[0426] Optionally, the processor 1710 is configured to implement one of the following:
[0427] If the slope range occupied by the V-axis component of the target node is smaller than the slope resolution of the point cloud acquisition device, it is determined that the first component can be encoded in a planar pattern.
[0428] If the slope range occupied by the V-axis component of the target node is greater than or equal to the slope resolution of the point cloud acquisition device, it is determined that the first component cannot be encoded in planar mode.
[0429] Optionally, the processor 1710 is configured to:
[0430] Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, the first indication information is determined. The first indication information is used to indicate the processing order of the planar mode of the S-axis component, T-axis component and V-axis component.
[0431] Based on the determination result and the first indication information, the processing order of the planar mode of the target node is determined;
[0432] According to the processing order, the S-axis component, T-axis component, and V-axis component of the target node are encoded.
[0433] Optionally, the processor 1710 is configured to:
[0434] If the result indicates that the first component of the target node can be encoded in a planar mode, the processing order of the planar modes of the S-axis component, T-axis component, and V-axis component of the target node is determined according to the first indication information.
[0435] Optionally, the processor 1710 is configured to implement at least one of the following:
[0436] If the multi-plane flag bit takes the first value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are respectively calculated according to the processing order, and the plane position information is encoded.
[0437] If the multi-plane flag bit takes the second value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are encoded with plane flag bits according to the processing order.
[0438] Optionally, the processor 1710 is configured to implement at least one of the following:
[0439] If, according to the processing order, the values of the plane flag bits corresponding to the first two components of the encoding are not all the first value, then the plane flag bits corresponding to the third component are encoded.
[0440] After encoding the planar flag bit of the target component, if the value of the planar flag bit corresponding to the target component is the first value, then the planar position information of the target component is calculated and the planar position information is encoded. The target component is any one of the S-axis component, T-axis component and V-axis component.
[0441] In another case, the processor 1710 is used for:
[0442] The target node is obtained by reconstructing the point cloud to be decoded using an octree.
[0443] Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, determine the processing order of the planar mode of the S-axis component, T-axis component and V-axis component of the target node;
[0444] According to the processing order, the S-axis component, T-axis component, and V-axis component of the target node are decoded.
[0445] Optionally, the processor 1710 is configured to implement at least one of the following:
[0446] The multi-plane flag is decoded. If the multi-plane flag value is the first value, the S-axis component, T-axis component, and V-axis component of the target node are decoded according to the processing order.
[0447] The multi-plane flag is decoded. If the multi-plane flag value is the second value, the S-axis component, T-axis component and V-axis component of the target node are decoded according to the processing order.
[0448] Optionally, the processor 1710 is configured to implement at least one of the following:
[0449] If, according to the processing order, the values of the plane flag bits corresponding to the first two components are not all the first value, then the plane flag bits corresponding to the third component are decoded.
[0450] After decoding the plane flag bit of the target component, if the plane flag bit corresponding to the target component takes the first value, then the plane position information corresponding to the target component is decoded. The target component is any one of the S-axis component, T-axis component, and V-axis component.
[0451] In another case, the processor 1710 is used for:
[0452] The point cloud to be encoded is partitioned into an octree to obtain the target nodes;
[0453] Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, determine the processing order of the planar mode of the S-axis component, T-axis component and V-axis component of the target node;
[0454] According to the processing order, the S-axis component, T-axis component, and V-axis component of the target node are encoded.
[0455] Optionally, the processor 1710 is configured to implement at least one of the following:
[0456] If the multi-plane flag bit takes the first value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are respectively calculated according to the processing order, and the plane position information is encoded.
[0457] If the multi-plane flag bit takes the second value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are encoded with plane flag bits according to the processing order.
[0458] Optionally, the processor 1710 is configured to implement at least one of the following:
[0459] If, according to the processing order, the values of the plane flag bits corresponding to the first two components of the encoding are not all the first value, then the plane flag bits corresponding to the third component are encoded.
[0460] After encoding the planar flag bit of the target component, if the value of the planar flag bit corresponding to the target component is the first value, then the planar position information of the target component is calculated and the planar position information is encoded. The target component is any one of the S-axis component, T-axis component and V-axis component.
[0461] It is understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description of the above method embodiments and achieve the same or corresponding technical effects. To avoid repetition, it will not be described again here.
[0462] This application also provides a readable storage medium on which a program or instructions are stored. When the program or instructions are executed by a processor, they implement the various processes of the above-described encoding or decoding method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0463] The computer-readable storage medium mentioned above includes, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0464] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described encoding or decoding method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0465] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0466] This application also provides a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described encoding or decoding method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0467] This application also provides an encoding / decoding system, including an encoding device and a decoding device. The encoding device can be used to perform the steps of the above-described encoding method, and the decoding device can be used to perform the steps of the above-described decoding method.
[0468] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0469] From the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of computer software products plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes several instructions to cause the terminal or network-side device to execute the methods described in the various embodiments of this application.
[0470] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other implementations under the guidance of this application without departing from the spirit and scope of the claims. All of these implementations are within the protection scope of this application.
Claims
1. A decoding method, characterized in that, include: The decoding device performs octree reconstruction on the point cloud to be decoded to obtain the target node; Based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device, the planar mode decoding qualification of the first component of the target node is determined, where the first component is the S-axis component or the T-axis component of the target node. Based on the determination result of the planar mode decoding qualification of the first component, the S-axis component, T-axis component and V-axis component of the target node are decoded.
2. The method according to claim 1, characterized in that, The step of determining the planar mode decoding qualification of the first component of the target node based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device includes: If at least one of the planar mode and angle mode of the target node is enabled, the planar mode decoding qualification of the first component of the target node is determined according to the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device.
3. The method according to claim 1 or 2, characterized in that, The determination of the planar mode decoding qualification of the first component of the target node based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device includes the following: If the slope range occupied by the V-axis component of the target node is smaller than the slope resolution of the point cloud acquisition device, it is determined that the first component can be decoded in planar mode. If the slope range occupied by the V-axis component of the target node is greater than or equal to the slope resolution of the point cloud acquisition device, it is determined that the first component cannot be decoded in planar mode.
4. The method according to any one of claims 1-3, characterized in that, The determination result of the planar mode decoding qualification based on the first component, and the decoding processing of the S-axis component, T-axis component and V-axis component of the target node, include: Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, the first indication information is determined. The first indication information is used to indicate the processing order of the planar mode of the S-axis component, T-axis component and V-axis component. Based on the determination result and the first indication information, the processing order of the planar mode of the target node is determined; According to the processing order, the S-axis component, T-axis component, and V-axis component of the target node are decoded.
5. The method according to claim 4, characterized in that, The step of determining the processing order of the planar mode of the target node based on the determination result and the first indication information includes: If the result indicates that the first component of the target node can be decoded in planar mode, the processing order of the planar mode of the S-axis component, T-axis component and V-axis component of the target node is determined according to the first indication information.
6. The method according to claim 4, characterized in that, The decoding process for the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following: The multi-plane flag is decoded. If the multi-plane flag value is the first value, the S-axis component, T-axis component, and V-axis component of the target node are decoded according to the processing order. The multi-plane flag is decoded. If the multi-plane flag value is the second value, the S-axis component, T-axis component and V-axis component of the target node are decoded according to the processing order.
7. The method according to claim 6, characterized in that, The decoding of the planar flag bits of the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following: If, according to the processing order, the values of the plane flag bits corresponding to the first two components are not all the first value, then the plane flag bits corresponding to the third component are decoded. After decoding the plane flag bit of the target component, if the plane flag bit corresponding to the target component takes the first value, then the plane position information corresponding to the target component is decoded. The target component is any one of the S-axis component, T-axis component, and V-axis component.
8. An encoding method, characterized in that, include: The encoding device performs an octree partitioning of the point cloud to be encoded to obtain the target node; Based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device, the planar pattern encoding qualification of the first component of the target node is determined, where the first component is the S-axis component or the T-axis component of the target node. Based on the determination result of the planar pattern encoding qualification of the first component, the S-axis component, T-axis component and V-axis component of the target node are encoded.
9. The method according to claim 8, characterized in that, The step of determining the planar pattern encoding qualification of the first component of the target node based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device includes: If at least one of the planar mode and angle mode of the target node is enabled, the planar mode encoding qualification of the first component of the target node is determined according to the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device.
10. The method according to claim 8 or 9, characterized in that, The determination of the planar pattern encoding qualification of the first component of the target node based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device includes the following: If the slope range occupied by the V-axis component of the target node is smaller than the slope resolution of the point cloud acquisition device, it is determined that the first component can be encoded in a planar pattern. If the slope range occupied by the V-axis component of the target node is greater than or equal to the slope resolution of the point cloud acquisition device, it is determined that the first component cannot be encoded in planar mode.
11. The method according to any one of claims 8-10, characterized in that, The determination result of the planar pattern encoding qualification based on the first component, and the encoding processing of the S-axis component, T-axis component and V-axis component of the target node, include: Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, the first indication information is determined. The first indication information is used to indicate the processing order of the planar mode of the S-axis component, T-axis component and V-axis component. Based on the determination result and the first indication information, the processing order of the planar mode of the target node is determined; According to the processing order, the S-axis component, T-axis component, and V-axis component of the target node are encoded.
12. The method according to claim 11, characterized in that, The step of determining the processing order of the planar mode of the target node based on the determination result and the first indication information includes: If the result indicates that the first component of the target node can be encoded in a planar mode, the processing order of the planar modes of the S-axis component, T-axis component, and V-axis component of the target node is determined according to the first indication information.
13. The method according to claim 11, characterized in that, The encoding process for the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following: If the multi-plane flag bit takes the first value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are respectively calculated according to the processing order, and the plane position information is encoded. If the multi-plane flag bit takes the second value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are encoded with plane flag bits according to the processing order.
14. The method according to claim 13, characterized in that, The encoding of multi-plane flags, which involves encoding the S-axis, T-axis, and V-axis components of the target node according to the processing order, includes at least one of the following: If, according to the processing order, the values of the plane flag bits corresponding to the first two components of the encoding are not all the first value, then the plane flag bits corresponding to the third component are encoded. After encoding the planar flag bit of the target component, if the value of the planar flag bit corresponding to the target component is the first value, then the planar position information of the target component is calculated and the planar position information is encoded. The target component is any one of the S-axis component, T-axis component and V-axis component.
15. A decoding method, characterized in that, include: The decoding device performs octree reconstruction on the point cloud to be decoded to obtain the target node; Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, determine the processing order of the planar mode of the S-axis component, T-axis component and V-axis component of the target node; According to the processing order, the S-axis component, T-axis component, and V-axis component of the target node are decoded.
16. The method according to claim 15, characterized in that, The decoding process for the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following: The multi-plane flag is decoded. If the multi-plane flag value is the first value, the S-axis component, T-axis component, and V-axis component of the target node are decoded according to the processing order. The multi-plane flag is decoded. If the multi-plane flag value is the second value, the S-axis component, T-axis component and V-axis component of the target node are decoded according to the processing order.
17. The method according to claim 16, characterized in that, The decoding of the planar flag bits of the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following: If, according to the processing order, the values of the plane flag bits corresponding to the first two components are not all the first value, then the plane flag bits corresponding to the third component are decoded. After decoding the plane flag bit of the target component, if the plane flag bit corresponding to the target component takes the first value, then the plane position information corresponding to the target component is decoded. The target component is any one of the S-axis component, T-axis component, and V-axis component.
18. An encoding method, characterized in that, include: The encoding device performs an octree partitioning of the point cloud to be encoded to obtain the target node; Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, determine the processing order of the planar mode of the S-axis component, T-axis component and V-axis component of the target node; According to the processing order, the S-axis component, T-axis component, and V-axis component of the target node are encoded.
19. The method according to claim 18, characterized in that, The encoding process for the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following: If the multi-plane flag bit takes the first value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are respectively calculated according to the processing order, and the plane position information is encoded. If the multi-plane flag bit takes the second value, the multi-plane flag bit is encoded, and the S-axis component, T-axis component and V-axis component of the target node are encoded with plane flag bits according to the processing order.
20. The method according to claim 19, characterized in that, The encoding of multi-plane flags, which involves encoding the S-axis, T-axis, and V-axis components of the target node according to the processing order, includes at least one of the following: If, according to the processing order, the values of the plane flag bits corresponding to the first two components of the encoding are not all the first value, then the plane flag bits corresponding to the third component are encoded. After encoding the planar flag bit of the target component, if the value of the planar flag bit corresponding to the target component is the first value, then the planar position information of the target component is calculated and the planar position information is encoded. The target component is any one of the S-axis component, T-axis component and V-axis component.
21. A decoding device, characterized in that, include: The first processing module is used to reconstruct the point cloud to be decoded into an octree to obtain the target node; The second processing module is used to determine the planar mode decoding qualification of the first component of the target node based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device. The first component is the S-axis component or the T-axis component of the target node. The third processing module is used to decode the S-axis component, T-axis component and V-axis component of the target node based on the determination result of the planar mode decoding qualification of the first component.
22. The apparatus according to claim 21, characterized in that, The second processing module is used for: If at least one of the planar mode and angle mode of the target node is enabled, the planar mode decoding qualification of the first component of the target node is determined according to the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device.
23. The apparatus according to claim 21 or 22, characterized in that, The second processing module is configured to implement at least one of the following: If the slope range occupied by the V-axis component of the target node is smaller than the slope resolution of the point cloud acquisition device, it is determined that the first component can be decoded in planar mode. If the slope range occupied by the V-axis component of the target node is greater than or equal to the slope resolution of the point cloud acquisition device, it is determined that the first component cannot be decoded in planar mode.
24. The apparatus according to any one of claims 21-23, characterized in that, The third processing module is used for: Based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device, the first indication information is determined. The first indication information is used to indicate the processing order of the planar mode of the S-axis component, T-axis component and V-axis component. Based on the determination result and the first indication information, the processing order of the planar mode of the target node is determined; According to the processing order, the S-axis component, T-axis component, and V-axis component of the target node are decoded.
25. The apparatus according to claim 24, characterized in that, The implementation of determining the processing order of the planar mode of the target node based on the determination result and the first indication information includes: If the result indicates that the first component of the target node can be decoded in planar mode, the processing order of the planar mode of the S-axis component, T-axis component and V-axis component of the target node is determined according to the first indication information.
26. The apparatus according to claim 24, characterized in that, The implementation of decoding the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following: The multi-plane flag is decoded. If the multi-plane flag value is the first value, the S-axis component, T-axis component, and V-axis component of the target node are decoded according to the processing order. The multi-plane flag is decoded. If the multi-plane flag value is the second value, the S-axis component, T-axis component and V-axis component of the target node are decoded according to the processing order.
27. The apparatus according to claim 24, characterized in that, The implementation of decoding the planar flag bits of the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following: If, according to the processing order, the values of the plane flag bits corresponding to the first two components are not all the first value, then the plane flag bits corresponding to the third component are decoded. After decoding the plane flag bit of the target component, if the plane flag bit corresponding to the target component takes the first value, then the plane position information corresponding to the target component is decoded. The target component is any one of the S-axis component, T-axis component, and V-axis component.
28. A decoding device, characterized in that, include: The fourth processing module is used to reconstruct the point cloud to be decoded into an octree to obtain the target node; The fifth processing module is used to determine the processing order of the planar modes of the S-axis, T-axis and V-axis components of the target node based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device. The sixth processing module is used to decode the S-axis component, T-axis component and V-axis component of the target node according to the processing order.
29. The apparatus according to claim 28, characterized in that, The sixth processing module is used to implement at least one of the following: The multi-plane flag is decoded. If the multi-plane flag value is the first value, the S-axis component, T-axis component, and V-axis component of the target node are decoded according to the processing order. The multi-plane flag is decoded. If the multi-plane flag value is the second value, the S-axis component, T-axis component and V-axis component of the target node are decoded according to the processing order.
30. The apparatus according to claim 29, characterized in that, The implementation of decoding the planar flag bits of the S-axis, T-axis, and V-axis components of the target node according to the processing order includes at least one of the following: If, according to the processing order, the values of the plane flag bits corresponding to the first two components are not all the first value, then the plane flag bits corresponding to the third component are decoded. After decoding the plane flag bit of the target component, if the plane flag bit corresponding to the target component takes the first value, then the plane position information corresponding to the target component is decoded. The target component is any one of the S-axis component, T-axis component, and V-axis component.
31. A decoding device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the decoding method as described in any one of claims 1 to 7, 15 to 17.
32. An encoding device, characterized in that, include: The seventh processing module is used to perform octree partitioning on the point cloud to be encoded to obtain the target nodes; The eighth processing module is used to determine the planar pattern encoding qualification of the first component of the target node based on the slope range occupied by the V-axis component of the target node and the slope resolution of the point cloud acquisition device. The first component is the S-axis component or the T-axis component of the target node. The ninth processing module is used to encode the S-axis component, T-axis component, and V-axis component of the target node based on the determination result of the planar pattern encoding qualification of the first component.
33. An encoding device, characterized in that, include: The tenth processing module is used to perform octree partitioning on the point cloud to be encoded to obtain the target node; The eleventh processing module is used to determine the processing order of the planar modes of the S-axis, T-axis and V-axis components of the target node based on the laser resolution calculated according to the parameter settings of the point cloud acquisition device. The twelfth processing module is used to encode the S-axis component, T-axis component and V-axis component of the target node according to the processing order.
34. An encoding device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the encoding method as described in any one of claims 8 to 14, 18 to 20.
35. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the decoding method as described in any one of claims 1 to 7, 15 to 17, or the steps of the encoding method as described in any one of claims 8 to 14, 18 to 20.
36. A chip, characterized in that, The chip includes a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the steps of the decoding method as described in any one of claims 1 to 7, 15 to 17 or the steps of the encoding method as described in any one of claims 8 to 14, 18 to 20.