Point cloud information decoding method and device, point cloud information coding method and device and related equipment
By agreeing on the prediction modes of K preset transformation layers in the point cloud tree data structure, the problem of high encoding resource consumption in the existing technology is solved, and a more efficient encoding scheme is achieved.
Patent Information
- Application Number
- CN202410457542.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-16
- Publication Date
- 2025-10-24
AI Technical Summary
Existing point cloud prediction information coding schemes consume large coding resources and are not conducive to improving coding efficiency, especially in the case of a large number of transform layers.
By determining the prediction modes of K preset transformation layers as pre-agreed or the same prediction mode in the tree data structure of the point cloud, only the prediction modes of the other NK transformation layers are encoded, thereby reducing coding bits and saving resources.
It effectively reduces the consumption of coding resources and improves coding efficiency.
Smart Images

Figure CN120835147A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of coding and decoding, and particularly relates to a method and device for decoding and encoding point cloud information and related equipment. BACKGROUND
[0002] In related technologies, when encoding attribute information of a point cloud, a Region Adaptive Hierarchical Transform (RAHT) encoding mode can be used. In this encoding mode, the prediction method of each RAHT transform layer needs to be encoded. In the case of a transform tree structure of a point cloud including a large number of transform layers, a large amount of encoding resources will be consumed, and the encoding efficiency is not improved. SUMMARY
[0003] The embodiments of the present application provide a method and device for decoding and encoding point cloud information, which can solve the problem that the current encoding scheme of prediction information of a point cloud consumes a large amount of encoding resources and is not conducive to improving the encoding efficiency.
[0004] In a first aspect, a method for decoding point cloud information is provided, comprising:
[0005] A decoding end decodes a code stream to obtain first information or to obtain the first information and second information. The first information is used to indicate a prediction mode of N-K transform layers in a tree-shaped data structure of a target point cloud, except for K preset transform layers in N transform layers, and the second information is used to indicate a same prediction mode corresponding to the K preset transform layers. N is a positive integer, 0
[0006] The decoding end obtains the prediction mode of the N transform layers according to the first information and a pre-agreed prediction mode corresponding to the K preset transform layers, or according to the first information and the second information.
[0007] In a second aspect, a method for encoding point cloud information is provided, comprising:
[0008] A coding end constructs a tree-shaped data structure of a target point cloud, and the tree-shaped data structure includes N transform layers, where N is a positive integer.
[0009] The coding end determines a prediction mode of the N transform layers of the tree-shaped data structure. The prediction mode of K preset transform layers of the tree-shaped data structure is a pre-agreed prediction mode, or the prediction mode of the K preset transform layers is a same prediction mode. 0
[0010] The encoding end encodes the prediction mode of the transform layer of the tree-shaped data structure to obtain a code stream, and the code stream includes first information or includes the first information and second information.
[0011] The first information is used to indicate the prediction mode of the N-K transform layers other than the K preset transform layers in the N transform layers; and the second information is used to indicate the same prediction mode.
[0012] In a third aspect, a decoding device for point cloud information is provided, including:
[0013] A first obtaining module is configured to decode a code stream to obtain first information or to obtain the first information and second information; the first information is used to indicate the prediction mode of N-K transform layers other than K preset transform layers in N transform layers included in a tree-shaped data structure of a target point cloud, and the second information is used to indicate the same prediction mode corresponding to the K preset transform layers; N is a positive integer, 0
[0014] A second obtaining module is configured to obtain the prediction mode of the N transform layers according to the first information and the pre-agreed prediction mode corresponding to the K preset transform layers or according to the first information and the second information.
[0015] In a fourth aspect, an encoding device for point cloud information is provided, including:
[0016] A second constructing module is configured to construct a tree-shaped data structure of a target point cloud, and the tree-shaped data structure includes N transform layers, N being a positive integer.
[0017] A determining module is configured to determine the prediction mode of the N transform layers of the tree-shaped data structure, wherein the prediction mode of K preset transform layers of the tree-shaped data structure is a pre-agreed prediction mode, or the prediction mode of the K preset transform layers is the same prediction mode; 0
[0018] A fourth obtaining module is configured to encode the prediction mode of the transform layer of the tree-shaped data structure to obtain a code stream, and the code stream includes first information or includes the first information and second information.
[0019] The first information is used to indicate the prediction mode of the N-K transform layers other than the K preset transform layers in the N transform layers; and the second information is used to indicate the same prediction mode.
[0020] In a fifth aspect, an electronic device is provided, which includes a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method according to the first aspect, or implementing the steps of the method according to the second aspect.
[0021] In a sixth aspect, an electronic device is provided, which includes a processor and a communication interface, wherein the processor is configured to perform decoding processing on a bitstream to obtain first information or to obtain first information and second information; wherein the first information is used to indicate prediction modes of N-K transform layers in a tree-shaped data structure of a target point cloud, except for K preset transform layers in N transform layers, the second information is used to indicate a same prediction mode corresponding to the K preset transform layers, N is a positive integer, 0
[0022] In a seventh aspect, an electronic device is provided, which includes a memory configured to store video data, and a processing circuit configured to implement the steps of the method according to the first aspect, or implement the steps of the method according to the second aspect.
[0023] In an eighth aspect, a readable storage medium is provided, the readable storage medium storing a program or instructions, the program or instructions, when executed by a processor, implementing the steps of the method according to the first aspect, or implementing the steps of the method according to the second aspect.
[0024] In a ninth aspect, a coding system is provided, which includes an encoding end device and a decoding end device, the decoding end device being configured to implement the steps of the method according to the first aspect, and the encoding end device being configured to implement the steps of the method according to the second aspect.
[0025] In a tenth aspect, a chip is provided, which includes a processor and a communication interface, the communication interface and the processor are coupled, the processor is configured to run programs or instructions, and implement steps of the method according to the first aspect or implement steps of the method according to the second aspect.
[0026] In an eleventh aspect, a computer program / program product is provided, which is stored in a storage medium, and the program / program product is executed by at least one processor to implement steps of the method according to the first aspect or implement steps of the method according to the second aspect.
[0027] In a twelfth aspect, a computer program product is provided, which includes computer instructions, and the computer instructions are executed by a processor to implement steps of the method according to the first aspect or implement steps of the method according to the second aspect.
[0028] In the embodiments of the present application, a decoding end decodes a code stream to obtain first information or to obtain the first information and second information; wherein the first information is used to indicate a prediction mode of N-K transform layers in a tree-shaped data structure of a target point cloud, the N-K transform layers being other than K preset transform layers in N transform layers; and the prediction mode of the N transform layers is obtained according to the first information and a pre-agreed prediction mode corresponding to the K preset transform layers, or according to the first information and the second information. In the above scheme, since the prediction modes of the K preset transform layers are pre-agreed or correspond to the same prediction mode, the encoding end can not need to encode the prediction modes of the K preset transform layers or only encode the above-mentioned one prediction mode for the K preset transform layers, thereby being able to reduce encoding bits, save encoding resources, and improve encoding efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0029] Figure 1 A schematic diagram of a coding system provided by the embodiments of the present application is shown;
[0030] Figure 2a An encoding flowchart executed by an encoder based on an AVS-PCC encoding framework is shown;
[0031] Figure 2b An encoding flowchart executed by an encoder based on an MPEG G-PCC encoding framework is shown;
[0032] Figure 3a A decoding flowchart executed by a decoder based on an AVS-PCC decoding framework is shown;
[0033] Figure 3b A decoding flowchart executed by a decoder based on an MPEG G-PCC decoding framework is shown;
[0034] Figure 4 a flowchart of a decoding method of point cloud information according to an embodiment of the present application;
[0035] Figure 5 a flowchart of an encoding method of point cloud information according to an embodiment of the present application;
[0036] Figure 6 a module diagram of a decoding device of point cloud information according to an embodiment of the present application;
[0037] Figure 7 a module diagram of an encoding device of point cloud information according to an embodiment of the present application;
[0038] Figure 8 a structural block diagram of an electronic device according to an embodiment of the present application;
[0039] Figure 9 a structural block diagram of a terminal according to an embodiment of the present application. DETAILED DESCRIPTION
[0040] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.
[0041] The terms “first”, “second”, and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by “first”, “second” are generally of a kind and do not limit the number of objects, for example, the first object can be one or more. In addition, “or” in the present application means at least one of the connected objects. For example, “A or B” covers three schemes, namely, scheme one: including A and not including B; scheme two: including B and not including A; scheme three: including A and including B. The character “ / ” generally represents that the objects before and after are in an “or” relationship.
[0042] Before introducing the technical solutions provided by the embodiments of the present application, the meanings of some terms therein will be introduced first.
[0043] Point Cloud: Point cloud refers to a set of discrete points that are irregularly distributed in space and express the spatial structure and surface attributes of a three-dimensional object or a three-dimensional scene. Point clouds can be classified into different categories according to different classification standards. For example, according to the acquisition method, point clouds can be divided into dense point clouds and sparse point clouds. For another example, according to the time sequence type, point clouds can be divided into static point clouds and dynamic point clouds.
[0044] Point Cloud Data: The geometric coordinate information and attribute information of each point in the point cloud together constitute the point cloud data. The geometric coordinate information can also be referred to as three-dimensional position information. The geometric coordinate information of a point in the point cloud refers to the spatial coordinates (x, y, z) of the point, which can include the coordinate values of the point in each coordinate axis direction of the three-dimensional coordinate system, such as the coordinate value x in the X-axis direction, the coordinate value y in the Y-axis direction, and the coordinate value z in the Z-axis direction. The attribute information of a point in the point cloud can include at least one of the following: color information, material information, and laser reflection intensity information (also referred to as reflectivity). Generally, each point in the point cloud has the same number of attribute information, such as each point in the point cloud having color information and laser reflection intensity, or each point in the point cloud having color information, material information, and laser reflection intensity information.
[0045] Point Cloud Compression (PCC): Point cloud compression refers to the process of encoding the geometric coordinate information and attribute information of each point in the point cloud to obtain a compressed bitstream. Point cloud compression can include two main processes: geometric coordinate information encoding and attribute information encoding. Currently, the point cloud compression framework that can compress point clouds can be the Geometry Point Cloud Compression (G-PCC) codec framework provided by the Moving Picture Experts Group (MPEG) or the Video Point Cloud Compression (V-PCC) codec framework, or the AVS-PCC codec framework provided by the Audio Video Standard (AVS).
[0046] Point cloud decoding: Point cloud decoding is the process of decoding the compressed bitstream generated by point cloud encoding to reconstruct the point cloud. Specifically, it refers to the process of reconstructing the geometric coordinate information and attribute information of each point in the point cloud based on the geometry bitstream and attribute bitstream in the compressed bitstream. After obtaining the compressed bitstream at the decoding end, the geometry bitstream is first entropy decoded to obtain the quantized information of each point in the point cloud. Then, dequantization is performed to reconstruct the geometric coordinate information of each point in the point cloud.
[0047] For the attribute bitstream, entropy decoding is first performed to obtain the quantized attribute residual information or quantized residual transform coefficients for each point in the point cloud. The quantized attribute residual information is then dequantized to obtain the reconstructed residual information, or the quantized residual transform coefficients are dequantized to obtain the reconstructed residual transform coefficients. The reconstructed residual transform coefficients are then inversely transformed to obtain the reconstructed residual information. The reconstructed residual information is then combined with the attribute prediction information to obtain the attribute information for each point in the reconstructed point cloud. The final result is a reconstructed point cloud.
[0048] Figure 1 Schematic diagram of a coding and decoding system 10 provided in an embodiment of the present application. The technical solution of the embodiment of the present application relates to coding and decoding (CODEC) (including encoding or decoding) of point cloud data.
[0049] like Figure 1 As shown, the codec system 10 includes a source device 100, which provides encoded point cloud data that is decoded and displayed by a destination device 110. Specifically, the source device 100 provides the point cloud data to the destination device 110 via a communication medium 120. The source device 100 and the destination device 110 may include any one or more of a desktop computer, a notebook (i.e., laptop) computer, a tablet computer, a set-top box, a mobile phone, a wearable device (e.g., a smart watch or a wearable camera), a television, a camera, a display device, an in-vehicle device, a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, a digital media player, a video game console, a video conferencing device, a video streaming device, a broadcast receiver device, a broadcast transmitter device, a spacecraft, an aircraft, a robot, a satellite, and the like.
[0050] exist Figure 1 In the example of FIG, the source device 100 includes a data source 101, a memory 102, an encoder 200, and an output interface 104. The destination device 110 includes an input interface 111, a decoder 300, a memory 113, and a display device 114. The source device 100 represents an example of an encoding device, and the destination device 110 represents an example of a decoding device. In other examples, the source device 100 and the destination device 110 may not includeFigure 1 part of the components in the source device 100, or can include other components not shown in FIG. 1. Figure 1 For example, the source device 100 can acquire point cloud data through an external capturing device. Likewise, the destination device 110 can interface with an external display device, rather than include an integrated display device. For another example, the memories 102, 113 can be external memories.
[0051] Although Figure 1 Although the source device 100 and the destination device 110 are illustrated as separate devices, both can be integrated in one device in some examples. In such embodiments, the source device 100 can use the same hardware or software, or separate hardware or software, or any combination thereof, to implement the corresponding functions of the source device 100 and the destination device 110.
[0052] In some examples, the source device 100 and the destination device 110 can perform unidirectional or bidirectional transmission of data. If bidirectional transmission of data, each of the source device 100 and the destination device 110 can operate in a substantially symmetrical manner. That is, each of the source device 100 and the destination device 110 can include an encoder and a decoder.
[0053] The data source 101 represents a source of point cloud data (i.e., raw, uncoded point cloud data) and provides the encoder 200 with point cloud data for encoding by the encoder 103. The source device 100 can include a capturing device (e.g., a camera device, a sensor device, or a scanning device), an archive including previously captured point cloud data, or a feed interface for receiving point cloud data from a data content provider. The camera device can include a conventional camera, a stereo camera, a light field camera, etc., the sensor device can include a laser device, a radar device, etc., and the scanning device can include a three-dimensional laser scanning device, etc. The point cloud data can be obtained by capturing a visual scene of a real world through the capturing device. Alternatively, the data source 101 can generate computer graphics based data as source data, or combine real-time data, archived data, and computer generated data. For example, the data source generates point cloud data based on a virtual object (e.g., a virtual three-dimensional object and a virtual three-dimensional scene obtained by three-dimensional modeling).
[0054] The encoder 200 encodes the captured, pre-captured, or computer generated data. The encoder 200 can rearrange the point cloud data from a received order (sometimes referred to as a “display order”) to an encoding order. The encoder 200 can generate a bitstream including the encoded point cloud data. The source device 100 can then output the encoded point cloud data via the output interface 104 onto the communication medium 120 for reception or retrieval by, for example, the input interface 111 of the destination device 110.
[0055] Memory 102 of source device 100 and memory 113 of destination device 110 represent general purpose memories. In some examples, memory 102 can store raw data from data source 101, and memory 113 can store decoded point cloud data from decoder 300. Additionally or alternatively, memory 102, 113 can store software instructions executable by, for example, encoder 200 and decoder 300, respectively. Although memory 102 and memory 113 are shown separately from encoder 200 and decoder 300 in this example, it should be understood that encoder 200 and decoder 300 can also include internal memories for functionally similar or equivalent purposes. If encoder 200 and decoder 300 are deployed on the same hardware device, memory 102 and memory 113 can be the same memory. Furthermore, memory 102, 113 can store encoded point cloud data that is output from, for example, encoder 200 and input to decoder 300. In some examples, portions of memory 102, 113 can be allocated as one or more point cloud buffers, for example, for storing raw, decoded, or encoded point cloud data.
[0056] In some examples, source device 100 can output encoded data from output interface 104 to memory 113. Similarly, destination device 110 can access encoded data from memory 113 via input interface 111. Memory 113 or memory 102 can comprise any of a variety of distributed or locally accessed data storage media such as a hard drive, Blu-ray discs, Digital Versatile Discs (DVDs), Compact Disc Read-Only Memory (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded point cloud data.
[0057] Output interface 104 can include any type of medium or device capable of transmitting encoded point cloud data from source device 100 to destination device 110. For example, output interface 104 can include a transmitter or a transceiver, such as an antenna, configured to transmit encoded point cloud data from source device 100 directly to destination device 110 in real-time. The encoded point cloud data can be modulated according to a communication standard of a wireless communication protocol and transmitted to destination device 110.
[0058] Communication medium 120 can include transient media, such as a wireless broadcast or wired network transmission. For example, communication medium 120 can include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cable). Communication medium 120 can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. Communication medium 120 can also be in a form of a storage medium, such as a hard disk, flash drive, compact disk, digital dot cloud disk, Blu-ray disk, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded point cloud data.
[0059] In some embodiments, communication medium 120 can include routers, switches, base stations, or any other devices that can be used to facilitate communication from source device 100 to destination device 110. For example, a server (not shown) can receive encoded point cloud data from source device 100 and provide to destination device 110, e.g., via network transmission to destination device 110. The server can include a web server (e.g., for a website), a server configured to provide a file transfer protocol service such as File Transfer Protocol (FTP) or File Delivery Over Unidirectional Transport (FLUTE) protocol, a content delivery network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Services (MBMS) or evolved Multimedia Broadcast Multicast Service (eMBMS) server, or a Network-attached storage (NAS) device, etc. The server can implement one or more HTTP streaming protocols such as MPEG Media Transport (MMT) protocol, Dynamic Adaptive Streaming over HTTP (DASH) protocol, HTTP Live Streaming (HLS) protocol, or Real Time Streaming Protocol (RTSP), etc.
[0060] Destination device 110 can access the encoded point cloud data from a server, e.g., through a wireless channel (e.g., a Wi-Fi connection) or a wired connection (e.g., a Digital subscriber line (DSL), a cable modem, etc.) for accessing encoded point cloud data stored on the server.
[0061] Output interface 104 and input interface 111 can represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components operating according to the IEEE 802.11 standards or the IEEE 802.15 standards (e.g., ZigBee™), the Bluetooth standard, etc., or other physical components. In examples where output interface 104 and input interface 111 comprise wireless components, output interface 104 and input interface 111 can be configured to transfer data, such as encoded point cloud data, according to WIFI, Ethernet, cellular networks (such as 4G, LTE (Long-Term Evolution), LTE-Advanced, 5G, 6G, etc.).
[0062] The technology provided by embodiments of the present application can be applied to support one or more of the following application scenarios: machine perception point cloud, which can be used in autonomous navigation systems, real-time inspection systems, geographic information systems, visual sorting robots, rescue robots, etc.; human eye perception point cloud, which can be used in digital cultural heritage, free-view broadcasting, three-dimensional immersive communication, three-dimensional immersive interaction, etc.
[0063] Input interface 111 of destination device 110 receives an encoded bitstream from communication medium 120. The encoded bitstream can include high-level syntax elements and encoded data units (e.g., sequences, groups of pictures, pictures, slices, blocks, etc.), where the high-level syntax elements are used to decode the encoded data units to obtain decoded point cloud data. Display device 114 displays the decoded point cloud data to a user. Display device 114 can include a Cathoderay tube (CRT), a liquid-crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices. In some examples, destination device 110 can not have display device 114, e.g., if the decoded point cloud data is used to determine a location of a physical object, display device 114 can be replaced with a processor.
[0064] The encoder 200 and the decoder 300 can be implemented as one or more of various processing circuitry, which can include one or more microprocessors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), discrete logic circuitry, hardware, or any combinations thereof. When the techniques are implemented partially in software, a device can store instructions for the software in a suitable, non- transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure.
[0065] The basic principles of the encoder 200 and the decoder 300 provided by the embodiments of the present application are introduced below taking the G-PCC and AVS-PCC coding framework as an example.
[0066] The coding framework of G-PCC and AVS-PCC is roughly the same. As shown in Figure 2a The encoding process performed by the encoder based on the AVS-PCC coding framework is shown in the encoding flowchart, as shown in Figure 2b The encoding process performed by the encoder based on the MPEG G-PCC coding framework is shown in the encoding flowchart, and the above-mentioned encoder can be Figure 1 The encoder 200 shown. The above-mentioned encoding framework can be roughly divided into a geometry coordinate information encoding process and an attribute information encoding process. In the geometry information encoding process, the geometry coordinate information of each point in the point cloud is encoded to obtain a geometry bitstream; in the attribute information encoding process, the attribute information of each point in the point cloud is encoded to obtain an attribute bitstream; the geometry bitstream and the attribute bitstream jointly constitute the compressed code stream of the point cloud.
[0067] For the geometry information encoding process, the encoding process performed by the encoder 200 is as follows:
[0068] 1. Pre-processing (Pre-Processing): It can include coordinate transformation (Transform Coordinates) and voxelization (Voxelize). Through the operation of scaling and translation, the pre-processing is to convert the point cloud data in three-dimensional space into an integer form and move its minimum geometry position to the coordinate origin. In some examples, the encoder 200 can not perform pre-processing.
[0069] 2、Geometry Coding: For AVS-PCC coding framework, geometry coding includes two modes, which are Octree-based geometry coding and Prediction Tree-based geometry coding. For G-PCC coding framework, geometry coding includes three modes, which are Octree-based geometry coding, Trisoup-based geometry coding and Prediction Tree-based geometry coding.
[0070] wherein:
[0071] Octree-based geometry coding: Octree is a tree data structure, which uniformly divides a pre-defined bounding box in three-dimensional space, and each node has eight child nodes. By using "1" and "0" to indicate whether each child node of the octree is occupied or not, the occupancy code information is obtained as the code stream of the point cloud geometry information.
[0072] Prediction Tree-based geometry coding: A prediction strategy is used to generate a prediction tree, and each node of the prediction tree is traversed from the root node. The residual coordinate values corresponding to each traversed node are encoded.
[0073] Trisoup-based geometry coding: The point cloud is divided into blocks of a certain size, and the intersection points (called vertices) of the edges of the block on the surface of the point cloud are located. The compression of the geometry information is realized by encoding whether there is an intersection point on each edge of the block and the position of the intersection point.
[0074] 3、Geometry Entropy Encoding: Statistical compression encoding is performed on the occupancy code information of the octree, the prediction residual information of the prediction tree and the vertex information of the trisoup, and finally the binary (0 or 1) compressed code stream is output. Statistical encoding is a lossless encoding method, which can effectively reduce the code rate required to express the same signal. The commonly used statistical encoding method is Content Adaptive Binary Arithmetic Coding (CABAC) based on context.
[0075] 4、Geometry Reconstruction: The geometry information after geometry coding is decoded and reconstructed.
[0076] For the attribute information encoding process, the encoder 200 performs the following encoding process:
[0077] 1、Color Transformation: A transformation is applied to transform the color information of the attribute to a different domain, for example, the color information can be transformed from the RGB color space to the YCbCr color space.
[0078] 2. Attribute Recoloring: In the case of lossy coding, the original geometry information of the point cloud is changed after the geometry information is encoded. At this time, the point cloud does not have any attribute information attached, and the attribute information needs to be encoded. The original attribute information of the point cloud is used to recolor the geometry-encoding-completed point cloud. The recoloring pursues the minimum distortion to the reconstructed point cloud after the geometry coding is completed. The attribute value of the recoloring is taken from the attribute value of the neighboring point of the original point cloud, and the weighted average method determined by the spatial distance is used.
[0079] In some examples, the encoder 200 can not perform color transformation or attribute recoloring.
[0080] 3. Attribute information processing: In AVS-PCC, attribute information processing can include three modes, namely prediction encoding, transform encoding, and prediction and transform encoding. The three encoding modes can be used under different conditions.
[0081] Among them, the prediction encoding refers to determining the neighbor point of the to-be-encoded point as the prediction point in the encoded point according to the distance or spatial relationship information, calculating the predicted attribute information of the to-be-encoded point based on the set criteria according to the attribute information of the prediction point, and calculating the difference between the real attribute information of the to-be-encoded point and the predicted attribute information as the attribute residual information. The attribute residual information is quantized, transformed (optional) and entropy encoded.
[0082] The transform encoding refers to using a transform method such as discrete cosine transform (DCT), Haar transform (Haar), etc. to group and transform the attribute information, and quantize the transform coefficients. The attribute reconstruction information is obtained by inverse quantization and inverse transform. The difference between the real attribute information and the attribute reconstruction information is calculated to obtain the attribute residual information, which is quantized. The quantized transform coefficients and the attribute residual are entropy encoded.
[0083] The prediction and transform encoding refers to using the attribute residual information obtained by prediction to perform transform, quantizing and entropy encoding the transform coefficients.
[0084] In MPEG G-PCC, attribute information processing can include three modes, namely prediction and transform (Prediction Transform) encoding, lifting transform (Lifting Transform) encoding, and region adaptive hierarchical transform (RegionAdaptive Hierarchical Transform, RAHT) encoding. The three encoding modes can be used under different conditions.
[0085] In the prediction transform coding, the point cloud is divided into multiple levels of detail (LoD) according to the distance selection sub-point set, and a multi-quality level point cloud representation from coarse to fine is realized. The prediction between adjacent levels can be realized from bottom to top, that is, the attribute information of the points introduced in the fine level is predicted from the adjacent points in the coarse level, and the corresponding attribute residual information is obtained. The points in the bottom layer are encoded as reference information.
[0086] The lifting transform coding refers to introducing a weight update strategy of the neighborhood points on the basis of the LoD adjacent layer prediction, and finally obtaining the predicted attribute information of each point and the corresponding attribute residual information.
[0087] The layered region adaptive transform coding refers to converting the signal into a transform domain by performing a RAHT transform on the attribute information, and the signal in the transform domain is referred to as a transform coefficient.
[0088] 4. Attribute quantization: The quantization precision is usually determined by a quantization parameter. The transform coefficient or the attribute residual information obtained by processing the attribute information is quantized, and the quantized result is entropy encoded. For example, in the prediction transform coding and the lifting transform coding, the attribute residual information after quantization is entropy encoded; in the RAHT, the residual transform coefficient after quantization is entropy encoded.
[0089] 5. Entropy coding: The attribute residual information and / or the residual transform coefficient after quantization are generally compressed by using run length coding and arithmetic coding. The corresponding coding mode, quantization parameter and other information are also encoded by using an entropy encoder.
[0090] The encoder 200 encodes the geometric coordinate information of each point in the point cloud to obtain a geometric bitstream, and encodes the attribute information of each point in the point cloud to obtain an attribute bitstream. The encoder 200 can transmit the encoded geometric bitstream and attribute bitstream to the decoder 300.
[0091] Figure 3a A decoding flowchart executed by a decoder based on the decoding framework of the AVS-PCC is shown in FIG. 6. Figure 3b A decoding flowchart executed by a decoder based on the decoding framework of the MPEG G-PCC is shown in FIG. 7. The decoder can be the decoder 300. Figure 1The decoder 300 is shown. After the decoder 300 receives the compressed bitstream (i.e., attribute bitstream and geometry bitstream) transmitted by the encoder 200, the geometry bitstream is decoded to reconstruct the geometry coordinate information of the points in the point cloud, and the attribute bitstream is decoded to reconstruct the attribute information of the points in the point cloud.
[0092] The decoding process performed by the decoder 300 is as follows:
[0093] 1. Entropy Decoding: The geometry bitstream and the attribute bitstream are respectively entropy decoded to obtain geometry syntax elements and attribute syntax elements.
[0094] 2. Geometry Decoding: For the AVS-PCC encoding framework, geometry decoding includes two modes, namely, octree-based geometry decoding and prediction tree-based geometry decoding. For the G-PCC encoding framework, geometry decoding includes three modes, namely, octree-based geometry decoding, trisoup-based geometry decoding, and prediction tree-based prediction decoding.
[0095] Octree-based geometry decoding: The octree is reconstructed based on the geometry syntax elements parsed from the geometry bitstream.
[0096] Prediction tree-based geometry decoding: The prediction tree is reconstructed based on the geometry syntax elements parsed from the geometry bitstream.
[0097] Trisoup-based geometry decoding: The triangle model is reconstructed based on the geometry syntax elements parsed from the geometry bitstream.
[0098] 3. Geometry Reconstruction: Reconstruction is performed to obtain the geometry coordinate information of the points in the point cloud.
[0099] 4. Coordinate Inverse Transformation: The reconstructed geometry coordinate information is inverse transformed to convert the reconstructed coordinates (positions) of the points in the point cloud from the transformed domain back to the original domain.
[0100] 5. Dequantization: The attribute syntax elements are dequantized.
[0101] 6. Attribute Information Processing: In AVS-PCC, attribute information processing determines the color information of the points in the point cloud by predicting or prediction-transforming the prediction residual or prediction residual transform coefficients after dequantization, or by transforming the residual transform coefficients after dequantization.
[0102] In MPEG G-PCC, attribute information processing determines the color information of the points in the point cloud by RAHT on the attribute information after dequantization, or by LOD and inverse lifting on the attribute information after dequantization.
[0103] 7. Color inverse transform: transform color information from YCbCr color space to RGB color space. In some examples, the color inverse transform operation can not be performed.
[0104] Based on the above description, it is known that the attribute information encoding mode includes RAHT encoding and prediction and lifting transform based on LOD division.
[0105] For the decoding end, the RAHT includes:
[0106] Step 1: First, entropy-decode and dequantize the quantized transform coefficients or coefficient residuals from the input bitstream.
[0107] Step 2: Then, build the transform tree structure from bottom to top.
[0108] Since the attribute decoding is after the geometry decoding, when decoding the attribute information, the position coordinates of each point that has been reconstructed are obtained. The point cloud is reordered in ascending order according to the Morton order, and then the octree structure is built from the bottom layer to the top layer. In the process of building the transform tree, the corresponding Morton code information and weight information are generated for the merged nodes.
[0109] Step 3: Then, process from top to bottom, which includes:
[0110] A flag is parsed from the bitstream for the current transform layer to identify whether the prediction method used by the current layer is layer-based or node-based. The flag is temporarily identified as layer_based_pred_mode_flag.
[0111] When layer_based_pred_mode_flag is 1, it indicates that the prediction is layer by layer, and another flag needs to be parsed from the bitstream to identify whether the current layer is an interlayer (inter layer) or an intra layer (intra layer). The flag is temporarily identified as layer_pred_mode. When layer_pred_mode is 1, it can indicate that the current layer is an inter layer; when layer_pred_mode is 0, it can indicate that the current layer is an intra layer.
[0112] When layer_based_pred_mode_flag is 0, it means that the current transform layer uses node by node prediction, in which the whole raht transform layer is divided into three layers, and different layers use different methods to determine the prediction mode and prediction value of the current node.
[0113] The specific decoding method is as follows:
[0114]
[0115] That is, if layer_based_pred_mode_flag is 0, it is determined that the prediction mode is node by node prediction; if layer_based_pred_mode_flag is 1 and the transform layer is in the upper layer, it is determined that the prediction mode is the intra-layer prediction, that is, layer_pred_mode = 0; if layer_based_pred_mode_flag is 1 and the transform layer is in the lower layer and the lower layer does not allow weighted prediction, it is determined that the prediction mode is the inter-layer prediction, that is, layer_pred_mode = 1; if layer_based_pred_mode_flag is 1 and the transform layer is in the middle layer or in the lower layer and the lower layer allows weighted prediction, continue to analyze layer_pred_mode to determine whether the prediction mode is the inter-layer prediction or the intra-layer prediction (also described as the prediction method).
[0116] When the transform layer is in the upper layer, the calculation process of the cost nodeByNodeCost of the node by node prediction method and the cost interLayerCost of the inter-layer prediction method is the same, that is, the final calculated values of nodeByNodeCost and interLayerCost are the same.
[0117] When the transform layer is in the lower layer and does not allow weighted prediction, the calculation process of the cost nodeByNodeCost of the node by node prediction method and the cost intraLayerCost of the intra-layer prediction method is the same, that is, the final calculated values of nodeByNodeCost and intraLayerCost are the same.
[0118] The above-mentioned division of the upper layer, the middle layer and the lower layer is divided by analyzing two syntax elements in the code stream for identifying the hierarchical structure: raht_prediction_signalling_upper_dist and raht_prediction_signalling_lower_dist.
[0119] raht_prediction_signalling_upper_dist: indicates the first layer of the middle layer from the layer with a distance equal to raht_prediction_signalling_upper_dist from the layer where the root node is located. Define the layer where the root node is located as RahtRootLvl, and the first layer of the middle layer can be identified as StartSignalPredictionLvlP:
[0120] StartSignalPredictionLvlP = RahtRootLvl - raht_prediction_signalling_upper_dist
[0121] raht_prediction_signalling_lower_dist: indicates the last layer of the middle layer.
[0122] Definition: EndSignalPredictionLvl = raht_prediction_signalling_lower_dist
[0123] When the current layer lvl is greater than StartSignalPredictionLvlP, it indicates that the current layer is in the upper layer; when the current layer lvl is less than EndSignalPredictionLvl, it indicates that the current layer is in the lower layer; when the current layer lvl is less than or equal to StartSignalPredictionLvlP and greater than or equal to EndSignalPredictionLvl, it indicates that the current layer is in the middle layer.
[0124] At the same time, optionally, a flag is transmitted in the code stream to indicate whether intra-inter frame weighted prediction is allowed, if the flag is 1, two relative quantities raht_average_prediction_upper_num and raht_average_prediction_lower_num are further transmitted in the code stream to further identify the number of layers that can use intra-inter frame weighted prediction values:
[0125] raht_average_prediction_upper_num: indicates the number of layers allowed to perform weighted prediction on EndSignalPredictionLvl, including EndSignalPredictionLvl layer.
[0126] raht_average_prediction_lower_num: indicates the number of layers allowed to do weighted prediction under EndSignalPredictionLvl. Does not include EndSignalPredictionLvl. (For example, when EndSignalPredictionLvl is 5, raht_average_prediction_upper_num and raht_average_prediction_lower_num are both 1, then the number of layers allowed to do weighted prediction is the 5th and 4th layers)
[0127] The prediction mode of each node can be divided into four types, for the convenience of distinction, the following are represented by inter, intra, average, null, respectively, inter prediction, intra prediction, intra inter weighted prediction and no prediction.
[0128] The following are respectively introduced to the prediction method with layer as a unit and the prediction method with node as a unit:
[0129] 1. The prediction method with node as a unit (layer_based_pred_mode_flag = 0)
[0130] Under this method, the entire raht transform layer is divided into upper, middle and lower three layers, and different layers use different methods to determine the prediction mode and prediction value of the current node.
[0131] (1) Upper layer: infer the prediction mode of the node (inter -> intra -> null).
[0132] First, determine whether the placeholder child node of the current frame node and the child node of the corresponding reference frame node are completely matched:
[0133] For a 2*2*2 size node block, select the node with the same geometric position in the reference frame as its reference frame node. Since some child nodes of the corresponding position of the reference frame node may be empty, there is no prediction value. Therefore, the parameter inter_pred_complete_match is used to represent whether all placeholder child nodes of the current frame node can be found in the corresponding reference frame node. If yes, it represents complete match, inter_pred_complete_match = 1; otherwise, 0.
[0134] If inter_pred_complete_match = 1, the prediction mode of the current node is Inter, and the prediction value is the inter prediction value;
[0135] Otherwise, we need to determine whether the current node is eligible for intra prediction: we use a flag RahtIntraPredEligible to indicate whether the current node is eligible for intra prediction. The following are the conditions for determining whether the current node is eligible for intra prediction:
[0136] If the current node is a root node, no intra prediction is performed, and RahtIntraPredEligible = 0.
[0137] If the current node is not a root node, we need to determine whether the current 2*2*2 node is predicted:
[0138] 1) Determine whether the number of placeholder child nodes is equal to 1: if the number of non-empty placeholder child nodes of the current node to be encoded is 1, set the number of neighbor parent nodes to 19, do not perform intra prediction, and RahtIntraPredEligible = 0. Otherwise, continue to determine 2:
[0139] 2) Determine whether the number of neighbor grandparent nodes is greater than or equal to 2: if the number of neighbor grandparent nodes (including grandparent nodes) of the current node to be encoded is less than 2, do not perform intra prediction, and RahtIntraPredEligible = 0. Otherwise, proceed to 3:
[0140] 3) Neighbor search: the search range is: the parent node of the current node to be encoded (1), the coplanar neighbor parent nodes of the parent node of the current node to be encoded (6), the collinear neighbor parent nodes of the parent node of the current node to be encoded (12), the coplanar neighbor child nodes of the current node to be encoded (6), and the collinear neighbor child nodes of the current node to be encoded (12). Search the neighbor nodes in the above order, and if the neighbor node exists, record the index information corresponding to the neighbor node, and record the number of neighbor parent nodes (including the parent node itself). Then proceed to 4.
[0141] 4) Determine whether the number of neighbor parent nodes is greater than or equal to 6: if the number of neighbor parent nodes of the current node to be encoded is less than the threshold value 6, do not perform intra prediction, and RahtIntraPredEligible = 0.
[0142] 5) Otherwise, intra prediction is enabled, and RahtIntraPredEligible = 1.
[0143] If intra prediction is enabled, the prediction mode of the current node is intra, and the prediction value is the intra prediction value; otherwise, the prediction mode of the current node is null, no prediction is performed, and there is no prediction value.
[0144] It should be noted that the above method of determining whether the inter node matches and whether the intra prediction is enabled is also applicable to other transform layers, and will not be described again in the following.
[0145] (2) Mid-level: According to the matching between the current node and the reference frame node, and the opening of Intra prediction, the prediction mode of the current node is inferred or decoded from the bitstream.
[0146] If inter_pred_complete_match = 0 and RahtIntraPredEligible = = 0, the prediction mode of the current node is directly inferred as Null. Otherwise, the prediction mode of the current node is decoded from the bitstream.
[0147] When the decoded prediction mode of the current node is inter, the determination of the prediction value is divided into two cases: 1. When the current layer allows Intra-Inter weighted prediction and inter_pred_complete_match = 1 and RahtIntraPredEligible = = 1, the prediction value is the Intra-Inter weighted prediction value; 2. When the current layer does not allow Intra-Inter weighted prediction, the prediction value is the Inter prediction value.
[0148] When the decoded prediction mode of the current node is intra, the prediction value is the Intra prediction value.
[0149] When the decoded prediction mode of the current node is null, no prediction is performed, and there is no prediction value.
[0150] (3) Lower level: Infer the prediction mode of the node (intra -> null)
[0151] First, check whether the current node opens Intra prediction. If it does, the determination of the prediction value is divided into two cases: 1. When the current layer allows Intra-Inter weighted prediction and inter_pred_complete_match = 1 and RahtIntraPredEligible = = 1, the prediction value is the Intra-Inter weighted prediction value; 2. When the current layer does not allow Intra-Inter weighted prediction, the prediction value is the Intra prediction value.
[0152] Otherwise, no prediction is performed, and there is no prediction value.
[0153] 2. Prediction method in units of layers (layer_based_pred_mode_flag = 1);
[0154] At this time, the prediction mode is used in layers, and layer_pred_mode is used to identify whether the current layer is an inter layer or an intra layer. When layer_pred_mode is 1, it can represent that the current layer is an inter layer; when layer_pred_mode is 0, it can represent that the current layer is an intra layer.
[0155] 2.1, inter layer (layer_pred_mode = 1);
[0156] When layer_pred_mode is 1, it represents that the current layer is an inter layer, and the determination method of the prediction value of each node of the current layer at this time is as follows:
[0157] When the current layer allows weighted prediction, if the current node matches the inter node and the intra prediction is enabled, the prediction value is the intra-inter weighted prediction value. Otherwise, if the intra prediction is enabled, the prediction value is the intra prediction value; otherwise, it is not predicted, and there is no prediction value.
[0158] When the current layer does not allow weighted prediction, if the current node matches the inter node, the prediction value is the inter prediction value. Otherwise, if the intra prediction is enabled, the prediction value is the intra prediction value; otherwise, it is not predicted, and there is no prediction value.
[0159] 2.2, intra layer (layer_pred_mode = 0);
[0160] When layer_pred_mode is 0, it represents that the current layer is an intra layer, and the determination method of the prediction value of each node of the current layer at this time is as follows:
[0161] When the current node intra prediction is enabled, the prediction value is the intra prediction value. Otherwise, it is not predicted, and there is no prediction value.
[0162] Then, the method for obtaining the intra prediction value, the inter prediction value and the intra-inter weighted prediction value described above is as follows:
[0163] (1) Intra prediction value:
[0164] When RahtIntraPredEligible is 1, the nearest neighbor found in the neighbor search is used to perform weighted prediction on each sub-node of the current node.
[0165] The prediction weight of the parent node is 9, the prediction weight of the neighbor node coplanar with the current to-be-encoded child node is 5, the prediction weight of the neighbor node collinear with the current to-be-encoded child node is 2, the prediction weight of the neighbor parent node coplanar with the current to-be-encoded child node is 3, and the prediction weight of the neighbor parent node collinear with the current to-be-encoded child node is 1.
[0166] Each placeholder child node of the current 2*2*2 node is weighted predicted by using the neighbor node and the weight, to obtain an intra-frame prediction value of each placeholder child node of the 2*2*2 node.
[0167] (2) Inter-frame prediction value:
[0168] The inter-frame prediction value of the current 2*2*2 node comes from the prediction value of the 2*2*2 node at the same position in the reference frame. When a certain child node of the node at the same position in the reference frame is empty (the prediction value is 0), the attribute average value of the current node is used to replace the inter-frame prediction value of the child node (at this time, the inter-frame prediction value can be called revision).
[0169] (3) Intra-inter weighted prediction value:
[0170] The intra-inter weighted prediction value is a weighted average of the intra-frame prediction value and the inter-frame prediction value of the current node, and the weight values of the two are determined by the prediction mode of the parent node, the neighbor parent node and the neighbor coded child node.
[0171] Step 4: Finally, the RAHT inverse transform is performed on the coefficient residual / coefficients after the inverse quantization of the current node, and then the corresponding prediction value (or not: the prediction mode is null) is added to obtain the final attribute reconstruction value.
[0172] For the encoding end, the RAHT includes:
[0173] Step 1: First, construct the transform tree structure. Using the reconstructed geometry, start from the bottom layer and construct the octree structure from bottom to top. In the process of constructing the transform tree, the corresponding Moltin code information, attribute information and weight information need to be generated for the merged node.
[0174] Step 2: Then, process from top to bottom, which includes:
[0175] The rate-distortion optimization method (RDO: rate-distortion optimization) calculates three costs for the prediction method of the current transform layer according to the node, the inter-layer inter-frame prediction layer and the intra-layer intra-frame prediction layer, which are represented by nodeByNodeCost, interLayerCost and intraLayerCost.
[0176] Similarly, we can continue to divide the whole raht transform layer into three layers: upper, middle and lower.
[0177] (1) When the current layer is in the upper layer:
[0178] (1.1) The calculation of the cost of the prediction method in node units:
[0179] First, determine whether the current node and the child nodes of the corresponding reference frame node are completely matched. If matched, the prediction value of the current node is the inter-frame prediction value, calculate the inter-frame prediction residual, perform RAHT transform and quantization, and calculate the cost of encoding the inter-frame prediction residual coefficient, and then add it to nodeByNodeCost;
[0180] Otherwise, determine whether the current node is enabled for intra-frame prediction. If intra-frame prediction is enabled, the prediction value of the current node is the intra-frame prediction value, calculate the intra-frame prediction residual, perform RAHT transform and quantization, and calculate the cost of encoding the intra-frame prediction residual coefficient, and then add it to nodeByNodeCost;
[0181] Otherwise, the current node has no prediction value, and the original value is directly subjected to RAHT transform and quantization, and the cost of encoding the original transform coefficient is calculated, and then added to nodeByNodeCost.
[0182] (1.2) The calculation of the cost of the Inter layer prediction method:
[0183] Not calculated (because it is the same as the node by node method), set to infinity
[0184] (1.3) The calculation of the cost of the Intra layer prediction method:
[0185] First, if intra-frame prediction is enabled, the prediction value of the current node is the intra-frame prediction value, calculate the cost of encoding the intra-frame prediction residual coefficient, and then add it to intraLayerCost;
[0186] Otherwise, the current node has no prediction value, and the cost of encoding the original transform coefficient is calculated, and then added to intraLayerCost.
[0187] (2) When the current layer is in the middle layer:
[0188] (2.1) The calculation of the cost of the prediction method in node units:
[0189] If inter_pred_complete_match = 1 and intra prediction is on, use RDO to select the best prediction mode among Inter, Intra, Null. Meanwhile, it should be noted that if the current layer allows weighted prediction, the value used in Inter mode is the weighted prediction value, if not, the value used in Inter mode is the inter prediction value.
[0190] If inter_pred_complete_match = 0 and intra prediction is on, use RDO to select the best prediction mode among Inter (in this case, the inter prediction value of some sub-nodes is replaced by the average value of the parent node, which can be regarded as the inter prediction value of revision), Intra, Null.
[0191] If inter_pred_complete_match = 1 and intra prediction is not on, use RDO to select the best prediction mode among Inter, Null.
[0192] If inter_pred_complete_match = 0 and intra prediction is not on, no RDO is needed at this time, and it is directly inferred as the Null prediction mode.
[0193] Table 1
[0194]
[0195] If the inter node does not match and the intra prediction is not on, the cost of encoding the original transform coefficient is directly calculated, and then added to nodeByNodeCost.
[0196] Otherwise, the mode with the minimum cost of encoding the current node is calculated by the method of RDO, and then the cost of encoding the mode information (because at this time the prediction mode information needs to be transmitted in the code stream, including inter, intra and null, and the cost of transmitting the prediction mode also needs to be added when calculating the cost) and the cost of encoding the residual transform coefficient / original transform coefficient is calculated and added to nodeByNodeCost.
[0197] (2.2) The calculation of the cost of the Inter layer prediction method:
[0198] First, determine whether the current layer allows weighted prediction;
[0199] Case 1: If weighted prediction is allowed:
[0200] First, if inter node matches and intra is on, the prediction value of current node is weighted prediction value, calculate the cost of encoding weighted prediction residual coefficients, then add to interLayerCost;
[0201] Otherwise, if intra prediction is on, the prediction value is intra prediction value, calculate the cost of encoding intra prediction residual coefficients, then add to interLayerCost;
[0202] Otherwise, there is no prediction value, calculate the cost of encoding original transform coefficients, then add to interLayerCost.
[0203] Case 2: If weighted prediction is not allowed:
[0204] If inter node matches, the prediction value is inter prediction value, calculate the cost of encoding inter prediction residual coefficients, then add to interLayerCost;
[0205] Otherwise, if intra prediction is on, the prediction value is intra prediction value, calculate the cost of encoding intra prediction residual coefficients, then add to interLayerCost;
[0206] Otherwise, no prediction, there is no prediction value, calculate the cost of encoding original transform coefficients, then add to interLayerCost.
[0207] (2.3) Calculation of the cost of Intra layer prediction method:
[0208] First, if intra prediction is on, the prediction value of current node is intra prediction value, calculate the cost of encoding intra prediction residual coefficients, then add to intraLayerCost;
[0209] Otherwise, the current node has no prediction value, calculate the cost of encoding original transform coefficients, then add to intraLayerCost.
[0210] (3) When the current layer is in the lower layer:
[0211] (3.1) Calculation of the cost of prediction method in node unit:
[0212] First, determine whether the current layer allows weighted prediction;
[0213] Case 1: Allow weighted prediction:
[0214] First, if inter node match and intra is on, the prediction value of current node is weighted prediction value, calculate the cost of encoding weighted prediction residual coefficients, then add to nodeByNodeCost;
[0215] Otherwise, no prediction value, calculate the cost of encoding original transform coefficients, then add to nodeByNodeCost.
[0216] Case 2: if not allow weighted prediction:
[0217] If intra prediction is on, the prediction value is intra prediction value, calculate the cost of encoding intra prediction residual coefficients, then add to nodeByNodeCost;
[0218] Otherwise, no prediction, no prediction value, calculate the cost of encoding original transform coefficients, then add to nodeByNodeCost;
[0219] (3.2) The calculation of cost of Inter layer prediction method:
[0220] First, judge whether the current layer allows weighted prediction;
[0221] Case 1: allow weighted prediction:
[0222] First, if inter node match and intra is on, the prediction value of current node is weighted prediction value, calculate the cost of encoding weighted prediction residual coefficients, then add to interLayerCost;
[0223] Otherwise, if intra prediction is on, the prediction value is intra prediction value, calculate the cost of encoding intra prediction residual coefficients, then add to interLayerCost;
[0224] Otherwise, no prediction value, calculate the cost of encoding original transform coefficients, then add to interLayerCost.
[0225] Case 2: if not allow weighted prediction:
[0226] If inter node match, the prediction value is inter prediction value, calculate the cost of encoding inter prediction residual coefficients, then add to interLayerCost;
[0227] Otherwise, if intra prediction is on, the prediction value is intra prediction value, calculate the cost of encoding intra prediction residual coefficients, then add to interLayerCost;
[0228] Otherwise, no prediction, no prediction value, calculate the cost of coding the original transform coefficients, then add to interLayerCost;
[0229] (3.3) Calculation of cost of Intra layer prediction method:
[0230] First, judge whether the current layer allows weighted prediction;
[0231] Case 1: Allow weighted prediction:
[0232] First, if intra prediction is enabled, the prediction value of the current node is the intra prediction value, calculate the cost of coding the intra prediction residual coefficients, then add to intraLayerCost;
[0233] Otherwise, the current node has no prediction value, calculate the cost of coding the original transform coefficients, then add to intraLayerCost.
[0234] Case 2: If not allow weighted prediction:
[0235] Do not calculate (because the same as node by node), set to infinity.
[0236] The above process can be represented by Table 2:
[0237] Table 2
[0238]
[0239]
[0240] It should be noted that Inter->intra->null represents that if inter match, calculate the inter residual cost, otherwise if intra is on, calculate the intra residual cost, otherwise calculate the original coefficient cost, add to the cost of the corresponding prediction method of the current layer
[0241] Step 3: Then, when the intraLayerCost, interLayerCost and nodeByNodeCost of the current layer are obtained, the following comparison is made:
[0242] Case 1: When (nodeByNodeCost<intraLayerCost)&&(nodeByNodeCost<interLayerCost)
[0243] layer_based_pred_mode_flag=0; / / is the prediction method based on node;
[0244] The transform coefficients / residual transform coefficients of all nodes in the current layer are calculated according to the node-based prediction method, and quantized and entropy coded.
[0245] Case 2: When (intraLayerCost <nodeByNodeCost)&&(intraLayerCost<interLayerCost)
[0246] layer_based_pred_mode_flag = 1; / / Prediction method based on layer;
[0247] layer_pred_mode=0; / / It is intra layer;
[0248] According to the prediction method of the intra layer, the transform coefficients / residual transform coefficients of all nodes in the current layer are calculated, and quantized and entropy coded.
[0249] Case 3: When (interLayerCost <intraLayerCost)&&(interLayerCost<nodeByNodeCost)
[0250] layer_based_pred_mode_flag = 1; / / Prediction method based on layer;
[0251] layer_pred_mode=1; / / It is inter layer;
[0252] According to the inter layer prediction method, the transform coefficients / residual transform coefficients of all nodes in the current layer are calculated, and quantized and entropy coded.
[0253] Step 4: Finally: Encode and transmit the resulting prediction mode information
[0254]
[0255] The following describes in detail the method for decoding point cloud information provided by the embodiments of the present application through some embodiments and their application scenarios in combination with the accompanying drawings.
[0256] like Figure 4 As shown, the embodiment of the present application provides a method for decoding point cloud information, including:
[0257] Step 401: a decoding end decodes a code stream to obtain first information or to obtain first information and second information; wherein the first information is used to indicate a prediction mode of N-K transform layers of a tree-shaped data structure of a target point cloud, except for K preset transform layers, the second information is used to indicate a same prediction mode corresponding to the K preset transform layers, N is a positive integer, 0
[0258] Optionally, the method of the embodiment of the present application further comprises:
[0259] The decoding end constructs a tree-shaped data structure of the target point cloud, and the tree-shaped data structure comprises N transform layers, N being a positive integer.
[0260] Optionally, the target point cloud is a point cloud sequence or a point cloud slice in a point cloud sequence.
[0261] As an implementation manner, the tree-shaped data structure is an octree structure (or described as a transform tree structure).
[0262] In the embodiment of the present application, the decoding end first entropy decodes and dequantizes the quantized transform coefficients or coefficient residuals from the input code stream, and then constructs the transform tree structure from top to bottom.
[0263] Since the attribute decoding is after the geometry decoding, when the attribute information is decoded, the position coordinates of each point that is reconstructed are obtained. The point cloud is reordered in ascending order according to the Morton order, and then the octree structure is constructed from bottom to top from the bottom layer. In the process of constructing the transform tree, the corresponding Morton code information and weight information are generated for the merged nodes.
[0264] The K preset transform layers can be K preset transform layers agreed in advance, for example, the K preset transform layers are the first K transform layers in the tree-shaped data structure, or the last K transform layers in the tree-shaped data structure, or the K transform layers at other positions in the tree-shaped data structure, and the present application does not make a specific limitation thereon.
[0265] The same prediction mode corresponding to the K preset transform layers can be a prediction mode in units of nodes, an inter-layer prediction mode or an intra-layer prediction mode.
[0266] The prediction modes of the N-K transform layers can be the same or different.
[0267] Step 402: the decoding end obtains the prediction modes of the N transform layers according to the first information and the prediction mode agreed in advance corresponding to the K preset transform layers, or according to the first information and the second information.
[0268] In the case that the K preset transform layers correspond to the pre-agreed prediction mode, the prediction modes of the K preset transform layers can be the same or different.
[0269] In the embodiments of the present application, the decoding end constructs a tree-shaped data structure of the target point cloud, and the tree-shaped data structure includes N transform layers; the code stream is decoded to obtain first information or to obtain the first information and second information; the first information is used to indicate the prediction mode of N-K transform layers except for K preset transform layers in the N transform layers, and the second information is used to indicate the same prediction mode corresponding to the K preset transform layers; the prediction mode of the N transform layers is obtained according to the first information and the pre-agreed prediction mode corresponding to the K preset transform layers, or according to the first information and the second information. In the above scheme, since the prediction mode of the K preset transform layers is pre-agreed or corresponds to the same prediction mode, the encoding end does not need to encode the prediction mode of the K preset transform layers or only encodes the above-mentioned prediction mode for the K preset transform layers, thereby reducing the encoding bits, saving the encoding resources, and improving the encoding efficiency.
[0270] Optionally, the decoding end decodes the code stream to obtain the first information or to obtain the first information and the second information, including:
[0271] In the case that the prediction mode corresponding to the K preset transform layers is the pre-agreed prediction mode, the code stream is decoded to obtain the first information.
[0272] Or, in the case that the prediction mode corresponding to the K preset transform layers is not the pre-agreed prediction mode, the code stream is decoded to obtain the first information and the second information.
[0273] In the embodiments of the present application, in the case that the prediction mode corresponding to the K preset transform layers is the pre-agreed prediction mode, the decoding end obtains the prediction mode of the K preset transform layers based on the pre-agreed method, and the encoding end does not need to encode the prediction mode of the K preset transform layers, and only needs to encode the prediction mode of the N-K transform layers. In this case, the decoding end obtains the prediction mode of the N-K transform layers, i.e., the above-mentioned first information, from the code stream. In this case, the encoding end does not need to encode the prediction mode of the K preset transform layers, and does not need to determine the prediction mode of each transform layer in the K preset transform layers through the RDO method, thereby effectively saving the encoding resources of the encoding end and improving the encoding efficiency.
[0274] In the case that the prediction modes of the K preset transform layers are not pre-agreed (i.e. the prediction modes corresponding to the K preset transform layers are the same prediction mode), the encoding end encodes the same prediction mode corresponding to the K preset transform layers and encodes the prediction modes of the N-K transform layers, so that the decoding end decodes the code stream to obtain the above-mentioned same prediction mode (i.e. the second information) and the prediction modes of the N-K transform layers (i.e. the above-mentioned first information). In this case, the encoding end only encodes the above-mentioned same prediction mode for the K preset transform layers, and does not need to encode the prediction mode of each preset transform layer in the K preset transform layers, thereby effectively saving the encoding resources and improving the encoding efficiency.
[0275] Optionally, the code stream is decoded to obtain the second information, including:
[0276] The code stream is decoded to obtain the first target identifier corresponding to the K preset transform layers, the first target identifier including at least one of the first identifier and the second identifier corresponding to the K preset transform layers, wherein the first identifier corresponding to the K preset transform layers is used to indicate whether the prediction modes of the K preset transform layers are the prediction modes in layer units, and the second identifier corresponding to the K preset transform layers is used to indicate whether the prediction modes of the K preset transform layers are the inter-layer prediction modes or the intra-layer prediction modes;
[0277] The second information is obtained according to the prediction mode indicated by at least one of the first identifier and the second identifier.
[0278] Optionally, the second information is obtained according to the prediction mode indicated by at least one of the first identifier and the second identifier, including:
[0279] In the case that the first target identifier includes the first identifier, and the first identifier indicates that the prediction modes of the K preset transform layers are not the prediction modes in layer units, the second information is obtained according to the prediction mode indicated by the first identifier;
[0280] Or, in the case that the first target identifier includes the first identifier and the second identifier, and the first identifier indicates that the prediction modes of the K preset transform layers are the prediction modes in layer units, the second information is obtained according to the prediction mode indicated by the second identifier;
[0281] Or, in the case that the first target identifier includes the second identifier, the second information is obtained according to the prediction mode indicated by the second identifier.
[0282] An exemplary decoding process of the code stream is as follows: a first identifier corresponding to the K preset transform layers is obtained, the first identifier corresponding to the K preset transform layers is used to indicate whether the prediction mode of the K preset transform layers is a layer-based prediction mode;
[0283] In a case where the first identifier corresponding to the K preset transform layers indicates that the prediction mode of the K preset transform layers is not a layer-based prediction mode, it is determined that the prediction mode of the K preset transform layers is a node-based prediction mode, i.e., the second information is determined to be a node-based prediction mode of the K preset transform layers;
[0284] In a case where the first identifier corresponding to the K preset transform layers indicates that the prediction mode of the K preset transform layers is a layer-based prediction mode, a second identifier corresponding to the K preset transform layers is obtained by decoding the code stream; and the prediction mode of the K preset transform layers is determined according to the prediction mode indicated by the second identifier corresponding to the K preset transform layers, wherein the second identifier corresponding to the K preset transform layers is used to indicate whether the prediction mode of the K preset transform layers is an inter-layer prediction mode or an intra-layer prediction mode.
[0285] Here, the first identifier can be represented by layer_based_pred_mode_flag, and when layer_based_pred_mode_flag is 1, it can be used to indicate that the prediction mode of the K preset transform layers is a layer-based prediction mode, and when layer_based_pred_mode_flag is 0, it can be used to indicate that the prediction mode of the K preset transform layers is not a layer-based prediction mode, i.e., a node-based prediction mode.
[0286] When layer_based_pred_mode_flag is 1, a second identifier needs to be further obtained, and the second identifier can be represented by layer_pred_mode, and when layer_pred_mode is 1, it is used to indicate that the prediction mode of the K preset transform layers is an inter-layer prediction mode, and when layer_pred_mode is 0, it is used to indicate that the prediction mode of the K preset transform layers is an intra-layer prediction mode.
[0287] The decoding process of the prediction mode of the K preset transform layers is as follows:
[0288]
[0289] The decoding process of the first identifier is performed only once, or the decoding process of the first identifier and the second identifier is performed only once, so that the prediction modes of the K preset transform layers are obtained, and the decoding efficiency is effectively improved.
[0290] Optionally, the code stream is decoded to obtain the first information, including:
[0291] The code stream is decoded to obtain a second target identifier corresponding to a target transform layer, the second target identifier including at least one of a first identifier and a second identifier corresponding to the target transform layer, the first identifier corresponding to the target transform layer being used to indicate whether the prediction mode of the target transform layer is a layer-based prediction mode, and the second identifier corresponding to the target transform layer being used to indicate whether the prediction mode of the target transform layer is an inter-layer prediction mode or an intra-layer prediction mode, the target transform layer being any one of the N-K transform layers.
[0292] The first information is obtained according to at least one of the first identifier and the second identifier corresponding to the target transform layer.
[0293] Optionally, as an implementation manner, the first information is obtained according to at least one of the first identifier and the second identifier corresponding to the target transform layer, including:
[0294] In a case where the second target identifier includes the first identifier corresponding to the target transform layer, and the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is not a layer-based prediction mode, the first information is obtained according to the prediction mode indicated by the first identifier corresponding to the target transform layer.
[0295] In a case where the second target identifier includes the first identifier corresponding to the target transform layer, and the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a layer-based prediction mode, the first information is obtained according to position information of the target transform layer in the tree-shaped data structure.
[0296] Or, in a case where the second target identifier includes the first identifier and the second identifier corresponding to the target transform layer, and the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a layer-based prediction mode, the first information is obtained according to the prediction mode indicated by the second identifier corresponding to the target transform layer.
[0297] Or, in a case where the second target identifier includes the second identifier corresponding to the target transform layer, the first information is obtained according to the prediction mode indicated by the second identifier corresponding to the target transform layer.
[0298] An exemplary decoding process is performed on the code stream to obtain a first identifier corresponding to a target transform layer, the target transform layer being any one of the N-K transform layers.
[0299] In a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is not a layer-based prediction mode, it is determined that the prediction mode of the target transform layer is a node-based prediction mode.
[0300] In a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a layer-based prediction mode, a second identifier corresponding to the target transform layer is obtained by decoding the code stream, and the prediction mode of the target transform layer is determined according to the prediction mode indicated by the second identifier corresponding to the target transform layer.
[0301] The first information is obtained according to the prediction mode of each of the N-K transform layers.
[0302] In this implementation, the first identifier can be represented by layer_based_pred_mode_flag. When layer_based_pred_mode_flag is 1, it can be used to indicate that the prediction mode of the target transform layer is a layer-based prediction mode, and when layer_based_pred_mode_flag is 0, it can be used to indicate that the prediction mode of the target transform layer is not a layer-based prediction mode, i.e., a node-based prediction mode.
[0303] When layer_based_pred_mode_flag is 1, a second identifier needs to be further obtained, which can be represented by layer_pred_mode. When layer_pred_mode is 1, it is used to indicate that the prediction mode of the target transform layer is an inter-layer prediction mode, and when layer_pred_mode is 0, it is used to indicate that the prediction mode of the target transform layer is an intra-layer prediction mode.
[0304] In this implementation, the decoding process of the first identifier (or the first identifier and the second identifier) is performed once for each of the N-K transform layers, so as to obtain the prediction mode of each of the N-K transform layers.
[0305] The decoding process of the prediction mode of the transform layer in this implementation is as follows:
[0306]
[0307]
[0308] Optionally, the first information is obtained according to position information of the target transform layer in the tree-shaped data structure, and the first information comprises:
[0309] In a case where the target transform layer is in an upper layer part of the tree-shaped data structure, the prediction mode of the target transform layer is determined as an intra-layer prediction mode;
[0310] In a case where the target transform layer is in a lower layer part of the tree-shaped data structure and the lower layer part does not allow weighted prediction, the prediction mode of the target transform layer is determined as an inter-layer prediction mode.
[0311] Optionally, the code stream is decoded to obtain a first identifier corresponding to the target transform layer, the target transform layer being any one of the N-K transform layers;
[0312] In a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is not a layer-unit prediction mode, the prediction mode of the target transform layer is determined as a node-unit prediction mode;
[0313] In a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a layer-unit prediction mode, if the target transform layer satisfies a first condition, the prediction mode of the target transform layer is determined as an intra-layer prediction mode;
[0314] In a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a layer-unit prediction mode, if the target transform layer satisfies a second condition, the prediction mode of the target transform layer is determined as an inter-layer prediction mode.
[0315] In a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a layer-unit prediction mode, if the target transform layer does not satisfy the first condition and the second condition, the code stream is decoded to obtain a second identifier corresponding to the target transform layer, and the prediction mode of the target transform layer is determined according to a prediction mode indicated by the second identifier, the second identifier corresponding to the target transform layer being used to indicate that the prediction mode of the target transform layer is an inter-layer prediction mode or an intra-layer prediction mode.
[0316] The first condition comprises that the target transform layer is in an upper layer part of the tree-shaped data structure.
[0317] The second condition comprises that the target transform layer is in a lower layer part of the tree-shaped data structure and the lower layer part does not allow weighted prediction.
[0318] In the implementation, the first identifier can be represented by layer_based_pred_mode_flag. When layer_based_pred_mode_flag is 1, it can be used to indicate that the prediction mode of the target transform layer is the prediction mode in units of layers. When layer_based_pred_mode_flag is 0, it can be used to indicate that the prediction mode of the target transform layer is not the prediction mode in units of layers, i.e., the prediction mode in units of nodes.
[0319] If the first identifier indicates that the prediction mode of the target transform layer is the prediction mode in units of layers, and the target transform layer is in the upper layer part of the tree data structure, it can be determined that the prediction mode of the target transform layer is the intra-layer prediction mode. Thus, for the transform layers in the upper layer part of the tree data structure, the prediction mode can be determined only by the first identifier, without the need to transmit the second identifier, thereby effectively reducing the encoding bits.
[0320] If the first identifier indicates that the prediction mode of the target transform layer is the prediction mode in units of layers, and the target transform layer is in the lower layer part of the tree data structure, and the lower layer part does not allow weighted prediction, it is determined that the prediction mode of the target transform layer is the inter-layer prediction mode. Thus, for the lower layer part that does not allow weighted prediction, the prediction mode can be determined only by the first identifier, without the need to transmit the second identifier, thereby effectively reducing the encoding bits.
[0321] In the implementation, if the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is the prediction mode in units of layers, if the target transform layer does not satisfy the first condition and the second condition, the second identifier, which can be represented by layer_pred_mode, needs to be further acquired to determine the prediction mode of the target transform layer.
[0322] The decoding process of the prediction mode of the transform layer in the implementation is as follows:
[0323]
[0324] It should be noted that, in the embodiments of the present application, the division of the upper layer part, the middle layer part and the lower layer part of the tree data structure is determined based on raht_prediction_signalling_upper_dist and raht_prediction_signalling_lower_dist, and the specific division process has been described in detail in the above description, which will not be repeated here.
[0325] Optionally, the method of the embodiments of the present application further comprises:
[0326] decoding the code stream to obtain the K;
[0327] Alternatively, the K is obtained according to a protocol agreement.
[0328] In the embodiments of the present application, the encoding end can encode the K value and transmit it to the decoding end, and the decoding end obtains the K value from the code stream; the decoding end can directly obtain the K value according to a protocol agreement.
[0329] In the scheme of the embodiments of the present application, by changing the value of K, different ways of obtaining the prediction mode can be used for different transform layers in the tree-shaped data structure, for example, when K is not 0, the prediction mode of the first K transform layers is specified or unified to a certain prediction mode, so that the way of obtaining the prediction mode is more flexible, and the complexity of encoding can be reduced.
[0330] The overall flow of decoding the decoding end is described as follows, and the above K preset transform layers are taken as the first K transform layers for example.
[0331] The flow includes:
[0332] Step 1: First, the quantized transform coefficients or coefficient residuals are entropy decoded and dequantized from the input code stream.
[0333] Step 2: Then, the transform tree structure is constructed from bottom to top.
[0334] Since the attribute decoding is after the geometry decoding, when the attribute information is decoded, the position coordinates of each point are reconstructed. The point cloud is reordered in ascending order according to the Morton order, and then the octree structure is constructed from bottom to top from the bottom layer. In the process of constructing the transform tree, the corresponding Morton code information and weight information are generated for the merged nodes.
[0335] Step 3: Then, it is processed from top to bottom, and the processing process includes:
[0336] The value of K is parsed from the code stream or obtained according to the method agreed by the encoding and decoding ends.
[0337] The prediction method information of each layer is obtained, which can be a node-based prediction method, an inter-layer prediction method, or an intra-layer prediction method.
[0338] The obtaining of the prediction method information of each layer includes:
[0339] For the first K layers, when the encoding end specifies the prediction method of each layer of the first K layers according to the agreed method, the decoding end also specifies the prediction method used by each layer of the first K layers according to the prediction method agreed by the encoding end.
[0340] For the first K layers, when the encoding end obtains the prediction method of the first K layers according to the RDO method, the decoding end needs to parse the prediction method used by the transform layer of the first K layers from the code stream, and the specific method is as follows (parse once):
[0341]
[0342] For each of the remaining layers, the relevant information needs to be parsed from the code stream to obtain the prediction method information of the layer, and the specific method is as follows:
[0343]
[0344] It should be noted that when K is equal to 0: the prediction method information of each transform layer needs to be parsed, and the specific method is as follows:
[0345]
[0346]
[0347] Otherwise, continue to parse layer_pred_mode.
[0348] Step 4: Then each layer applies the corresponding prediction method according to the obtained layer_based_pred_mode_flag and layer_pred_mode.
[0349] Assumption 1: If (layer_based_pred_mode_flag = 0);
[0350] Under this assumption, different layers use different methods to determine the prediction mode and prediction value of the current node. The current node:
[0351] a) If in the upper layer: infer the prediction mode of the node (inter->intra->null).
[0352] First, judge whether the placeholder node of the current frame node and the node of the corresponding reference frame node are completely matched:
[0353] If matched, the prediction mode of the current node is Inter, and the prediction value is the inter-frame prediction value;
[0354] Otherwise, if the intra-frame prediction is enabled, the prediction mode of the current node is intra, and the prediction value is the intra-frame prediction value;
[0355] Otherwise, the prediction mode of the current node is null, and no prediction is performed, and there is no prediction value.
[0356] b) If in middle layer: infer or decode from bitstream the prediction mode of the current node according to the matching between the current node and the reference frame node, the enabling of Intra prediction.
[0357] If inter_pred_complete_match = 0 and RahtIntraPredEligible = = 0, directly infer the prediction mode of the current node as Null. Otherwise, decode the prediction mode of the current node from bitstream:
[0358] (1) When the decoded prediction mode of the current node is inter, the determination of the prediction value is divided into two cases:
[0359] The first case: When the current layer allows Intra-Inter weighted prediction and inter_pred_complete_match = 1 and RahtIntraPredEligible = 1, the prediction value is Intra-Inter weighted prediction value.
[0360] The second case: When the current layer does not allow Intra-Inter weighted prediction, the prediction value is Inter prediction value.
[0361] (2) When the decoded prediction mode of the current node is intra, the prediction value is Intra prediction value.
[0362] (3) When the decoded prediction mode of the current node is null, no prediction, no prediction value.
[0363] c) If in lower layer: infer the prediction mode of the node (intra -> null);
[0364] First, check whether Intra prediction is enabled. If enabled, the determination of the prediction value is divided into two cases:
[0365] The first case: When the current layer allows Intra-Inter weighted prediction and inter_pred_complete_match = 1 and RahtIntraPredEligible = 1, the prediction value is Intra-Inter weighted prediction value.
[0366] The second case: When the current layer does not allow Intra-Inter weighted prediction, the prediction value is Intra prediction value.
[0367] Otherwise, no prediction, no prediction value.
[0368] Assumption 2: if (layer_based_pred_mode_flag = 1 && layer_pred_mode = 1);
[0369] When layer_based_pred_mode_flag = 1 and layer_pred_mode is 1, it indicates that the current layer is an inter layer, and the determination method of the prediction value of each node of the current layer is as follows:
[0370] When the current layer allows weighted prediction, if the current node matches the inter node and the intra prediction is enabled, the prediction value is the intra-inter weighted prediction value. Otherwise, if the intra prediction is enabled, the prediction value is the intra prediction value; otherwise, it is not predicted, and there is no prediction value.
[0371] When the current layer does not allow weighted prediction, if the current node matches the inter node, the prediction value is the inter prediction value. Otherwise, if the intra prediction is enabled, the prediction value is the intra prediction value; otherwise, it is not predicted, and there is no prediction value.
[0372] Assumption 3: if (layer_based_pred_mode_flag = 1 && layer_pred_mode = 0);
[0373] When layer_based_pred_mode_flag = 1 and layer_pred_mode is 0, it indicates that the current layer is an intra layer, and the determination method of the prediction value of each node of the current layer is as follows:
[0374] When the intra prediction is enabled, the prediction value is the intra prediction value. Otherwise, it is not predicted, and there is no prediction value.
[0375] Step 5: Finally, the RAHT inverse transform is performed on the dequantized coefficients / coefficient residuals of the current node, and then the corresponding prediction value is added to obtain the final attribute reconstruction value.
[0376] In the above scheme, since the prediction modes of the K preset transformation layers are predetermined or correspond to the same prediction mode, the encoding end can not need to encode the prediction modes of the K preset transformation layers or only encode the above-mentioned one prediction mode for the K preset transformation layers, thereby being able to reduce the encoding bits, save the encoding resources, and improve the encoding efficiency. And by changing the value of K, different ways can be used to obtain the prediction mode for different transformation layers in the tree-shaped data structure, for example, when K is not 0, the prediction modes of the first K transformation layers are specified or are uniformly one prediction mode, so that the way of obtaining the prediction mode is more flexible, and the complexity of encoding can be reduced.
[0377] As shown in FIG. 1, the present application also provides a method for encoding point cloud information, comprising: Figure 5
[0378] Step 501: a target point cloud is constructed by an encoding end to form a tree data structure, and the tree data structure includes N transform layers, where N is a positive integer.
[0379] Optionally, the target point cloud is a point cloud sequence or a point cloud slice in a point cloud sequence.
[0380] As an implementation manner, the tree data structure is an octree structure (or described as a transform tree structure).
[0381] The encoding end starts from the bottom layer to construct the octree structure from bottom to top using the reconstruction geometry, and in the process of constructing the transform tree, the corresponding Morton code information, attribute information and weight information need to be generated for the merged node.
[0382] Step 502: the encoding end determines the prediction mode of the N transform layers of the tree data structure, where the prediction mode of the K preset transform layers of the tree data structure is a pre-agreed prediction mode, or the prediction mode of the K preset transform layers is the same prediction mode, 0 < K ≤ N, and K is an integer.
[0383] The K preset transform layers can be pre-agreed K preset transform layers, for example, the K preset transform layers are the first K transform layers in the tree data structure, or the last K transform layers in the tree data structure, or the K transform layers in other positions of the tree data structure, which is not limited in the present application.
[0384] The same prediction mode corresponding to the K preset transform layers can be a node-based prediction mode, an inter-layer prediction mode or an intra-layer prediction mode.
[0385] The prediction modes of the N-K transform layers can be the same or different. The prediction mode of the N-K transform layers can be selected by comparing the RD-cost (cost) of multiple prediction modes through RDO (Rate Distortion Optimization).
[0386] Step 503: the encoding end encodes the prediction mode of the transform layer of the tree data structure to obtain a bitstream, and the bitstream includes first information or includes the first information and second information.
[0387] The first information is used to indicate the prediction mode of the N-K transform layers other than the K preset transform layers in the N transform layers, and the second information is used to indicate the same prediction mode.
[0388] In the embodiments of the present application, the encoding end constructs a tree-shaped data structure of a target point cloud, the tree-shaped data structure comprising N transform layers; the encoding end determines the prediction modes of the N transform layers of the tree-shaped data structure, wherein the prediction modes of K preset transform layers of the tree-shaped data structure are a pre-agreed prediction mode, or the prediction modes of the K preset transform layers are the same prediction mode; the encoding end encodes the prediction modes of the transform layers of the tree-shaped data structure to obtain a bitstream, the bitstream comprising first information or comprising first information and second information; wherein the first information is used to indicate the prediction modes of N-K transform layers other than the K preset transform layers in the N transform layers; and the second information is used to indicate the same prediction mode. In the above scheme, since the prediction modes of the K preset transform layers are pre-agreed or correspond to the same prediction mode, the encoding end can not need to encode the prediction modes of the K preset transform layers or only encode the above-mentioned prediction mode for the K preset transform layers, thereby being able to reduce the encoding bits, save the encoding resources, and improve the encoding efficiency.
[0389] Optionally, the encoding of the prediction modes of the transform layers of the tree-shaped data structure by the encoding end comprises:
[0390] In the case where the prediction modes of the K preset transform layers are the same prediction mode, the encoding end encodes the same prediction mode and the prediction mode of each transform layer in the N-K transform layers;
[0391] Or, in the case where the prediction modes of the K preset transform layers are a pre-agreed prediction mode, the encoding end encodes the prediction mode of each transform layer in the N-K transform layers.
[0392] In the embodiments of the present application, in the case where the prediction modes corresponding to the K preset transform layers are a pre-agreed prediction mode, the encoding end does not need to encode the prediction modes of the K preset transform layers, and the encoding end only needs to encode the prediction modes of the N-K transform layers, in which case the decoding end obtains the prediction modes of the N-K transform layers, i.e. the above-mentioned first information, from the bitstream. In this case, the encoding end does not need to encode the prediction modes of the K preset transform layers, and does not need to determine the prediction mode of each transform layer in the K preset transform layers by the RDO method, thereby effectively saving the encoding resources of the encoding end and improving the encoding efficiency.
[0393] In the case that the prediction modes of the K preset transform layers are not pre-agreed (i.e. the prediction modes corresponding to the K preset transform layers are the same prediction mode), the encoding end encodes the same prediction mode corresponding to the K preset transform layers and encodes the prediction modes of the N-K transform layers, so that the decoding end decodes the code stream to obtain the same prediction mode (i.e. the second information) and the prediction modes of the N-K transform layers (i.e. the first information). In this case, the encoding end only encodes the same prediction mode for the K preset transform layers, and does not need to encode the prediction mode of each preset transform layer in the K preset transform layers, thereby effectively saving the encoding resources and improving the encoding efficiency.
[0394] Optionally, the encoding end encodes the prediction modes of the transform layers of the tree-shaped data structure to obtain the second information, including:
[0395] The encoding end encodes the first target identifier corresponding to the K preset transform layers to obtain the second information, and the first target identifier includes at least one of a first identifier and a second identifier corresponding to the K preset transform layers, wherein the first identifier corresponding to the K preset transform layers is used to indicate whether the prediction modes of the K preset transform layers are the prediction modes in layer units, and the second identifier corresponding to the K preset transform layers is used to indicate whether the prediction modes of the K preset transform layers are the inter-layer prediction modes or the intra-layer prediction modes.
[0396] Optionally, as an implementation manner, the encoding end encodes the first target identifier corresponding to the K preset transform layers to obtain the second information, and the first target identifier includes at least one of a first identifier and a second identifier corresponding to the K preset transform layers, wherein the first identifier corresponding to the K preset transform layers is used to indicate whether the prediction modes of the K preset transform layers are the prediction modes in layer units, and the second identifier corresponding to the K preset transform layers is used to indicate whether the prediction modes of the K preset transform layers are the inter-layer prediction modes or the intra-layer prediction modes.
[0397] For example, the encoding end encodes the first identifier corresponding to the K preset transform layers to obtain the first encoding information;
[0398] In the case that the first identifier corresponding to the K preset transform layers indicates that the prediction modes of the K preset transform layers are not the prediction modes in layer units, the encoding end obtains the second information according to the first encoding information;
[0399] The encoding end encodes the second identifier to obtain second encoding information in the case that the first identifier indicates that the prediction mode of the K preset transform layers is a layer-based prediction mode, and obtains the second information according to the second encoding information and the first encoding information.
[0400] Here, the first identifier can be represented by layer_based_pred_mode_flag. When layer_based_pred_mode_flag is 1, it can be used to indicate that the prediction mode of the K preset transform layers is a layer-based prediction mode. When layer_based_pred_mode_flag is 0, it can be used to indicate that the prediction mode of the K preset transform layers is not a layer-based prediction mode, but a node-based prediction mode.
[0401] When layer_based_pred_mode_flag is 1, the second identifier needs to be further encoded. The second identifier can be represented by layer_pred_mode. When layer_pred_mode is 1, it is used to indicate that the prediction mode of the K preset transform layers is an inter-layer prediction mode. When layer_pred_mode is 0, it is used to indicate that the prediction mode of the K preset transform layers is an intra-layer prediction mode.
[0402] The encoding process of the prediction mode of the K preset transform layers is as follows:
[0403] For the K preset transform layers (transmitted once):
[0404] First, encode and transmit layer_based_pred_mode_flag.
[0405] If (layer_based_pred_mode_flag == 1)
[0406] Continue to encode and transmit layer_pred_mode.
[0407] For the K preset transform layers, the encoding process of the first identifier is only performed once, or the encoding process of the first identifier and the second identifier is only performed once, which effectively improves the encoding efficiency.
[0408] Optionally, the encoding end encodes the prediction mode of the transform layer of the tree-shaped data structure to obtain first information, including:
[0409] encoding the second target identifier corresponding to the target transform layer to obtain the first information, the second target identifier including at least one of a first identifier and a second identifier corresponding to the target transform layer, the first identifier corresponding to the target transform layer being used to indicate whether the prediction mode of the target transform layer is a prediction mode in a layer unit, the second identifier corresponding to the target transform layer being used to indicate whether the prediction mode of the target transform layer is an inter-layer prediction mode or an intra-layer prediction mode, the target transform layer being any one of the N-K transform layers.
[0410] Optionally, as an implementation manner, the encoding the second target identifier corresponding to the target transform layer to obtain the first information includes:
[0411] in a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is not a prediction mode in a layer unit, obtaining the first information according to the encoding information of the first identifier corresponding to the target transform layer;
[0412] or, in a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a prediction mode in a layer unit, obtaining the first information according to the encoding information of the first identifier and the second identifier corresponding to the target transform layer, or obtaining the first information according to the encoding information of the second identifier corresponding to the target transform layer;
[0413] or, in a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a prediction mode in a layer unit, if the target transform layer is in an upper layer part of the tree-shaped data structure, or the target transform layer is in a lower layer part of the tree-shaped data structure and the lower layer part does not allow weighted prediction, obtaining the first information according to the encoding information of the first identifier, if the target transform layer is in a middle layer part of the tree-shaped data structure, obtaining the first information according to the encoding information of the first identifier and the second identifier.
[0414] For example, the encoding end encodes the first identifier corresponding to the target transform layer to obtain third encoding information;
[0415] in a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is not a prediction mode in a layer unit, obtaining the first information according to the third encoding information;
[0416] In a case where the first identification corresponding to the target transform layer indicates that the prediction mode of the target transform layer is the layer-based prediction mode, a second identification corresponding to the target transform layer is encoded to obtain fourth encoding information, and the first information is obtained according to the third encoding information and the fourth encoding information, wherein the second identification corresponding to the target transform layer is used to indicate that the prediction mode of the target transform layer is the inter-layer prediction mode or the intra-layer prediction mode.
[0417] In the example, the first identification can be represented by layer_based_pred_mode_flag. When layer_based_pred_mode_flag is 1, it can be used to indicate that the prediction mode of the target transform layer is the layer-based prediction mode. When layer_based_pred_mode_flag is 0, it can be used to indicate that the prediction mode of the target transform layer is not the layer-based prediction mode, i.e., the node-based prediction mode.
[0418] When layer_based_pred_mode_flag is 1, the second identification needs to be further encoded. The second identification can be represented by layer_pred_mode. When layer_pred_mode is 1, it is used to indicate that the prediction mode of the target transform layer is the inter-layer prediction mode. When layer_pred_mode is 0, it is used to indicate that the prediction mode of the target transform layer is the intra-layer prediction mode.
[0419] In the example, the encoding process of the prediction mode of the N-K transform layers is as follows (the process is performed once for each layer):
[0420] First, layer_based_pred_mode_flag is encoded and transmitted.
[0421] If (layer_based_pred_mode_flag == 1)
[0422] Continue to encode and transmit layer_pred_mode.
[0423] In an example, the first identification corresponding to the target transform layer is encoded to obtain third encoding information.
[0424] In a case where the first identification corresponding to the target transform layer indicates that the prediction mode of the target transform layer is not the layer-based prediction mode, the first information is obtained according to the third encoding information.
[0425] In a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is the layer-based prediction mode, if the target transform layer satisfies the first condition or the second condition, the first information is obtained according to the third coding information;
[0426] In a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is the layer-based prediction mode, if the target transform layer does not satisfy the first condition and the second condition, a second identifier corresponding to the target transform layer is coded to obtain fourth coding information, and the first information is obtained according to the third coding information and the fourth coding information, wherein the second identifier corresponding to the target transform layer is used to indicate that the prediction mode of the target transform layer is the inter-layer prediction mode or the intra-layer prediction mode.
[0427] The first condition comprises that the target transform layer is in the upper layer part of the tree-shaped data structure.
[0428] The second condition comprises that the target transform layer is in the lower layer part of the tree-shaped data structure, and the lower layer part does not allow weighted prediction.
[0429] In this example, the first identifier can be represented by layer_based_pred_mode_flag. When layer_based_pred_mode_flag is 1, it can be used to indicate that the prediction mode of the target transform layer is the layer-based prediction mode. When layer_based_pred_mode_flag is 0, it can be used to indicate that the prediction mode of the target transform layer is not the layer-based prediction mode, but the node-based prediction mode.
[0430] In a case where the first identifier indicates that the prediction mode of the target transform layer is the layer-based prediction mode, and the target transform layer is in the upper layer part of the tree-shaped data structure, the decoding end can determine that the prediction mode of the target transform layer is the intra-layer prediction mode, without the need to code the second identifier, thereby effectively reducing the coding bits.
[0431] In a case where the first identifier indicates that the prediction mode of the target transform layer is the layer-based prediction mode, and the target transform layer is in the lower layer part of the tree-shaped data structure, and the lower layer part does not allow weighted prediction, the decoding end can determine that the prediction mode of the target transform layer is the inter-layer prediction mode, without the need to transmit the second identifier, thereby effectively reducing the coding bits.
[0432] In the example, if the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a layer-based prediction mode, if the target transform layer does not satisfy the first condition and the second condition, it is necessary to further encode a second identifier, which can be represented by layer_pred_mode, and the prediction mode of the target transform layer is determined through the second identifier.
[0433] The encoding process of the prediction mode of each transform layer in the example is as follows:
[0434] First, encode the transmission layer_based_pred_mode_flag,
[0435] If (middle layer || (lower layer && allow weighting))
[0436] If (layer_based_pred_mode_flag == 1)
[0437] Continue to encode the transmission layer_pred_mode.
[0438] It should be noted that in the embodiments of the present application, the division of the upper part, the middle part and the lower part of the tree data structure is determined based on raht_prediction_signalling_upper_dist and raht_prediction_signalling_lower_dist, and the specific division process has been described in detail in the above description, which will not be repeated here.
[0439] Optionally, the method of the embodiments of the present application further comprises:
[0440] Obtain the target rate-distortion optimization cost of each of the at least two preset prediction modes, and the target rate-distortion optimization cost of the preset prediction mode is obtained based on the sum of the rate-distortion optimization costs of the K preset transform layers under the preset prediction mode;
[0441] Among the at least two preset prediction modes, the preset prediction mode with the minimum target rate-distortion optimization cost (RD-cost) is selected as the prediction mode of the K preset transform layers.
[0442] For example, the at least two preset prediction modes include a node-based prediction mode, an inter-layer prediction mode and an intra-layer prediction mode, and for the first K preset transform layers, the RD-cost of the K preset transform layers is determined by RDO comparison according to the three prediction modes respectively, and the prediction method with the minimum RD-cost is selected as the prediction method of the K preset transform layers. Assuming that the above K preset transform layers are the first K transform layers, the specific method is as follows:
[0443] The method based on rate distortion optimization (RDO) calculates three costs for these layers according to the inter layer and intra layer prediction methods in node units, denoted as K_nodeByNodeCost, K_interLayerCost and K_intraLayerCost. Take K_nodeByNodeCost as an example, the final value is the RD cost of the first K layers encoded according to the prediction method in node units.
[0444] The following operations are performed on each of the first K transformed layers:
[0445] 1) When the current transformed layer is in the upper layer:
[0446] 1.1) Calculation of the cost of the prediction method in node units:
[0447] The calculation method is consistent with the encoding end calculation method described above, which is not repeated here, and is added to K_nodeByNodeCost;
[0448] 1.2) Calculation of the cost of the inter layer prediction method:
[0449] The calculation method is consistent with the encoding end calculation method described above, which is not repeated here, and is added to K_interLayerCost;
[0450] 1.3) Calculation of the cost of the intra layer prediction method:
[0451] The calculation method is consistent with the encoding end calculation method described above, which is not repeated here, and is added to K_intraLayerCost.
[0452] 2) When the current transformed layer is in the middle layer:
[0453] 2.1) Calculation of the cost of the prediction method in node units:
[0454] The calculation method is consistent with the encoding end calculation method described above, which is not repeated here, and is added to K_nodeByNodeCost;
[0455] 2.2) Calculation of the cost of the inter layer prediction method:
[0456] The calculation method is consistent with the encoding end calculation method described above, which is not repeated here, and is added to K_interLayerCost;
[0457] 2.3) Calculation of cost of Intra layer prediction method:
[0458] In line with the encoding side calculation method described above, it is not repeated here, and is added to K_intraLayerCost.
[0459] 3) When the current transform layer is in the lower layer:
[0460] 3.1) Calculation of cost of prediction method in node unit:
[0461] In line with the encoding side calculation method described above, it is not repeated here, and is added to K_nodeByNodeCost;
[0462] 3.2) Calculation of cost of Intra layer prediction method:
[0463] In line with the encoding side calculation method described above, it is not repeated here, and is added to K_intraLayerCost;
[0464] 3.3) Calculation of cost of Inter layer prediction method:
[0465] In line with the encoding side calculation method described above, it is not repeated here, and is added to K_interLayerCost.
[0466] Then, when the encoding cost of different prediction methods of the first K layers: K_nodeByNodeCost, K_intraLayerCost and K_interLayerCost are obtained
[0467] Comparison is made as follows:
[0468] Case 1: When (K_nodeByNodeCost
[0469] layer_based_pred_mode_flag = 0; / / It is the prediction method in node unit
[0470] The transform coefficients / residual transform coefficients of all nodes of the first K layers are calculated according to the prediction method in node unit, and quantization and entropy encoding are performed.
[0471] Case 2: When (K_intraLayerCost
[0472] layer_based_pred_mode_flag = 1; / / is layer-based prediction method
[0473] layer_pred_mode = 0; / / is intra layer
[0474] According to the intra layer prediction method, the transform coefficients / residual transform coefficients of all nodes in the first K layers are calculated, quantized and entropy encoded.
[0475] Case 3: when (K_interLayerCost < K_intraLayerCost) && (K_interLayerCost < K_nodeByNodeCost)
[0476] layer_based_pred_mode_flag = 1; / / is layer-based prediction method
[0477] layer_pred_mode = 1; / / is inter layer
[0478] According to the inter layer prediction method, the transform coefficients / residual transform coefficients of all nodes in the first K layers are calculated, quantized and entropy encoded.
[0479] Optionally, the method of the embodiment of the application further comprises:
[0480] Obtaining the rate-distortion optimization cost of each transform layer in the N-K transform layers under at least two preset prediction modes;
[0481] Among the at least two preset prediction modes, the preset prediction mode with the minimum rate-distortion optimization cost is selected as the prediction mode of each transform layer.
[0482] For example, the at least two preset prediction modes include a node-based prediction mode, an inter layer prediction mode and an intra layer prediction mode. For each transform layer in the N-K transform layers, the RD-cost of the three prediction methods is compared by RDO to determine the prediction method with the minimum cost as the prediction method of the layer. The specific method is as follows:
[0483] The rate-distortion optimization method (RDO: rate-distortion optimization) calculates three costs for the current layer according to the node-based, inter layer and intra layer prediction methods, which are represented by nodeByNodeCost, interLayerCost and intraLayerCost.
[0484] 1) When the current transform layer is in the upper layer:
[0485] 1.1) The calculation of the cost of the prediction method in the node unit:
[0486] In line with the encoding end calculation method described above, it will not be repeated here;
[0487] 1.2) The calculation of the cost of the Inter layer prediction method:
[0488] In line with the encoding end calculation method described above, it will not be repeated here;
[0489] 1.3) The calculation of the cost of the Intra layer prediction method:
[0490] In line with the encoding end calculation method described above, it will not be repeated here.
[0491] 2) When the current transform layer is in the middle layer:
[0492] 2.1) The calculation of the cost of the prediction method in the node unit:
[0493] In line with the encoding end calculation method described above, it will not be repeated here;
[0494] 2.2) The calculation of the cost of the Inter layer prediction method:
[0495] In line with the encoding end calculation method described above, it will not be repeated here;
[0496] 2.3) The calculation of the cost of the Intra layer prediction method:
[0497] In line with the encoding end calculation method described above, it will not be repeated here.
[0498] 3) When the current transform layer is in the lower layer:
[0499] 3.1) The calculation of the cost of the prediction method in the node unit:
[0500] In line with the encoding end calculation method described above, it will not be repeated here;
[0501] 3.2) The calculation of the cost of the Intra layer prediction method:
[0502] In line with the encoding end calculation method described above;
[0503] 3.3) The calculation of the cost of the Inter layer prediction method:
[0504] Consistent with the encoding end calculation method described above.
[0505] Then, when the intraLayerCost, interLayerCost and nodeByNodeCost of the current layer are obtained, the following comparison is made:
[0506] Case 1: when (nodeByNodeCost < intraLayerCost) && (nodeByNodeCost < interLayerCost);
[0507] layer_based_pred_mode_flag = 0; / / is the node-based prediction method
[0508] According to the node-based prediction method, the transform coefficients / residual transform coefficients of all nodes of the current layer are calculated, quantized and entropy encoded.
[0509] Case 2: when (intraLayerCost < nodeByNodeCost) && (intraLayerCost < interLayerCost);
[0510] layer_based_pred_mode_flag = 1; / / is the layer-based prediction method
[0511] layer_pred_mode = 0; / / is intra layer
[0512] According to the intra layer prediction method, the transform coefficients / residual transform coefficients of all nodes of the current layer are calculated, quantized and entropy encoded.
[0513] Case 3: when (interLayerCost < intraLayerCost) && (interLayerCost < nodeByNodeCost);
[0514] layer_based_pred_mode_flag = 1; / / is the layer-based prediction method
[0515] layer_pred_mode = 1; / / is inter layer
[0516] According to the inter layer prediction method, the transform coefficients / residual transform coefficients of all nodes of the current layer are calculated, quantized and entropy encoded.
[0517] Optionally, the method of the embodiment of the application further comprises encoding information of the K in the code stream.
[0518] In the embodiments of the present application, the value of K can be agreed upon at the encoding and decoding end, or can be transmitted in the code stream. If the value of K is transmitted, the value of K is encoded. The encoding method can be context-based adaptive arithmetic coding, variable-length coding, exponential Golomb coding, etc.
[0519] The overall process of encoding at the encoding end is described as follows, which is described by taking the first K preset transform layers as the first K transform layers.
[0520] The process includes:
[0521] Step 1: First, build the transform tree structure. Use the reconstructed geometry to build the octree structure from the bottom layer to the top layer. In the process of building the transform tree, the corresponding Morden code information, attribute information and weight information need to be generated for the merged nodes.
[0522] Step 2: Then, process from top to bottom, which includes:
[0523] Select a prediction method for all transform layers. The prediction method can be a node-based prediction method, an inter-layer prediction method, or an intra-layer prediction method.
[0524] When a prediction method is selected for all transform layers, the prediction method of the first K transform layers can be specified as one of the three, and the remaining layers can be determined by comparing the RD-cost of the three prediction modes through RDO.
[0525] When a prediction method is selected for all transform layers, the first K transform layers can be compared through RDO to determine the RD-cost of the three prediction modes, that is, the prediction method with the minimum cost required to encode the first K layers is selected as the prediction method of the transform layers of the first K layers, and the remaining layers can be determined by comparing the RD-cost of the three prediction modes through RDO.
[0526] (1) When a prediction method is selected for all transform layers, the prediction method of the first K transform layers is specified as one of the three, and the remaining layers can be determined by comparing the RD-cost of the three prediction methods through RDO. The method is as follows:
[0527] For the first K layers, their prediction methods are respectively specified as one of the three prediction methods, for example, the rootlvl layer can be specified as the node-by-node prediction method, the (rootlvl-1) layer can be specified as the inter layer prediction method, and the (rootlvl-K-1) layer can be specified as the intra layer prediction method. The first K layers can have the same prediction method, or different prediction methods, or some of the first K layers have the same prediction method. The encoding and decoding ends can encode and decode each of the first K layers according to the specified prediction method according to the agreed method.
[0528] For the remaining layers, the RD-costs of the three prediction methods are compared by RDO to determine the prediction method with the minimum cost as the prediction method of the layer. The specific method is as follows:
[0529] First, the rate-distortion optimization (RDO) method is used to calculate the costs of the node-by-node, inter layer and intra layer prediction methods for the current layer, which are represented by nodeByNodeCost, interLayerCost and intraLayerCost.
[0530] 1) When the current layer is in the upper layer:
[0531] 1.1) Calculation of the cost of the node-by-node prediction method:
[0532] The calculation method is consistent with the encoding end calculation method described above, and will not be described again;
[0533] 1.2) Calculation of the cost of the inter layer prediction method:
[0534] The calculation method is consistent with the encoding end calculation method described above, and will not be described again;
[0535] 1.3) Calculation of the cost of the intra layer prediction method:
[0536] The calculation method is consistent with the encoding end calculation method described above, and will not be described again.
[0537] 2) When the current layer is in the middle layer:
[0538] 2.1) Calculation of the cost of the node-by-node prediction method:
[0539] The calculation method is consistent with the encoding end calculation method described above, and will not be described again;
[0540] 2.2) Calculation of the cost of the inter layer prediction method:
[0541] In accordance with the encoding end calculation method described above, no further description is given.
[0542] 2.3) Calculation of the cost of the Intra layer prediction method:
[0543] In accordance with the encoding end calculation method described above, no further description is given.
[0544] 3) When the current layer is in the lower layer:
[0545] 3.1) Calculation of the cost of the prediction method in node units:
[0546] In accordance with the encoding end calculation method described above, no further description is given.
[0547] 3.2) Calculation of the cost of the Intra layer prediction method:
[0548] In accordance with the encoding end calculation method described above, no further description is given.
[0549] 3.3) Calculation of the cost of the Inter layer prediction method:
[0550] In accordance with the encoding end calculation method described above, no further description is given.
[0551] Then, when the intraLayerCost, interLayerCost and nodeByNodeCost of the current layer are obtained, the following comparison is made:
[0552] Case 1: When (nodeByNodeCost < intraLayerCost) && (nodeByNodeCost < interLayerCost);
[0553] layer_based_pred_mode_flag = 0; / / It is the prediction method in node units
[0554] The transform coefficients / residual transform coefficients of all nodes of the current layer are calculated according to the prediction method in node units, and quantization and entropy encoding are performed.
[0555] Case 2: When (intraLayerCost < nodeByNodeCost) && (intraLayerCost < interLayerCost);
[0556] layer_based_pred_mode_flag = 1; / / It is the prediction method in layer units
[0557] layer_pred_mode = 0; / / is intra layer
[0558] Calculate transform coefficients / residual transform coefficients of all nodes in current layer according to the prediction method of intra layer, and then quantize and entropy encode.
[0559] Case 3: When (interLayerCost < intraLayerCost) && (interLayerCost < nodeByNodeCost)
[0560] layer_based_pred_mode_flag = 1; / / is the prediction method based on layer
[0561] layer_pred_mode = 1; / / is inter layer
[0562] Calculate transform coefficients / residual transform coefficients of all nodes in current layer according to the prediction method of inter layer, and then quantize and entropy encode.
[0563] Finally, encode the prediction method information. At the same time, the value of K can be agreed upon at the encoding and decoding end, using the same value, or transmitted in the code stream.
[0564] If the value of K is transmitted, encode the value of K. The encoding method can be context-based adaptive arithmetic coding, variable length coding, exponential Golomb coding, etc.
[0565] In addition, encode and transmit the prediction method information of each remaining layer, as follows:
[0566] First, encode and transmit layer_based_pred_mode_flag,
[0567] If (middle layer || (lower layer && allow weighting))
[0568] If (layer_based_pred_mode_flag == 1)
[0569] Continue to encode and transmit layer_pred_mode.
[0570] (2) When selecting a prediction method for all transform layers, compare the RD-cost of the three prediction modes for the first K transform layers together through RDO to determine, i.e. select the prediction method with the minimum cost required to encode the first K layers as the prediction method for the first K transform layers, and the remaining layers can be determined by comparing the RD-cost of the three prediction modes through RDO. The method is as follows:
[0571] For the first K layers, the RD-cost of encoding these layers by three prediction methods is compared, and the prediction method with the minimum cost is selected as the prediction method of these layers. The specific method is as follows:
[0572] The method based on rate-distortion optimization (RDO) calculates three costs for these layers according to the inter layer and intra layer prediction methods in units of nodes, denoted as K_nodeByNodeCost, K_interLayerCost and K_intraLayerCost. Take K_nodeByNodeCost as an example. The final value is the RD cost of encoding the first K layers according to the prediction method in units of nodes.
[0573] The following operations are performed on each of the first K transform layers:
[0574] 1) When the current transform layer is in the upper layer:
[0575] 1.1) Calculation of the cost of the prediction method in units of nodes:
[0576] The calculation method is consistent with the encoding end calculation method described above, which is not repeated here, and is added to K_nodeByNodeCost;
[0577] 1.2) Calculation of the cost of the inter layer prediction method:
[0578] The calculation method is consistent with the encoding end calculation method described above, which is not repeated here, and is added to K_interLayerCost;
[0579] 1.3) Calculation of the cost of the intra layer prediction method:
[0580] The calculation method is consistent with the encoding end calculation method described above, which is not repeated here, and is added to K_intraLayerCost.
[0581] 2) When the current transform layer is in the middle layer:
[0582] 2.1) Calculation of the cost of the prediction method in units of nodes:
[0583] The calculation method is consistent with the encoding end calculation method described above, which is not repeated here, and is added to K_nodeByNodeCost;
[0584] 2.2) Calculation of the cost of the inter layer prediction method:
[0585] In line with the encoding side calculation method described above, the calculation is not repeated here and is added to K_intraLayerCost;
[0586] 2.3) Calculation of cost of Intra layer prediction method:
[0587] In line with the encoding side calculation method described above, the calculation is not repeated here and is added to K_intraLayerCost.
[0588] 3) When the current transform layer is in the lower layer:
[0589] 3.1) Calculation of cost of prediction method in node unit:
[0590] In line with the encoding side calculation method described above, the calculation is not repeated here and is added to K_nodeByNodeCost;
[0591] 3.2) Calculation of cost of Intra layer prediction method:
[0592] In line with the encoding side calculation method described above, the calculation is not repeated here and is added to K_intraLayerCost.
[0593] 3.3) Calculation of cost of Inter layer prediction method:
[0594] In line with the encoding side calculation method described above, the calculation is not repeated here and is added to K_interLayerCost.
[0595] Then, when the encoding cost of different prediction methods of the previous K layers: K_nodeByNodeCost, K_intraLayerCost and K_interLayerCost are obtained
[0596] The following comparison is made:
[0597] Case 1: When (K_nodeByNodeCost
[0598] layer_based_pred_mode_flag = 0; / / It is a prediction method in node unit
[0599] According to the prediction method in node unit, the transform coefficients / residual transform coefficients of all nodes of the previous K layers are calculated, quantized and entropy encoded.
[0600] Case 2: when (K_intraLayerCost < K_nodeByNodeCost) && (K_intraLayerCost < K_interLayerCost)
[0601] layer_based_pred_mode_flag = 1; / / is layer-based prediction method
[0602] layer_pred_mode = 0; / / is intra layer
[0603] Calculate transform coefficients / residual transform coefficients of all nodes in the first K layers according to the intra layer prediction method, and perform quantization and entropy coding.
[0604] Case 3: when (K_interLayerCost < K_intraLayerCost) && (K_interLayerCost < K_nodeByNodeCost)
[0605] layer_based_pred_mode_flag = 1; / / is layer-based prediction method
[0606] layer_pred_mode = 1; / / is inter layer
[0607] Calculate transform coefficients / residual transform coefficients of all nodes in the first K layers according to the inter layer prediction method, and perform quantization and entropy coding.
[0608] Finally, encode the prediction method information. At the same time, the value of K can be agreed upon at the encoding and decoding ends, using the same value, or transmitted in the code stream.
[0609] If the value of K is transmitted, encode the value of K. The encoding method can be context-based adaptive arithmetic coding, variable-length coding, exponential Golomb coding, etc.
[0610] In addition, the prediction method information of the first K layers and the prediction method information of the remaining layers are encoded and transmitted as follows:
[0611] For the first K layers: (execute the following process once)
[0612] First, encode and transmit layer_based_pred_mode_flag,
[0613] If (layer_based_pred_mode_flag == 1)
[0614] Continue to encode the transmission layer_pred_mode
[0615] For the rest of the layers: (each layer performs the following process once)
[0616] First encode the transmission layer_based_pred_mode_flag,
[0617] If (middle layer || (lower layer && allow weighting))
[0618] If (layer_based_pred_mode_flag == 1)
[0619] Continue to encode the transmission layer_pred_mode.
[0620] It should be noted that when the K value described above is 0, it means that the prediction method of each transform layer is determined by RDO comparing the RD-cost of the three prediction modes of the current layer, that is, the prediction method with the minimum cost required to encode the current layer is selected as the prediction method of the current layer.
[0621] That is, the rate-distortion optimization method (RDO: rate-distortion optimization) calculates three costs for the inter layer and intra layer prediction methods of the current layer according to the node unit, which are represented by nodeByNodeCost, interLayerCost and intraLayerCost.
[0622] 1) When the current layer is in the upper layer:
[0623] 1.1) Calculation of the cost of the node-based prediction method:
[0624] Consistent with the encoding end calculation method described above, it will not be repeated
[0625] 1.2) Calculation of the cost of the inter layer prediction method:
[0626] Consistent with the encoding end calculation method described above, it will not be repeated
[0627] 1.3) Calculation of the cost of the intra layer prediction method:
[0628] Consistent with the encoding end calculation method described above, it will not be repeated.
[0629] 2) When the current layer is in the middle layer:
[0630] 2.1) Calculation of the cost of the node-based prediction method:
[0631] In accordance with the encoding end calculation method described above, no further description is given
[0632] 2.2) Calculation of the cost of the inter layer prediction method:
[0633] In accordance with the encoding end calculation method described above, no further description is given
[0634] 2.3) Calculation of the cost of the intra layer prediction method:
[0635] In accordance with the encoding end calculation method described above, no further description is given
[0636] 3) When the current layer is in the lower layer:
[0637] 3.1) Calculation of the cost of the prediction method in node units:
[0638] In accordance with the encoding end calculation method described above, no further description is given
[0639] 3.2) Calculation of the cost of the intra layer prediction method:
[0640] In accordance with the encoding end calculation method described above, no further description is given
[0641] 3.3) Calculation of the cost of the inter layer prediction method:
[0642] In accordance with the encoding end calculation method described above, no further description is given
[0643] Then, when the intraLayerCost, interLayerCost and nodeByNodeCost of the current layer are obtained, the following comparison is made:
[0644] Case 1: When (nodeByNodeCost < intraLayerCost) && (nodeByNodeCost < interLayerCost);
[0645] layer_based_pred_mode_flag = 0; / / It is the prediction method in node units
[0646] According to the prediction method in node units, the transform coefficients / residual transform coefficients of all nodes of the current layer are calculated, quantized and entropy encoded.
[0647] Case 2: When (intraLayerCost < nodeByNodeCost) && (intraLayerCost < interLayerCost);
[0648] layer_based_pred_mode_flag = 1; / / is layer-based prediction method
[0649] layer_pred_mode = 0; / / is intra layer
[0650] Calculate transform coefficients / residual transform coefficients of all nodes in current layer according to the prediction method of intra layer, and perform quantization and entropy coding.
[0651] Case 3: when (interLayerCost < intraLayerCost) && (interLayerCost < nodeByNodeCost)
[0652] layer_based_pred_mode_flag = 1; / / is layer-based prediction method
[0653] layer_pred_mode = 1; / / is inter layer
[0654] Calculate transform coefficients / residual transform coefficients of all nodes in current layer according to the prediction method of inter layer, and perform quantization and entropy coding.
[0655] Finally, encode and transmit the prediction method information of each layer, as follows:
[0656] First, encode and transmit layer_based_pred_mode_flag,
[0657] If (middle layer || (lower layer && allow weighting))
[0658] If (layer_based_pred_mode_flag == 1)
[0659] Continue to encode and transmit layer_pred_mode.
[0660] In the scheme, since the prediction modes of the K preset transform layers are pre-agreed or correspond to the same prediction mode, the encoding end can not encode the prediction modes of the K preset transform layers or only encode the prediction mode of the K preset transform layers, thereby reducing the encoding bits, saving the encoding resources, and improving the encoding efficiency. By changing the value of K, different ways can be used to obtain the prediction mode for different transform layers in the tree-shaped data structure. For example, when K is not 0, the prediction modes of the first K transform layers are specified or are uniformly set to one prediction mode, so that the way of obtaining the prediction mode is more flexible, and the encoding complexity can be reduced.
[0661] The point cloud information decoding method provided in the embodiments of the present application can be executed by a point cloud information decoding device. In the embodiments of the present application, the point cloud information decoding method is executed by a point cloud information decoding device, and the point cloud information decoding device provided in the embodiments of the present application is described.
[0662] As shown in Figure 6 The embodiments of the present application further provide a point cloud information decoding device 600, which comprises:
[0663] A first obtaining module 601 is configured to decode a bitstream to obtain first information or to obtain the first information and second information, wherein the first information is used to indicate the prediction modes of N-K transform layers in a tree-shaped data structure of a target point cloud, the N-K transform layers being different from K preset transform layers in N transform layers, and the second information is used to indicate the same prediction mode corresponding to the K preset transform layers, N is a positive integer, 0
[0664] A second obtaining module 602 is configured to obtain the prediction modes of the N transform layers according to the first information and the pre-agreed prediction mode corresponding to the K preset transform layers, or according to the first information and the second information.
[0665] Optionally, the first obtaining module is configured to:
[0666] In a case where the prediction mode corresponding to the K preset transform layers is a pre-agreed prediction mode, the bitstream is decoded to obtain the first information.
[0667] Or, in a case where the prediction mode corresponding to the K preset transform layers is not a pre-agreed prediction mode, the bitstream is decoded to obtain the first information and the second information.
[0668] Optionally, the first obtaining module comprises:
[0669] The first obtaining sub-module is configured to decode the code stream and obtain a first target identifier corresponding to the K preset transform layers, the first target identifier comprising at least one of a first identifier and a second identifier corresponding to the K preset transform layers, wherein the first identifier corresponding to the K preset transform layers is used to indicate whether the prediction mode of the K preset transform layers is a prediction mode in units of layers, and the second identifier corresponding to the K preset transform layers is used to indicate whether the prediction mode of the K preset transform layers is an inter-layer prediction mode or an intra-layer prediction mode.
[0670] The second obtaining sub-module is configured to obtain the second information according to the prediction mode indicated by at least one of the first identifier and the second identifier.
[0671] Optionally, the second obtaining sub-module is configured to:
[0672] in a case where the first target identifier comprises the first identifier and the first identifier indicates that the prediction mode of the K preset transform layers is not a prediction mode in units of layers, obtain the second information according to the prediction mode indicated by the first identifier;
[0673] or, in a case where the first target identifier comprises the first identifier and the second identifier and the first identifier indicates that the prediction mode of the K preset transform layers is a prediction mode in units of layers, obtain the second information according to the prediction mode indicated by the second identifier;
[0674] or, in a case where the first target identifier comprises the second identifier, obtain the second information according to the prediction mode indicated by the second identifier.
[0675] Optionally, the first obtaining module comprises:
[0676] The third obtaining sub-module is configured to decode the code stream and obtain a second target identifier corresponding to a target transform layer, the second target identifier comprising at least one of a first identifier and a second identifier corresponding to the target transform layer, wherein the first identifier corresponding to the target transform layer is used to indicate whether the prediction mode of the target transform layer is a prediction mode in units of layers, and the second identifier corresponding to the target transform layer is used to indicate whether the prediction mode of the target transform layer is an inter-layer prediction mode or an intra-layer prediction mode, and the target transform layer is any one of the N-K transform layers.
[0677] The fourth obtaining sub-module is configured to obtain the first information according to at least one of the first identifier and the second identifier corresponding to the target transform layer.
[0678] Optionally, the fourth obtaining sub-module is configured to:
[0679] When the second target identifier includes the first identifier corresponding to the target transform layer, and the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is not a layer-based prediction mode, obtaining the first information according to the prediction mode indicated by the first identifier corresponding to the target transform layer;
[0680] When the second target identifier includes the first identifier corresponding to the target transform layer, and the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a layer-based prediction mode, obtaining the first information according to position information of the target transform layer in the tree data structure;
[0681] Alternatively, when the second target identifier includes a first identifier and a second identifier corresponding to the target transform layer, and the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a prediction mode in units of layers, obtaining the first information according to the prediction mode indicated by the second identifier corresponding to the target transform layer;
[0682] Alternatively, in a case where the second target identifier includes a second identifier corresponding to the target transformation layer, the first information is obtained according to a prediction mode indicated by the second identifier corresponding to the target transformation layer.
[0683] Optionally, the device of the embodiment of the present application further includes:
[0684] The third acquisition module is used to decode the code stream to obtain K; or obtain K according to a protocol agreement.
[0685] Optionally, the K preset transformation layers are the first K transformation layers in the tree data structure.
[0686] In an embodiment of the present application, since the prediction modes of the K preset transformation layers are pre-agreed or correspond to the same prediction mode, the encoding end may not need to encode the prediction modes of the K preset transformation layers or only encode the above-mentioned one prediction mode for the K preset transformation layers, thereby reducing the coding bits, saving coding resources, and improving coding efficiency.
[0687] The point cloud information encoding method provided in the embodiment of the present application can be executed by a point cloud information encoding device. In the embodiment of the present application, the point cloud information encoding method performed by the point cloud information encoding device is used as an example to illustrate the point cloud information encoding device provided in the embodiment of the present application.
[0688] The decoding device for point cloud information provided in the embodiment of the present application can achieve Figure 4 The various processes implemented by the method embodiment achieve the same technical effect and are not described here again to avoid repetition.
[0689] As Figure 7 described above, the application further provides a point cloud information encoding device 700, comprising:
[0690] a construction module 701, configured to construct a tree-shaped data structure of a target point cloud, the tree-shaped data structure comprising N transform layers, N being a positive integer;
[0691] a determination module 702, configured to determine a prediction mode of the N transform layers of the tree-shaped data structure, wherein the prediction mode of K preset transform layers of the tree-shaped data structure is a pre-agreed prediction mode, or the prediction mode of the K preset transform layers is a same prediction mode, 0
[0692] a fourth acquisition module 703, configured to encode the prediction mode of the transform layers of the tree-shaped data structure to obtain a code stream, the code stream comprising first information or comprising the first information and second information;
[0693] wherein the first information is used to indicate the prediction mode of N-K transform layers other than the K preset transform layers in the N transform layers; and the second information is used to indicate the same prediction mode.
[0694] Optionally, the fourth acquisition module is configured to:
[0695] encode a first target identifier corresponding to the K preset transform layers to obtain the second information, the first target identifier comprising at least one of a first identifier and a second identifier corresponding to the K preset transform layers, wherein the first identifier corresponding to the K preset transform layers is used to indicate whether the prediction mode of the K preset transform layers is a layer-based prediction mode, and the second identifier corresponding to the K preset transform layers is used to indicate whether the prediction mode of the K preset transform layers is an inter-layer prediction mode or an intra-layer prediction mode.
[0696] Optionally, the fourth acquisition module is configured to:
[0697] in a case where the first identifier corresponding to the K preset transform layers indicates that the prediction mode of the K preset transform layers is not a layer-based prediction mode, the second information is obtained according to the encoding information of the first identifier;
[0698] or, in a case where the first identifier corresponding to the K preset transform layers indicates that the prediction mode of the K preset transform layers is a layer-based prediction mode, the second information is obtained according to the encoding information of the first identifier and the encoding information of the second identifier, or the second information is obtained according to the encoding information of the second identifier.
[0699] Optionally, the fourth acquisition module is configured to:
[0700] encode the second target identifier corresponding to the target transform layer to obtain the first information, the second target identifier including at least one of a first identifier and a second identifier corresponding to the target transform layer, the first identifier corresponding to the target transform layer being used to indicate whether the prediction mode of the target transform layer is a prediction mode in a layer unit, the second identifier corresponding to the target transform layer being used to indicate whether the prediction mode of the target transform layer is an inter-layer prediction mode or an intra-layer prediction mode, the target transform layer being any one of the N-K transform layers.
[0701] Optionally, the fourth obtaining module is configured to:
[0702] in a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is not a prediction mode in a layer unit, obtaining the first information according to the encoding information of the first identifier corresponding to the target transform layer;
[0703] or, in a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a prediction mode in a layer unit, obtaining the first information according to the encoding information of the first identifier and the second identifier corresponding to the target transform layer, or obtaining the first information according to the encoding information of the second identifier corresponding to the target transform layer;
[0704] or, in a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a prediction mode in a layer unit, if the target transform layer is in the upper layer part of the tree-shaped data structure, or the target transform layer is in the lower layer part of the tree-shaped data structure and the lower layer part does not allow weighted prediction, obtaining the first information according to the encoding information of the first identifier, if the target transform layer is in the middle layer part of the tree-shaped data structure, obtaining the first information according to the encoding information of the first identifier and the second identifier.
[0705] Optionally, the apparatus of the embodiment of the present application further includes:
[0706] the fifth obtaining module is configured to obtain a target rate-distortion optimization cost of each preset prediction mode in at least two preset prediction modes, the target rate-distortion optimization cost of the preset prediction mode being obtained based on a sum of rate-distortion optimization costs of the K preset transform layers in the preset prediction mode;
[0707] the first selecting module is configured to select, in the at least two preset prediction modes, a preset prediction mode with a minimum target rate-distortion optimization cost as the prediction mode of the K preset transform layers.
[0708] Optionally, the apparatus of the embodiment of the present application further comprises:
[0709] The sixth obtaining module is configured to obtain rate-distortion optimization costs of each of the N-K transform layers under at least two preset prediction modes.
[0710] The second selecting module is configured to select, from the at least two preset prediction modes, a preset prediction mode with the minimum rate-distortion optimization cost as the prediction mode of each of the transform layers.
[0711] Optionally, the code stream further comprises coding information of the K.
[0712] In the embodiment of the present application, since the prediction modes of the K preset transform layers are pre-agreed or correspond to the same prediction mode, the encoding end can not need to encode the prediction modes of the K preset transform layers or only encode the above-mentioned one prediction mode for the K preset transform layers, thereby being able to reduce coding bits, save coding resources, and improve coding efficiency.
[0713] The encoding apparatus for point cloud information provided in the embodiment of the present application can implement each process of the method embodiment and achieve the same technical effects, and for the sake of avoiding repetition, details are not described herein. Figure 5
[0714] As shown in Figure 8 , the embodiment of the present application further provides an electronic device 800 comprising a processor 801 and a memory 802, wherein the memory 802 has stored programs or instructions executable on the processor 801. For example, when the electronic device 800 is an encoding end device, the programs or instructions are executed by the processor 801 to implement each step of the above-mentioned encoding method for point cloud information, and achieve the same technical effects. When the electronic device 800 is a decoding end device, the programs or instructions are executed by the processor 801 to implement each step of the above-mentioned decoding method for point cloud information, and achieve the same technical effects. For the sake of avoiding repetition, details are not described herein. Optionally, the memory 802 can be the memory 102 or the memory 113 in the embodiment shown in Figure 1 , Figure 1 , Figure 2a , Figure 2b , Figure 3a and Figure 3b the embodiment shown in
[0715] The embodiment of the present application further provides an electronic device comprising a memory configured to store video data and a processing circuit configured to implement each step of the above-mentioned encoding method or decoding method for point cloud information. Optionally, the memory can be the memory 102 or the memory 113 in the embodiment shown in Figure 1 In the embodiment shown, the memory 102 or the memory 113, the processing circuit can implement Figure 1 、 Figure 2a 、 Figure 2b 、 Figure 3a as well as Figure 3b Functionality of the encoder 200 or decoder 300 in the illustrated embodiment.
[0716] The embodiment of the present application also provides an electronic device, including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the following Figure 4 or Figure 5 The device embodiment corresponds to the above method embodiment, and each implementation process and implementation method of the above method embodiment are applicable to the terminal embodiment and can achieve the same technical effect.
[0717] The electronic device may be a terminal, or may be other devices other than a terminal, such as a server, a network attached storage (NAS), etc.
[0718] The terminal can be a mobile phone, a tablet personal computer, a laptop computer, a notebook computer, a personal digital assistant (PDA), a palm computer, a netbook, an ultra-mobile personal computer (UMPC), a mobile Internet device (MID), an augmented reality (AR) device, a virtual reality (VR) device, a mixed reality (MR) device, a robot, a wearable device, a flight vehicle, a vehicle user equipment (VUE), a shipboard device, a pedestrian user equipment (PUE), a smart home (a home device with a wireless communication function, such as a refrigerator, a television, a washing machine, or furniture), a game console, a personal computer (PC), a teller machine, or a self-service machine, and the like. The wearable device includes a smart watch, a smart bracelet, a smart earphone, smart glasses, smart jewelry (a smart bracelet, a smart necklace, a smart ring, a smart necklace, a smart anklet, a smart necklace, and the like), a smart wristband, smart clothing, and the like. The vehicle-mounted device can also be referred to as a vehicle-mounted terminal, a vehicle-mounted controller, a vehicle-mounted module, a vehicle-mounted component, a vehicle-mounted chip, or a vehicle-mounted unit, and the like. It should be noted that the specific type of the terminal is not limited in the embodiments of the present application.
[0719] The server can be a stand-alone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server. The cloud server can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), or cloud computing services based on big data and artificial intelligence platforms.
[0720] For example, the electronic device can include, but is not limited to Figure 1 The type of the source device 100 or the destination device 110 is shown.
[0721] Taking the electronic device as an example, Figure 9 A hardware structure schematic diagram of a terminal for implementing the embodiments of the present application.
[0722] The terminal 900 includes, but is not limited to, at least part of components such as a radio frequency unit 901, a network module 902, an audio output unit 903, an input unit 904, a sensor 905, a display unit 906, a user input unit 907, an interface unit 908, a memory 909, and a processor 910.
[0723] Those skilled in the art can understand that the terminal 900 can further include a power supply (such as a battery) for supplying power to each component, and the power supply can be logically connected to the processor 910 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 9 The terminal structure shown in the figure does not constitute a limitation on the terminal, and the terminal can include more or fewer components than the figure, or combine certain components, or different component arrangements, which are not described here.
[0724] It should be understood that in the embodiments of the present application, the input unit 904 can include a graphics processing unit (GPU) 9041 and a microphone 9042. The graphics processor 9041 processes image data of a still picture or a video obtained by an image acquisition device (such as a camera) in a video capture mode or an image capture mode, or can process obtained point cloud data. The display unit 906 can include a display panel 9061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 907 includes at least one of a touch panel 9071 and other input devices 9072. The touch panel 9071 is also called a touch screen. The touch panel 9071 can include two parts of a touch detection device and a touch controller. The other input devices 9072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), trackballs, mice, joysticks, etc., which are not described here.
[0725] In the embodiments of the present application, the radio frequency unit 901 can transmit downlink data from the network side device to the processor 910 for processing after receiving the downlink data. In addition, the radio frequency unit 901 can send uplink data to the network side device. Generally, the radio frequency unit 901 includes, but is not limited to, an antenna, an amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, etc.
[0726] The memory 909 can be used to store software programs or instructions and various data. The memory 909 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), and the like. In addition, the memory 909 can include a volatile memory or a non-volatile memory. The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 909 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.
[0727] The processor 910 can include one or more processing units; optionally, the processor 910 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 910.
[0728] The processor 910 is configured to perform decoding processing on the code stream to obtain first information or to obtain the first information and second information, wherein the first information is used to indicate a prediction mode of N-K transform layers of a tree-shaped data structure of a target point cloud, the N-K transform layers being different from K preset transform layers in N transform layers, the second information is used to indicate a same prediction mode corresponding to the K preset transform layers, N is a positive integer, 0
[0729] The first information is used to indicate the prediction mode of the N-K transform layers different from the K preset transform layers in the N transform layers, and the second information is used to indicate the same prediction mode.
[0730] In the above scheme, since the prediction modes of the K preset transform layers are predetermined or correspond to the same prediction mode, the encoding end can not need to encode the prediction modes of the K preset transform layers or only encode the above-mentioned one prediction mode for the K preset transform layers, thereby being able to reduce the encoding bits, save the encoding resources, and improve the encoding efficiency.
[0731] It can be understood that the implementation processes of the implementation manners mentioned in the embodiment can refer to the related descriptions of the encoding method or the decoding method of the point cloud information, and achieve the same or corresponding technical effects. To avoid repetition, they will not be described here again.
[0732] The embodiment of the present application further provides a readable storage medium, the readable storage medium stores a program or instructions, the program or instructions are executed by a processor to realize each process of the encoding method or the decoding method of the point cloud information, and the same technical effects can be achieved. To avoid repetition, they will not be described here again.
[0733] The processor is the processor in the terminal in the above embodiments. The readable storage medium includes a computer readable storage medium, such as a ROM, a RAM, a magnetic disk, or an optical disk, etc. In some examples, the readable storage medium can be a non-transitory readable storage medium.
[0734] The chip provided by the embodiments of the present application includes a processor and a communication interface, the communication interface is coupled with the processor, the processor is used to run programs or instructions, realizes each process of the encoding method or the decoding method of the point cloud information, and can achieve the same technical effects. To avoid repetition, details are not described here.
[0735] It should be understood that the chip mentioned in the embodiments of the present application can include a system-on-chip (also known as a system chip, a chip system, or a system-on-chip), and can also include a separate display chip, etc.
[0736] The embodiments of the present application further provide a computer program / program product stored in a storage medium, which is executed by at least one processor to realize each process of the encoding method or the decoding method of the point cloud information, and can achieve the same technical effects. To avoid repetition, details are not described here.
[0737] The embodiments of the present application further provide a point cloud information encoding and decoding system, including an encoding end device and a decoding end device. The encoding end device can be used to execute the steps of the encoding method of the point cloud information as described above. The decoding end device can be used to execute the steps of the decoding method of the point cloud information as described above.
[0738] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles, or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or further include elements inherent to such processes, methods, articles, or devices. Without more limitations, the element defined by the statement "including a" does not exclude the presence of additional identical elements in the process, method, article, or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to the order of performing the functions shown or discussed, and can also include performing the functions in a substantially simultaneous manner or in a reverse order, for example, the described method can be performed in an order different from the described order, and various steps can also be added, omitted, or combined. In addition, the features described with reference to some examples can be combined in other examples.
[0739] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned example methods can be realized by means of a computer software product and a general hardware platform as necessary, and of course can also be realized by hardware. The computer software product is stored in a storage medium (such as a ROM, a RAM, a magnetic disc, an optical disc, etc.), and includes a plurality of instructions for enabling a terminal or a network side device to execute the method described in each embodiment of the present application.
[0740] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the specific embodiments described above, and the specific embodiments described above are merely illustrative rather than limiting. Those skilled in the art can make many forms of embodiments under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims, and these embodiments all belong to the protection of the present application.
Claims
1. A method of decoding point cloud information, characterized by, The decoding end decodes the code stream to obtain first information or to obtain the first information and second information; wherein the first information is used to indicate a prediction mode of N-K transform layers of a tree-shaped data structure of a target point cloud, except for K preset transform layers, the second information is used to indicate a same prediction mode corresponding to the K preset transform layers, N is a positive integer, 0 The decoding end obtains the prediction mode of the N transform layers according to the first information and the pre-agreed prediction mode corresponding to the K preset transform layers, or according to the first information and the second information. The decoding end decodes the code stream to obtain first information or to obtain the first information and second information, comprising:
2. The method of claim 1, wherein, In a case where the prediction mode corresponding to the K preset transform layers is a pre-agreed prediction mode, decoding the code stream to obtain the first information; Or, in a case where the prediction mode corresponding to the K preset transform layers is not a pre-agreed prediction mode, decoding the code stream to obtain the first information and the second information. Decoding the code stream to obtain the second information, comprising:
3. The method according to claim 1 or 2, characterized in that, Decoding the code stream to obtain a first target identifier corresponding to the K preset transform layers, the first target identifier comprising at least one of a first identifier and a second identifier corresponding to the K preset transform layers, wherein the first identifier corresponding to the K preset transform layers is used to indicate whether the prediction mode of the K preset transform layers is a layer-based prediction mode, and the second identifier corresponding to the K preset transform layers is used to indicate whether the prediction mode of the K preset transform layers is an inter-layer prediction mode or an intra-layer prediction mode; Obtaining the second information according to the prediction mode indicated by at least one of the first identifier and the second identifier. The obtaining the second information according to the prediction mode indicated by at least one of the first identifier and the second identifier, comprising:
4. The method of claim 3, wherein, In a case where the first target identifier comprises the first identifier and the first identifier indicates that the prediction mode of the K preset transform layers is not a layer-based prediction mode, obtaining the second information according to the prediction mode indicated by the first identifier; Or, in a case where the first target identifier comprises the first identifier and the second identifier, and the first identifier indicates that the prediction mode of the K preset transform layers is a layer-based prediction mode, obtaining the second information according to the prediction mode indicated by the second identifier; Or, in a case where the first target identifier comprises the second identifier, obtaining the second information according to the prediction mode indicated by the second identifier. Decoding the code stream to obtain the first information, comprising:
5. The method according to any one of claims 1 to 4, characterized in that, decoding the code stream to obtain a second target identifier corresponding to a target transform layer, the second target identifier comprising at least one of a first identifier and a second identifier corresponding to the target transform layer, the first identifier corresponding to the target transform layer being used to indicate whether a prediction mode of the target transform layer is a layer-based prediction mode, the second identifier corresponding to the target transform layer being used to indicate whether the prediction mode of the target transform layer is an inter-layer prediction mode or an intra-layer prediction mode, the target transform layer being any one of N-K transform layers; obtaining the first information according to at least one of the first identifier and the second identifier corresponding to the target transform layer.
6. The method of claim 5, wherein, obtaining the first information according to at least one of the first identifier and the second identifier corresponding to the target transform layer, comprising: in a case where the second target identifier comprises the first identifier corresponding to the target transform layer and the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is not the layer-based prediction mode, obtaining the first information according to the prediction mode indicated by the first identifier corresponding to the target transform layer; in a case where the second target identifier comprises the first identifier corresponding to the target transform layer and the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is the layer-based prediction mode, obtaining the first information according to position information of the target transform layer in the tree-shaped data structure; or, in a case where the second target identifier comprises the first identifier and the second identifier corresponding to the target transform layer and the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is the layer-based prediction mode, obtaining the first information according to the prediction mode indicated by the second identifier corresponding to the target transform layer; or, in a case where the second target identifier comprises the second identifier corresponding to the target transform layer, obtaining the first information according to the prediction mode indicated by the second identifier corresponding to the target transform layer.
7. The method of claim 6, wherein, obtaining the first information according to the position information of the target transform layer in the tree-shaped data structure, comprising: in a case where the target transform layer is in an upper layer part of the tree-shaped data structure, determining that the prediction mode of the target transform layer is the intra-layer prediction mode; in a case where the target transform layer is in a lower layer part of the tree-shaped data structure and the lower layer part does not allow weighted prediction, determining that the prediction mode of the target transform layer is the inter-layer prediction mode.
8. The method according to any one of claims 1 to 7, characterized in that, further comprising: decoding the code stream to obtain the K; or, obtaining the K according to a protocol agreement.
9. The method according to any one of claims 1 to 8, characterized in that, the K preset transform layers are the first K transform layers in the tree-shaped data structure.
10. A method of encoding point cloud information, characterized by, comprising: constructing, at an encoding end, a tree-shaped data structure of a target point cloud, the tree-shaped data structure comprising N transform layers, N being a positive integer; The encoding end determines the prediction modes of N transform layers of the tree-shaped data structure, wherein the prediction modes of K preset transform layers of the tree-shaped data structure are a pre-agreed prediction mode, or the prediction modes of the K preset transform layers are a same prediction mode, 0 The encoding end encodes the prediction modes of the transform layers of the tree-shaped data structure to obtain a code stream, wherein the code stream comprises first information or comprises the first information and second information. The first information is used to indicate the prediction modes of N-K transform layers among the N transform layers except the K preset transform layers; and the second information is used to indicate the same prediction mode.
11. The method of claim 10, wherein, The encoding end encodes the prediction modes of the transform layers of the tree-shaped data structure, comprising: The encoding end encodes the same prediction mode and the prediction mode of each transform layer among the N-K transform layers in the case that the prediction modes of the K preset transform layers are the same prediction mode. Or, the encoding end encodes the prediction mode of each transform layer among the N-K transform layers in the case that the prediction modes of the K preset transform layers are the pre-agreed prediction mode.
12. The method according to claim 10 or 11, characterized in that, The encoding end encodes the prediction modes of the transform layers of the tree-shaped data structure to obtain the second information, comprising: The encoding end encodes the first target identifier corresponding to the K preset transform layers to obtain the second information, wherein the first target identifier comprises at least one of a first identifier and a second identifier corresponding to the K preset transform layers, wherein the first identifier corresponding to the K preset transform layers is used to indicate whether the prediction modes of the K preset transform layers are layer-based prediction modes, and the second identifier corresponding to the K preset transform layers is used to indicate whether the prediction modes of the K preset transform layers are inter-layer prediction modes or intra-layer prediction modes.
13. The method of claim 12, wherein, The encoding end encodes the first target identifier corresponding to the K preset transform layers to obtain the second information, comprising: In the case that the first identifier corresponding to the K preset transform layers indicates that the prediction modes of the K preset transform layers are not layer-based prediction modes, the second information is obtained according to the encoding information of the first identifier; Or, in the case that the first identifier corresponding to the K preset transform layers indicates that the prediction modes of the K preset transform layers are layer-based prediction modes, the second information is obtained according to the encoding information of the first identifier and the encoding information of the second identifier, or the second information is obtained according to the encoding information of the second identifier.
14. The method according to any one of claims 10 to 13, characterized in that, The encoding end encodes the prediction modes of the transform layers of the tree-shaped data structure to obtain the first information, comprising: The second target identifier corresponding to the target transform layer is encoded to obtain the first information, the second target identifier includes at least one of a first identifier and a second identifier corresponding to the target transform layer, the first identifier corresponding to the target transform layer is used to indicate whether the prediction mode of the target transform layer is a prediction mode in a layer unit, and the second identifier corresponding to the target transform layer is used to indicate whether the prediction mode of the target transform layer is an inter-layer prediction mode or an intra-layer prediction mode, and the target transform layer is any one of N-K transform layers.
15. The method of claim 14, wherein, The encoding of the second target identifier corresponding to the target transform layer to obtain the first information includes: In a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is not a prediction mode in a layer unit, the first information is obtained according to the encoding information of the first identifier corresponding to the target transform layer; Or, in a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a prediction mode in a layer unit, the first information is obtained according to the encoding information of the first identifier and the second identifier corresponding to the target transform layer, or the first information is obtained according to the encoding information of the second identifier corresponding to the target transform layer; Or, in a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a prediction mode in a layer unit, if the target transform layer is in an upper layer part of the tree-shaped data structure, or the target transform layer is in a lower layer part of the tree-shaped data structure and the lower layer part does not allow weighted prediction, the first information is obtained according to the encoding information of the first identifier, and if the target transform layer is in a middle layer part of the tree-shaped data structure, the first information is obtained according to the encoding information of the first identifier and the second identifier.
16. The method according to any one of claims 10 to 15, characterized in that, Further comprising: Obtaining a target rate-distortion optimization cost of each preset prediction mode in at least two preset prediction modes, the target rate-distortion optimization cost of the preset prediction mode being obtained based on a sum of rate-distortion optimization costs of K preset transform layers in the preset prediction mode; In the at least two preset prediction modes, a preset prediction mode with a minimum target rate-distortion optimization cost is selected as a prediction mode of the K preset transform layers.
17. The method according to any one of claims 10 to 16, characterized in that, Further comprising: Obtaining a rate-distortion optimization cost of each transform layer in the N-K transform layers in at least two preset prediction modes; In the at least two preset prediction modes, a preset prediction mode with a minimum rate-distortion optimization cost is selected as a prediction mode of each transform layer.
18. The method according to any one of claims 10 to 17, characterized in that, The code stream further includes encoding information of the K.
19. An apparatus for decoding point cloud information, the apparatus comprising: Comprising: The first obtaining module is configured to decode a code stream to obtain first information or to obtain the first information and second information, the first information is used to indicate a prediction mode of N-K transform layers except K preset transform layers in a tree-shaped data structure of a target point cloud, and the second information is used to indicate a same prediction mode corresponding to the K preset transform layers, N is a positive integer, 0 The second obtaining module is configured to obtain the prediction mode of the N transform layers according to the first information and the K preset transform layers corresponding to the pre-agreed prediction mode, or according to the first information and the second information.
20. The apparatus of claim 19, wherein, The first obtaining module is configured to: decode the bitstream to obtain the first information in a case where the prediction mode corresponding to the K preset transform layers is a pre-agreed prediction mode; or decode the bitstream to obtain the first information and the second information in a case where the prediction mode corresponding to the K preset transform layers is not a pre-agreed prediction mode.
21. The apparatus of claim 19 or 20, wherein, The first obtaining module includes: a first obtaining submodule configured to decode the bitstream to obtain a first target identifier corresponding to the K preset transform layers, the first target identifier including at least one of a first identifier and a second identifier corresponding to the K preset transform layers, wherein the first identifier corresponding to the K preset transform layers is used to indicate whether the prediction mode of the K preset transform layers is a layer-based prediction mode, and the second identifier corresponding to the K preset transform layers is used to indicate whether the prediction mode of the K preset transform layers is an inter-layer prediction mode or an intra-layer prediction mode; a second obtaining submodule configured to obtain the second information according to the prediction mode indicated by at least one of the first identifier and the second identifier.
22. The apparatus of claim 21, wherein, The second obtaining submodule is configured to: obtain the second information according to the prediction mode indicated by the first identifier in a case where the first target identifier includes the first identifier and the first identifier indicates that the prediction mode of the K preset transform layers is not a layer-based prediction mode; or obtain the second information according to the prediction mode indicated by the second identifier in a case where the first target identifier includes the first identifier and the second identifier and the first identifier indicates that the prediction mode of the K preset transform layers is a layer-based prediction mode; or obtain the second information according to the prediction mode indicated by the second identifier in a case where the first target identifier includes the second identifier.
23. The apparatus of any one of claims 19 to 22, wherein, The first obtaining module includes: a third obtaining submodule configured to decode the bitstream to obtain a second target identifier corresponding to a target transform layer, the second target identifier including at least one of a first identifier and a second identifier corresponding to the target transform layer, wherein the first identifier corresponding to the target transform layer is used to indicate whether the prediction mode of the target transform layer is a layer-based prediction mode, and the second identifier corresponding to the target transform layer is used to indicate whether the prediction mode of the target transform layer is an inter-layer prediction mode or an intra-layer prediction mode, and the target transform layer is any one of the N-K transform layers; a fourth obtaining submodule configured to obtain the first information according to at least one of the first identifier and the second identifier corresponding to the target transform layer.
24. The apparatus of claim 23, wherein, The fourth obtaining submodule is configured to: in a case where the second target identifier comprises the first identifier corresponding to the target transform layer and the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is not the prediction mode in a layer unit, obtaining the first information according to the prediction mode indicated by the first identifier corresponding to the target transform layer; in a case where the second target identifier comprises the first identifier corresponding to the target transform layer and the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is the prediction mode in a layer unit, obtaining the first information according to the position information of the target transform layer in the tree-shaped data structure; or, in a case where the second target identifier comprises the first identifier and the second identifier corresponding to the target transform layer and the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is the prediction mode in a layer unit, obtaining the first information according to the prediction mode indicated by the second identifier corresponding to the target transform layer; or, in a case where the second target identifier comprises the second identifier corresponding to the target transform layer, obtaining the first information according to the prediction mode indicated by the second identifier corresponding to the target transform layer.
25. The apparatus of any one of claims 19 to 24, wherein, Further comprising: a third obtaining module configured to decode the code stream to obtain the K; or obtain the K according to a protocol agreement.
26. An apparatus for encoding point cloud information, the apparatus comprising: Comprising: a constructing module configured to construct a tree-shaped data structure of a target point cloud, the tree-shaped data structure comprising N transform layers, N being a positive integer; a determining module configured to determine prediction modes of the N transform layers of the tree-shaped data structure, wherein the prediction modes of K preset transform layers of the tree-shaped data structure are a pre-agreed prediction mode, or the prediction modes of the K preset transform layers are a same prediction mode, 0 < K ≤ N, and K is an integer; a fourth obtaining module configured to encode the prediction modes of the transform layers of the tree-shaped data structure to obtain a code stream, the code stream comprising first information or comprising first information and second information; wherein the first information is used to indicate the prediction modes of N-K transform layers other than the K preset transform layers in the N transform layers; and the second information is used to indicate the same prediction mode.
27. The apparatus of claim 26, wherein, The fourth obtaining module is configured to: in a case where the prediction modes of the K preset transform layers are the same prediction mode, encode the same prediction mode and the prediction mode of each of the N-K transform layers; or, in a case where the prediction modes of the K preset transform layers are the pre-agreed prediction mode, encode the prediction mode of each of the N-K transform layers.
28. The apparatus of claim 26 or 27, wherein, The fourth obtaining module is configured to: The first target identifier corresponding to the K preset transform layers is encoded to obtain second information, the first target identifier includes at least one of a first identifier and a second identifier corresponding to the K preset transform layers, wherein the first identifier corresponding to the K preset transform layers is used to indicate whether the prediction mode of the K preset transform layers is a prediction mode in a layer unit, and the second identifier corresponding to the K preset transform layers is used to indicate whether the prediction mode of the K preset transform layers is an inter-layer prediction mode or an intra-layer prediction mode.
29. The apparatus of claim 28, wherein, The fourth acquisition module is used to: In a case where the first identifier corresponding to the K preset transform layers indicates that the prediction mode of the K preset transform layers is not a prediction mode in a layer unit, the second information is obtained according to the encoding information of the first identifier; Or, in a case where the first identifier corresponding to the K preset transform layers indicates that the prediction mode of the K preset transform layers is a prediction mode in a layer unit, the second information is obtained according to the encoding information of the first identifier and the encoding information of the second identifier, or the second information is obtained according to the encoding information of the second identifier.
30. The apparatus of any one of claims 26 to 29, wherein, The fourth acquisition module is used to: The second target identifier corresponding to the target transform layer is encoded to obtain the first information, the second target identifier includes at least one of a first identifier and a second identifier corresponding to the target transform layer, the first identifier corresponding to the target transform layer is used to indicate whether the prediction mode of the target transform layer is a prediction mode in a layer unit, and the second identifier corresponding to the target transform layer is used to indicate whether the prediction mode of the target transform layer is an inter-layer prediction mode or an intra-layer prediction mode, and the target transform layer is any one of the N-K transform layers.
31. The apparatus of claim 30, wherein, The fourth acquisition module is used to: In a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is not a prediction mode in a layer unit, the first information is obtained according to the encoding information of the first identifier corresponding to the target transform layer; Or, in a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a prediction mode in a layer unit, the first information is obtained according to the encoding information of the first identifier and the second identifier corresponding to the target transform layer, or the first information is obtained according to the encoding information of the second identifier corresponding to the target transform layer; Or, in a case where the first identifier corresponding to the target transform layer indicates that the prediction mode of the target transform layer is a prediction mode in a layer unit, if the target transform layer is in the upper layer part of the tree-shaped data structure, or the target transform layer is in the lower layer part of the tree-shaped data structure and the lower layer part does not allow weighted prediction, the first information is obtained according to the encoding information of the first identifier, and if the target transform layer is in the middle layer part of the tree-shaped data structure, the first information is obtained according to the encoding information of the first identifier and the second identifier.
32. The apparatus of any one of claims 26 to 31, wherein, Further comprising: The fifth obtaining module is configured to obtain a target rate-distortion optimization cost of each of at least two preset prediction modes, wherein the target rate-distortion optimization cost of the preset prediction mode is obtained based on a sum of rate-distortion optimization costs of the K preset transform layers in the preset prediction mode. The first selecting module is configured to select, from the at least two preset prediction modes, a preset prediction mode with a minimum target rate-distortion optimization cost as the prediction mode of the K preset transform layers.
33. The apparatus of any one of claims 26-32, wherein, Further comprising: The sixth obtaining module is configured to obtain a rate-distortion optimization cost of each of the N-K transform layers in at least two preset prediction modes. The second selecting module is configured to select, from the at least two preset prediction modes, a preset prediction mode with a minimum rate-distortion optimization cost as the prediction mode of each of the transform layers.
34. The apparatus of any one of claims 26-33, wherein, The code stream further comprises coding information of the K.
35. An electronic device, comprising: The processor and the memory, the memory stores programs or instructions that can be run on the processor, and the programs or instructions are executed by the processor to implement the steps of the point cloud information decoding method of any one of claims 1-9, or implement the steps of the point cloud information encoding method of any one of claims 10-18.
36. A readable storage medium characterized by, The readable storage medium stores programs or instructions, and the programs or instructions are executed by the processor to implement the steps of the point cloud information decoding method of any one of claims 1-9, or implement the steps of the point cloud information encoding method of any one of claims 10-18.
37. A chip, characterized by The chip comprises a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the steps of the point cloud information decoding method of any one of claims 1-9, or implement the steps of the point cloud information encoding method of any one of claims 10-18.
38. A computer program product, characterised in that, The computer instructions are executed by the processor to implement the steps of the point cloud information decoding method of any one of claims 1-9, or implement the steps of the point cloud information encoding method of any one of claims 10-18.