Coding and decoding method and related equipment

Selecting inter prediction mode or non-prediction mode through rate distortion optimization solves the problem of poor point cloud encoding effect when intra prediction mode is not turned on, and improves encoding quality and efficiency.

CN120343259APending Publication Date: 2025-07-18VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410072679.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, in point cloud encoding, if the intra prediction mode is not turned on, directly adopting the non-prediction mode may lead to poor prediction encoding effect and affect the encoding quality.

Method used

Select the prediction mode of the current node from the inter prediction mode and the non-prediction mode through rate distortion optimization (RDO), and determine the appropriate prediction mode in combination with entropy decoding to encode and decode point cloud data.

Benefits of technology

It improves the prediction effect of point cloud encoding and decoding, improves encoding quality and efficiency, and reduces computing complexity and memory overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343259A_ABST
    Figure CN120343259A_ABST
Patent Text Reader

Abstract

The invention discloses a coding and decoding method and related equipment, and belongs to the technical field of coding and decoding, and the coding method comprises the steps that a coding end determines a transformation tree structure of a to-be-coded point cloud based on geometric reconstruction information, and the transformation tree structure comprises at least one first node layer; determining a node with the same geometric position as the to-be-coded node in a reference frame of a current node in the at least one first node layer as a reference frame node; if the current node is completely matched with the reference frame node and the intra-frame prediction mode is not started, selecting a prediction mode adopted by the current node from an inter-frame prediction mode and a non-prediction mode through rate-distortion optimization; predicting and transforming the to-be-coded node according to the prediction mode to obtain an attribute transformation coefficient; the node to be coded is a child node of the current node; and coding the attribute transformation coefficient to obtain a code stream. According to the embodiment of the invention, a more appropriate prediction mode can be selected for the current node, so that the prediction coding effect is improved, and the coding quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of encoding and decoding technologies, and particularly relates to an encoding and decoding method and related devices. Background Art

[0002] With the continuous development of point cloud technology, the compression and encoding of point cloud data have become an important research issue. At present, both the Audio Video coding Standard Workgroup of China (AVS) and the Moving Picture Experts Group (MPEG) in the International Organization for Standardization are formulating standards for point cloud coding, such as Geometry-based Point Cloud Compression (G-PCC). How to further improve the performance of point cloud encoding and decoding is an urgent problem to be solved. Summary of the Invention

[0003] Embodiments of this application provide an encoding and decoding method and related devices, which can select a more suitable prediction mode for the current node, thereby facilitating the improvement of the prediction coding effect and the encoding quality.

[0004] In a first aspect, an encoding method is provided, which is executed by an encoding end. The method includes:

[0005] The encoding end determines a transform tree structure of the point cloud to be encoded based on geometric reconstruction information, and the transform tree structure includes at least one first node layer;

[0006] A node with the same geometric position as the current node in the reference frame of the current node in the at least one first node layer is determined as a reference frame node;

[0007] If the current node completely matches the reference frame node and the intra prediction mode is not enabled, the prediction mode adopted by the current node is selected from the inter prediction mode and the non-prediction mode through rate-distortion optimization;

[0008] The node to be encoded is predicted and transformed according to the prediction mode to obtain attribute transform coefficients; the node to be encoded is a child node of the current node;

[0009] The attribute transform coefficients are encoded to obtain a bitstream.

[0010] In a second aspect, a decoding method is provided, which is executed by a decoding end. The method includes:

[0011] The decoding end analyzes the bitstream to obtain the reconstructed value of the transform coefficient of the point cloud to be decoded;

[0012] Determine the transform tree structure of the to-be-decoded point cloud based on geometric reconstruction information; the transform tree structure includes at least one first node layer;

[0013] Determine a node with the same geometric position as the current node in the reference frame of the current node in the at least one first node layer as the reference frame node;

[0014] If the current node completely matches the reference frame node and the intra prediction mode is not enabled, determine the prediction mode adopted by the current node through entropy decoding; wherein, the prediction mode is an inter prediction mode or a non-prediction mode;

[0015] Inverse-transform the reconstructed value of the transform coefficient of the to-be-decoded node according to the prediction mode to obtain a reconstructed attribute value; wherein, the to-be-decoded node is a child node of the current node.

[0016] In a third aspect, an encoding device is provided, including:

[0017] A determination module, configured to determine the transform tree structure of the to-be-encoded point cloud based on geometric reconstruction information, where the transform tree structure includes at least one first node layer;

[0018] The determination module is further configured to determine a node with the same geometric position as the current node in the reference frame of the current node in the at least one first node layer as the reference frame node;

[0019] The determination module is further configured to, if the current node completely matches the reference frame node and the intra prediction mode is not enabled, select the prediction mode adopted by the current node from the inter prediction mode and the non-prediction mode through rate-distortion optimization;

[0020] A prediction and transform module, configured to perform prediction and transform on the to-be-encoded node according to the prediction mode to obtain attribute transform coefficients; the to-be-encoded node is a child node of the current node;

[0021] An encoding module, configured to perform encoding processing on the attribute transform coefficients to obtain a bitstream.

[0022] In a fourth aspect, a decoding device is provided, including:

[0023] A parsing module, configured to parse the bitstream to obtain the reconstructed value of the transform coefficient of the to-be-decoded point cloud;

[0024] A determination module, configured to determine the transform tree structure of the to-be-decoded point cloud based on geometric reconstruction information; the transform tree structure includes at least one first node layer;

[0025] The determining module is further configured to determine, in a reference frame of a current node in the at least one first node layer, a node having the same geometric position as that of the current node as a reference frame node;

[0026] The decoding module is configured to, if the current node exactly matches the reference frame node and the intra prediction mode is not enabled, determine, by entropy decoding, a prediction mode adopted by the current node; wherein the prediction mode is an inter prediction mode or a non-prediction mode;

[0027] The inverse transformation module is configured to perform an inverse transformation on a reconstructed value of transform coefficients of a node to be decoded according to the prediction mode, to obtain a reconstructed attribute value; wherein the node to be decoded is a child node of the current node.

[0028] In a fifth aspect, an electronic device is provided. The terminal includes a processor and a memory. The memory stores a program or instructions that can run on the processor. When the program or instructions are executed by the processor, the steps of the method described in the first aspect are implemented, or the steps of the method described in the second aspect are implemented.

[0029] In a sixth aspect, an electronic device is provided, including: a memory configured to store video data, and a processing circuit configured to implement the steps of the method described in the first aspect, or implement the steps of the method described in the second aspect.

[0030] In a seventh aspect, a readable storage medium is provided. A program or instructions are stored on the readable storage medium. When the program or instructions are executed by a processor, the steps of the method described in the first aspect are implemented, or the steps of the method described in the second aspect are implemented.

[0031] In an eighth aspect, an encoding / decoding system is provided, including: an encoding end device and a decoding end device. The encoding end device can be used to execute the steps of the method described in the first aspect, and the decoding end device can be used to execute the steps of the method described in the second aspect.

[0032] In a ninth aspect, a chip is provided. The chip includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instructions to implement the steps of the method described in the first aspect, or implement the steps of the method described in the second aspect.

[0033] In a tenth aspect, a computer program / program product is provided. The computer program / program product is stored in a storage medium. The program / program product is executed by at least one processor to implement the steps of the method described in the first aspect, or implement the steps of the method described in the second aspect.

[0034] In an embodiment of the present application, when the current node is exactly matched with its corresponding reference frame node and the intra prediction mode is not enabled, the encoding end selects the prediction mode adopted by the current node from the inter prediction mode and the non-prediction mode through rate-distortion optimization, and then performs transformation, prediction, and encoding processing on the attribute information of the child nodes of the current node, that is, the nodes to be encoded, to obtain a bitstream. Compared with directly adopting the non-prediction mode as the prediction mode when the current node is exactly matched with its corresponding reference frame node and the intra prediction mode is not enabled, the embodiment of the present application can select a more suitable prediction mode for the current node, which is beneficial to improving the prediction encoding effect and the encoding quality.

[0035] When the current node is exactly matched with its corresponding reference frame node and the intra prediction mode is not enabled, the decoding end determines the prediction mode of the current node through entropy decoding, which is the inter prediction mode or the non-prediction mode, and then performs inverse transformation on the reconstructed value of the transform coefficient of the child nodes of the current node, that is, the nodes to be decoded, according to the prediction mode to obtain the reconstructed attribute value. Compared with directly adopting the non-prediction mode as the prediction mode when the current node is exactly matched with its corresponding reference frame node and the intra prediction mode is not enabled, the embodiment of the present application can select a more suitable prediction mode for the current node, which is beneficial to improving the prediction decoding effect and the decoding quality. Description of the Drawings

[0036] Figure 1 is a schematic diagram of an encoding and decoding system provided by an embodiment of the present application;

[0037] Figure 2a is a flowchart of encoding executed by an encoder based on the AVS-PCC encoding framework;

[0038] Figure 2b is a flowchart of encoding executed by an encoder based on the MPEG G-PCC encoding framework;

[0039] Figure 3a is a flowchart of decoding executed by a decoder based on the AVS-PCC decoding framework;

[0040] Figure 3b is a flowchart of decoding executed by a decoder based on the MPEG G-PCC decoding framework;

[0041] Figure 4 is a schematic flowchart of an encoding method provided by an embodiment of the present application;

[0042] Figure 5 is a schematic flowchart of another encoding method provided by an embodiment of the present application;

[0043] Figure 6It is a schematic flowchart of another encoding method provided by an embodiment of the present application;

[0044] Figure 7 It is a schematic flowchart of a decoding method provided by an embodiment of the present application;

[0045] Figure 8 It is a schematic flowchart of another decoding method provided by an embodiment of the present application;

[0046] Figure 9 It is a schematic diagram of a decoding process provided by an embodiment of the present application;

[0047] Figure 10 It is a schematic block diagram of an encoding device provided by an embodiment of the present application;

[0048] Figure 11 It is a schematic block diagram of a decoding device provided by an embodiment of the present application;

[0049] Figure 12 It is a schematic block diagram of an electronic device further provided by an embodiment of the present application;

[0050] Figure 13 It is a schematic diagram of the hardware structure of a terminal implementing the embodiments of the present application. Detailed implementation manners

[0051] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0052] The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "or" in the present application means at least one of the connected objects. For example, "A or B" covers three scenarios, namely, Scenario 1: including A and not including B; Scenario 2: including B and not including A; Scenario 3: including both A and B. The character " / " generally indicates an "or" relationship between the associated objects before and after.

[0053] Before introducing the technical solutions provided by the embodiments of the present application, the meanings of some terms therein are first introduced.

[0054] Point Cloud: A point cloud refers to a set of irregularly distributed discrete points in space that represent the spatial structure and surface properties of a three-dimensional object or three-dimensional scene. Point clouds can be divided into different categories according to different classification criteria. For example, according to the acquisition method of the point cloud, it can be divided into dense point clouds and sparse point clouds; or according to the temporal type of the point cloud, it can be divided into static point clouds and dynamic point clouds.

[0055] Point Cloud Data: The geometric coordinate information and attribute information of each point in the point cloud together constitute the point cloud data. Among them, the geometric coordinate information can also be called three-dimensional position information. The geometric coordinate information of a certain point in the point cloud refers to the spatial coordinates (x, y, z) of the point, which can include the coordinate values of the point in the directions of each coordinate axis in the three-dimensional coordinate system. For example, the coordinate value x in the X-axis direction, the coordinate value y in the Y-axis direction, and the coordinate value z in the Z-axis direction. The attribute information of a certain point in the point cloud can include at least one of the following: color information, material information, and laser reflection intensity information (which can also be called reflectivity). Usually, each point in the point cloud has the same number of attribute information. For example, each point in the point cloud can have two attribute information, namely color information and laser reflection intensity; or each point in the point cloud can have three attribute information, namely color information, material information, and laser reflection intensity information.

[0056] Point Cloud Compression (PCC): Point cloud compression refers to the process of encoding the geometric coordinate information and attribute information of each point in the point cloud to obtain a compressed bitstream. Point cloud compression can include two main processes: geometric coordinate information encoding and attribute information encoding. Currently, the point cloud compression framework that can compress point clouds can be the geometry-based point cloud compression (G-PCC) codec framework provided by the Moving Picture Experts Group (MPEG), or the video-based point cloud compression (V-PCC) codec framework, or the AVS-PCC codec framework provided by the Audio Video Standard (AVS).

[0057] Point cloud decoding: Point cloud decoding refers to the process of decoding the compressed bitstream obtained by point cloud encoding to reconstruct the point cloud. Specifically, it refers to the process of reconstructing the geometric coordinate information and attribute information of each point in the point cloud based on the geometric bitstream and attribute bitstream in the compressed bitstream. After obtaining the compressed bitstream at the decoding end, for the geometric bitstream, first perform entropy decoding to obtain the quantized information of each point in the point cloud, and then perform inverse quantization to reconstruct the geometric coordinate information of each point in the point cloud. For the attribute bitstream, first perform entropy decoding to obtain the quantized attribute residual information or quantized transform coefficients of each point in the point cloud; then perform inverse quantization on the quantized attribute residual information to obtain the reconstructed residual information, perform inverse quantization on the quantized transform coefficients to obtain the reconstructed transform coefficients, and the reconstructed transform coefficients are inversely transformed to obtain the reconstructed residual information. According to the reconstructed residual information of each point in the point cloud, the attribute information of each point in the point cloud can be reconstructed. The reconstructed attribute information of each point in the point cloud is sequentially corresponding to the reconstructed geometric coordinate information one by one to reconstruct the point cloud.

[0058] Figure 1 FIG. is a schematic diagram of the codec system 10 provided by an embodiment of the present application. The technical solution of the embodiment of the present application relates to encoding and decoding (CODEC) (including encoding or decoding) point cloud data.

[0059] As Figure 1 shown, the codec system 10 includes a source device 100, and the source device 100 provides encoded point cloud data to be decoded and displayed by a destination device 110. Specifically, the source device 100 provides point cloud data to the destination device 110 via a communication medium 120. The source device 100 and the destination device 110 may include any one or more of a desktop computer, a notebook (i.e., laptop) computer, a tablet computer, a set-top box, a mobile phone, a wearable device (such as a smart watch or a wearable camera), a television, a camera, a display device, a vehicle-mounted device, a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, a digital media player, a video game console, a video conferencing device, a video streaming device, a broadcast receiver device, a broadcast transmitter device, a spacecraft, an airplane, a robot, a satellite, etc.

[0060] In Figure 1 the example of, the source device 100 includes a data source 101, a memory 102, an encoder 200, and an output interface 104. The destination device 110 includes an input interface 111, a decoder 300, a memory 113, and a display device 114. The source device 100 represents an example of an encoding device, and the destination device 110 represents an example of a decoding device. In other examples, the source device 100 and the destination device 110 may not include Figure 1Some components in, or may also include Figure 1 Other components than. For example, the source device 100 can obtain point cloud data through an external capture device. Similarly, the destination device 110 can be interfaced with an external display device without including an integrated display device. For another example, the memories 102 and 113 can be external memories.

[0061] Although Figure 1 The source device 100 and the destination device 110 are depicted as separate devices, but in some examples, they can also be integrated into one device. In such embodiments, the functions corresponding to the source device 100 and the functions corresponding to the destination device 110 can be implemented using the same hardware or software, or using separate hardware or software, or any combination thereof.

[0062] In some examples, the source device 100 and the destination device 110 can perform unidirectional data transmission or bidirectional data transmission. If it is bidirectional data transmission, the source device 100 and the destination device 110 can operate in a substantially symmetric manner, that is, each of the source device 100 and the destination device 110 includes an encoder and a decoder.

[0063] The data source 101 represents the source of the point cloud data (i.e., the original, uncoded point cloud data) and provides the point cloud data to the encoder 200, and the encoder 103 encodes the point cloud data. The source device 100 can include a capture device (such as a camera device, a sensing device, or a scanning device), an archive including previously captured point cloud data, or a feed interface for receiving point cloud data from a data content provider. Among them, the camera device can include an ordinary camera, a stereo camera, a light field camera, etc., the sensing device can include a laser device, a radar device, etc., and the scanning device can include a three-dimensional laser scanning device, etc. Point cloud data can be obtained by collecting the visual scene of the real world through the capture device. As an alternative, the data source 101 can generate computer graphics-based data as the source data, or combine real-time data, archived data, and computer-generated data. For example, the data source generates point cloud data according to virtual objects (such as virtual three-dimensional objects and virtual three-dimensional scenes obtained through three-dimensional modeling).

[0064] The encoder 200 encodes the captured, pre-captured, or computer-generated data. The encoder 200 can rearrange the point cloud data from the received order (sometimes referred to as the "display order") in the encoding order. The encoder 200 can generate a bitstream including the encoded point cloud data. The source device 100 can then output the encoded point cloud data to the communication medium 120 via the output interface 104 for reception or retrieval by, for example, the input interface 111 of the destination device 110.

[0065] The memory 102 of the source device 100 and the memory 113 of the destination device 110 represent general memories. In some examples, the memory 102 may store the original data from the data source 101, and the memory 113 may store the decoded point cloud data from the decoder 300. Additionally or alternatively, the memories 102, 113 may store software instructions that can be executed by, for example, the encoder 200 and the decoder 300, respectively. Although the memory 102 and the memory 113 are shown separately from the encoder 200 and the decoder 300 in this example, it should be understood that the encoder 200 and the decoder 300 may also include internal memories for functionally similar or equivalent purposes. If the encoder 200 and the decoder 300 are deployed on the same hardware device, the memory 102 and the memory 113 may be the same memory. Further, the memories 102, 113 may store, for example, the encoded point cloud data output from the encoder 200 and input to the decoder 300. In some examples, portions of the memories 102, 113 may be allocated as one or more point cloud buffers, such as for storing raw, decoded, or encoded point cloud data.

[0066] In some examples, the source device 100 may output the encoded data from the output interface 104 to the memory 113. Similarly, the destination device 110 may access the encoded data from the memory 113 via the input interface 111. The memory 113 or the memory 102 may include any of a variety of distributed or local access data storage media, such as hard drives, Blu-ray discs, Digital Versatile Discs (DVDs), Compact Disc Read-Only Memories (CD-ROMs), flash memories, volatile or non-volatile memories, or any other suitable digital storage media for storing the encoded point cloud data.

[0067] The output interface 104 may include any type of medium or device capable of sending the encoded point cloud data from the source device 100 to the destination device 110. For example, the output interface 104 may include a transmitter or transceiver, such as an antenna, configured to send the encoded point cloud data from the source device 100 directly and in real time to the destination device 110. The encoded point cloud data may be modulated according to the communication standards of a wireless communication protocol and sent to the destination device 110.

[0068] The communication medium 120 may include a transient medium, such as a wireless broadcast or a wired network transmission. For example, the communication medium 120 may include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). The communication medium 120 may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium 120 may also take the form of a storage medium (e.g., a non-transitory storage medium), such as a hard disk, a flash drive, a compact disc, a digital versatile disc, a Blu-ray disc, a volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded point cloud data.

[0069] In some embodiments, the communication medium 120 may include a router, a switch, a base station, or any other device that may be used to facilitate communication from the source device 100 to the destination device 110. For example, a server (not shown) may receive the encoded point cloud data from the source device 100 and provide it to the destination device 110, e.g., via a network transmission to the destination device 110. The server may include, for example, a web server (such as for a website), a server configured to provide file transfer protocol services (such as File Transfer Protocol (FTP) or File Delivery Over Unidirectional Transport (FLUTE) protocol), a content delivery network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Services (MBMS) or evolved Multimedia Broadcast Multicast Service (eMBMS) server, or a Network-attached storage (NAS) device, etc. The server may implement one or more HTTP streaming protocols, such as the MPEG Media Transport (MMT) protocol, the Dynamic Adaptive Streaming over HTTP (DASH) protocol, the HTTP Live Streaming (HLS) protocol, or the Real Time Streaming Protocol (RTSP), etc.

[0070] The destination device 110 can access the encoded point cloud data from the server, for example, via a wireless channel (e.g., Wi-Fi connection) or a wired connection (e.g., Digital Subscriber Line (DSL), cable modem, etc.) for accessing the encoded point cloud data stored on the server.

[0071] The output interface 104 and the input interface 111 can represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to the IEEE 802.11 standard or the IEEE 802.15 standard (e.g., ZigBeeTM), the Bluetooth standard, etc.), or other physical components. In an example where the output interface 104 and the input interface 111 include wireless components, the output interface 104 and the input interface 111 can be configured to transmit data, such as encoded point cloud data, according to WIFI, Ethernet, a cellular network (such as 4G, LTE (Long Term Evolution), Advanced LTE, 5G, 6G, etc.).

[0072] The technology provided by the embodiments of this application can be applied to support one or more of the following application scenarios: machine perception of point clouds, which can be used in scenarios such as autonomous navigation systems, real-time inspection systems, geographic information systems, vision sorting robots, disaster relief robots, etc.; human eye perception of point clouds, which can be used in point cloud application scenarios such as digital cultural heritage, free viewpoint broadcasting, three-dimensional immersive communication, three-dimensional immersive interaction, etc.

[0073] The input interface 111 of the destination device 110 receives the encoded bitstream from the communication medium 120. The encoded bitstream can include high-level syntax elements and encoded data units (e.g., sequences, groups of pictures, pictures, slices, blocks, etc.), where the high-level syntax elements are used to decode the encoded data units to obtain the decoded point cloud data. The display device 114 displays the decoded point cloud data to the user. The display device 114 can include a cathode ray tube (CRT), a liquid-crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices. In some examples, the destination device 110 may not have the display device 114. For example, if the decoded point cloud data is used to determine the position of a physical object, the display device 114 can be replaced by a processor.

[0074] The encoder 200 and the decoder 300 may be implemented as one or more of various processing circuits, which may include a microprocessor, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), discrete logic, hardware, or any combination thereof. When the technologies are implemented in software in whole or in part, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the technologies provided by the embodiments of the present application.

[0075] Taking the G-PCC and AVS-PCC encoding and decoding frameworks as examples, the basic principles of the encoder 200 and the decoder 300 provided by the embodiments of the present application are introduced below.

[0076] The encoding and decoding frameworks of G-PCC and AVS-PCC are substantially the same. As Figure 2a shows the encoding flowchart executed by the encoder of the encoding framework based on AVS-PCC. As Figure 2b shows the encoding flowchart executed by the encoder of the encoding framework based on MPEG G-PCC. The above encoder may be Figure 1 the encoder 200 shown. Generally, the above encoding frameworks can be divided into a geometric coordinate information encoding process and an attribute information encoding process. In the geometric information encoding process, the geometric coordinate information of each point in the point cloud is encoded to obtain a geometric bitstream; in the attribute information encoding process, the attribute information of each point in the point cloud is encoded to obtain an attribute bitstream; the geometric bitstream and the attribute bitstream together form the compressed bitstream of the point cloud.

[0077] For the geometric information encoding process, the encoding flow executed by the encoder 200 is as follows:

[0078] 1. Pre-Processing: It may include Transform Coordinates and Voxelize. Through operations of scaling and translation, the pre-processing converts the point cloud data in three-dimensional space into an integer form and moves its minimum geometric position to the origin of coordinates. In some examples, the encoder 200 may not perform pre-processing.

[0079] 2. Geometric Encoding: For the AVS-PCC coding framework, geometric encoding includes two modes, namely octree-based geometric encoding and prediction-tree-based geometric encoding. For the G-PCC coding framework, geometric encoding includes three modes, namely octree-based geometric encoding, triangle representation (Trisoup)-based geometric encoding, and prediction-tree-based prediction encoding. Among them:

[0080] Octree-based geometric encoding: An octree is a tree-shaped data structure that evenly divides a pre-set bounding box in three-dimensional space, and each node has eight child nodes. By using "1" and "0" to indicate whether each child node of the octree is occupied or not, occupancy code information (Occupancy Code) is obtained as the bitstream of point cloud geometric information.

[0081] Prediction-tree-based geometric encoding: A prediction tree is generated using a prediction strategy. Starting from the root node of the prediction tree, each node is traversed, and the residual coordinate values corresponding to each traversed node are encoded.

[0082] Triangle representation (Trisoup)-based geometric encoding: The point cloud is divided into blocks of a certain size, and the intersection points (referred to as vertices) of the point cloud surface at the edges of the blocks are located. Geometric information compression is achieved by encoding whether there are intersection points on each edge of the block and the positions of the intersection points.

[0083] 3. Geometry Entropy Encoding: Statistical compression encoding is performed on the occupancy code information of the octree, the prediction residual information of the prediction tree, and the vertex information of the triangle representation, and finally a binary (0 or 1) compressed bitstream is output. Statistical encoding is a lossless encoding method that can effectively reduce the bitrate required to represent the same signal. The commonly used statistical encoding method is context-based binary arithmetic coding (Content Adaptive Binary Arithmetic Coding, CABAC).

[0084] 4. Geometric Reconstruction: The geometric information after geometric encoding is decoded and reconstructed.

[0085] For the attribute information encoding process, the encoding process executed by the encoder 200 is as follows:

[0086] 1. Color Transformation: A transformation is applied to transform the color information of the attribute to a different domain. For example, the color information can be transformed from the RGB color space to the YCbCr color space.

[0087] 2. Attribute Recoloring: In the case of lossy coding, after encoding the geometric coordinate information, the encoder needs to decode and reconstruct the geometric information, that is, restore the geometric information of each point in the point cloud. Find the attribute information corresponding to one or more neighboring points in the original point cloud as the attribute information of the reconstructed point.

[0088] In some examples, the encoder 200 may not perform color transformation or attribute recoloring.

[0089] 3. Attribute Information Processing: In AVS-PCC, attribute information processing can include three modes, namely Prediction coding, Transform coding, and Prediction&Transform coding. These three coding modes can be used under different conditions.

[0090] Among them, Prediction coding means that according to information such as distance or spatial relationship, determine the neighboring points of the point to be encoded as the prediction points among the encoded points. Based on the set criteria, calculate the predicted attribute information of the point to be encoded according to the attribute information of the prediction points. Calculate the difference between the true attribute information and the predicted attribute information of the point to be encoded as the attribute residual information, and perform quantization, transformation (optional), and entropy coding on the attribute residual information.

[0091] Transform coding means using transformation methods such as Discrete Cosine Transform (DCT) and Haar Transform (Haar) to group and transform the attribute information, quantize the transform coefficients; obtain the attribute reconstruction information through inverse quantization and inverse transformation; calculate the difference between the true attribute information and the attribute reconstruction information to obtain the attribute residual information and quantize it; perform entropy coding on the quantized transform coefficients and the attribute residual.

[0092] Prediction&Transform coding means using the attribute residual information obtained by prediction for transformation, and performing quantization and entropy coding on the transform coefficients.

[0093] In MPEG G-PCC, attribute information processing can include three modes, namely PredictionTransform coding, Lifting Transform coding, and RegionAdaptive Hierarchical Transform (RAHT) coding. These three coding modes can be used under different conditions.

[0094] Among them, predictive transform coding refers to selecting a subset of sub-points according to distance, dividing the point cloud into multiple different levels of detail (LoD), and realizing a multi-quality level point cloud representation from rough to refined. Prediction can be achieved from bottom to top between adjacent layers, that is, the attribute information of the points introduced in the fine layer is predicted by the neighboring points in the rough layer to obtain the corresponding attribute residual information. Among them, the points in the bottom layer are used as reference information for encoding.

[0095] Enhanced transform coding refers to introducing a weight update strategy for neighboring points on the basis of LoD adjacent layer prediction, and finally obtaining the predicted attribute information of each point to obtain the corresponding attribute residual information.

[0096] Hierarchical region adaptive transform coding means that the attribute information is transformed by RAHT, and the signal is converted into the transform domain, which is called the transform coefficient.

[0097] 4. Attribute Quantization: The fineness of quantization is usually determined by quantization parameters. The transform coefficients or attribute residual information obtained by processing the attribute information are quantized, and the quantized results are entropy encoded. For example, in predictive transform coding and enhanced transform coding, the quantized attribute residual information is entropy encoded; in RAHT, the quantized transform coefficients are entropy encoded.

[0098] 5. Entropy Coding: The quantized attribute residual information and / or transform coefficients generally use run length coding (RLC) and arithmetic coding (AC) to achieve final compression. The corresponding coding mode, quantization parameter and other information are also encoded by the entropy encoder.

[0099] The encoder 200 encodes the geometric coordinate information of each point in the point cloud to obtain a geometric bitstream, and encodes the attribute information of each point in the point cloud to obtain an attribute bitstream. The encoder 200 can transmit the encoded geometric bitstream and attribute bitstream to the decoder 300 together.

[0100] Figure 3a shows the decoding flowchart executed by the decoder of the decoding framework based on AVS-PCC, as Figure 3b shows the decoding flowchart executed by the decoder of the decoding framework based on MPEG G-PCC. The above decoder can be Figure 1The decoder 300 shown. After receiving the compressed bitstreams (i.e., the attribute bitstream and the geometry bitstream) transmitted by the encoder 200, the decoder 300 decodes the geometry bitstream to reconstruct the geometric coordinate information of each point in the point cloud, and decodes the attribute bitstream to reconstruct the attribute information of each point in the point cloud.

[0101] The decoding process executed by the decoder 300 is as follows:

[0102] 1. Entropy Decoding: Entropy decode the geometry bitstream and the attribute bitstream respectively to obtain geometric syntax elements and attribute syntax elements.

[0103] 2. Geometry Decoding: For the AVS-PCC coding framework, geometry decoding includes two modes, namely octree-based geometry decoding and prediction tree-based geometry decoding. For the G-PCC coding framework, geometry coding includes three modes, namely octree-based geometry decoding, triangle representation (Trisoup)-based geometry decoding, and prediction decoding based on a prediction tree.

[0104] Octree-based Geometry Decoding: Reconstruct an octree based on the geometric syntax elements parsed from the geometry bitstream.

[0105] Prediction Tree-based Geometry Decoding: Reconstruct a prediction tree based on the geometric syntax elements parsed from the geometry bitstream.

[0106] Triangle Representation-based Geometry Decoding: Reconstruct a triangle model based on the geometric syntax elements parsed from the geometry bitstream.

[0107] 3. Geometry Reconstruction: Perform reconstruction to obtain the geometric coordinate information of the points in the point cloud.

[0108] 4. Coordinate Inverse Transformation: Perform an inverse transformation on the reconstructed geometric coordinate information to transform the reconstructed coordinates (positions) of the points in the point cloud from the transformed domain back to the initial domain.

[0109] 5. Inverse Quantization: Perform inverse quantization on the attribute syntax elements.

[0110] 6. Attribute Information Processing: In AVS-PCC, attribute information processing determines the color information of the points in the point cloud for the inverse quantized prediction residuals or prediction residual transform coefficients through prediction or prediction transformation, or determines the color information of the points in the point cloud for the inverse quantized transform coefficients through transformation.

[0111] In MPEG G-PCC, attribute information processing determines the color information of the points in the point cloud for the inverse quantized attribute information through RAHT, or determines the color information of the points in the point cloud for the inverse quantized attribute information through LOD and inverse lifting.

[0112] 7. Color inverse transformation: Transform color information from the YCbCr color space to the RGB color space. In some examples, the color inverse transformation operation may not be performed.

[0113] In the related art, in the RAHT process based on upsampling prediction, different coding modes are adopted for the nodes in the transform tree structure. For example, for some nodes, if intra prediction is enabled, RDO is used to select an appropriate prediction coding mode among the intra prediction mode, the inter prediction mode, and the non - prediction mode; if the intra prediction mode is not enabled, the non - prediction mode is directly adopted. However, when the intra prediction mode is not enabled, if the non - prediction mode is directly used, the prediction coding effect may be poor. Therefore, when the intra prediction mode is not enabled, how to select the prediction mode urgently needs to be solved.

[0114] In view of this, the embodiments of the present application provide an encoding and decoding method and related devices, which can, when the intra prediction mode is not enabled, select the prediction mode adopted by the node to be encoded from the inter prediction mode and the non - prediction mode through rate - distortion optimization, thereby being beneficial to improving the prediction coding effect and the coding quality.

[0115] The following introduces the encoding method and decoding method provided by the embodiments of the present application with reference to the accompanying drawings. The encoding method provided by the embodiments of the present application can be executed by an encoding end, such as Figure 1 、 Figure 2a or Figure 2b the encoder 200 shown. The decoding method provided by the embodiments of the present application can be executed by a decoding end, such as Figure 1 、 Figure 3a or Figure 3b the decoder 300 described. Among them, the encoding end and the decoding end can be implemented by software, hardware, or a combination thereof. When implemented by hardware, the encoding end can be referred to as an encoding end device or an encoding device, and the decoding end can be referred to as a decoding end device or a decoding device.

[0116] Figure 4 shows a schematic flowchart of an encoding method 400 provided by the embodiments of the present application. As Figure 4 shown, the method 400 includes steps 410 to 450.

[0117] 410. The encoding end determines the transform tree structure of the point cloud to be encoded based on geometric reconstruction information, and the transform tree structure includes at least one first - node layer.

[0118] Exemplarily, the encoding end may reorder the point cloud to be encoded, and construct an N-layer transformation tree structure for the reordered point cloud based on the geometric reconstruction information. Wherein N is a positive integer greater than 2. As a method, a bottom-up construction method may be used to construct the transformation tree structure. Wherein, the first node layer may be any one or more node layers in the N-layer transformation tree structure, which is not limited in the embodiment of the present application.

[0119] In some embodiments, in the process of constructing the transformation tree structure, corresponding Morton code information, attribute information and weight information may be generated for the merged nodes.

[0120] Optionally, for a node to be encoded in the transform tree structure, its prediction mode may include an inter-frame prediction (Inter) mode, an intra-frame prediction (Intra) mode or a non-prediction (Null) mode. For different nodes in the transform tree structure, it is possible to determine which prediction mode to use for prediction in different ways.

[0121] In some embodiments, the first node layer is a middle node layer in the transform tree structure.

[0122] Exemplarily, the encoder can perform upsampling prediction and RAHT on each node layer by layer from the root node from top to bottom based on the transform tree structure. As an implementable method, the constructed transform tree structure can be layered into three layers: upper, middle, and lower, and different prediction modes can be used respectively. Among them, the first node layer is the middle node layer.

[0123] It should be noted that when encoding the transform tree structure from top to bottom, when encoding the child nodes of the current node, the current node serves as the parent node, and its attribute value has been encoded.

[0124] Optionally, for upper node layers in the transform tree structure, the prediction mode of the node may be inferred.

[0125] Exemplarily, first, it can be determined whether the inter-frame prediction mode is enabled. If enabled, the inter-frame prediction mode is selected; if not enabled, it is determined whether the intra-frame prediction mode is enabled. If enabled, the intra-frame prediction mode is selected; otherwise, the non-prediction mode is selected.

[0126] Optionally, for the middle node layer in the transform tree structure, a suitable prediction mode may be selected from three prediction modes through Rate-Distortion Optimization (RDO).

[0127] 420 , determine, in a reference frame of a current node in at least one first node layer, a node having the same geometric position as the current node as a reference frame node.

[0128] Exemplarily, by performing an operation of building a transform tree from bottom to top on the reference frame of the current node that is the same as the current frame where the current node is located, the transform tree structure of the reference frame can be obtained, so that inter-frame prediction mode can be used for prediction. Among them, the reference frame is an encoded frame. After that, a node in the reference frame with the same geometric position as the current node can be determined as the reference frame node.

[0129] In some embodiments, for a current node that needs to adopt the inter-frame prediction method, the operation in step 420 can be performed to determine its reference frame node. In some embodiments, for a current node that needs to select a suitable prediction mode from multiple prediction modes (such as the above three prediction modes) through RDO, the operation in step 420 can be performed to determine its reference frame node.

[0130] 430. If the current node and the reference frame node are completely matched and the intra-frame prediction mode is not enabled, then through rate-distortion optimization, select the prediction mode adopted by the current node from the inter-frame prediction mode and the non-prediction mode.

[0131] Specifically, for a node in the first node layer, if the current node and its corresponding reference frame node are completely matched and the intra-frame prediction mode is not enabled, then through RDO, select a prediction mode from the inter-frame prediction and the non-prediction mode as the prediction mode adopted by the current node.

[0132] Among them, the current node and the reference frame node being completely matched means that all the child nodes of the current node can all find child nodes at the same positions in the corresponding reference frame node. Specifically, since the current node block (i.e., the current node) in the transform tree structure is not completely occupied, when a certain child node of the current node exists, the child node of the reference frame node at the same position may not exist. Therefore, for all the child nodes of the current node can all find child nodes at the same positions in the corresponding reference frame node, it can be said that the current node and the reference frame node are completely matched. For all the child nodes of the current node cannot all find child nodes at the same positions in the corresponding reference frame node, it is said that the current node and the reference frame node are not completely matched.

[0133] Exemplarily, a parameter inter_pred_complete_match can be used to represent whether all the child nodes of the current node can all find child nodes at the same positions in the corresponding reference frame node. If so, it represents complete match, inter_pred_complete_match = 1; otherwise it is 0.

[0134] Therefore, when the current node completely matches the reference frame node and the intra prediction mode is not enabled, RDO is used to select a suitable prediction mode from the inter prediction mode and the non - prediction mode as the prediction mode adopted by the current node. Compared with directly using the non - prediction mode as the prediction mode, it can ensure a low bit rate while ensuring a low distortion rate, improving the coding performance.

[0135] In some other embodiments, if the current node completely matches the reference frame node (such as inter_pred_complete_match = 1) and the intra prediction mode is enabled, RDO is used to select the prediction mode adopted by the current node from the intra prediction mode, the inter prediction mode, and the non - prediction mode.

[0136] In some other embodiments, if the current node does not completely match the reference frame node (such as inter_pred_complete_match = 0) and the intra prediction mode is enabled, RDO is used to select the prediction mode adopted by the current node from the intra prediction mode, the inter prediction mode, and the non - prediction mode.

[0137] In some other embodiments, if the current node does not completely match the reference frame node (such as inter_pred_complete_match = 0) and the intra prediction mode is not enabled, the prediction mode adopted by the current node is the non - prediction mode.

[0138] Exemplarily, for the middle - layer node layer in the transform tree structure, such as the first node layer above, a suitable prediction mode can be selected from the above three prediction modes according to Table 1 below.

[0139] Table 1

[0140]

[0141] 440. Predict and transform the node to be encoded according to the prediction mode to obtain the attribute transform coefficients. The node to be encoded is a child node of the current node.

[0142] Specifically, according to the prediction mode of the current node determined in step 430, the attribute information (value) of the child node of the current node, that is, the node to be encoded, can be predicted and transformed. For example, based on RATH of up - sampling prediction, the attribute transform coefficients can be obtained. According to the above step 430, the prediction mode of the current node can be the inter prediction mode or the non - prediction mode.

[0143] In some embodiments, referring to Figure 5 , the attribute transform coefficients of the node to be encoded can be obtained according to the following steps 441 to 443.

[0144] 441. If the prediction mode of the current node is the inter-frame prediction mode, determine the prediction attribute value of the node to be encoded.

[0145] Optionally, the reconstruction attribute value of the child node at the corresponding position of the reference frame node can be determined as the prediction attribute value of the node to be encoded at the corresponding position of the current node.

[0146] Specifically, since the current node exactly matches the reference frame node, that is, each node in the current node finds the child node at the same position in the corresponding reference frame node. At this time, the reconstruction attribute value of the child node at the corresponding position in the reference frame node can be determined as the prediction attribute value of the child node (i.e., the node to be encoded) at the corresponding position of the current node. It can be understood that the attribute values of each child node in the reference frame node have been reconstructed and encoded.

[0147] 442. Transform the original attribute value of the node to be encoded to obtain the first transformation coefficient, and transform the prediction attribute value of the node to be encoded to obtain the second transformation coefficient.

[0148] Specifically, when the prediction mode of the current node is the inter-frame prediction mode, transform the original attribute value and the prediction attribute value of the child node of the current node to be encoded respectively to obtain the corresponding transformation coefficients. Exemplarily, the RAHT transformation can be performed on the original attribute value and the prediction attribute value respectively to obtain the corresponding AC transformation coefficients.

[0149] 443. Obtain the attribute transformation coefficient according to the first transformation coefficient and the second transformation coefficient.

[0150] Specifically, the difference can be made between the first transformation coefficient and the second transformation coefficient to obtain the attribute transformation coefficient. The attribute transformation system can also be called the residual transformation coefficient, the AC residual transformation coefficient, etc., without limitation.

[0151] Therefore, in the embodiment of the present application, when the prediction mode of the current node is the inter-frame prediction mode, by respectively transforming the original attribute value and the prediction attribute value of the node to be encoded and making the difference between the obtained transformation coefficients, the attribute transformation coefficient can be obtained.

[0152] Optionally, continue to refer to Figure 5 , and the attribute transformation coefficient of the node to be encoded can also be obtained according to step 444 as follows.

[0153] 444. If the prediction mode of the current node is the non-prediction mode, transform the original attribute value of the node to be encoded to obtain the attribute transformation coefficient.

[0154] Specifically, since the prediction mode of the current node is the non-prediction mode, only the original attribute values of the children nodes of the current node to be encoded need to be transformed to obtain the corresponding attribute transformation coefficients. Exemplarily, the attribute transformation coefficients can be referred to as AC transformation coefficients.

[0155] Therefore, in the embodiment of the present application, when the prediction mode of the current node is the non-prediction mode, the corresponding attribute transformation coefficients are obtained by transforming the original attribute values of the node to be encoded.

[0156] In some embodiments, for the lower node layer in the transformation tree structure, the prediction mode of the nodes can be inferred.

[0157] Optionally, inter-frame prediction is not performed on the nodes in the lower node layer. Exemplarily, for the nodes in the lower layer, it can first be determined whether intra-frame prediction is enabled. If it is enabled, the intra-frame prediction mode is selected; otherwise, the non-prediction mode is selected.

[0158] In some embodiments, for the intra-frame prediction mode and the non-prediction mode, the upsampling prediction and the RAHT process are as follows:

[0159] When the current node has only one occupied child node, no prediction is performed;

[0160] When the number of neighbor parent nodes of the current node (i.e., the number of grandparent neighbors of the children nodes of the current node) is less than threshold 1, no prediction is performed, and the RAHT transformation is directly performed on the original attribute values of the children nodes of the current node to obtain the AC transformation coefficients;

[0161] If threshold 1 is satisfied, neighbors are searched for the children nodes of the current node. The neighbor search range includes: the current node, the neighbor parent nodes coplanar and collinear with the children nodes of the current node, and the neighbor children nodes coplanar and collinear with the children nodes of the current node;

[0162] When the number of found neighbor parent nodes is less than threshold 2, no prediction is performed, and the RAHT transformation is directly performed on the original attribute values of the children nodes of the current node to obtain the AC transformation coefficients;

[0163] If threshold 2 is satisfied, intra-frame prediction is performed according to the neighbor nodes to obtain the attribute prediction value of the current child node; the RAHT transformation is respectively performed on the original attribute value and the attribute prediction value, and the obtained AC transformation coefficients are subtracted to obtain the AC residual transformation coefficients.

[0164] Exemplarily, threshold 1 can be configured as 2, and threshold 2 can be configured as 6.

[0165] In some embodiments, for the inter-frame prediction mode, the upsampling prediction and the RAHT process are as follows:

[0166] Perform the same bottom-up operation of constructing a transform tree on the reference frame, and generate the attribute values of the reference frame nodes corresponding to the current frame nodes, which can be used as the inter-frame prediction values.

[0167] Exemplarily, for a node block of size 2*2*2, select the node with the same geometric position in the reference frame as its reference frame node. Since not all of each 2*2*2 block is occupied, when a certain child node of the current frame node exists, the child node of the reference frame node at the same position may not exist. Therefore, use the parameter inter_pred_complete_match to indicate whether all child nodes of the current frame node can find child nodes at the same position in the corresponding reference frame node. If so, it represents a complete match, inter_pred_complete_match = 1; otherwise, it is 0.

[0168] When inter_pred_complete_match = 1, the attribute values of the child nodes at the corresponding positions of the reference frame nodes can be used as the inter-frame prediction values of each child node of the current frame node; when inter_pred_complete_match = 0, that is, when a certain child node of the current frame node cannot find a child node at the same position in the corresponding reference frame node, use the reconstructed attribute values of the current frame node to determine the inter-frame prediction values of the child nodes. For example, the average of the reconstructed attribute values of the current frame node can be used as its inter-frame prediction value. This inter-frame prediction mode is called the revised frame inter-prediction mode and can be expressed as the Inter(revision) mode. Since the RAHT transform is encoded from top to bottom, when encoding the child nodes of the current node, the current node is used as the parent node, and its attribute values have been encoded and reconstructed, so they can be used.

[0169] 450, perform encoding processing on the attribute transform coefficients to obtain the bitstream.

[0170] Exemplarily, when the prediction mode is the inter-frame prediction mode, the AC residual transform coefficients obtained in step 440 can be encoded to obtain the bitstream. When the prediction mode is the non-prediction mode, the AC transform coefficients obtained according to the original attribute values of the node to be encoded in step 440 can be encoded to obtain the bitstream.

[0171] Therefore, in the embodiments of the present application, when the encoding end completely matches the current node with its corresponding reference frame node and the intra prediction mode is not enabled, the prediction mode adopted by the current node is selected from the inter prediction mode and the non-prediction mode through rate-distortion optimization, and then the attribute information of the child node of the current node, that is, the node to be encoded, is transformed, predicted, and encoded according to the selected appropriate prediction mode to obtain a bitstream. Compared with directly using the non-prediction mode as the prediction mode when the current node completely matches its corresponding reference frame node and the intra prediction mode is not enabled, the embodiments of the present application can select a more appropriate prediction mode for the current node, which is beneficial to improving the prediction coding effect and the coding quality.

[0172] In some embodiments, when using RDO to select the best prediction mode, it is necessary to perform entropy coding on the selected prediction mode and pass the coding result into the bitstream. In the related RAHT-based coding process, 2 bits are used to encode the above three prediction modes, and these 2 bits are isNullFlag and isIntraFlag respectively. Specifically, each bit uses 108 contexts to perform arithmetic coding on these prediction modes. Specifically, the 108 contexts can be jointly determined according to the following four context information a) to d):

[0173] a) Using whether intra prediction is enabled and whether inter prediction matches, it can be divided into 3 states:

[0174] 1) Intra prediction mode is enabled and inter prediction matches;

[0175] 2) Intra prediction mode is enabled and inter prediction does not match;

[0176] 3) Intra prediction mode is not enabled.

[0177] b) Using the prediction mode of the current node, it can be divided into 3 states:

[0178] 1) The current node uses the inter prediction mode;

[0179] 2) The current node uses the intra prediction mode;

[0180] 3) The current node uses the non-prediction mode.

[0181] c) Using the most frequently occurring prediction mode among the neighbor parent node and the decoded same-layer child nodes (i.e., neighbor nodes), it can be divided into 4 states:

[0182] 1) The neighbor nodes most frequently use the intra prediction mode;

[0183] 2) The neighbor nodes most frequently use the inter prediction mode;

[0184] 3) The neighbor nodes most frequently use the non-prediction mode;

[0185] 4) Neighbor nodes use the non-predictive mode.

[0186] d) Classify into three states according to the number of occupied child nodes of the current node:

[0187] 1) The number of occupied child nodes is 2 or 3;

[0188] 2) The number of occupied child nodes is 4 or 5;

[0189] 3) The number of occupied child nodes is 6, 7, or 8.

[0190] Among them, inter-frame prediction matching means that the current node completely matches the reference frame node, such as inter_pred_complete_match = 1; inter-frame prediction mismatch means that the current node does not completely match the reference frame node, such as inter_pred_complete_match = 0.

[0191] Specifically, according to the above four context information of a), b), c), and d), the context index (modeIdx) corresponding to the node to be encoded can be determined. As an example, the value range of the context index is {0, 107}. Then, the bits representing the prediction mode (such as isNullFlag and isIntraFlag) can be encoded according to the context index. However, in the related art, too many contexts are used, and each context needs to maintain and update its state, which will increase the computational complexity and memory overhead during the encoding process and affect the encoding efficiency.

[0192] In view of this, when entropy encoding the best prediction mode selected by RDO in the embodiments of the present application, a new method for determining context information is proposed to determine the context, which can help reduce the types of contexts, and thus is beneficial to reducing the computational complexity during the encoding process, reducing the memory overhead, and improving the encoding efficiency.

[0193] In some embodiments, referring to Figure 6 , the above encoding method 400 may further include the following steps 460 to 490:

[0194] 460. Determine the first context information by using whether the intra-frame prediction mode is enabled and whether the current node completely matches the reference frame node.

[0195] Exemplarily, the first context information corresponds to one of the following three states:

[0196] 1) The intra-frame prediction mode is enabled and the current node completely matches the reference frame node;

[0197] 2) Intra prediction mode is enabled and the current node does not exactly match the reference frame node;

[0198] 3) Intra prediction mode is not enabled.

[0199] Exemplarily, it can be determined whether the current node exactly matches the reference frame node according to whether the above parameter inter_pred_complete_match is 1. For example, when inter_pred_complete_match = 1, it is an exact match, and when inter_pred_complete_match = 0, it is not an exact match.

[0200] 470. Determine the second context information by using the most frequently used prediction mode among the current node, the neighboring parent node, and the decoded same-layer child nodes.

[0201] Specifically, since the current node has been encoded and reconstructed when encoding the child nodes of the current node, the current node can also be regarded as a kind of neighboring parent node for predicting the child nodes of the current node (i.e., the nodes to be encoded). Based on this, the context information corresponding to the current node and the neighboring parent node can be combined together to reduce the types of contexts.

[0202] Exemplarily, the second context information corresponds to one of the following three states:

[0203] The most frequently used intra prediction mode among the current node, the neighboring parent node, and the decoded same-layer child nodes;

[0204] The most frequently used inter prediction mode among the current node, the neighboring parent node, and the decoded same-layer child nodes;

[0205] The most frequently used non-prediction mode among the current node, the neighboring parent node, and the decoded same-layer child nodes.

[0206] In some other embodiments, the second context information can also correspond to four contexts. In addition to including the above three states, it can also include the state corresponding to the non-prediction mode used among the current node, the neighboring parent node, and the decoded same-layer child nodes.

[0207] As a specific example, the context can be determined jointly according to the following two context information a) and b):

[0208] a) Using whether intra prediction is enabled and whether inter prediction matches, it can be divided into 3 states:

[0209] 1) Intra prediction mode is enabled and inter prediction matches;

[0210] 2) Intra prediction mode is enabled and inter prediction does not match;

[0211] 3) Intra prediction mode is not enabled.

[0212] b) Using the most frequently occurring prediction mode among the current node, the neighbor parent node, and the decoded children nodes of the same layer, it is divided into 3 states:

[0213] 1) The neighbor node most frequently uses the intra prediction mode;

[0214] 2) The neighbor node most frequently uses the inter prediction mode;

[0215] 3) The neighbor node most frequently uses the non - prediction mode.

[0216] Among them, inter - prediction matching means that the current node completely matches the reference frame node, such as inter_pred_complete_match = 1; inter - prediction non - matching means that the current node does not completely match the reference frame node, such as inter_pred_complete_match = 0.

[0217] According to the above two types of context information a) and b), a total of 9 contexts can be determined. Therefore, compared with the 108 contexts in the related art, the embodiments of the present application can greatly reduce the number of context types.

[0218] In some embodiments, the third context information can also be determined by using the number of occupied children nodes in the current node; among them, the third context information corresponds to two contexts. Compared with the solution in the related art where the context information corresponding to d) corresponds to three states, the embodiments of the present application can help reduce the number of context types by corresponding to one of the two states according to the context information of the occupied children nodes in the current node.

[0219] Exemplarily, the two states corresponding to the third context information can include the following two:

[0220] The number range of the occupied children nodes of the current node is the first interval;

[0221] The number range of the occupied children nodes of the current node is the second interval, where the value range of the number of children nodes of the current node is the first interval and the second interval.

[0222] Exemplarily, the value range of the number n of the current children nodes is [2, 8], n is a positive integer, the first interval can include [2, 4], that is, the number of occupied children nodes is 2, 3, 4, and the second interval can include [5, 8], that is, 5, 6, 7, 8; or the first interval can include [2, 5], that is, the number of occupied children nodes is 2, 3, 4, 5, and the second interval can include [6, 8], that is, 6, 7, 8, etc. The present application does not make a limitation on this.

[0223] As a specific example, the context can be jointly determined according to the following three types of context information in a) and c):

[0224] a) Using whether intra prediction is enabled and whether inter prediction matches, it can be divided into 3 states:

[0225] 1) Intra prediction mode is enabled and inter prediction matches;

[0226] 2) Intra prediction mode is enabled and inter prediction does not match;

[0227] 3) Intra prediction mode is not enabled.

[0228] b) Using the most frequently occurring prediction mode among the current node, the neighbor parent node, and the decoded same-layer child nodes, it is divided into 3 states:

[0229] 1) The neighbor node most frequently uses the intra prediction mode;

[0230] 2) The neighbor node most frequently uses the inter prediction mode;

[0231] 3) The neighbor node most frequently uses the non-prediction mode.

[0232] c) Using the number of child nodes occupied by the current node, it is divided into 2 states:

[0233] 1) The number of current child nodes occupied is 2, 3, 4;

[0234] 2) The number of current child nodes occupied is 5, 6, 7, 8;

[0235] According to the above three types of context information in a), b), and c), a total of 18 contexts can be determined. Therefore, compared with the 108 contexts in the related art, the embodiments of the present application can greatly reduce the number of context types.

[0236] 480, according to the first context information and the second context information, determine the index of the context corresponding to the prediction mode of the current node.

[0237] Specifically, the index of the context corresponding to the prediction model of the current node can be determined according to the prediction method of the current node and the state matching the first context information and the second context information.

[0238] In some embodiments, if the third context information is further determined by using the number of child nodes occupied by the current node, the index of the context corresponding to the prediction mode can be determined according to the first context information, the second context information, and the third context information.

[0239] Exemplarily, for the above 9 contexts, the value range of the index of the context corresponding to the prediction mode of the current node can be {0, 8}; for the above 18 contexts, the value range of the index of the context corresponding to the prediction mode of the current node can be {0, 17}.

[0240] 490, entropy-encode the prediction mode of the current node according to the index of the context to obtain the prediction mode bit, and add the prediction mode bit to the bitstream.

[0241] Exemplarily, the bits (such as isNullFlag and isIntraFlag) representing the prediction mode can be encoded according to the context index. For example, first, it can be determined whether it is a non-prediction mode according to the prediction mode. If so, encode isNullFlag using the context index (modeIdx), and isNullFlag = 1, and end the current encoding; if not, encode isNullFlag using the context index (modeIdx), and isNullFlag = 0, and continue the encoding. Then, if the inter prediction mode is not enabled or the intra prediction mode is not enabled, end the current encoding; otherwise, continue the encoding. Finally, determine whether it is an intra prediction mode according to the prediction mode. If so, encode isIntraFlag using the context index (modeIdx), and isNullFlag = 1; if not, encode isIntraFlag using the context index (modeIdx), and isNullFlag = 0, and end the current encoding.

[0242] Therefore, when entropy-encoding the best prediction mode selected by RDO in the embodiments of the present application, by using the current node as one of the determination context information of the neighbor nodes, or determining one of the two states of the context information according to the number of child nodes occupied in the current node, it is possible to help reduce the types of contexts, thereby facilitating reducing the computational complexity in the encoding process, reducing the memory overhead, and improving the encoding efficiency.

[0243] The encoding method provided by the embodiments of the present application has been described above in conjunction with the accompanying drawings. Next, the decoding method provided by the embodiments of the present application will be described in conjunction with the accompanying drawings.

[0244] Figure 7 The schematic flowchart of a decoding method 500 provided by the embodiments of the present application is shown. As Figure 8 shown, method 500 includes steps 510 to 550.

[0245] 510, the decoding end parses the bitstream to obtain the reconstructed value of the transform coefficient of the point cloud to be decoded.

[0246] Specifically, the data to be decoded can be parsed from the bitstream, and the entropy decoding and inverse quantization are performed on the data to be decoded to obtain the reconstructed transform coefficient values. The geometric decoding is performed on the data to be decoded to obtain the geometric reconstruction information of the point cloud to be decoded.

[0247] 520, determine the transform tree structure of the point cloud to be decoded based on the geometric reconstruction information; the transform tree structure includes at least one first node layer.

[0248] Specifically, the process of constructing the transform tree structure of the point cloud to be decoded is similar to the process of constructing the transform tree structure of the point cloud to be encoded, and reference can be made to Figure 4 the description in step 410 therein, which will not be elaborated here.

[0249] 530, determine the node with the same geometric position as the current node in the reference frame of the current node in at least one first node layer as the reference frame node.

[0250] Specifically, the process of determining the reference frame node of the current node can be referred to Figure 4 the relevant description in step 420 therein, which will not be elaborated here.

[0251] 540, if the current node completely matches the reference frame node and the intra prediction mode is not enabled, determine the prediction mode adopted by the current node through entropy decoding; wherein, the prediction mode is an inter prediction mode or a non - prediction mode.

[0252] Specifically, for the nodes in the first node layer, if the current node completely matches its corresponding reference frame node and the intra prediction mode is not enabled, the prediction mode adopted by the current node is determined through entropy decoding. Specifically, the complete match between the current node and its corresponding reference frame node can be referred to Figure 4 the relevant description therein, which will not be elaborated here.

[0253] Therefore, when the current node completely matches the reference frame node and the intra prediction mode is not enabled, determining the prediction mode adopted by the current node through entropy decoding can ensure a low bit rate while ensuring a low distortion rate, improving the coding performance compared to directly using the non - prediction mode as the prediction mode.

[0254] In some other embodiments, if the current node completely matches the reference frame node (such as inter_pred_complete_match = 1) and the intra prediction mode is enabled, determine the prediction mode of the current node through entropy decoding, and the prediction mode includes an intra prediction mode, an inter prediction mode, or a non - prediction mode.

[0255] In some other embodiments, if the current node does not exactly match the reference frame node (e.g., inter_pred_complete_match = 0) and the intra prediction mode is enabled, the prediction mode of the current node is determined by entropy decoding, and the prediction mode includes an intra prediction mode, an inter prediction mode, or a non-prediction mode.

[0256] In some other embodiments, if the current node does not exactly match the reference frame node (e.g., inter_pred_complete_match = 0) and the intra prediction mode is not enabled, the prediction mode adopted by the current node is the non-prediction mode, that is, the non-prediction mode is directly selected.

[0257] Optionally, for the upper node layer in the transform tree structure, the prediction mode of the node can be inferred.

[0258] Exemplarily, first, it can be determined whether the inter prediction mode is enabled. If it is enabled, the inter prediction mode is selected; if not, it is determined whether the intra prediction mode is enabled. If it is enabled, the intra prediction mode is selected; otherwise, the non-prediction mode is selected.

[0259] In some embodiments, for the lower node layer in the transform tree structure, the prediction mode of the node can be inferred.

[0260] Optionally, inter prediction is not performed on the nodes in the lower node layer. Exemplarily, for the nodes in the lower layer, first, it can be determined whether intra prediction is enabled. If it is enabled, the intra prediction mode is selected; otherwise, the non-prediction mode is selected.

[0261] Specifically, the intra prediction mode, the inter prediction mode, and the non-prediction mode can refer to Figure 4 the relevant descriptions therein, which will not be elaborated here.

[0262] 550, perform an inverse transform on the reconstructed value of the transform coefficient of the node to be decoded according to the prediction mode to obtain a reconstructed attribute value; wherein, the node to be decoded is a child node of the current node.

[0263] In some embodiments, if the prediction mode is the inter prediction mode, perform a transform on the predicted attribute value of the node to be decoded to obtain a second transform coefficient; according to the reconstructed value of the transform coefficient of the node to be decoded and the second transform coefficient, obtain a first transform coefficient; and perform an inverse transform on the first transform coefficient to obtain the reconstructed attribute value.

[0264] Exemplarily, if the prediction mode is the inter-frame prediction mode, the prediction attribute value of the child node (i.e., the node to be decoded) of the current node can be obtained, the AC transform coefficient of the prediction attribute value is obtained through transformation, and it is added to the reconstructed value of the transform coefficient obtained by decoding (i.e., the reconstructed value of the AC residual coefficient) to obtain the reconstructed value of the AC coefficient. Optionally, the DC coefficient of the node to be decoded can be inherited from the parent node, and the reconstructed value of the node to be decoded is obtained through the RAHT inverse transform according to the reconstructed value of the AC coefficient and the DC coefficient.

[0265] In some embodiments, the reconstructed attribute value of the child node at the corresponding position of the reference frame node can be determined as the prediction attribute value of the node to be decoded at the corresponding position of the current node.

[0266] In some embodiments, if the prediction mode is the non-prediction mode, the inverse transform is performed on the reconstructed value of the transform coefficient of the node to be decoded to obtain the reconstructed attribute value.

[0267] Exemplarily, if the prediction mode is the non-prediction mode, the RAHT inverse transform can be performed on the reconstructed value of the transform coefficient obtained by decoding to obtain the reconstructed attribute value of the node to be decoded.

[0268] When traversing to the bottom layer of the transform tree structure, the reconstructed attribute values of all nodes can be obtained, thereby completing the attribute decoding.

[0269] Therefore, in the embodiments of the present application, when the decoding end determines that the current node and its corresponding reference frame node are completely matched and the intra-frame prediction mode is not enabled, the prediction mode of the current node is determined through entropy decoding, which is the inter-frame prediction mode or the non-prediction mode. Then, according to the prediction mode, the inverse transform is performed on the reconstructed value of the transform coefficient of the child node of the current node, that is, the node to be decoded, to obtain the reconstructed attribute value. Compared with directly using the non-prediction mode as the prediction mode when the current node and its corresponding reference frame node are completely matched and the intra-frame prediction mode is not enabled, the embodiments of the present application can select a more appropriate prediction mode for the current node, which is beneficial to improving the prediction decoding effect and the decoding quality.

[0270] In some embodiments, refer to Figure 8 , the prediction mode adopted by the current node can be determined through entropy decoding according to the following steps 560 to 590.

[0271] 560. Determine the first context information by using whether the intra-frame prediction mode is enabled and whether the current node and the reference frame node are completely matched.

[0272] Optionally, the first context information corresponds to one of the following three states:

[0273] The intra prediction mode is enabled and the current node exactly matches the reference frame node;

[0274] The intra prediction mode is enabled and the current node does not exactly match the reference frame node;

[0275] The intra prediction mode is not enabled.

[0276] 570. Using the prediction mode most frequently used among the current node, the neighbor parent node, and the decoded children nodes of the same layer, determine the second context information.

[0277] Optionally, the second context information corresponds to one of the following three states:

[0278] The intra prediction mode is most frequently used among the current node, the neighbor parent node, and the decoded children nodes of the same layer;

[0279] The inter prediction mode is most frequently used among the current node, the neighbor parent node, and the decoded children nodes of the same layer;

[0280] The non - prediction mode is most frequently used among the current node, the neighbor parent node, and the decoded children nodes of the same layer.

[0281] Specifically, for the process of determining the first context information and the second context information, reference can be made to Figure 6 the relevant descriptions in steps 460 and 470 therein, which will not be elaborated here.

[0282] Optionally, the number of occupied children nodes in the current node can also be used to determine the third context information; among them, the third context information corresponds to one of two states. Exemplarily, these two states can include:

[0283] The number of occupied children nodes is within the first interval;

[0284] The number of occupied children nodes is within the second interval; among them, the range of the number of children nodes of the current node is divided into the first interval and the second interval.

[0285] Specifically, for the process of determining the third context information, reference can be made to Figure 6 the relevant descriptions therein, which will not be elaborated here.

[0286] 580. According to the first context information and the second context information, determine the index of the context corresponding to the prediction mode of the current node.

[0287] In some embodiments, if the number of occupied children nodes of the current node is also used to determine the third context information, then the index of the context corresponding to the prediction mode can be determined according to the first context information, the second context information, and the third context information.

[0288] Specifically, step 580 can be referred toFigure 6 The relevant description in step 480 will not be elaborated here.

[0289] 590, Parse the bitstream, perform entropy decoding according to the context index, and obtain the prediction mode.

[0290] Exemplarily, the bitstream can be parsed to obtain the bits representing the prediction mode (such as isNullFlag and isIntraFlag). For example, see Figure 9 , First, the isNullFlag can be entropy decoded using the context index (modeIdx). If isNullFlag = 1, it means the prediction mode is the non - prediction mode (Null), and the current decoding ends; if isNullFlag = 0, continue decoding. Then, if inter - frame prediction (Inter) is not enabled, the prediction mode is the intra - frame prediction mode (Intra), and the current decoding ends; otherwise, continue decoding. Next, determine whether the intra - frame prediction mode (Intra) is enabled. If it is enabled, determine the prediction mode as the inter - frame prediction mode, and the current decoding ends. If it is not enabled, further entropy decode the isIntraFlag using the context (modeIdx). If isNullFlag = 1, it means the prediction mode is the intra - frame prediction (Intra) mode, and the current decoding ends; if isNullFlag = 0, it means the prediction mode is the inter - frame prediction (Inter) mode, and the current decoding ends.

[0291] Therefore, when determining the prediction mode of the current node through entropy decoding in the embodiments of the present application, by using the current node as one of the neighbor nodes to determine the context information, or determining one of the two states of the context information according to the number of child nodes occupied in the current node, it is possible to help reduce the types of contexts, thereby facilitating reducing the computational complexity in the encoding process, reducing the memory overhead, and improving the encoding efficiency.

[0292] Exemplarily, compared with the related art, since the embodiments of the present application use a better RDO algorithm when selecting the prediction mode of the current node and effectively reduce the number of contexts during the encoding and decoding of the prediction mode, the purpose of improving the encoding efficiency can be achieved. Compared with the performance of GPCC ges - tm_4.0, performance gains of 0.2, 0.5, and 0.5 are obtained for the three color attribute channels Luma, Chrome Cb, and Cr respectively.

[0293] The encoding method provided by the embodiments of the present application may be executed by an encoding device. In the embodiments of the present application, taking the encoding device executing the encoding method as an example, the encoding device provided by the embodiments of the present application is described.

[0294] Figure 10The schematic block diagram of an encoding device 1000 provided by an embodiment of the present application is shown. As Figure 10 shown, the encoding device 1000 includes a determination module 1010, a prediction and transformation module 1020, and an encoding module 1030.

[0295] The determination module 1010 is configured to determine a transformation tree structure of a point cloud to be encoded based on geometric reconstruction information, where the transformation tree structure includes at least one first node layer;

[0296] The determination module 1010 is further configured to determine a node having the same geometric position as the current node in the reference frame of the current node in the at least one first node layer as a reference frame node;

[0297] The determination module 1010 is further configured to, if the current node completely matches the reference frame node and the intra prediction mode is not enabled, select a prediction mode adopted by the current node from an inter prediction mode and a non-prediction mode through rate-distortion optimization;

[0298] The prediction and transformation module 1020 is configured to perform prediction and transformation on a node to be encoded according to the prediction mode to obtain attribute transformation coefficients; the node to be encoded is a child node of the current node;

[0299] The encoding module 1030 is configured to perform encoding processing on the attribute transformation coefficients to obtain a bitstream.

[0300] Optionally, the prediction and transformation module 1020 is specifically configured to:

[0301] If the prediction mode is the inter prediction mode, determine a predicted attribute value of the node to be encoded;

[0302] Perform transformation on the original attribute value of the node to be encoded to obtain a first transformation coefficient, and perform transformation on the predicted attribute value to obtain a second transformation coefficient;

[0303] Obtain the attribute transformation coefficients according to the first transformation coefficient and the second transformation coefficient.

[0304] Optionally, the prediction and transformation module 1020 is specifically configured to:

[0305] Determine the reconstructed attribute value of the child node at the corresponding position of the reference frame node as the predicted attribute value of the node to be encoded at the corresponding position of the current node.

[0306] Optionally, the prediction and transformation module 1020 is specifically configured to:

[0307] If the prediction mode is the non-prediction mode, perform transformation on the original attribute value of the node to be encoded to obtain the attribute transformation coefficients.

[0308] Optionally, the determining module 1010 is further configured to:

[0309] If the current node does not exactly match the reference frame node and the intra prediction mode is not enabled, determine that the prediction mode adopted by the current node is the non-prediction mode.

[0310] Optionally, the first node layer is the middle node layer of the transform tree structure.

[0311] Optionally, the determining module 1010 is further configured to:

[0312] Determine first context information by using whether the intra prediction mode is enabled and whether the current node exactly matches the reference frame node;

[0313] Determine second context information by using the prediction mode most frequently used among the current node, the neighboring parent node, and the decoded children nodes of the same layer;

[0314] Determine the index of the context corresponding to the prediction mode according to the first context information and the second context information;

[0315] The encoding module 1030 is further configured to:

[0316] Perform entropy encoding on the prediction mode according to the index to obtain prediction mode bits, and add the prediction mode bits to the bitstream.

[0317] Optionally, the first context information corresponds to one of the following three states:

[0318] The intra prediction mode is enabled and the current node exactly matches the reference frame node;

[0319] The intra prediction mode is enabled and the current node does not exactly match the reference frame node;

[0320] The intra prediction mode is not enabled.

[0321] Optionally, the second context information corresponds to one of the following three states:

[0322] The intra prediction mode is most frequently used among the current node, the neighboring parent node, and the decoded children nodes of the same layer;

[0323] The inter prediction mode is most frequently used among the current node, the neighboring parent node, and the decoded children nodes of the same layer;

[0324] The non-prediction mode is most frequently used among the current node, the neighboring parent node, and the decoded children nodes of the same layer.

[0325] Optionally, the determining module 1010 is further configured to:

[0326] Determine the third context information by using the number of occupied child nodes in the current node; wherein, the third context information corresponds to one of two states.

[0327] Determine the index of the context corresponding to the prediction mode according to the first context information, the second context information, and the third context information.

[0328] Optionally, the two states include:

[0329] The range of the number of occupied child nodes is a first interval.

[0330] The range of the number of occupied child nodes is a second interval; wherein, the range of the possible values of the number of child nodes of the current node is divided into the first interval and the second interval.

[0331] The encoding device 1000 provided by the embodiments of the present application can implement Figures 4 to 6 each process implemented by the method embodiments and achieve the same technical effects. To avoid repetition, details are not described here again.

[0332] In the embodiments of the present application, when the current node completely matches its corresponding reference frame node and the intra prediction mode is not enabled, the encoding end selects the prediction mode adopted by the current node from the inter prediction mode and the non-prediction mode through rate-distortion optimization, and then performs transformation, prediction, and encoding processing on the attribute information of the child nodes of the current node, that is, the nodes to be encoded, to obtain a bitstream. Compared with directly using the non-prediction mode as the prediction mode when the current node completely matches its corresponding reference frame node and the intra prediction mode is not enabled, the embodiments of the present application can select a more appropriate prediction mode for the current node, which is beneficial to improving the prediction coding effect and the encoding quality.

[0333] Furthermore, when entropy encoding the best prediction mode selected by using RDO in the embodiments of the present application, by using the current node as one of the neighbor nodes to determine the context information, or determining that one of the two states of the context information according to the number of occupied child nodes in the current node, it is helpful to reduce the types of contexts, which is beneficial to reducing the computational complexity in the encoding process, reducing the memory overhead, and improving the encoding efficiency.

[0334] Figure 11 Fig. shows a schematic block diagram of a decoding device 1100 provided by the embodiments of the present application. As Figure 11 shown, the decoding device 1100 includes a parsing module 1110, a determining module 1120, a decoding module 1130, and an inverse transformation module 1140.

[0335] A parsing module 1110, configured to parse a bitstream to obtain a reconstructed value of a transform coefficient of a point cloud to be decoded;

[0336] A determining module 1120, configured to determine a transform tree structure of the point cloud to be decoded based on geometric reconstruction information; the transform tree structure includes at least one first node layer;

[0337] The determining module 1120 is further configured to determine, in a reference frame of a current node in the at least one first node layer, a node having the same geometric position as the current node as a reference frame node;

[0338] A decoding module 1130, configured to, if the current node completely matches the reference frame node and the intra prediction mode is not enabled, determine, through entropy decoding, a prediction mode adopted by the current node; wherein the prediction mode is an inter prediction mode or a non-prediction mode;

[0339] An inverse transform module 1140, configured to perform an inverse transform on the reconstructed value of the transform coefficient of a node to be decoded according to the prediction mode to obtain a reconstructed attribute value; wherein the node to be decoded is a child node of the current node.

[0340] Optionally, the inverse transform module 1140 is specifically configured to:

[0341] If the prediction mode is the inter prediction mode, perform a transform on a predicted attribute value of the node to be decoded to obtain a second transform coefficient;

[0342] According to the reconstructed value of the transform coefficient of the node to be decoded and the second transform coefficient, obtain a first transform coefficient;

[0343] Perform an inverse transform on the first transform coefficient to obtain the reconstructed attribute value.

[0344] Optionally, the inverse transform module 1140 is specifically configured to:

[0345] Determine the reconstructed attribute value of a child node at a corresponding position of the reference frame node as the predicted attribute value of the node to be decoded at the corresponding position of the current node.

[0346] Optionally, the inverse transform module 1140 is specifically configured to:

[0347] If the prediction mode is the non-prediction mode, perform an inverse transform on the reconstructed value of the transform coefficient of the node to be decoded to obtain the reconstructed attribute value.

[0348] Optionally, the determining module 1120 is further configured to:

[0349] If the current node does not exactly match the reference frame node and the intra prediction mode is not enabled, determine that the prediction mode adopted by the current node is the non-prediction mode.

[0350] Optionally, the first node layer is the middle node layer of the transform tree structure.

[0351] Optionally, the decoding module 1130 is specifically configured to:

[0352] Determine first context information by using whether the intra prediction mode is enabled and whether the current node exactly matches the reference frame node;

[0353] Determine second context information by using the prediction mode most frequently used among the current node, the neighbor parent node, and the decoded children nodes of the same layer;

[0354] Determine the index of the context corresponding to the prediction mode according to the context information and the second context information;

[0355] Parse the code stream, perform entropy decoding according to the index, and obtain the prediction mode.

[0356] Optionally, the first context information corresponds to one of the following three states:

[0357] The intra prediction mode is enabled and the current node exactly matches the reference frame node;

[0358] The intra prediction mode is enabled and the current node does not exactly match the reference frame node;

[0359] The intra prediction mode is not enabled.

[0360] Optionally, the second context information corresponds to one of the following three states:

[0361] The intra prediction mode is most frequently used among the current node, the neighbor parent node, and the decoded children nodes of the same layer;

[0362] The inter prediction mode is most frequently used among the current node, the neighbor parent node, and the decoded children nodes of the same layer;

[0363] The non-prediction mode is most frequently used among the current node, the neighbor parent node, and the decoded children nodes of the same layer.

[0364] Optionally, the determination module 1120 is further configured to:

[0365] Determine third context information by using the number of occupied children nodes in the current node; wherein, the third context information corresponds to one of two states;

[0366] Determine the index of the context corresponding to the prediction mode according to the first context information, the second context information, and the third context information.

[0367] Optionally, the two states include:

[0368] The number range of the occupied child nodes is the first interval;

[0369] The number range of the occupied child nodes is the second interval; wherein, the range of the available values of the number of child nodes of the current node is divided into the first interval and the second interval.

[0370] In the embodiment of the present application, at the decoding end, when the current node completely matches its corresponding reference frame node and the intra prediction mode is not enabled, the prediction mode of the current node is determined by entropy decoding, which is an inter prediction mode or a non-prediction mode. Then, according to the prediction mode, the inverse transform is performed on the reconstructed values of the transform coefficients of the child nodes of the current node, that is, the nodes to be decoded, to obtain the reconstructed attribute values. Compared with directly using the non-prediction mode as the prediction mode when the current node completely matches its corresponding reference frame node and the intra prediction mode is not enabled, the embodiment of the present application can select a more appropriate prediction mode for the current node, which is beneficial to improving the prediction decoding effect and the decoding quality.

[0371] Further, when determining the prediction mode of the current node by entropy decoding, regarding the current node as one of the neighbor nodes to determine the context information, or determining one of the two states of the context information according to the number of occupied child nodes in the current node, can help reduce the types of contexts, which is beneficial to reducing the computational complexity in the encoding process, reducing the memory overhead, and improving the encoding efficiency.

[0372] The decoding device 1100 provided in the embodiment of the present application can implement Figures 7 to 8 each process implemented by the method embodiment and achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0373] As Figure 12 shown, the embodiment of the present application further provides an electronic device 1200, including a processor 1201 and a memory 1202. A program or instruction that can run on the processor 1201 is stored on the memory 1202. For example, when the electronic device 1200 is an encoding end device, when the program or instruction is executed by the processor 1201, it implements each step of the above encoding method embodiment and can achieve the same technical effect. When the electronic device 1200 is a decoding end device, when the program or instruction is executed by the processor 1201, it implements each step of the above decoding method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Optionally, the memory 1202 may be Figure 1In the memory 102 or memory 113 in the illustrated embodiment, the processor 1201 may implement Figure 1 - the functions of the encoder 200 or decoder 300 in the embodiment shown in FIG. 3.

[0374] An embodiment of the present application further provides an electronic device, including: a memory configured to store video data; and a processing circuit configured to implement each step of the method embodiment as described above. Optionally, the memory may be Figure 1 the memory 102 or memory 113 in the illustrated embodiment, and the processing circuit may implement Figure 1 - the functions of the encoder 200 or decoder 300 in the embodiment shown in FIG. 3.

[0375] An embodiment of the present application further provides an electronic device, including a processor and a communication interface, the communication interface being coupled to the processor, and the processor being configured to run a program or an instruction to implement Figures 4 to 8 the steps in the method embodiment as shown. This device embodiment corresponds to the above method embodiment, and each implementation process and implementation manner of the above method embodiment can be applied to this terminal embodiment, and the same technical effects can be achieved.

[0376] The above electronic device may be a terminal or other devices other than a terminal, such as a server, a Network Attached Storage (NAS), etc.

[0377] Among them, the terminal can be a mobile phone, a tablet personal computer, a laptop computer, a notebook computer, a personal digital assistant (PDA), a handheld computer, a netbook, an ultra-mobile personal computer (UMPC), a mobile internet device (MID), an augmented reality (AR) device, a virtual reality (VR) device, a mixed reality (MR) device, a robot, a wearable device, a flight vehicle, a vehicle user equipment (VUE), a shipborne device, a pedestrian user equipment (PUE), a smart home (home devices with wireless communication functions, such as refrigerators, TVs, washing machines or furniture, etc.), a game console, a personal computer (PC), a teller machine or a self-service machine, etc., which are terminal-side devices. Wearable devices include: smart watches, smart bracelets, smart earphones, smart glasses, smart jewelry (such as smart bracelets, smart bracelets, smart rings, smart necklaces, smart anklets, smart ankle chains, etc.), smart wristbands, smart clothing, etc. Among them, vehicle user equipment can also be referred to as vehicle terminal, vehicle controller, vehicle module, vehicle component, vehicle chip or vehicle unit, etc. It should be noted that the specific type of the terminal is not limited in the embodiments of the present application.

[0378] The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server. The cloud server can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), or cloud computing services based on big data and artificial intelligence platforms, etc.

[0379] Exemplarily, the above-mentioned electronic devices can include, but are not limited to Figure 1 the types of the source device 100 or the destination device 110 shown.

[0380] Taking the electronic device as the terminal as an example, Figure 13 is a schematic diagram of the hardware structure of a terminal for implementing an embodiment of the present application.

[0381] The terminal 1300 includes, but is not limited to, at least some components such as a radio frequency unit 1301, a network module 1302, an audio output unit 1303, an input unit 1304, a sensor 1305, a display unit 1306, a user input unit 1307, an interface unit 1308, a memory 1309, and a processor 1310.

[0382] Those skilled in the art can understand that the terminal 1300 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 1310 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 13 The terminal structure shown does not limit the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0383] It should be understood that in the embodiments of the present application, the input unit 1304 may include a graphics processing unit (GPU) 13041 and a microphone 13042. The graphics processor 13041 processes the image data of a static picture or video obtained by an image acquisition device (such as a camera) in a video acquisition mode or an image acquisition mode, or may process the obtained point cloud data. The display unit 1306 may include a display panel 13061, and the display panel 13061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1307 includes at least one of a touch panel 13071 and other input devices 13072. The touch panel 13071 is also called a touch screen. The touch panel 13071 may include two parts: a touch detection device and a touch controller. The other input devices 13072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.

[0384] In the embodiments of the present application, after receiving downlink data from a network-side device, the radio frequency unit 1301 can transmit it to the processor 1310 for processing; in addition, the radio frequency unit 1301 can send uplink data to the network-side device. Generally, the radio frequency unit 1301 includes, but is not limited to, an antenna, an amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, etc.

[0385] The memory 1309 can be used to store software programs or instructions and various data. The memory 1309 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area may store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1309 may include volatile memory or non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 1309 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.

[0386] The processor 1310 may include one or more processing units; optionally, the processor 1310 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1310 either.

[0387] In some embodiments, the processor 1310 is configured to determine a transform tree structure of a point cloud to be encoded based on geometric reconstruction information, and the transform tree structure includes at least one first node layer;

[0388] Determine a node having the same geometric position as the current node in the reference frame of the current node in the at least one first node layer as a reference frame node;

[0389] If the current node exactly matches the reference frame node and the intra prediction mode is not enabled, then rate-distortion optimization selects the prediction mode adopted by the current node from the inter prediction mode and the non-prediction mode;

[0390] Perform prediction and transformation on the node to be encoded according to the prediction mode to obtain attribute transformation coefficients; the node to be encoded is a child node of the current node;

[0391] Encode the attribute transformation coefficients to obtain a bitstream.

[0392] Therefore, when the current node exactly matches its corresponding reference frame node and the intra prediction mode is not enabled, the encoder selects the prediction mode adopted by the current node from the inter prediction mode and the non-prediction mode through rate-distortion optimization, and then performs transformation, prediction, and encoding processing on the attribute information of the child node of the current node, that is, the node to be encoded, according to the selected appropriate prediction mode to obtain a bitstream. Compared with directly using the non-prediction mode as the prediction mode when the current node exactly matches its corresponding reference frame node and the intra prediction mode is not enabled, the embodiment of the present application can select a more appropriate prediction mode for the current node, which is beneficial to improving the prediction coding effect and the coding quality.

[0393] In some embodiments, the processor 1310 is configured to parse the bitstream to obtain the reconstructed value of the transformation coefficient of the point cloud to be decoded;

[0394] Determine the transformation tree structure of the point cloud to be decoded based on the geometric reconstruction information; the transformation tree structure includes at least one first node layer;

[0395] In the reference frame of the current node in the at least one first node layer, determine the node with the same geometric position as the current node as the reference frame node;

[0396] If the current node exactly matches the reference frame node and the intra prediction mode is not enabled, then determine the prediction mode adopted by the current node through entropy decoding; wherein, the prediction mode is the inter prediction mode or the non-prediction mode;

[0397] Inverse-transform the reconstructed value of the transformation coefficient of the node to be decoded according to the prediction mode to obtain the reconstructed attribute value; wherein, the node to be decoded is a child node of the current node.

[0398] In the embodiments of the present application, when the decoding end determines that the current node completely matches its corresponding reference frame node and the intra prediction mode is not enabled, the prediction mode of the current node is determined by entropy decoding, which is an inter prediction mode or a non-prediction mode. Then, according to the prediction mode, the inverse transform is performed on the reconstruction value of the transform coefficient of the child node of the current node, that is, the node to be decoded, to obtain the reconstructed attribute value. Compared with directly using the non-prediction mode as the prediction mode when the current node completely matches its corresponding reference frame node and the intra prediction mode is not enabled, the embodiments of the present application can select a more appropriate prediction mode for the current node, which is beneficial to improving the prediction decoding effect and the decoding quality.

[0399] It can be understood that the implementation processes of the various implementation manners mentioned in this embodiment can refer to the relevant descriptions of the method embodiments and achieve the same or corresponding technical effects. To avoid repetition, they will not be elaborated here.

[0400] The embodiments of the present application further provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0401] Wherein, the processor is the processor in the terminal described in the above embodiment. The readable storage medium includes computer-readable storage media, such as ROM, RAM, magnetic disks, or optical discs, etc. In some examples, the readable storage medium can be a non-transitory readable storage medium.

[0402] The embodiments of the present application further provide a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement each process of the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0403] It should be understood that the chip mentioned in the embodiments of the present application may include a system-on-chip (also referred to as a system chip, chip system, or system-on-chip), or may include an independent display chip, etc.

[0404] The embodiments of the present application further provide a computer program / program product, which is stored in a storage medium. The computer program / program product is executed by at least one processor to implement each process of the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0405] The embodiments of the present application further provide an encoding and decoding system, including: an encoding end device and a decoding end device. The encoding end device can be used to execute the steps of the above encoding method, and the decoding end device can be used to execute the steps of the above decoding method.

[0406] It should be noted that in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0407] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of a computer software product plus a necessary general hardware platform, and of course, can also be implemented by hardware. This computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes several instructions for causing a terminal or a network-side device to execute the methods described in various embodiments of the present application.

[0408] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms of embodiments without departing from the purpose of the present application and the scope protected by the claims. These embodiments are all within the protection scope of the present application.

Claims

1. A coding method, characterized in that, Including: The encoding end determines a transformation tree structure of the point cloud to be encoded based on geometric reconstruction information, and the transformation tree structure includes at least one first node layer; In the reference frame of the current node in the at least one first node layer, a node having the same geometric position as the current node is determined as a reference frame node; If the current node completely matches the reference frame node and the intra prediction mode is not enabled, then through rate-distortion optimization, a prediction mode adopted by the current node is selected from an inter prediction mode and a non-prediction mode; Predicting and transforming the node to be encoded according to the prediction mode to obtain attribute transformation coefficients; the node to be encoded is a child node of the current node; Encoding the attribute transformation coefficients to obtain a bitstream.

2. The method according to claim 1, wherein The predicting and transforming the node to be encoded according to the prediction mode to obtain attribute transformation coefficients includes: If the prediction mode is the inter prediction mode, determining a predicted attribute value of the node to be encoded; Transforming an original attribute value of the node to be encoded to obtain a first transformation coefficient, and transforming the predicted attribute value to obtain a second transformation coefficient; According to the first transformation coefficient and the second transformation coefficient, obtaining the attribute transformation coefficients.

3. The method according to claim 2, wherein The determining the predicted attribute value of the node to be encoded includes: Determining a reconstructed attribute value of a child node at a corresponding position of the reference frame node as a predicted attribute value of the node to be encoded at a corresponding position of the current node.

4. The method according to claim 1, characterized in that The predicting and transforming the node to be encoded according to the prediction mode to obtain attribute transformation coefficients includes: If the prediction mode is the non-prediction mode, transforming an original attribute value of the node to be encoded to obtain the attribute transformation coefficients.

5. The method according to any one of claims 1-4, characterized in that, Also including: If the current node does not completely match the reference frame node and the intra prediction mode is not enabled, determining that the prediction mode adopted by the current node is the non-prediction mode.

6. The method according to any one of claims 1-5, characterized in that, The first node layer is a middle node layer of the transformation tree structure.

7. The method according to any one of claims 1-6, characterized in that, Also including: Determining first context information by using whether the intra prediction mode is enabled and whether the current node completely matches the reference frame node; Determining second context information by using a prediction mode most frequently used among the current node, a neighbor parent node, and decoded child nodes of the same layer; Determining an index of a context corresponding to the prediction mode according to the first context information and the second context information; Entropy encoding the prediction mode according to the index to obtain prediction mode bits, and adding the prediction mode bits to the bitstream.

8. The method according to claim 7, wherein The first context information corresponds to one of the following three states: The intra prediction mode is enabled and the current node completely matches the reference frame node; The intra prediction mode is enabled and the current node does not completely match the reference frame node; The intra prediction mode is not enabled.

9. The method according to claim 7, wherein The second context information corresponds to one of the following three states: The intra prediction mode is most frequently used among the current node, a neighbor parent node, and decoded child nodes of the same layer; The inter prediction mode is most frequently used among the current node, a neighbor parent node, and decoded child nodes of the same layer; The most frequently used non - prediction mode is used among the current node, the neighbor parent node, and the decoded same - layer child nodes.

10. The method according to any one of claims 7-9, characterized in that It further includes: Determining third context information by using the number of occupied child nodes in the current node; wherein, the third context information corresponds to one of two states. Among them, the determining the index of the context corresponding to the prediction mode according to the first context information and the second context information includes: Determining the index of the context corresponding to the prediction mode according to the first context information, the second context information, and the third context information.

11. The method according to claim 10, wherein The two states include: The range of the number of occupied child nodes is the first interval. The range of the number of occupied child nodes is the second interval; wherein, the range of the possible values of the number of child nodes of the current node is divided into the first interval and the second interval.

12. A decoding method, characterized in that, It includes: The decoding end parses the code stream to obtain the reconstructed value of the transform coefficient of the point cloud to be decoded. Determining the transform tree structure of the point cloud to be decoded based on the geometric reconstruction information. The transform tree structure includes at least one first - node layer. Determining a node with the same geometric position as the current node in the reference frame of the current node in the at least one first - node layer as the reference - frame node. If the current node completely matches the reference - frame node and the intra - prediction mode is not enabled, then determining the prediction mode adopted by the current node through entropy decoding; wherein, the prediction mode is an inter - prediction mode or a non - prediction mode. Performing an inverse transform on the reconstructed value of the transform coefficient of the node to be decoded according to the prediction mode to obtain the reconstructed attribute value; wherein, the node to be decoded is a child node of the current node.

13. The method according to claim 12, characterized in that, The performing an inverse transform on the reconstructed value of the transform coefficient of the node to be decoded according to the prediction mode to obtain the reconstructed attribute value includes: If the prediction mode is the inter - prediction mode, then performing a transform on the predicted attribute value of the node to be decoded to obtain a second transform coefficient. Obtaining a first transform coefficient according to the reconstructed value of the transform coefficient of the node to be decoded and the second transform coefficient. Performing an inverse transform on the first transform coefficient to obtain the reconstructed attribute value.

14. The method according to claim 13, wherein It further includes: Determining the reconstructed attribute value of the child node at the corresponding position of the reference - frame node as the predicted attribute value of the node to be decoded at the corresponding position of the current node.

15. The method according to claim 12, wherein The performing an inverse transform on the reconstructed value of the transform coefficient of the node to be decoded according to the prediction mode to obtain the reconstructed attribute value includes: If the prediction mode is the non - prediction mode, then performing an inverse transform on the reconstructed value of the transform coefficient of the node to be decoded to obtain the reconstructed attribute value.

16. The method according to any one of claims 12-15, characterized in that, It further includes: If the current node does not completely match the reference - frame node and the intra - prediction mode is not enabled, then determining that the prediction mode adopted by the current node is the non - prediction mode.

17. The method according to any one of claims 12-16, characterized in that, The first - node layer is the middle - node layer of the transform tree structure.

18. The method according to any one of claims 12 - 17, characterized in that, The determining the prediction mode adopted by the current node through entropy decoding includes: Determining first context information by using whether the intra - prediction mode is enabled and whether the current node completely matches the reference - frame node. Determine the second context information by using the most frequently used prediction mode among the current node, the neighbor parent node, and the decoded children nodes of the same layer; Determine the index of the context corresponding to the prediction mode according to the context information and the second context information; Parse the code stream, perform entropy decoding according to the index, and obtain the prediction mode.

19. The method according to claim 18, wherein The first context information corresponds to one of the following three states: The intra prediction mode is enabled and the current node exactly matches the reference frame node; The intra prediction mode is enabled and the current node does not exactly match the reference frame node; The intra prediction mode is not enabled.

20. The method according to claim 18, wherein The second context information corresponds to one of the following three states: The most frequently used prediction mode among the current node, the neighbor parent node, and the decoded children nodes of the same layer is the intra prediction mode; The most frequently used prediction mode among the current node, the neighbor parent node, and the decoded children nodes of the same layer is the inter prediction mode; The most frequently used prediction mode among the current node, the neighbor parent node, and the decoded children nodes of the same layer is the non - prediction mode.

21. The method according to any one of claims 18 - 20, characterized in that, Further included: Determine the third context information by using the number of occupied children nodes in the current node; wherein, the third context information corresponds to one of two states; Wherein, the determining the index of the context corresponding to the prediction mode according to the first context information and the second context information includes: Determine the index of the context corresponding to the prediction mode according to the first context information, the second context information, and the third context information.

22. The method according to claim 21, wherein The two states include: The number range of the occupied children nodes is in the first interval; The number range of the occupied children nodes is in the second interval; wherein, the range of the possible values of the number of children nodes of the current node is divided into the first interval and the second interval.

23. An encoding device, characterized in that, Including: A determination module, configured to determine a transformation tree structure of the point cloud to be encoded based on geometric reconstruction information, where the transformation tree structure includes at least one first node layer; The determination module is further configured to determine, in the reference frame of the current node in the at least one first node layer, a node with the same geometric position as the current node as the reference frame node; The determination module is further configured to, if the current node exactly matches the reference frame node and the intra prediction mode is not enabled, select the prediction mode adopted by the current node from the inter prediction mode and the non - prediction mode through rate - distortion optimization; A prediction and transformation module, configured to perform prediction and transformation on the node to be encoded according to the prediction mode to obtain attribute transformation coefficients; the node to be encoded is a child node of the current node; An encoding module, configured to perform encoding processing on the attribute transformation coefficients to obtain a code stream.

24. The device according to claim 23, wherein The determination module is further configured to: Determine the first context information by using whether the intra prediction mode is enabled and whether the current node exactly matches the reference frame node; determine the second context information by using the most frequently used prediction mode among the current node, the neighbor parent node, and the decoded children nodes of the same layer; determine the index of the context corresponding to the prediction mode according to the first context information and the second context information; The encoding module is further configured to perform entropy encoding on the prediction mode according to the index to obtain prediction mode bits, and add the prediction mode bits to the bitstream.

25. The device according to claim 24, characterized in that, The second context information corresponds to one of the following three states: The state corresponding to the most frequently used intra prediction mode among the current node, the neighboring parent node, and the decoded co-layer child nodes; The state corresponding to the most frequently used inter prediction mode among the current node, the neighboring parent node, and the decoded co-layer child nodes; The state corresponding to the most frequently used non-prediction mode among the current node, the neighboring parent node, and the decoded co-layer child nodes.

26. A decoding device, characterized in that, Comprising: A parsing module, configured to parse the bitstream to obtain the reconstructed value of the transform coefficient of the point cloud to be decoded; A determination module, configured to determine the transform tree structure of the point cloud to be decoded based on the geometric reconstruction information; the transform tree structure includes at least one first node layer; The determination module is further configured to determine, in the reference frame of the current node in the at least one first node layer, a node having the same geometric position as the current node as the reference frame node; The decoding module is configured to, if the current node exactly matches the reference frame node and the intra prediction mode is not enabled, determine the prediction mode adopted by the current node through entropy decoding; wherein the prediction mode is an inter prediction mode or a non-prediction mode; An inverse transform module, configured to perform an inverse transform on the reconstructed value of the transform coefficient of the node to be decoded according to the prediction mode to obtain a reconstructed attribute value; wherein the node to be decoded is a child node of the current node.

27. The device according to claim 26, wherein Specifically, the decoding module is configured to: Determine first context information by using whether the intra prediction mode is enabled and whether the current node exactly matches the reference frame node; determine second context information by using the most frequently used prediction mode among the current node, the neighboring parent node, and the decoded co-layer child nodes; Determine an index of the context corresponding to the prediction mode according to the context information and the second context information; Parse the bitstream, and perform entropy decoding according to the index to obtain the prediction mode.

28. The device according to claim 27, wherein The second context information corresponds to one of the following three states: The state corresponding to the most frequently used intra prediction mode among the current node, the neighboring parent node, and the decoded co-layer child nodes; The state corresponding to the most frequently used inter prediction mode among the current node, the neighboring parent node, and the decoded co-layer child nodes; The state corresponding to the most frequently used non-prediction mode among the current node, the neighboring parent node, and the decoded co-layer child nodes.

29. An electronic device, characterized in that, Comprising a processor and a memory, the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented, or the steps of the method according to any one of claims 12 to 22 are implemented.

30. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the method according to any one of claims 1 to 11 is implemented, or the steps of the method according to any one of claims 12 to 22 are implemented.

31. A chip, characterized in that, The chip includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is configured to run programs or instructions to implement the steps of the method according to any one of claims 1 to 11, or to implement the steps of the method according to any one of claims 12 to 22.