Encoding method, decoding method, and related device
By constructing a transformation tree structure and performing cross-attribute prediction residual processing, the problem of low attribute encoding efficiency in point cloud coding is solved, achieving more efficient encoding and decoding performance.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2025-10-29
- Publication Date
- 2026-05-07
AI Technical Summary
Among existing point cloud coding technologies, attribute coding has low coding efficiency and is difficult to meet the requirements of efficient compression and decoding.
By constructing a transformation tree structure for the point cloud, the original attribute value and reference attribute value of the current node are obtained. The scaling parameter is used to transform, quantize and encode the cross-attribute prediction residual, thereby realizing the encoding of the cross-attribute prediction residual.
It improves the encoding efficiency of point cloud data, increases the compression rate of the bitstream, and enhances the encoding performance.
Smart Images

Figure CN2025130786_07052026_PF_FP_ABST
Abstract
Description
Encoding and decoding methods and related equipment
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese Patent Application No. 202411532908.7, filed on October 30, 2024, entitled "Encoding / Decoding Method and Related Equipment", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application belongs to the field of encoding and decoding technology, specifically relating to an encoding and decoding method and related equipment. Background Technology
[0004] With the continuous development of point cloud technology, the compression and encoding of point cloud data has become an important research issue. Currently, both the Audio Video Coding Standard Workgroup of China (AVS) and the Moving Picture Experts Group (MPEG) of the international standardization organization are developing standards for point cloud encoding, such as Geometry-based Point Cloud Compression (G-PCC). How to further improve the performance of point cloud encoding and decoding is an urgent problem to be solved. Summary of the Invention
[0005] This application provides an encoding / decoding method and related equipment that can solve the problem of low encoding efficiency in point cloud attribute encoding.
[0006] Firstly, an encoding method is provided, executed by the encoding end, the method comprising:
[0007] The encoding end constructs a transformation tree structure of the current attributes of the point cloud to be encoded;
[0008] Obtain the first original attribute value and the reference attribute value of the current attribute of the placeholder child node of the current node in the transformed tree structure;
[0009] Based on the first original attribute value, the reference attribute value, and the scaling parameter, the cross-attribute prediction residuals of the current attribute and the reference attribute are obtained; the scaling parameter is used to characterize the correlation between the current attribute and the reference attribute.
[0010] The cross-attribute prediction residuals are transformed, quantized, and encoded to obtain a bitstream.
[0011] Secondly, a decoding method is provided, executed by the decoding end, the method comprising:
[0012] The decoding end decodes and dequantizes the bitstream to obtain the reconstructed value of the first transform coefficient residual of the point cloud to be decoded;
[0013] Construct a transformation tree structure of the current attributes of the point cloud to be decoded;
[0014] Obtain the reference attribute value of the placeholder child node of the current node in the transformed tree structure;
[0015] Cross-attribute prediction is performed based on the reference attribute value and scaling parameters to obtain the first attribute prediction value; the scaling parameters are used to characterize the correlation between the current attribute and the reference attribute.
[0016] The attribute reconstruction values of the placeholder child nodes of the current node are obtained by performing an inverse transformation based on the first attribute prediction value and the reconstructed value of the first transformation coefficient residual.
[0017] Thirdly, an encoding apparatus is provided, comprising:
[0018] The building module is used by the encoding end to construct the transformation tree structure of the current attributes of the point cloud to be encoded;
[0019] The acquisition module is used to acquire the first original attribute value and the reference attribute value of the current attribute of the placeholder child node of the current node of the transformed tree structure;
[0020] The prediction module is used to obtain the cross-attribute prediction residuals of the current attribute and the reference attribute based on the first original attribute value, the reference attribute value, and the scaling parameter; the scaling parameter is used to characterize the correlation between the current attribute and the reference attribute.
[0021] The encoding module is used to transform, quantize, and encode the cross-attribute prediction residuals to obtain a bitstream.
[0022] Fourthly, a decoding apparatus is provided, comprising:
[0023] The parsing module is used by the decoding end to decode and dequantize the bitstream to obtain the reconstructed value of the first transform coefficient residual of the point cloud to be decoded;
[0024] A construction module is used to construct the transformation tree structure of the current attributes of the point cloud to be decoded;
[0025] The acquisition module is used to acquire the reference attribute values of the placeholder child nodes of the current node in the transformed tree structure.
[0026] The prediction module is used to perform cross-attribute prediction based on the reference attribute value and the scaling parameter to obtain a first attribute prediction value; the scaling parameter is used to characterize the correlation between the current attribute and the reference attribute.
[0027] The inverse transformation module is used to perform an inverse transformation based on the first attribute prediction value and the reconstructed value of the first transformation coefficient residual to obtain the attribute reconstruction value of the placeholder child node of the current node.
[0028] Fifthly, an encoding apparatus is provided, the apparatus being configured to perform the steps of the method described in the first aspect.
[0029] In a sixth aspect, a decoding apparatus is provided, the apparatus being configured to perform the steps of the method described in the second aspect.
[0030] In a seventh aspect, an electronic device is provided, the terminal including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first aspect, or implementing the steps of the method as described in the second aspect.
[0031] Eighthly, an electronic device is provided, including a processor and a communication interface, wherein the processor is used for:
[0032] The encoding end constructs a transformation tree structure of the current attributes of the point cloud to be encoded;
[0033] Obtain the first original attribute value and the reference attribute value of the current attribute of the placeholder child node of the current node in the transformed tree structure;
[0034] Based on the first original attribute value, the reference attribute value, and the scaling parameter, the cross-attribute prediction residuals of the current attribute and the reference attribute are obtained; the scaling parameter is used to characterize the correlation between the current attribute and the reference attribute.
[0035] The cross-attribute prediction residuals are transformed, quantized, and encoded to obtain a bitstream.
[0036] Alternatively, the processor is used to:
[0037] The bitstream is decoded and dequantized to obtain the reconstructed value of the first transform coefficient residual of the point cloud to be decoded;
[0038] Construct a transformation tree structure of the current attributes of the point cloud to be decoded;
[0039] Obtain the reference attribute value of the placeholder child node of the current node in the transformed tree structure;
[0040] Cross-attribute prediction is performed based on the reference attribute value and scaling parameters to obtain the first attribute prediction value; the scaling parameters are used to characterize the correlation between the current attribute and the reference attribute.
[0041] The attribute reconstruction values of the placeholder child nodes of the current node are obtained by performing an inverse transformation based on the first attribute prediction value and the reconstructed value of the first transformation coefficient residual.
[0042] A ninth aspect provides an electronic device comprising: a memory configured to store video data, and processing circuitry configured to implement the steps of the method described in the first aspect, or the steps of the method described in the second aspect.
[0043] In a tenth aspect, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect, or implement the steps of the method described in the second aspect.
[0044] Eleventhly, an encoding / decoding system is provided, comprising: an encoding end device and a decoding end device, wherein the decoding end device can be used to perform the steps of the method described in the first aspect, and the encoding end device can be used to perform the steps of the method described in the second aspect.
[0045] In a twelfth aspect, a chip is provided, the chip including a processor and a communication interface coupled to the processor, the processor being configured to run a program or instructions to implement the steps of the method described in the first aspect, or to implement the steps of the method described in the second aspect.
[0046] In a thirteenth aspect, a computer program / program product is provided, which is stored in a storage medium and is executed by at least one processor to implement the steps of the method as described in the first aspect, or to implement the steps of the method as described in the second aspect.
[0047] In this embodiment, the encoding end constructs a transformation tree structure of the current attributes of the point cloud to be encoded, obtains the first original attribute value and the reference attribute value of the current attribute of the placeholder child node of the current node in the transformation tree structure, and obtains the cross-attribute prediction residual of the current specification attribute and the reference attribute based on the first original attribute value, the reference attribute value, and the scaling parameter. Then, the cross-attribute prediction residual is transformed, quantized, and encoded to obtain the bitstream. This embodiment can utilize the correlation between the attribute types of the current node to perform cross-attribute prediction, obtain the cross-attribute prediction residual, and then transform and encode the cross-attribute prediction residual to obtain the bitstream, which can help improve the compression ratio of the bitstream and thus improve encoding efficiency.
[0048] The decoding end decodes and dequantizes the bitstream to obtain the reconstructed value of the first transform coefficient residual of the point cloud to be decoded, and constructs a transform tree structure of the current attribute of the point cloud to be decoded. It then obtains the reference attribute value of the reference attribute of the placeholder child node of the current node in the transform tree structure, and performs cross-attribute prediction based on the reference attribute value and scaling parameters to obtain the predicted value of the first attribute. Finally, it performs an inverse transform based on the predicted value of the first attribute and the reconstructed value of the first transform coefficient residual to obtain the attribute reconstructed value of the placeholder child node of the current node. This embodiment of the application can utilize the correlation between the attribute types of the current node to perform cross-attribute prediction and obtain the reconstructed attribute value of the placeholder child node of the current node, which can help improve the compression rate of the bitstream. Attached Figure Description
[0049] Figure 1 is a schematic diagram of an encoding / decoding system provided in an embodiment of this application;
[0050] Figure 2a is a flowchart of the encoding process performed by an encoder based on the AVS-PCC encoding framework;
[0051] Figure 2b is a flowchart of the encoding process performed by the encoder based on the MPEG G-PCC encoding framework;
[0052] Figure 3a is a flowchart of the decoding process performed by the decoder based on the AVS-PCC decoding framework;
[0053] Figure 3b is a flowchart of the decoding process performed by the decoder based on the MPEG G-PCC decoding framework;
[0054] Figure 4 is a schematic flowchart for determining whether to perform intra-frame prediction;
[0055] Figure 5 is a schematic flowchart for determining whether to perform weighted forecasting;
[0056] Figure 6 is a schematic diagram of the transformation;
[0057] Figure 7 is a schematic flowchart of an encoding method provided in an embodiment of this application;
[0058] Figure 8 is a schematic flowchart of a decoding method provided in an embodiment of this application;
[0059] Figure 9 is a schematic block diagram of an encoding device provided in an embodiment of this application;
[0060] Figure 10 is a schematic block diagram of a decoding device provided in an embodiment of this application;
[0061] Figure 11 is a schematic block diagram of an electronic device provided in an embodiment of this application;
[0062] Figure 12 is a schematic diagram of the hardware structure of a terminal implementing an embodiment of this application. Detailed Implementation
[0063] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0064] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, the scope of protection for "A or B" covers at least three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. In addition, the terms "A and / or B," "at least one of A and B," and "at least one of A or B" also cover at least the above three scenarios. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0065] Before introducing the technical solutions provided in the embodiments of this application, the meanings of some terms will be explained first.
[0066] Point cloud: A point cloud is a set of discrete points in space that are randomly distributed and represent the spatial structure and surface properties of a three-dimensional object or scene. Point clouds can be classified into different categories according to different classification criteria. For example, according to the method of acquiring the point cloud, it can be divided into dense point clouds and sparse point clouds; or according to the temporal type of the point cloud, it can be divided into static point clouds and dynamic point clouds.
[0067] Point cloud data: Point cloud data is composed of the geometric coordinates and attribute information of each point. Geometric coordinate information, also known as 3D position information, refers to the spatial coordinates (x, y, z) of a point in the point cloud. This can include the coordinate values of the point along each coordinate axis of a 3D coordinate system, such as the coordinate value x along the X-axis, the coordinate value y along the Y-axis, and the coordinate value z along the Z-axis. The attribute information of a point in the point cloud can include at least one of the following: color information, material information, and laser reflection intensity information (also known as reflectivity). Typically, each point in the point cloud has the same number of attribute information. For example, each point in the point cloud can have both color information and laser reflection intensity information, or it can have color information, material information, and laser reflection intensity information.
[0068] Point cloud compression (PCC) refers to the process of encoding the geometric coordinates and attribute information of each point in a point cloud to obtain a compressed bitstream. Point cloud compression includes two main processes: geometric coordinate information encoding and attribute information encoding. Currently, point cloud compression frameworks that can compress point clouds include the Geometry Point Cloud Compression (G-PCC) or Video Point Cloud Compression (V-PCC) framework provided by the Moving Picture Experts Group (MPEG), or the AVS-PCC framework provided by the Audio Video Standard (AVS).
[0069] Point cloud decoding: Point cloud decoding refers to decoding the compressed bitstream obtained from point cloud encoding to reconstruct the point cloud. More specifically, it refers to the process of reconstructing the geometric coordinates and attribute information of each point in the point cloud based on the geometric bitstream and attribute bitstream in the compressed bitstream. After obtaining the compressed bitstream at the decoding end, for the geometric bitstream, entropy decoding is first performed to obtain the quantized information of each point in the point cloud, and then inverse quantization is performed to reconstruct the geometric coordinates of each point in the point cloud. For the attribute bitstream, entropy decoding is first performed to obtain the quantized attribute residual information or quantized transform coefficients of each point in the point cloud; then, inverse quantization is performed on the quantized attribute residual information to obtain the reconstructed residual information, and inverse quantization is performed on the quantized transform coefficients to obtain the reconstructed transform coefficients. The reconstructed transform coefficients are then inversely transformed to obtain the reconstructed residual information. Based on the reconstructed residual information of each point in the point cloud, the attribute information of each point in the point cloud can be reconstructed. The reconstructed attribute information of each point in the point cloud is then matched one-to-one with the reconstructed geometric coordinate information in sequence to reconstruct the point cloud.
[0070] Figure 1 is a schematic diagram of the encoding / decoding system 10 provided in an embodiment of this application. The technical solution of this application embodiment relates to encoding / decoding (CODEC) point cloud data (including encoding or decoding).
[0071] As shown in Figure 1, the encoding / decoding system 10 includes a source device 100, which provides encoded point cloud data to be decoded and displayed by the destination device 110. Specifically, the source device 100 provides the point cloud data to the destination device 110 via a communication medium 120. The source device 100 and the destination device 110 may include any one or more of the following: desktop computer, laptop computer, tablet computer, set-top box, mobile phone, wearable device (e.g., smartwatch or wearable camera), television, camera, display device, in-vehicle device, virtual reality (VR) device, augmented reality (AR) device, mixed reality (MR) device, digital media player, video game console, video conferencing equipment, video streaming equipment, broadcast receiver equipment, broadcast transmitter equipment, spacecraft, aircraft, robot, satellite, etc.
[0072] In the example of Figure 1, source device 100 includes a data source 101, a memory 102, an encoder 200, and an output interface 104. Destination device 110 includes an input interface 111, a decoder 300, a memory 113, and a display device 114. Source device 100 represents an example of an encoding device, while destination device 110 represents an example of a decoding device. In other examples, source device 100 and destination device 110 may not include some of the components shown in Figure 1, or they may include components other than those shown in Figure 1. For example, source device 100 may acquire point cloud data through an external capture device. Similarly, destination device 110 may interface with an external display device instead of including an integrated display device. Furthermore, memory 102 and memory 113 may be external memories.
[0073] Although Figure 1 illustrates the source device 100 and the destination device 110 as separate devices, in some examples, they may be integrated into a single device. In such embodiments, the same hardware or software, separate hardware or software, or any combination thereof may be used to implement the functionality corresponding to the source device 100 and the functionality corresponding to the destination device 110.
[0074] In some examples, source device 100 and destination device 110 can perform unidirectional or bidirectional data transmission. In the case of bidirectional data transmission, source device 100 and destination device 110 can operate in a substantially symmetrical manner, i.e., each of source device 100 and destination device 110 includes an encoder and a decoder.
[0075] Data source 101 represents the source of point cloud data (i.e., raw, unencoded point cloud data) and provides the point cloud data to encoder 200, which encodes the point cloud data. Source device 100 may include capture devices (e.g., camera devices, sensing devices, or scanning devices), archives containing previously captured point cloud data, or feed interfaces for receiving point cloud data from data content providers. Camera devices may include ordinary cameras, stereo cameras, and light field cameras; sensing devices may include laser devices, radar devices, etc.; and scanning devices may include 3D laser scanning devices, etc. Point cloud data can be obtained by capturing real-world visual scenes using capture devices. Alternatively, data source 101 may generate computer graphics-based data as source data, or combine real-time data, archived data, and computer-generated data. For example, the data source may generate point cloud data based on virtual objects (e.g., virtual 3D objects and virtual 3D scenes obtained through 3D modeling).
[0076] Encoder 200 encodes captured, pre-captured, or computer-generated data. Encoder 200 can rearrange point cloud data from the received order (sometimes referred to as the "display order") according to the encoded order. Encoder 200 can generate a bitstream including the encoded point cloud data. Source device 100 can then output the encoded point cloud data to communication medium 120 via output interface 104 for reception or retrieval, for example, by input interface 111 of destination device 110.
[0077] The memory 102 of the source device 100 and the memory 113 of the destination device 110 represent general-purpose memory. In some examples, memory 102 may store raw data from data source 101, and memory 113 may store decoded point cloud data from decoder 300. Additionally or alternatively, memories 102 and 113 may respectively store software instructions executable by, for example, encoder 200 and decoder 300. Although memories 102 and 113 are shown separately from encoder 200 and decoder 300 in this example, it should be understood that encoder 200 and decoder 300 may also include internal memory for functionally similar or equivalent purposes. If encoder 200 and decoder 300 are deployed on the same hardware device, memories 102 and 113 may be the same memory. Furthermore, memories 102 and 113 may store, for example, encoded point cloud data output from encoder 200 and input to decoder 300. In some examples, portions of memories 102 and 113 may be allocated as one or more point cloud buffers, for example, to store raw, decoded, or encoded point cloud data.
[0078] In some examples, source device 100 can output encoded data from output interface 104 to memory 113. Similarly, destination device 110 can access encoded data from memory 113 via input interface 111. Memory 113 or memory 102 can include any of a variety of distributed or locally accessed data storage media, such as hard drives, Blu-ray discs, digital versatile discs (DVDs), compact disc read-only memory (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded point cloud data.
[0079] Output interface 104 may include any type of medium or device capable of transmitting encoded point cloud data from source device 100 to destination device 110. For example, output interface 104 may include a transmitter or transceiver, such as an antenna, configured to transmit encoded point cloud data directly from source device 100 to destination device 110 in real time. The encoded point cloud data may be modulated according to the communication standards of a wireless communication protocol and transmitted to destination device 110.
[0080] Communication medium 120 may include transient media, such as wireless broadcasting or wired network transmission. For example, communication medium 120 may include radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). Communication medium 120 may form part of a packet-based network (such as a local area network, a wide area network, or a global network such as the Internet). Communication medium 120 may also take the form of a storage medium (e.g., a non-transitory storage medium), such as a hard disk, flash drive, compact disk, digital point cloud disk, Blu-ray disc, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded point cloud data.
[0081] In some implementations, the communication medium 120 may include a router, switch, base station, or any other device that can be used to facilitate communication from source device 100 to destination device 110. For example, a server (not shown) may receive encoded point cloud data from source device 100 and provide it to destination device 110, for example, via network transmission. The server may include (e.g., a web server for a website), a server configured to provide file transfer protocol services (such as File Transfer Protocol (FTP) or File Delivery Over Unidirectional Transport (FLUTE) protocol), a content delivery network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Services (MBMS) or Evolved Multimedia Broadcast Multicast Service (eMBMS) server, or a Network-attached storage (NAS) device, etc. The server can implement one or more HTTP streaming protocols, such as MPEG Media Transport (MMT), Dynamic Adaptive Streaming over HTTP (DASH), HTTP Live Streaming (HLS), or Real Time Streaming Protocol (RTSP).
[0082] Destination device 110 can access encoded point cloud data from a server, for example, via a wireless channel (e.g., Wi-Fi connection) or a wired connection (e.g., Digital subscriber line (DSL), cable modem, etc.) for accessing encoded point cloud data stored on the server.
[0083] Output interface 104 and input interface 111 can represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to the IEEE 802.11 or IEEE 802.15 standard (e.g., ZigBee™), Bluetooth standard, or other physical components. In an example where output interface 104 and input interface 111 include wireless components, output interface 104 and input interface 111 can be configured to transmit data, such as encoded point cloud data, via Wi-Fi, Ethernet, or cellular networks (such as 4G, LTE (Long Term Evolution), Advanced LTE, 5G, 6G, etc.).
[0084] The technology provided in this application can be applied to support one or more of the following application scenarios: machine-perceived point clouds, which can be used in autonomous navigation systems, real-time inspection systems, geographic information systems, visual sorting robots, disaster relief robots, and other scenarios; human-perceived point clouds, which can be used in point cloud application scenarios such as digital cultural heritage, free-viewpoint broadcasting, 3D immersive communication, and 3D immersive interaction.
[0085] The input interface 111 of the destination device 110 receives an encoded bitstream from the communication medium 120. The encoded bitstream may include high-level syntax elements and encoded data units (e.g., sequences, image groups, images, slices, blocks, etc.), where the high-level syntax elements are used to decode the encoded data units to obtain decoded point cloud data. The display device 114 displays the decoded point cloud data to the user. The display device 114 may include a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices. In some examples, the destination device 110 may not have a display device 114; for example, if the decoded point cloud data is used to determine the location of a physical object, the display device 114 may be replaced by a processor.
[0086] The encoder 200 and decoder 300 can be implemented as one or more of various processing circuits, which may include microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. When the technology is implemented wholly or partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and use one or more processors to execute the instructions in hardware to perform the technology provided in the embodiments of this application.
[0087] The basic principles of the encoder 200 and decoder 300 provided in this application embodiment are introduced below, taking the G-PCC and AVS-PCC codec frameworks as examples.
[0088] The encoding and decoding frameworks of G-PCC and AVS-PCC are largely the same. Figure 2a shows the encoding flowchart executed by the encoder based on the AVS-PCC encoding framework, and Figure 2b shows the encoding flowchart executed by the encoder based on the MPEG G-PCC encoding framework. The encoder mentioned above can be the encoder 200 shown in Figure 1. The above encoding frameworks can generally be divided into a geometric coordinate information encoding process and an attribute information encoding process. In the geometric information encoding process, the geometric coordinate information of each point in the point cloud is encoded to obtain a geometric bitstream; in the attribute information encoding process, the attribute information of each point in the point cloud is encoded to obtain an attribute bitstream; the geometric bitstream and the attribute bitstream together constitute the compressed bitstream of the point cloud.
[0089] For the geometric information encoding process, the encoding flow executed by encoder 200 is as follows:
[0090] 1. Pre-processing: This can include coordinate transformation and voxelization. Through scaling and translation operations, pre-processing converts the point cloud data in 3D space into integer form and moves its smallest geometric position to the origin. In some examples, encoder 200 may not perform pre-processing.
[0091] 2. Geometric Coding: For the AVS-PCC coding framework, geometric coding includes two modes: octree-based geometric coding and prediction tree-based geometric coding. For the G-PCC coding framework, geometric coding includes three modes: octree-based geometric coding, trisoup-based geometric coding, and prediction tree-based prediction coding. Among them:
[0092] Octree-based geometric encoding: An octree is a tree-like data structure that uniformly divides a predefined bounding box in three-dimensional space, with each node having eight child nodes. By using "1" and "0" to indicate whether each child node of the octree is occupied, occupancy code information is obtained as the bitstream of point cloud geometric information.
[0093] Geometric coding based on prediction trees: A prediction tree is generated using a prediction strategy. Starting from the root node of the prediction tree, each node is traversed, and the residual coordinate value corresponding to each traversed node is encoded.
[0094] Geometric encoding based on triangulation: The point cloud is divided into blocks of a certain size, and the intersection points (called vertices) of the point cloud surface with the edges of the blocks are located. Geometric information is compressed by encoding whether there are intersection points on the edges of the blocks and the positions of the intersection points.
[0095] 3. Geometric Entropy Encoding: This method uses statistical compression encoding on the occupancy code information of the octree, the prediction residual information of the prediction tree, and the vertex information of the triangular representation, finally outputting a binary (0 or 1) compressed bitstream. Statistical coding is a lossless coding method that can effectively reduce the bit rate required to represent the same signal. A commonly used statistical coding method is Content Adaptive Binary Arithmetic Coding (CABAC).
[0096] 4. Geometric Reconstruction: Decoding and reconstructing the geometric information after geometric encoding.
[0097] For the attribute information encoding process, the encoding flow executed by encoder 200 is as follows:
[0098] 1. Color Transformation: Apply transformations to change the color information of an attribute to a different domain. For example, color information can be transformed from the RGB color space to the YCbCr color space.
[0099] 2. Attribute Recoloring: In lossy encoding, after encoding the geometric coordinate information, the encoding end needs to decode and reconstruct the geometric information, that is, restore the geometric information of each point in the point cloud. Attribute information corresponding to one or more neighboring points in the original point cloud is used as the attribute information for the reconstructed point.
[0100] In some examples, encoder 200 may not perform color transformation or attribute recoloring.
[0101] 3. Attribute information processing: In AVS-PCC, attribute information processing can include three modes: prediction coding, transformation coding, and prediction & transformation coding. These three coding modes can be used under different conditions.
[0102] Predictive coding refers to determining the neighboring points of the point to be coded as prediction points among the already coded points based on information such as distance or spatial relationships. Based on set criteria, the predicted attribute information of the point to be coded is calculated according to the attribute information of the prediction points. The difference between the actual attribute information and the predicted attribute information of the point to be coded is calculated as attribute residual information. This attribute residual information is then quantized, transformed (optional), and entropy encoded.
[0103] Transform coding refers to using transformation methods such as Discrete Cosine Transform (DCT) and Haar Transform (Haar) to group and transform attribute information, quantize the transformation coefficients, obtain attribute reconstruction information through inverse quantization and inverse transformation, calculate the difference between the real attribute information and the attribute reconstruction information to obtain attribute residual information and quantize it, and then entropy-encode the quantized transformation coefficients and attribute residuals.
[0104] Predictive transform coding refers to using the attribute residual information obtained from prediction to perform transformation, and then quantizing and entropy coding the transform coefficients.
[0105] In MPEG G-PCC, attribute information processing can include three modes: Prediction Transform coding, Lifting Transform coding, and Region Adaptive Hierarchical Transform (RAHT) coding. These three coding modes can be used under different conditions.
[0106] Predictive transform coding refers to dividing the point cloud into multiple different levels of detail (LoD) based on distance-selected subsets of points, achieving a multi-quality, hierarchical point cloud representation from coarse to fine. Bottom-up prediction is possible between adjacent layers, where neighboring points in the coarse layer predict the attribute information of points introduced in the fine layer, obtaining the corresponding attribute residual information. The points at the lowest level are encoded as reference information.
[0107] Lift transform coding refers to introducing a weight update strategy for neighboring points on the basis of LoD neighboring layer prediction, and finally obtaining the predicted attribute information of each point and the corresponding attribute residual information.
[0108] Hierarchical region adaptive transform coding refers to the process of transforming attribute information into the transform domain, which is called the transform coefficient.
[0109] 4. Attribute Quantization: The fineness of quantization is usually determined by the quantization parameters. The transformation coefficients or attribute residuals obtained from attribute information processing are quantized, and the quantized results are entropy-coded. For example, in predictive transform coding and boost transform coding, entropy coding is performed on the quantized attribute residuals; in RAHT, entropy coding is performed on the quantized transform coefficients.
[0110] 5. Entropy Coding: The quantized attribute residual information and / or transform coefficients are generally compressed using run-length coding and arithmetic coding. The corresponding coding mode, quantization parameters, and other information are also encoded using an entropy encoder.
[0111] The encoder 200 encodes the geometric coordinate information of each point in the point cloud to obtain a geometric bitstream, and encodes the attribute information of each point in the point cloud to obtain an attribute bitstream. The encoder 200 can transmit the encoded geometric bitstream and attribute bitstream together to the decoder 300.
[0112] Figure 3a shows a decoding flowchart executed by the decoder in the AVS-PCC-based decoding framework, and Figure 3b shows a decoding flowchart executed by the decoder in the MPEG G-PCC-based decoding framework. The decoder can be the decoder 300 shown in Figure 1. After receiving the compressed bitstream (i.e., attribute bitstream and geometric bitstream) transmitted by the encoder 200, the decoder 300 decodes the geometric bitstream to reconstruct the geometric coordinate information of each point in the point cloud, and decodes the attribute bitstream to reconstruct the attribute information of each point in the point cloud.
[0113] The decoding process performed by decoder 300 is as follows:
[0114] 1. Entropy Decoding: Perform entropy decoding on the geometric bitstream and attribute bitstream respectively to obtain geometric syntax elements and attribute syntax elements.
[0115] 2. Geometric Decoding: For the AVS-PCC coding framework, geometric decoding includes two modes: octree-based geometric decoding and prediction tree-based geometric decoding. For the G-PCC coding framework, geometric decoding includes three modes: octree-based geometric decoding, trisoup-based geometric decoding, and prediction tree-based prediction decoding.
[0116] Octree-based geometric decoding: reconstructing the octree based on the geometric syntax elements obtained from the geometric bitstream.
[0117] Geometric Decoding Based on Prediction Trees: Reconstructing the prediction tree based on the geometric syntax elements obtained from parsing the geometric bitstream.
[0118] Geometric Decoding Based on Triangle Representation: Reconstructing the triangular model based on the geometric syntax elements obtained from parsing the geometric bitstream.
[0119] 3. Geometric Reconstruction: Perform reconstruction to obtain the geometric coordinate information of the points in the point cloud.
[0120] 4. Inverse coordinate transformation: Perform an inverse transformation on the reconstructed geometric coordinate information to convert the reconstructed coordinates (positions) of points in the point cloud from the transformation domain back to the initial domain.
[0121] 5. Dequantization: Dequantizes attribute syntax elements.
[0122] 6. Attribute Information Processing: In AVS-PCC, attribute information processing determines the color information of points in the point cloud by predicting or predicting the transformation of the inverse-quantized prediction residual or prediction residual transformation coefficients, or by transforming the transformation coefficients of the inverse-quantized transformation.
[0123] In MPEG G-PCC, attribute information processing determines the color information of points in the point cloud by using RAHT to invert the attribute information, or by using LOD and inverse boosting to determine the color information of points in the point cloud.
[0124] 7. Inverse Color Transformation: Transforms color information from the YCbCr color space to the RGB color space. In some examples, the inverse color transformation operation may not be necessary.
[0125] This application relates to the RATH process. Based on the prediction method used, it can be divided into intra-frame prediction RAHT and inter-frame prediction RAHT. Intra-frame prediction uses only the nodes already encoded in the current frame to predict the current node, while inter-frame prediction uses the nodes of a reference frame to predict the current node. This application may involve intra-frame prediction RAHT or inter-frame prediction RAHT. Intra-frame prediction RAHT is described in detail below.
[0126] Intra-frame RAHT technology, also known as Region Adaptive Transform based on upsampling prediction, performs the transformation layer by layer from top to bottom. In each layer, it is applied to each cell node containing 2×2×2 child nodes. When encoding a cell node with 8 child nodes in the current layer, it uses its parent node (upper layer), coplanar neighbor parent node (upper layer), collinear neighbor parent node (upper layer), and already encoded coplanar and collinear neighbor child nodes (same layer) to perform weighted prediction of the attributes of the placeholder child node. The predicted attribute values are then subtracted from the original values to obtain the prediction residual, which is then subjected to RAHT transformation to obtain the transform coefficient residual. This residual is then quantized and encoded. However, when the number of child nodes to be encoded, the number of their grandparent nodes, and the number of their parent nodes do not meet certain conditions, intra-frame prediction is not enabled, and the original attribute values of the current node to be encoded are directly subjected to RAHT transformation, followed by quantization entropy encoding of the transform coefficients.
[0127] The following is a specific example of the intra-frame RAHT technique:
[0128] The first step is to construct the transformation tree structure. Starting from the bottom layer, the transformation tree structure is built layer by layer from the bottom up. During the construction of the transformation tree structure, it is necessary to generate corresponding Morton code information, attribute information, and weight information for the merged nodes. Here, the transformation tree structure can be an octree structure.
[0129] The second step is to analyze the transformation tree structure from top to bottom. Starting from the root node, upsampling prediction and Region Adaptive Hierarchical Transform (RAHT) are performed layer by layer from top to bottom.
[0130] If the current node is the root node, no intra-frame prediction is performed. Instead, the node's attribute information is directly transformed using RAHT to obtain the DC coefficient and AC coefficient.
[0131] If it is not the root node, it is necessary to determine whether to perform intra-frame prediction on the current eight child nodes. For example, this can be done according to the process shown in Figure 4.
[0132] 401. Determine if the number of placeholder child nodes (NumValidC) is equal to 1. If the number of placeholder child nodes (not empty) of the current node to be encoded is 1, then proceed to step 405, set the number of its neighboring parent nodes (NumValidP) to Val, disable prediction, directly perform RAHT transformation on the original attribute information of the current node, and then quantize and entropy encode the obtained AC coefficients. Optionally, Val can be configured to 19. Otherwise, continue to step 402.
[0133] 402. Determine if the number of neighboring grandparent nodes (NumValidP) is greater than or equal to TH1. If the number of neighboring grandparent nodes (including grandparent nodes) of the current child node to be encoded is less than TH1, no prediction is performed. Instead, the original attribute information of the current node is directly subjected to RAHT transformation, and then the resulting AC coefficients are quantized and entropy encoded. Optionally, TH1 can be configured to 2. Otherwise, proceed to step 403.
[0134] 403. Neighbor Search. Optionally, the search scope includes: the parent node of the current child node to be encoded (1 node), the coplanar neighbors of the parent node of the current child node to be encoded (6 nodes), the collinear neighbors of the parent node of the current child node to be encoded (12 nodes), the coplanar neighbors of the current child node to be encoded (6 nodes), and the collinear neighbors of the current child node to be encoded (12 nodes). The above neighbors are searched sequentially. If a neighbor exists, its corresponding index information is recorded, along with the number of neighbors of the parent node (including the parent node itself). Then proceed to step 404.
[0135] 404. Determine if the number of neighboring parent nodes is greater than or equal to TH2. If the number of neighboring parent nodes of the current child node to be encoded is less than the threshold of 6, no prediction is performed; instead, the original attribute information of the current node to be encoded is directly subjected to RAHT transformation, and then the resulting AC coefficients are quantized and entropy encoded. Otherwise, weighted prediction is performed. Optionally, TH2 can be configured to 6.
[0136] The third step is weighted prediction. The nearest neighbors found in the neighbor search are used to perform weighted prediction on each child node of the node to be encoded. (See Figure 5.)
[0137] The prediction weight of the parent node is set to 9, the prediction weight of the neighboring child node that is coplanar with the current child node to be encoded is 5, the prediction weight of the neighboring child node that is collinear with the current child node to be encoded is 2, the prediction weight of the neighboring parent node that is coplanar with the current child node to be encoded is 3, and the prediction weight of the neighboring parent node that is collinear with the current child node to be encoded is 1.
[0138] In this process, the parent node can be directly used to predict each child node of the current block to be encoded. For other neighbor nodes (neighbor parent nodes and encoded neighbor child nodes), a step is needed to determine whether they can be used to predict the child nodes of the current block. The steps for this determination are shown in the figure below:
[0139] 501. Determine if the parent node's neighbor node Parent[i] (1-7) exists. If it does not exist, continue to determine the next parent neighbor node. If it exists, determine in turn whether the current parent neighbor node Parent can be used to predict the child node j (0-7) in the current parent node.
[0140] Specifically, for neighboring parent nodes with indices 1-7, since their child nodes have not yet been encoded, prediction can only be made using the neighboring parent nodes. First, it's determined whether a neighboring parent node exists at that position. If not, the next neighboring parent node is checked. If it exists, it's necessary to determine whether the current neighboring parent node can be used for prediction, thus filtering neighboring parent nodes. The specific process is as follows:
[0141] 1) Two prediction thresholds are set based on the attribute values of the parent node to further filter nearest neighbors, eliminating unreasonable neighbor parent nodes to improve prediction accuracy. Let these two thresholds be limitLow and limitHigh, and let the attribute value of the parent node be attrpar, then: limitLow = attrpar * 2 limitHigh = attrpar * 25
[0142] Suppose the attribute value of the current node is attrNei, and perform the following judgment on it: limitLow <attrnei*10<limitHigh
[0143] If the condition is not met, the current neighboring parent node cannot be used to predict the child nodes of the current block to be encoded; if the condition is met, the following judgment is performed.
[0144] 2) Determine whether the current neighboring parent node satisfies the condition that it is coplanar and collinear with each of the current child nodes to be encoded. If this condition is not met, the current neighboring nodes cannot be used to perform weighted prediction on the current child nodes to be encoded; if this condition is met, the current neighboring nodes are used to perform weighted prediction on the current child nodes to be encoded.
[0145] 502. Continue to check the parent node's neighbors (8-19). If the parent node's neighbors exist, check if there are any child nodes at the same level in the same position of the current parent node's neighbors. If they exist, use the child nodes for prediction and do not use the parent nodes for prediction. That is, the parent neighbor node is in the same position relative to the parent node as the child neighbor node is in the same position relative to the child node.
[0146] Specifically, for neighboring parent nodes with indices 8-19, since their child nodes have already been encoded, if there are neighboring child nodes at the same position, they can be used to replace their respective neighboring parent nodes for prediction, resulting in better prediction performance.
[0147] First, determine if a neighboring parent node exists at this location. If not, continue checking the next neighboring parent node. If a neighboring parent node exists, determine if it can be used for prediction, thus filtering neighboring parent nodes. The specific process is as follows:
[0148] 1) Two prediction thresholds are set based on the attribute values of the parent node to further filter nearest neighbors, eliminating unreasonable neighbor parent nodes to improve prediction accuracy. Let the attribute value of the current node be attrNei, and the following judgment be made on it: limitLow <attrnei*10<limitHigh
[0149] If the condition is not met, the current neighboring parent node cannot be used to predict the child nodes of the current block to be encoded; if the condition is met, the following judgment is performed.
[0150] 2) Determine if there are any child nodes on the same floor at the same position of the current parent node's neighbors: For coplanar neighbor parent nodes, if one of their child nodes is also coplanar with the current child node to be predicted, then use that coplanar neighbor child node instead of the coplanar neighbor parent node for weighted prediction of the child node to be predicted; for collinear neighbor parent nodes, if one of their child nodes is also collinear with the current child node to be predicted, then use that collinear neighbor child node instead of the collinear neighbor parent node for weighted prediction. If no such child node exists, continue using the current neighbor parent node for weighted prediction.
[0151] 503. The child nodes of the current node to be encoded use the neighboring nodes that meet the conditions as a set of reference points to perform weighted prediction and obtain the attribute prediction value.
[0152] The fourth step is the RAHT transformation, and the process is as follows:
[0153] If no prediction is performed on the current block, only the original attribute values need to be transformed using RAHT; if intra-frame prediction is performed on the current block, then the intra-frame prediction residuals of the attributes need to be transformed using RAHT.
[0154] Before performing the RAHT transformation, the original attribute values and the predicted attribute values (if prediction is performed) need to be normalized first, and then the RAHT transformation is performed on the processed values.
[0155] First, the original attribute values are normalized. Let the original attribute value of the current child node be A. i The weight is w i(The size is the number of points contained in the current node), then
[0156] If prediction is used, let the predicted attribute value of the current child node be A. pi ,but:
[0157] The predicted value of the processed attribute A′ pi Or the original attribute value A′ i After subtraction, a RAHT transformation is performed. For a 2x2x2 node block, the transformation is performed along three directions, with four transformations in each direction. Therefore, as shown in Figure 6, each 2x2x2 block undergoes a total of:
[0158] 1) Perform the transformation along the first direction to obtain the low-frequency L node and the high-frequency H node;
[0159] 2) Transform the L and H nodes along the second direction to obtain the low-frequency LL node and the high-frequency LH, HL, and HH nodes;
[0160] 3) Transform the LL, LH, HL, and HH nodes along the third direction to obtain the low-frequency LLL node and the high-frequency LLH, LHL, LHH, HLL, HLH, HHL, and HHH nodes.
[0161] Where LLL is the DC coefficient, and LLH, LHL, LHH, HLL, HLH, HHL, and HHH are the AC coefficients.
[0162] The DC and AC coefficients of the root node are quantized and entropy encoded in the order of (LLL)0, LLH(4), LHL(2), HLL(1), LHH(6), HLH(5), HHL(3), HHH(7).
[0163] For the remaining nodes in the remaining layers, only the AC coefficients are quantized and entropy encoded in the order of LLH(4), LHL(2), HLL(1), LHH(6), HLH(5), HHL(3), HHH(7).
[0164] Due to the sparsity of point clouds, each 2*2*2 node block typically occupies fewer than 8 child nodes. Therefore, not all AC coefficients will exist, and non-existent AC coefficients will not be encoded.
[0165] Specifically, when performing a two-point transformation, assume the input attribute values are T. 01 T 11 Transformation coefficients T1, The transformation formula used is:
[0166] Where a and b are calculated from the weights of the current node, and the weights of the current node are obtained from the weight array:
[0167] If the current node is the root node and no prediction is performed, the DC and AC coefficients of the original attribute values after transformation are quantized and entropy encoded. If the current node is not the root node and no prediction is performed, the AC coefficients of the original attribute values after transformation are quantized and entropy encoded. If the current node is not the root node and prediction is performed, the AC coefficient residuals are quantized and entropy encoded.
[0168] The coefficients are decoded and dequantized at the decoding end. Then, an inverse RAHT transform is performed. The inverse RAHT transform is the reverse process of the RAHT transform, and the inverse transform formula is as follows:
[0169] Among them, T1, T represents the transformation coefficients. 01 T 11 The calculation methods for a and b are the same during the reconstruction of attribute values.
[0170] If the current node is the root node, the inverse transform is performed directly to obtain the attribute reconstruction value of each child node. If the current node is not the root node and no prediction is performed, the DC coefficients inherit the attribute reconstruction value of the parent node, and then the inverse transform is performed with the reconstructed AC coefficients to obtain the attribute reconstruction value of each child node of the current node. If the current node is not the root node and prediction is performed, the DC coefficients inherit the attribute reconstruction value of the parent node, and the difference between the DC coefficient residual and the intra-frame predicted parent node attribute value is obtained. Then, the inverse transform is performed together with the reconstructed AC coefficient residual to obtain the reconstructed attribute prediction residual. Finally, the attribute prediction value is added to obtain the attribute reconstruction value of each child node.
[0171] In the existing GPCC attribute encoding process, the compression efficiency is improved by utilizing the correlation between the reconstructed temporal and spatial domains of the current attribute. When point cloud data contains multiple categories of attribute information, there is a certain correlation between different types of attribute information.
[0172] The relevant scheme utilizes this correlation to propose a RHATT cross-attribute prediction scheme.
[0173] Specifically, firstly, an enabling flag `cross_attr_prediction_enabled_flag` for inter-attribute type prediction is introduced into the sequence parameter sets (sps) to control whether the encoding and decoding process enables the cross-attribute type prediction method. When this flag is enabled (i.e., the flag is 1), `cross_attr_type_prediction_enabled_flag` is introduced into the attribute parameter sets (aps) corresponding to each attribute to be encoded. This controls whether the attribute can use the cross-attribute prediction method. If there are already encoded attributes of different types before encoding the current attribute, then the current attribute to be encoded can use the cross-attribute prediction method. That is, the attribute prediction incorporates the information of the already encoded attributes of different types, utilizing their correlation to improve the attribute encoding efficiency.
[0174] When `cross_attr_type_prediction_enabled_flag` is enabled (i.e., the flag is 1), `attrRefIdx` is introduced into the attribute parameter set of the current attribute to be encoded. This indicates that the current attribute to be encoded should be predicted using the attrRefIdx-th encoded attribute information of a different type, where `attrRefIdx` represents the attrRefIdx-th encoded attribute information appearing in the bitstream. The current attribute to be encoded must be in a higher order than `attrRefIdx` among all attributes.
[0175] However, the encoding efficiency of attribute encoding still needs to be further improved.
[0176] The encoding and decoding methods provided in the embodiments of this application are described below with reference to the accompanying drawings. The encoding method provided in the embodiments of this application can be executed by an encoding end, such as the encoder 200 shown in Figure 1, Figure 2a, or Figure 2b. The decoding method provided in the embodiments of this application can be executed by a decoding end, such as the decoder 300 shown in Figure 1, Figure 3a, or Figure 3b. The encoding end and decoding end can be implemented by software, hardware, or a combination thereof. When implemented by hardware, the encoding end can be referred to as an encoding end device or encoding device, and the decoding end can be referred to as a decoding end device or decoding device.
[0177] Figure 7 shows a schematic flowchart of an encoding method 700 provided in an embodiment of this application. As shown in Figure 7, method 700 includes steps 710 to 740.
[0178] 710, The encoding end constructs the transformation tree structure of the current attributes of the point cloud to be encoded.
[0179] 720, retrieves the first original attribute value and the reference attribute value of the current attribute of the placeholder child node of the current node in the transformed tree structure.
[0180] 730. Based on the first original attribute value, the reference attribute value, and the scaling parameter, obtain the cross-attribute prediction residuals of the current attribute and the reference attribute; the scaling parameter is used to characterize the correlation between the current attribute and the reference attribute.
[0181] 740. The cross-attribute prediction residuals are transformed, quantized, and encoded to obtain the bitstream.
[0182] Therefore, in this embodiment, the encoding end constructs a transformation tree structure of the current attributes of the point cloud to be encoded, obtains the first original attribute value and the reference attribute value of the current attribute of the placeholder child node of the current node in the transformation tree structure, and obtains the cross-attribute prediction residual of the current specification attribute and the reference attribute based on the first original attribute value, the reference attribute value, and the scaling parameter. Then, the cross-attribute prediction residual is transformed, quantized, and encoded to obtain the bitstream. This embodiment can utilize the correlation between the attribute types of the current node to perform cross-attribute prediction, obtain the cross-attribute prediction residual, and then transform and encode the cross-attribute prediction residual to obtain the bitstream, which can help improve the compression rate of the bitstream and thus improve encoding efficiency.
[0183] It should be understood that Figure 7 illustrates the steps or operations of the encoding method, but these steps or operations are merely examples, and other operations or variations of the operations in Figure 7 may also be performed in this application.
[0184] Here, the current attribute is either the currently encoded attribute or the attribute to be encoded, while the reference attribute is the attribute that has already been encoded. The current attribute and the reference attribute are two different types of attributes for the current node. For example, the current attribute could be reflectivity R, and the reference attribute could be luminance (luma, L).
[0185] For example, when both `cross_attr_prediction_enabled_flag` and `cross_attr_type_prediction_enabled_flag` are enabled (i.e., their values are 1), `attrRefIdx` is introduced into the attribute parameter set of the current attribute to be encoded. Here, `attrRefIdx` represents the `attrRefIdx`th encoded attribute information appearing in the bitstream, i.e., the reference attribute information. The current attribute's order among all attributes must be greater than `attrRefIdx`, meaning it must be greater than the reference attribute's order.
[0186] In some embodiments, step 710 above can be implemented as follows:
[0187] The point cloud to be encoded is reordered, and a P-layer transformation tree structure for the current attributes is constructed based on the reordered point cloud data according to geometric distance. Here, P is a positive integer. Optionally, a bottom-up construction method can be used to construct the transformation tree structure. The bottom layer of the transformation tree structure can contain all nodes, while the top layer is the root node layer, containing only one node.
[0188] In some embodiments, a transformation tree can also be constructed for the reference attribute. Specifically, the transformation tree is constructed in a manner similar to step 710. It can be understood that since the transformation tree structures of the current attribute and the reference attribute are based on the same geometric structural components, the transformation tree structures of the current attribute and the reference attribute are identical, and the Morton code information and weight information of each merged node are the same, but the attribute information is different.
[0189] Optionally, during the construction of the transformation tree structure, corresponding Morton code information, current attribute information, and weight information can be generated for the node after merging the current attributes. Optionally, corresponding Morton code information, reference attribute information, and weight information can also be generated for the node after merging the reference attributes.
[0190] In some embodiments, in step 710, the transformation tree structure can be parsed from top to bottom. For example, the transformation tree structure of the current attribute can be upsampled and predicted layer by layer and node by node from top to bottom, that is, starting from the root node, and the transformation tree structure of the reference attribute can be upsampled and predicted layer by layer and node by node from the root node.
[0191] In some embodiments, the current node belongs to the nodes of the first A layers of the transformation tree structure of the current attribute.
[0192] In other words, for nodes in the first A layers of the transformation tree structure of the current attribute, cross-attribute prediction is permitted using the technical solution provided in the embodiments of this application. Here, A is a positive integer.
[0193] As one implementation approach, the parameter `num_layers_CAP_enabled` can be used to indicate how many layers are allowed to enable cross-attribute prediction. For example, `num_layers_CAP_enabled` = A. Optionally, the parameter `num_layers_CAP_enabled` can be in the Attribute Parameter Set (APS), Attribute Data Unit (ADU), or other parameter sets; there are no restrictions here.
[0194] In some embodiments, a first parameter may also be passed to the bitstream; the first parameter is used to instruct the nodes of the first A layer to perform cross-attribute prediction.
[0195] For example, the first parameter can be the parameter num_layers_CAP_enabled mentioned above. Optionally, the first parameter can be the bitstream passed in from APS, ADU, or other parameter sets; there are no restrictions here.
[0196] Optionally, the first parameter can be transmitted using adaptive arithmetic entropy coding or exponential Golomb coding methods, without any restrictions.
[0197] In some embodiments, for nodes that are not in the first A layer, relevant schemes can be used for prediction, transformation, and entropy coding, without limitation.
[0198] In some embodiments, in step 720 above, the number of placeholder child nodes of the current node ranges from 1 to 8.
[0199] In some embodiments, in step 720 above, the reference attribute value of the reference attribute of each placeholder child node of the current node is the reconstructed value of the reference attribute.
[0200] In some embodiments, step 730 above can be achieved by the following steps 731-733:
[0201] 731. If the current node makes a prediction, then obtain the first predicted value of the current attribute and the second predicted value of the reference attribute of the current node.
[0202] Here, the first predicted value of the current attribute can refer to the attribute prediction value of the current attribute, which can be obtained according to relevant schemes, and there are no restrictions here. The second predicted value of the reference attribute can refer to the attribute prediction value of the reference attribute, which can also be obtained according to relevant schemes, and there are no restrictions here.
[0203] Optionally, a threshold value for the parent node of the current attribute can be used to determine whether to use the predicted value of the neighboring parent node of the reference attribute. Specifically, if the neighboring parent node of the current attribute meets the threshold condition, then that neighboring parent node will naturally be used to predict the current node, and the reference attribute value of that neighboring parent node will also be used to predict the reference attribute predicted value of the current node; otherwise, it will not be used.
[0204] It should be noted that prediction for the current node can include intra-frame prediction or inter-frame prediction, without limitation. For example, when the test sequence is a single-frame sequence, the prediction can refer to intra-frame prediction. As another example, when the test sequence is a multi-frame sequence, there are inter-frame prediction values, so the prediction can include at least one of intra-frame prediction and inter-frame prediction.
[0205] 732, obtain the first prediction residual of the current attribute based on the first original attribute value and the first prediction value; and obtain the second prediction residual of the reference attribute based on the reference attribute value and the second prediction value.
[0206] For example, the first prediction residual is the difference between the first original attribute value and the first predicted value, and the second prediction residual is the difference between the reference attribute value and the second predicted value.
[0207] For example, assuming the current attribute is reflectance, the first original attribute value of the current attribute is R, the reference attribute is luminance, and the reference attribute value of the reference attribute is L, if the current node makes a prediction, then the first prediction residual R_res of the current attribute and the second prediction residual L_res of the reference attribute can be obtained according to the following formula (1): R_res=R-R_pred L_res=L-L_pred (1)
[0208] Where R_pred represents the first predicted value of the current attribute, and L_pred represents the second predicted value of the reference attribute.
[0209] 733. Obtain the cross-attribute prediction residual based on the first prediction residual, the second prediction residual, and the scaling parameter.
[0210] For example, the cross-attribute prediction residual is the difference between the first prediction residual and the product of the second prediction residual and the scaling parameter.
[0211] For example, the cross-attribute prediction residual Res_cross can be obtained according to the following formula (2): Res_cross=R_res-s·L_res (2)
[0212] Where s represents the scaling parameter.
[0213] In some embodiments, the scaling parameter s is a scalar value that characterizes the correlation between the current attribute and the reference attribute. For example, s can be positive or negative when the two attribute values are correlated. s can be 0 when the two attributes may not be significantly correlated.
[0214] Therefore, the embodiments of this application can obtain the first predicted value of the current attribute and the second predicted value of the reference attribute when the current node is allowed to make predictions, and obtain the first prediction residual of the current attribute based on the first original attribute value and the first predicted value, and obtain the second prediction residual of the reference attribute based on the reference attribute value and the second predicted value, and then obtain the cross-attribute prediction residual based on the first prediction residual, the second prediction residual and the scaling parameter, thereby realizing the prediction of the current attribute by the reference attribute.
[0215] In some embodiments, step 730 above can be implemented by step 734:
[0216] If the current node does not make a prediction, the cross-attribute prediction residual is obtained based on the first original attribute value, the reference attribute value, and the scaling parameter.
[0217] For example, the cross-attribute prediction residual is the difference between the first original attribute value and the product of the reference attribute value and the scaling parameter.
[0218] For example, assuming the current attribute is reflectance, the first original attribute value of the current attribute is R, the reference attribute is luminance, and the reference attribute value of the reference attribute is L, if the current node does not make a prediction, then the cross-attribute prediction residual Res_cross can be obtained according to the following formula (3): Res_cross=Rs·L (3)
[0219] Therefore, the embodiments of this application can obtain cross-attribute prediction residuals based on the first original attribute value of the current attribute of the current node, the reference attribute value of the reference attribute, and the scaling parameters when prediction is not allowed at the current node, thereby realizing the prediction of the current attribute by the reference attribute.
[0220] It should be noted that the embodiments of this application do not limit the method of obtaining the scaling parameters. For example, the scaling reference can be obtained in at least one of the following ways:
[0221] Method 1: Transform the nodes in each level of the tree structure to have the same scaling parameter; or transform the nodes in the tree structure to have the same scaling parameter.
[0222] For example, a fixed scaling parameter s value can be set for each layer in the transformation tree structure, or the same scaling parameter s value can be set for each node in the transformation tree structure.
[0223] Therefore, by setting fixed scaling parameters for each layer or each node, it is possible to reduce the amount of computation at the encoding end and save computational resources at the encoding end.
[0224] Method 2: Calculate a scaling parameter s for every n nodes in the transformation tree structure. Here, n is an integer greater than or equal to 1. See steps S21-S23 below for the specific calculation method.
[0225] S21, for the current layer of the transformed tree structure, obtain the n nodes of the kth group; k and n are integers greater than or equal to 1.
[0226] Specifically, we can traverse the transformation tree structure through the transformation layers that have enabled cross-attribute prediction, calculating an s for every n nodes. For example, for the current lvl layer, we obtain the k-th group of nodes, where each group contains n nodes. Here, k represents the index order of each group of nodes.
[0227] S22, obtain the second original attribute value of the current attribute and the third original attribute value of the reference attribute of the i-th node of the current layer; where i∈kn~(k+1)n-1 and are integers.
[0228] For example, the second original attribute value of the current attribute (R) of the i-th node in the current level (lvl) is... The third primitive attribute value of the reference attribute (L) is
[0229] Where i represents the index order of nodes in each layer, and i∈kn~(k+1)n-1 represents all nodes in the kth group.
[0230] S23, determine the scaling parameter based on the second original attribute value and the third original attribute value of the i-th node; wherein the scaling parameter is shared by the n nodes.
[0231] It should be noted that the embodiments of this application do not limit the method of determining the scaling parameters based on the second and third original attribute values of the i-th node.
[0232] For example, the scaling parameters can be determined according to the following steps S231-S233:
[0233] S231, determine the product of the second original attribute value and the third original attribute value of the i-th node, and sum the products corresponding to the nodes within the range of the i-th node.
[0234] S232, determine the sum of squares of the third original attribute values of the nodes within the range of the i-th node.
[0235] S233, the scaling parameter is obtained based on the ratio of the sum of the products of the nodes within the range of the i-th node to the sum of squares.
[0236] In this method 2, the range of the i-th node is the range of the k-th group of nodes. That is, the product of the second and third original attribute values of all nodes in the k-th group is summed, and the sum of the squares of the third original attribute values of all nodes in the k-th group is calculated. The ratio of the sum of the products to the sum of the squares is then used to obtain the scaling parameter.
[0237] For example, the scaling parameter s of the k-th group of nodes in the current lvl layer. lvl,k It can be obtained according to the following formula (4):
[0238] In some embodiments, in this method 2, the scaling parameter s can be... lvl,kSend the bitstream. Optionally, the value of n can be fixed in the encoding / decoding segment, or sent into the bitstream for transmission; there are no restrictions here.
[0239] Therefore, by calculating a scaling parameter s value for each of the n nodes in the transformation tree structure based on the second original attribute value of the current attribute and the third original attribute value of the reference attribute of each of the n nodes in the current layer, it is possible to obtain the correlation between the attributes of each of the n nodes using the reference attribute and the current attribute of each of the n nodes in the current layer, which is beneficial to improving the accuracy of the scaling factor.
[0240] Method 3: Calculate a scaling parameter value for each node in the transformation tree structure. See steps S31-S32 below for the specific calculation method.
[0241] S31, for the current layer in the transformation tree structure, obtain the second original attribute value of the current attribute and the third original attribute value of the reference attribute of the i-th node of the current layer; wherein, i is less than or equal to the number of nodes in the current layer and is a positive integer;
[0242] Specifically, we can traverse the transformation layers in the transformation tree structure that enable cross-attribute prediction and compute an s for each layer. For example, for the current level (lvl), the second original attribute value of the current attribute (R) of the i-th node is... The third primitive attribute value of the reference attribute (L) is
[0243] Here, i represents the index order of nodes in each level. Unlike method 2, in method 3, i belongs to the range of all nodes in the lvl level.
[0244] S32, determine the scaling parameter based on the second original attribute value and the third original attribute value of the i-th node; wherein the scaling parameter is shared by the nodes of the current layer.
[0245] It should be noted that the embodiments of this application do not limit the method of determining the scaling parameters based on the second and third original attribute values of the i-th node. For example, the scaling parameters can be determined according to the above steps S231-S233.
[0246] Unlike method 2, in method 3, the range of the i-th node is the range of nodes in the lvl level. That is, the product of the second and third original attribute values of all nodes in the lvl level is summed, and the sum of the squares of the third original attribute values of all nodes in the lvl level is calculated. The ratio of the sum of the products to the sum of the squares is then used to obtain the scaling parameter.
[0247] For example, the scaling parameter s of the current LVL layer lvlIt can be obtained according to the following formula (5):
[0248] In some embodiments, in this method 3, the scaling parameter s can be... lvl Send the bitstream. For example, the scaling parameter s can be used. lvl Store it in CrossAttrPredCoeff[lvl] and send it to the bitstream.
[0249] Therefore, by calculating a scaling parameter s value for all nodes in the current layer of the transformation tree structure based on the second original attribute value of the current attribute and the third original attribute value of the reference attribute of all nodes in the current layer, it is possible to obtain the correlation between the attributes of the nodes in the current layer using the reference attribute and the current attribute of each node in the current layer, which is beneficial to improving coding efficiency while reducing the bit rate.
[0250] Method 4: Determine the scaling parameters according to the following steps S41 and S42.
[0251] S41, for the current layer of the transformation tree structure, obtain the first attribute reconstruction data of the current attribute and the second attribute reconstruction data of the reference attribute of the j-th node of the current layer; the j-th node belongs to the node that has been encoded and reconstructed; j is a positive integer;
[0252] Specifically, we can traverse the transformation layers in the transformation tree structure that have cross-attribute prediction enabled. Assuming that the current node j of the current lvl layer consists of 2*2*2 child nodes, each node can encode and reconstruct multiple attribute reconstruction values of the current attribute when prediction is not enabled, and can encode and reconstruct multiple attribute reconstruction residuals of the current attribute, as well as the attribute reconstruction value or attribute reconstruction residual corresponding to the reference attribute when prediction is enabled.
[0253] For example, for the current LVL layer, the first attribute reconstruction data of the current attribute (R) of the j-th node whose predecessor has been encoded and reconstructed can be obtained. Second reconstructed data and reference attribute (L) Specifically, when prediction is not enabled at the j-th node, the first attribute is reconstructed data. Reconstruct the attribute value for the current attribute, and then reconstruct the data. The attribute reconstruction value is used as the reference attribute; when prediction is enabled, the first attribute is reconstructed. Reconstruct residuals for the current attribute, and then reconstruct the second set of data. Reconstruct residuals for the reference attribute.
[0254] Where j represents the index order of nodes in each layer. Here, the j-th node is the encoded and reconstructed node used to calculate the scaling parameter s.
[0255] S42, determine the scaling parameter based on the first attribute reconstruction data and the second attribute reconstruction data of the j-th node.
[0256] It should be noted that the embodiments of this application do not limit the method of determining the scaling parameters based on the reconstruction data of the first attribute and the reconstruction data of the second attribute of the j-th node.
[0257] For example, step S42 may include the following steps S421-S423.
[0258] S421, determine the product of the first attribute reconstruction data and the second attribute reconstruction data of the j-th node, and sum the products corresponding to the nodes within the range of the j-th node.
[0259] S422, determine the sum of squares of the second attribute reconstructed data of the nodes within the range of the j-th node.
[0260] S423, the scaling parameter is obtained based on the ratio of the sum of the products of the nodes within the range of the j-th node to the sum of squares.
[0261] In this method 4, the range of the j-th node is the nodes that have already been encoded and reconstructed.
[0262] For example, the scaling parameter s of the current LVL layer is determined based on the j-th node. lvl,j It can be obtained according to the following formula (6):
[0263] In some embodiments, in method 4, n nodes can be grouped together, and all nodes in a group share a single node s. n can be 1 or an integer greater than 1.
[0264] In some embodiments, the j-th node satisfies at least one of the following:
[0265] The j-th node belongs to the pre-encoded and reconstructed nodes of the current layer; or
[0266] The j-th node belongs to the first m child nodes whose preceding order has been reconstructed; or
[0267] The j-th node belongs to the first k nodes of the current layer that have been encoded and reconstructed in the preceding sequence;
[0268] m and k are positive integers.
[0269] For example, the j-th node belongs to the preceding encoded and reconstructed nodes of the current layer, that is, the preceding encoded and reconstructed nodes of the current node in the current layer can be used to calculate the scaling parameters.
[0270] Optionally, the number of child nodes of the preceding encoded and reconstructed nodes in the current layer is greater than or equal to q, where q ∈ [0, 7] and is an integer. That is, it is possible to require that the number of child nodes of the preceding encoded and reconstructed nodes in the current layer is greater than or equal to q before they participate in the calculation of the scaling parameters. Specifying nodes with a number of child nodes greater than q for determining the scaling coefficient helps improve the accuracy of the scaling coefficient.
[0271] The j-th node belongs to the first m encoded and reconstructed child nodes of the current node. That is, the child nodes of the encoded and reconstructed child nodes of the current node can be used to calculate the scaling parameter. As an example, when m is 5, the first encoded and reconstructed child node before the current node has 3 child nodes, the second encoded and reconstructed child node has 1 child node, and the third encoded and reconstructed child node has 1 child node. Therefore, the child nodes of the first, second, and third child nodes before the current node are the first 5 encoded and reconstructed child nodes of the previous node.
[0272] The j-th node belongs to the first k encoded and reconstructed nodes of the current layer. That is, the k encoded and reconstructed nodes before the current node can be used to calculate the scaling parameters. As an example, when k is 3, the first, second, and third nodes before the current node are the first three encoded and reconstructed nodes of the previous layer.
[0273] Optionally, the number of child nodes of the first k nodes is greater than or equal to q; where q ∈ [0, 7] and is an integer. That is, it is possible to require that the number of child nodes of the first k nodes of the current layer that have been encoded and reconstructed in the preceding order is greater than or equal to q before they participate in the calculation of scaling parameters. By specifying nodes with a number of child nodes greater than q to determine the scaling factor, it is beneficial to improve the accuracy of the scaling factor.
[0274] Optionally, in method 4, the scaling parameters are calculated using the encoded and reconstructed information, so the scaling parameters do not need to be passed into the bitstream.
[0275] Optionally, either m or k can be passed into the bitstream.
[0276] Optionally, q can be passed into the bitstream.
[0277] Optionally, m, k, or q can be fixed in the encoding / decoding segment and do not need to be transmitted.
[0278] Optionally, in any of the above methods 1 to 4, the scaling parameter can be passed to the bitstream, and there is no limitation here.
[0279] After obtaining the cross-attribute prediction residuals of the current attribute and the reference attribute, step 740 can be performed, which involves transforming, quantizing, and encoding the cross-attribute prediction residuals to obtain the bitstream.
[0280] For example, the cross-attribute prediction residuals can be transformed using RAHT, and the resulting transform coefficient residuals can be quantized and entropy encoded to obtain a bitstream.
[0281] Therefore, by calculating a scaling parameter 's' value for the current node in the transform tree structure based on the attribute reconstruction data of the current attribute and the reference attribute of the encoded reconstructed node in the current layer, the correlation between the attributes of the current node can be obtained using the reference and current attributes of the encoded reconstructed node in the current layer, which helps improve the accuracy of the scaling parameter. Furthermore, calculating the scaling factor using the information of the encoded reconstructed node eliminates the need to transmit the scaling factor to the bitstream, thus saving bitstream data.
[0282] Figure 8 shows a schematic flowchart of a decoding method 800 provided in an embodiment of this application. As shown in Figure 8, method 800 includes steps 810 to 860.
[0283] 810, the decoding end decodes and dequantizes the bitstream to obtain the reconstructed value of the first transform coefficient residual of the point cloud to be decoded.
[0284] 820, Construct the transformation tree structure of the current attributes of the point cloud to be decoded.
[0285] 830, obtain the reference attribute value of the placeholder child node of the current node of the transformed tree structure.
[0286] 840. Perform cross-attribute prediction based on the reference attribute value and scaling parameters to obtain a first attribute prediction value; the scaling parameters are used to characterize the correlation between the current attribute and the reference attribute.
[0287] 850, perform an inverse transformation based on the first attribute prediction value and the reconstructed value of the first transformation coefficient residual to obtain the attribute reconstruction value of the placeholder child node of the current node.
[0288] Therefore, in this embodiment, the decoding end obtains the reconstructed value of the first transform coefficient residual of the point cloud to be decoded by decoding and inverse quantizing the bitstream, constructs a transform tree structure of the current attribute of the point cloud to be decoded, obtains the reference attribute value of the reference attribute of the placeholder child node of the current node in the transform tree structure, and performs cross-attribute prediction based on the reference attribute value and scaling parameters to obtain the first attribute prediction value. Then, it performs inverse transformation based on the first attribute prediction value and the reconstructed value of the first transform coefficient residual to obtain the attribute reconstruction value of the placeholder child node of the current node. This embodiment can utilize the correlation between the attribute types of the current node to perform cross-attribute prediction and obtain the reconstructed attribute value of the placeholder child node of the current node, which can help improve the compression rate of the bitstream.
[0289] It should be understood that Figure 8 illustrates the steps or operations of the decoding method, but these steps or operations are merely examples, and other operations or variations of the operations in Figure 8 may also be performed in this application.
[0290] Specifically, the current attribute and reference attribute can be found in the relevant descriptions of the encoding end in Figure 7, which will not be repeated here.
[0291] For example, in step 810 above, the decoder first parses the acquired input bitstream to obtain the cross_attr_prediction_enabled_flag, a multi-attribute type prediction enable flag, from the sequence parameter set. When this flag is enabled (i.e., the flag is 1), the cross_attr_type_prediction_enabled_flag flag corresponding to each attribute to be decoded is parsed out. When cross_attr_type_prediction_enabled_flag is enabled (i.e., the flag is 1), the attrRefIdx of the attribute type that the current attribute needs to reference is parsed out.
[0292] In some embodiments, the bitstream can also be parsed to obtain a first parameter; the first parameter is used to instruct the nodes of the first A layers in the transform tree structure to perform cross-attribute prediction. A is a positive integer. Specifically, the first parameter can be referred to the relevant description of the encoding end in Figure 7, and will not be repeated here.
[0293] One implementation approach involves parsing the input bitstream to obtain the `num_layers_CAP_enabled` parameter, which indicates how many layers preceding the current attribute are enabled for cross-attribute prediction. Optionally, A = `num_layers_CAP_enabled`.
[0294] Next, entropy decoding and dequantization are performed on the quantized transform coefficients or coefficient residuals from the input bitstream to obtain the reconstructed transform coefficients or transform residual coefficients. These reconstructed transform residual coefficients include the reconstructed values of the first transform coefficient residuals mentioned above. As one implementation, the first transform coefficient residual is the AC coefficient residual.
[0295] It should be understood that the reconstructed value of the first transform coefficient residual corresponds to the cross-attribute prediction residual obtained by cross-attribute prediction at the encoding end.
[0296] In some embodiments, step 720 above can be implemented as follows:
[0297] The point cloud to be decoded is reordered, and a P-layer transformation tree structure for the current attributes is constructed based on the reordered point cloud data according to geometric distance. Here, P is a positive integer. Optionally, a bottom-up construction method can be used to construct the transformation tree structure. The bottom layer of the transformation tree structure can contain all nodes, while the top layer is the root node layer, containing only one node.
[0298] In some embodiments, a transformation tree can also be constructed for the reference attribute. Specifically, the transformation tree is constructed in a manner similar to step 820. It can be understood that since the transformation tree structures of the current attribute and the reference attribute are constructed based on the same geometric structure, the transformation tree structure of the current attribute and the transformation tree structure of the reference attribute are the same, and the Morton code information and weight information of each merged node are the same.
[0299] Optionally, during the construction of the transformation tree structure, corresponding Morton code information and weight information can be generated for the nodes after merging the current attributes. Optionally, corresponding Morton code information, reference attribute information, and weight information can also be generated for the nodes after merging the reference attributes.
[0300] In some embodiments, in step 820, the transformation tree structure can be parsed from top to bottom. For example, the transformation tree structure of the current attribute can be upsampled and predicted layer by layer from the root node, and the transformation tree structure of the reference attribute can be upsampled and predicted layer by layer from the root node, node by node, to finally obtain the reconstructed attribute value of each node.
[0301] In some embodiments, the current node belongs to the nodes of the first A layer of the transformation tree structure.
[0302] In other words, for nodes in the first A layers of the transformation tree structure of the current attribute, cross-attribute prediction is permitted using the technical solution provided in the embodiments of this application. Here, A is a positive integer.
[0303] In some embodiments, for nodes that are not in the first A layer, relevant schemes can be used for prediction and inverse transformation, and no restrictions are imposed here.
[0304] For example, for nodes not in the first A layer, if the current node is the root node, an inverse transformation is directly performed to obtain the attribute reconstruction values of each child node. Alternatively, if the current node is not the root node and no prediction is performed, its DC coefficients inherit the attribute reconstruction values of its parent node, and are then combined with the reconstructed AC coefficients to perform a RAHT inverse transformation to obtain the attribute reconstruction values of each child node of the current node. Or, if the current node is not the root node but prediction is performed, its DC coefficients inherit the attribute reconstruction values of its parent node, and the difference between these and the parent node's attribute prediction values is used to obtain the DC coefficient residuals. These residuals are then combined with the reconstructed AC coefficient residuals to perform an inverse transformation to obtain the reconstructed attribute prediction residuals, which are then added to the attribute prediction values to obtain the attribute reconstruction values of each child node.
[0305] In some embodiments, the number of placeholder child nodes of the current node in step 830 above ranges from 1 to 8.
[0306] In some embodiments, in step 830 above, the reference attribute value of the reference attribute of each placeholder child node of the current node is the reconstructed value of the reference attribute. For example, when the reference attribute is brightness, the reference attribute value is L.
[0307] In some embodiments, step 840 above can be achieved by the following steps 841-843:
[0308] 841. If the current node allows prediction, then obtain the second predicted value of the reference attribute of the current node.
[0309] Here, the second predicted value of the reference attribute can refer to the attribute predicted value of the reference attribute, which can be obtained according to the relevant scheme, and there are no restrictions here.
[0310] Optionally, a threshold value for the parent node of the current attribute can be used to determine whether to use the predicted value of the neighboring parent node of the reference attribute. Specifically, if the neighboring parent node of the current attribute meets the threshold condition, then that neighboring parent node will naturally be used to predict the current node, and the reference attribute value of that neighboring parent node will also be used to predict the reference attribute predicted value of the current node; otherwise, it will not be used.
[0311] It should be noted that prediction for the current node can include intra-frame prediction or inter-frame prediction, without limitation. For example, when the test sequence is a single-frame sequence, the prediction can refer to intra-frame prediction. As another example, when the test sequence is a multi-frame sequence, there are inter-frame prediction values, so the prediction can include at least one of intra-frame prediction and inter-frame prediction.
[0312] 842. Based on the second predicted value and the reference attribute value, obtain the second predicted residual of the reference attribute.
[0313] For example, the second prediction residual is the difference between the reference attribute value and the second prediction value.
[0314] For example, assuming the current attribute is reflectivity, the reference attribute is luminance, and the reference attribute value of the reference attribute is L, if the current node makes a prediction, the second prediction residual L_res of the reference attribute can be obtained according to the following formula (7): L_res=L-L_pred (7)
[0315] Where L_pred represents the second predicted value of the reference attribute.
[0316] 843. Obtain the first attribute prediction value based on the second prediction residual and the scaling parameter.
[0317] For example, the predicted value of the first attribute is the product of the second prediction residual and the scaling parameter. For example, the predicted value L of the first attribute can be obtained according to the following formula (8). cross L cross =s·L_res (8)
[0318] The term 's' can be found in the description of the encoding end in Figure 7, and will not be repeated here.
[0319] It should be noted that the name of the predicted value of the first attribute is not limited in the embodiments of this application. For example, the predicted value of the first attribute may also be called a cross-attribute predicted value, or other names.
[0320] Therefore, the embodiments of this application can obtain the second predicted value of the reference attribute of the current node when the current node is allowed to make predictions, and obtain the second predicted residual of the reference attribute based on the second predicted value and the reference attribute value of the reference attribute, and then obtain the first attribute predicted value based on the second predicted residual and the scaling parameter, thereby realizing the prediction of the current attribute by the reference attribute.
[0321] In some embodiments, step 840 above can be implemented by the following step 844:
[0322] 844. If the current node does not make a prediction, then the first attribute prediction value is obtained based on the reference attribute value and the scaling parameter.
[0323] For example, the predicted value of the first attribute is the product of the reference attribute value and the scaling parameter.
[0324] For example, assuming the current attribute is reflectance and the reference attribute is luminance, with a reference attribute value of L, if the current node is not allowed to predict, then the predicted value L of the first attribute can be obtained according to the following formula (9). cross L cross =s·L (9)
[0325] Therefore, the embodiments of this application can obtain the first attribute prediction value based on the reference attribute value and scaling parameter of the reference attribute of the current node when prediction is not allowed at the current node, thereby realizing the prediction of the current attribute by the reference attribute.
[0326] It should be noted that the embodiments of this application do not limit the method of obtaining the scaling parameters.
[0327] Optionally, the decoding end can parse the bitstream to obtain the scaling parameters. Alternatively, the decoding end can be configured with the same scaling parameters as the encoding segment. Alternatively, the decoding end can calculate the scaling parameters used locally, without limitation.
[0328] Optionally, each node in the transformation tree structure corresponds to the same scaling parameter; or
[0329] The nodes in the current layer share the scaling parameters; or
[0330] The scaling parameter is shared by every n nodes in the current layer, where n is an integer greater than or equal to 1.
[0331] For example, the scaling reference can be obtained in at least one of the following ways:
[0332] Method 1: On a layer-by-layer basis, use (configure) the same fixed scaling parameter s value as the encoding end for each layer of the transform tree structure, or use (configure) the same scaling parameter s value as the encoding end for each node in the transform tree structure.
[0333] By setting fixed scaling parameters for each layer or node, it is possible to reduce the amount of computation at the encoding end and save computing resources.
[0334] Method 2: Traverse the transform tree structure and enable the transform layer for cross-attribute prediction, resolving a scaling parameter s from the bitstream for every n nodes. Optionally, n can be resolved from the bitstream or can use the same value as the encoder.
[0335] Optionally, the decoding end can also calculate the shared scaling parameters for every n nodes, following method 2 of the encoding end, without limitation.
[0336] Method 3: Traverse the transformation layers that enable cross-attribute prediction in the transformation tree structure and parse the scaling parameter s corresponding to each transformation layer from the bitstream.
[0337] Optionally, the decoding end can also calculate the shared scaling parameters for every n nodes, following method 3 of the encoding end, without limitation.
[0338] Method 4: The decoding end can calculate the scaling parameters in a similar manner to Method 4 of the encoding end. Specifically, this includes the following steps:
[0339] For the current layer of the transformed tree structure, obtain the first attribute reconstruction data of the current attribute and the second attribute reconstruction data of the reference attribute of the j-th node of the current layer; the j-th node belongs to the nodes that have been decoded and reconstructed; j is a positive integer;
[0340] The scaling parameter is determined based on the first attribute reconstruction data and the second attribute reconstruction data of the j-th node.
[0341] The step of determining the scaling parameter based on the first attribute reconstruction data and the second attribute reconstruction data of the j-th node includes:
[0342] Determine the product of the first attribute reconstruction data and the second attribute reconstruction data of the j-th node, and sum the products corresponding to the nodes within the range of the j-th node;
[0343] Determine the sum of squares of the second attribute reconstructed data of the nodes within the range of the j-th node;
[0344] The scaling parameter is obtained by comparing the sum of the products of the nodes within the range of the j-th node with the sum of squares.
[0345] Optionally, the j-th node belongs to the preceding decoded and reconstructed nodes of the current layer; or
[0346] The j-th node belongs to the first m child nodes that have been decoded and reconstructed before the current node; or
[0347] The j-th node belongs to the previous k nodes that have been decoded and reconstructed in the current layer;
[0348] m and k are positive integers.
[0349] Optionally, the number of child nodes of the previously decoded and reconstructed nodes in the current layer is greater than or equal to q; or
[0350] The number of child nodes of the first k nodes is greater than or equal to q; where q∈[0,7] and is an integer.
[0351] Optionally, the decoding end can also parse the bitstream to obtain m or k.
[0352] Optionally, the decoding end can also parse the bitstream to obtain q.
[0353] Optionally, m, q, or k can also use the same values as the encoding end, without restriction.
[0354] Specifically, the process of obtaining scaling parameters by the decoding end according to method 4 can refer to the description in method 4 of the encoding end, or some simple adaptations can be made. This application embodiment does not limit this.
[0355] In some embodiments, step 850 above can be specifically implemented through the following steps 851 to 853:
[0356] 851. Based on the DC coefficient value inherited by the current node and the DC coefficient value corresponding to the first attribute prediction value of the current node, the second transformation coefficient residual is obtained.
[0357] Specifically, the DC coefficient value of the current node can be inherited, that is, inherited from the corresponding inverse transform coefficient of its parent node. Specifically, in the transform tree structure, except for the root node, other nodes only transmit AC coefficients, and their DC coefficients can be obtained through inheritance. Therefore, the obtained DC coefficient is the attribute reconstruction value of the current node obtained from the previous layer. That is, the DC coefficient value of the current node is obtained by inheriting the attribute reconstruction value of the parent node. The corresponding attribute reconstruction value in the parent node is the attribute reconstruction value of the current node obtained when decoding the attribute of each placeholder child node of the parent node corresponding to the current node. The DC coefficient value corresponding to the first attribute prediction value of the current node is the cross-attribute predicted DC coefficient of the current node.
[0358] Optionally, if the current node does not perform prediction, the DC coefficient is inherited and the difference between the DC coefficient value of the first attribute prediction value is calculated to obtain the second transformation coefficient residual.
[0359] Optionally, if the current node makes a prediction, the second attribute prediction value of the current node is obtained; and the DC coefficient is inherited, and the difference between the DC coefficient value corresponding to the second attribute prediction value and the DC coefficient value corresponding to the first attribute prediction value is successively calculated to obtain the second transformation coefficient residual.
[0360] For example, the second attribute prediction value can be an attribute prediction value obtained from intra-frame prediction or an attribute prediction value obtained from inter-frame prediction. The DC coefficient value corresponding to the second attribute value is the predicted DC coefficient of the current node.
[0361] Specifically, since a prediction has been made for the current node, the first transformation coefficient is the AC coefficient prediction residual. The decoder needs to perform an inverse transformation on the AC coefficient residual and the DC coefficient corresponding to the current node to obtain the reconstructed attribute residual. At this time, the AC coefficient is the prediction residual coefficient, and the DC coefficient should also be of the same order of magnitude, and should also be the DC coefficient prediction residual, that is, the DC coefficient obtained by inheritance is subtracted from the predicted DC coefficient.
[0362] 852, Perform an inverse transformation based on the reconstructed values of the second transformation coefficient residual and the first transformation coefficient residual to obtain the reconstructed attribute residual value.
[0363] For example, the second transformation coefficient residual can be combined with the reconstructed value of the first transformation coefficient parameter to perform an inverse RAHT transformation to obtain the reconstructed attribute residual of the placeholder child node of the current node, which can be represented as Attr_res_rec.
[0364] 853. Based on the reconstructed attribute residual value and the first attribute prediction value, the attribute reconstruction value of the placeholder child node of the current node is obtained.
[0365] Optionally, if the current node is not predicted, the attribute reconstruction value of the placeholder child node of the current node is obtained by summing the reconstructed attribute residual value and the first attribute prediction value.
[0366] For example, the reconstructed attribute residual Attr_res_rec of each placeholder child node can be compared with the first attribute prediction value L of each placeholder child node. cross Add them together to get the attribute reconstruction values of each placeholder child node of the current node.
[0367] Optionally, if the current node makes a prediction, the attribute reconstruction value of the placeholder child node of the current node is obtained by summing the reconstructed attribute residual value, the second attribute prediction value and the first attribute prediction value.
[0368] For example, the reconstructed attribute residual Attr_res_rec for each placeholder child node, the predicted second attribute value for each placeholder child node, and the predicted first attribute value L for each placeholder child node can be calculated. cross Add them together to get the attribute reconstruction values of each placeholder child node of the current node.
[0369] In summary, according to the decoding method provided in the embodiments of this application, the transformation tree structure is parsed layer by layer and node by node from top to bottom. When the bottom layer of the transformation tree structure is reached, the attribute reconstruction values of all nodes can be obtained.
[0370] The encoding method provided in this application can be executed by an encoding device. As an example, the device can be an electronic device or a component within the electronic device, such as a chip or circuit. This application uses an encoding device executing the encoding method as an example to illustrate the encoding device provided in this application.
[0371] Figure 9 shows a schematic block diagram of the encoding device. As shown in Figure 9, the encoding device 900 includes:
[0372] Module 910 is used to construct the transformation tree structure of the current attributes of the point cloud to be encoded at the encoding end;
[0373] The acquisition module 920 is used to acquire the first original attribute value and the reference attribute value of the current attribute of the placeholder child node of the current node of the transformed tree structure;
[0374] The prediction module 930 is used to obtain the cross-attribute prediction residuals of the current attribute and the reference attribute based on the first original attribute value, the reference attribute value, and the scaling parameter; the scaling parameter is used to characterize the correlation between the current attribute and the reference attribute.
[0375] The encoding module 940 is used to transform, quantize, and encode the cross-attribute prediction residual to obtain a bitstream.
[0376] Optionally, the prediction module 930 is specifically used for:
[0377] If the current node makes a prediction, then obtain the first predicted value of the current attribute and the second predicted value of the reference attribute of the current node;
[0378] The first prediction residual of the current attribute is obtained based on the first original attribute value and the first predicted value; and the second prediction residual of the reference attribute is obtained based on the reference attribute value and the second predicted value.
[0379] The cross-attribute prediction residual is obtained based on the first prediction residual, the second prediction residual, and the scaling parameter.
[0380] Optionally, the prediction module 930 is specifically used for:
[0381] If the current node does not make a prediction, the cross-attribute prediction residual is obtained based on the first original attribute value, the reference attribute value, and the scaling parameter.
[0382] Optionally, nodes in each level of the transformation tree structure correspond to the same scaling parameter; or
[0383] Each node in the transformation tree structure corresponds to the same scaling parameter.
[0384] Optionally, the acquisition unit is further configured to:
[0385] For the current layer of the transformed tree structure, obtain the n nodes of the k-th group; k and n are integers greater than or equal to 1.
[0386] Obtain the second original attribute value of the current attribute and the third original attribute value of the reference attribute of the i-th node of the current layer; where i∈kn~(k+1)n-1 and are integers;
[0387] The scaling parameter is determined based on the second original attribute value and the third original attribute value of the i-th node; wherein the scaling parameter is shared by the n nodes.
[0388] Optionally, the acquisition unit 920 is further configured to:
[0389] For the current layer in the transformation tree structure, obtain the second original attribute value of the current attribute and the third original attribute value of the reference attribute of the i-th node of the current layer; where i is less than or equal to the number of nodes in the current layer and is a positive integer;
[0390] The scaling parameter is determined based on the second original attribute value and the third original attribute value of the i-th node; wherein the scaling parameter is shared by all nodes in the current layer.
[0391] Optionally, the acquisition unit 920 is specifically used for:
[0392] Determine the product of the second original attribute value and the third original attribute value of the i-th node, and sum the products corresponding to the nodes within the range of the i-th node;
[0393] Determine the sum of squares of the third original attribute values of the nodes within the range of the i-th node;
[0394] The scaling parameter is obtained by comparing the sum of the products of the nodes within the range of the i-th node with the sum of squares.
[0395] Optionally, the acquisition unit 920 is further configured to:
[0396] For the current layer of the transformed tree structure, obtain the first attribute reconstruction data of the current attribute and the second attribute reconstruction data of the reference attribute of the j-th node of the current layer; the j-th node belongs to the already encoded and reconstructed node; j is a positive integer;
[0397] The scaling parameter is determined based on the first attribute reconstruction data and the second attribute reconstruction data of the j-th node.
[0398] Optionally, the acquisition unit 920 is specifically used for:
[0399] Determine the product of the first attribute reconstruction data and the second attribute reconstruction data of the j-th node, and sum the products corresponding to the nodes within the range of the j-th node;
[0400] Determine the sum of squares of the second attribute reconstructed data of the nodes within the range of the j-th node;
[0401] The scaling parameter is obtained by comparing the sum of the products of the nodes within the range of the j-th node with the sum of squares.
[0402] Optionally, the j-th node belongs to the pre-encoded and reconstructed nodes of the current layer; or
[0403] The j-th node belongs to the first m child nodes whose preceding order has been reconstructed; or
[0404] The j-th node belongs to the first k nodes of the current layer that have been encoded and reconstructed in the preceding sequence;
[0405] m and k are positive integers.
[0406] Optionally, the number of child nodes of the previously encoded and reconstructed nodes in the current layer is greater than or equal to q; or
[0407] The number of child nodes of the first k nodes is greater than or equal to q;
[0408] Where q∈[0,7] and is an integer.
[0409] Optionally, the encoding module 910 is further configured to:
[0410] The m or the k is passed into the bitstream.
[0411] Optionally, the encoding module 940 is further configured to:
[0412] The q is passed into the bitstream.
[0413] Optionally, the encoding module 940 is further configured to:
[0414] The scaling parameter is passed into the bitstream.
[0415] Optionally, the current node belongs to the nodes of the first A levels of the transformed tree structure. A is a positive integer.
[0416] Optionally, the encoding module 940 is further configured to:
[0417] The first parameter is passed into the bitstream; the first parameter is used to instruct the nodes of the first A layer to perform cross-attribute prediction.
[0418] In this embodiment, the encoding end constructs a transformation tree structure of the current attributes of the point cloud to be encoded, obtains the first original attribute value and the reference attribute value of the current attribute of the placeholder child node of the current node in the transformation tree structure, and obtains the cross-attribute prediction residual of the current specification attribute and the reference attribute based on the first original attribute value, the reference attribute value, and the scaling parameter. Then, the cross-attribute prediction residual is transformed, quantized, and encoded to obtain the bitstream. This embodiment can utilize the correlation between the attribute types of the current node to perform cross-attribute prediction, obtain the cross-attribute prediction residual, and then transform and encode the cross-attribute prediction residual to obtain the bitstream, which can help improve the compression rate of the bitstream and thus improve encoding efficiency.
[0419] The encoding device 900 provided in this application embodiment can implement the various processes implemented in the method embodiment of FIG7 and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0420] The decoding method provided in this application can be executed by a decoding device. As an example, the device can be an electronic device or a component within the electronic device, such as a chip or circuit. This application uses the example of a decoding device executing the decoding method to illustrate the decoding device provided in this application.
[0421] Figure 10 shows a schematic block diagram of the decoding apparatus. As shown in Figure 10, the decoding apparatus 1000 includes:
[0422] The parsing module 1010 is used by the decoding end to decode and dequantize the bit stream to obtain the reconstructed value of the first transform coefficient residual of the point cloud to be decoded;
[0423] Construction module 1020 is used to construct the transformation tree structure of the current attributes of the point cloud to be decoded;
[0424] The acquisition module 1030 is used to acquire the reference attribute value of the reference attribute of the placeholder child node of the current node of the transformed tree structure;
[0425] The prediction module 1040 is used to perform cross-attribute prediction based on the reference attribute value and the scaling parameter to obtain a first attribute prediction value; the scaling parameter is used to characterize the correlation between the current attribute and the reference attribute.
[0426] The inverse transformation module 1050 is used to perform an inverse transformation based on the first attribute prediction value and the reconstructed value of the first transformation coefficient residual to obtain the attribute reconstruction value of the placeholder child node of the current node.
[0427] Optionally, the prediction module is specifically used for:
[0428] If the current node allows prediction, then obtain the second predicted value of the reference attribute of the current node;
[0429] Based on the second predicted value and the reference attribute value, obtain the second predicted residual of the reference attribute;
[0430] The predicted value of the first attribute is obtained based on the second prediction residual and the scaling parameter.
[0431] Optionally, the prediction module 1040 is specifically used for:
[0432] If the current node does not make a prediction, then the predicted value of the first attribute is obtained based on the reference attribute value and the scaling parameter.
[0433] Optionally, each node in the transformation tree structure corresponds to the same scaling parameter; or
[0434] The nodes in the current layer share the scaling parameters; or
[0435] The scaling parameter is shared by every n nodes in the current layer, where n is an integer greater than or equal to 1.
[0436] Optionally, the parsing module 1010 is further configured to:
[0437] The scaling parameters are obtained by parsing the bitstream.
[0438] Optionally, the acquisition module 1030 is further configured to:
[0439] For the current layer of the transformed tree structure, obtain the first attribute reconstruction data of the current attribute and the second attribute reconstruction data of the reference attribute of the j-th node of the current layer; the j-th node belongs to the nodes that have been decoded and reconstructed; j is a positive integer;
[0440] The scaling parameter is determined based on the first attribute reconstruction data and the second attribute reconstruction data of the j-th node.
[0441] Optionally, the acquisition module 1030 is specifically used for:
[0442] Determine the product of the first attribute reconstruction data and the second attribute reconstruction data of the j-th node, and sum the products corresponding to the nodes within the range of the j-th node;
[0443] Determine the sum of squares of the second attribute reconstructed data of the nodes within the range of the j-th node;
[0444] The scaling parameter is obtained by comparing the sum of the products of the nodes within the range of the j-th node with the sum of squares.
[0445] Optionally, the j-th node belongs to the preceding decoded and reconstructed nodes of the current layer; or
[0446] The j-th node belongs to the first m child nodes that have been decoded and reconstructed before the current node; or
[0447] The j-th node belongs to the previous k nodes that have been decoded and reconstructed in the current layer;
[0448] m and k are positive integers.
[0449] Optionally, the number of child nodes of the previously decoded and reconstructed nodes in the current layer is greater than or equal to q; or
[0450] The number of child nodes of the first k nodes is greater than or equal to q;
[0451] Where q∈[0,7] and is an integer.
[0452] Optionally, the parsing module 1010 is further configured to:
[0453] Parse the bitstream to obtain either m or k.
[0454] Optionally, the parsing module 1010 is further configured to:
[0455] The bitstream is parsed to obtain q.
[0456] Optionally, the current node belongs to the nodes of the first A layers of the transformed tree structure, where A is a positive integer.
[0457] Optionally, the parsing module 1010 is further configured to:
[0458] The bitstream is parsed to obtain a first parameter; the first parameter is used to instruct the nodes of the first A layer to perform cross-attribute prediction.
[0459] Optionally, the inverse transformation module 1050 is specifically used for:
[0460] The second transformation coefficient residual is obtained based on the DC coefficient value inherited by the current node and the DC coefficient value corresponding to the first attribute prediction value of the current node;
[0461] The reconstructed attribute residual value is obtained by performing an inverse transformation based on the reconstructed values of the second transformation coefficient residual and the first transformation coefficient residual;
[0462] Based on the reconstructed attribute residual value and the first attribute prediction value, the attribute reconstruction value of the placeholder child node of the current node is obtained.
[0463] Optionally, the inverse transformation module 1050 is specifically used for:
[0464] If the current node does not make a prediction, the DC coefficient is inherited and the difference between the DC coefficient value of the first attribute prediction value is calculated to obtain the second transformation coefficient residual.
[0465] Optionally, the inverse transformation module 1050 is specifically used for:
[0466] The attribute reconstruction value of the placeholder child node of the current node is obtained by summing the reconstructed attribute residual value and the first attribute prediction value.
[0467] Optionally, the inverse transformation module 1050 is specifically used for:
[0468] If the current node is to be predicted, then the predicted value of the second attribute of the current node is obtained;
[0469] The DC coefficient is inherited, and the difference between the DC coefficient value corresponding to the second attribute prediction value and the DC coefficient value corresponding to the first attribute prediction value is successively calculated to obtain the second transformation coefficient residual.
[0470] Optionally, the inverse transformation module 1050 is specifically used for:
[0471] The attribute reconstruction values of the placeholder child nodes of the current node are obtained by summing the reconstructed attribute residual value, the second attribute prediction value, and the first attribute prediction value.
[0472] In this embodiment, the decoding end decodes and dequantizes the bitstream to obtain the reconstructed value of the first transform coefficient residual of the point cloud to be decoded, constructs a transform tree structure of the current attribute of the point cloud to be decoded, obtains the reference attribute value of the reference attribute of the placeholder child node of the current node in the transform tree structure, and performs cross-attribute prediction based on the reference attribute value and scaling parameters to obtain the first attribute prediction value. Then, based on the first attribute prediction value and the reconstructed value of the first transform coefficient residual, an inverse transform is performed to obtain the attribute reconstruction value of the placeholder child node of the current node. This embodiment can utilize the correlation between the attribute types of the current node to perform cross-attribute prediction and obtain the reconstructed attribute value of the placeholder child node of the current node, which can help improve the compression rate of the bitstream.
[0473] The decoding apparatus 1000 provided in this application embodiment can implement the various processes implemented in the method embodiment of FIG8 and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0474] As shown in Figure 11, this application embodiment also provides an electronic device 1100, including a processor 1101 and a memory 1102. The memory 1102 stores a program or instructions that can run on the processor 1101. For example, when the electronic device 1100 is an encoding device, the program or instructions executed by the processor 1101 implement the various steps of the above-described encoding method embodiment and achieve the same technical effect. When the electronic device 1100 is a decoding device, the program or instructions executed by the processor 1101 implement the various steps of the above-described decoding method embodiment and achieve the same technical effect. To avoid repetition, this will not be described again here. Optionally, the memory 1102 can be the memory 102 or memory 113 in the embodiment shown in Figure 1, and the processor 1101 can implement the functions of the encoder 200 or decoder 300 in the embodiments shown in Figures 1-3.
[0475] This application also provides an electronic device, including: a memory configured to store video data; and a processing circuit configured to implement the steps of the encoding method embodiment described above, or the steps of the decoding method embodiment described above. Optionally, the memory may be memory 102 or memory 113 in the embodiment shown in FIG1, and the processing circuit may implement the functions of encoder 200 or decoder 300 in the embodiments shown in FIG1-3.
[0476] This application also provides an electronic device, including a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the steps in the method embodiments shown in FIG7 or 8. This device embodiment corresponds to the above method embodiments, and all implementation processes and methods of the above method embodiments can be applied to this terminal embodiment and can achieve the same technical effect.
[0477] The processor or processing circuit in this application embodiment may include general-purpose processors, special-purpose processors, etc., such as central processing units (CPUs), microprocessors, digital signal processors (DSPs), artificial intelligence (AI) processors, graphics processing units (GPUs), application-specific integrated circuits (ASICs), network processors (NPs), field-programmable gate arrays (FPGAs), or other programmable logic devices, gate circuits, transistors, discrete hardware components, etc. The communication interface in this application embodiment may include transceivers, pins, circuits, buses, etc.
[0478] The aforementioned electronic devices can be terminals or other devices besides terminals, such as servers, network attached storage (NAS), etc.
[0479] The terminal can be a mobile phone, tablet computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile internet device (MID), augmented reality (AR), virtual reality (VR) device, mixed reality (MR) device, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipboard equipment, pedestrian user equipment (PUE), smart home (home devices with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), game console, personal computer (PC), ATM or self-service machine, etc. Wearable devices include: smartwatches, smart bracelets, smart earphones, smart glasses, smart jewelry (smart bracelets, smart chains, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. Among these, in-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units, etc. It should be noted that the embodiments in this application do not limit the specific type of terminal.
[0480] A server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server. A cloud server can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), or cloud computing services based on big data and artificial intelligence platforms.
[0481] For example, the aforementioned electronic device may include, but is not limited to, the type of source device 100 or destination device 110 shown in FIG1.
[0482] Taking an electronic device as an example, Figure 12 is a schematic diagram of the hardware structure of a terminal implementing an embodiment of this application.
[0483] The terminal 1200 includes, but is not limited to, at least some of the following components: radio frequency unit 1201, network module 1202, audio output unit 1203, input unit 1204, sensor 1205, display unit 1206, user input unit 1207, interface unit 1208, memory 1209, and processor 1210.
[0484] Those skilled in the art will understand that terminal 1200 may also include a power supply (such as a battery) for powering various components. The power supply can be logically connected to processor 1210 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The terminal structure shown in Figure 12 does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0485] It should be understood that, in this embodiment, the input unit 1204 may include a graphics processor 12041 and a microphone 12042. The graphics processor 12041 processes image data of still images or videos obtained by an image acquisition device (such as a camera) in video acquisition mode or image acquisition mode, or it may process the obtained point cloud data. The display unit 1206 may include a display panel 12061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1207 includes at least one of a touch panel 12071 and other input devices 12072. The touch panel 12071 is also called a touch screen. The touch panel 12071 may include a touch detection device and a touch controller. Other input devices 12072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0486] In this embodiment, after receiving downlink data from the network-side device, the radio frequency unit 1201 can transmit it to the processor 1210 for processing; in addition, the radio frequency unit 1201 can send uplink data to the network-side device. Typically, the radio frequency unit 1201 includes, but is not limited to, antennas, amplifiers, transceivers, couplers, low-noise amplifiers, duplexers, etc.
[0487] The memory 1209 can be used to store software programs or instructions, as well as various data. The memory 1209 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1209 may include volatile memory or non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1209 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0488] Processor 1210 may include one or more processing units; optionally, processor 1210 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1210.
[0489] When terminal 1200 executes the encoding method, processor 1210 is used to:
[0490] Construct a transformation tree structure for the current attributes of the point cloud to be encoded;
[0491] Obtain the first original attribute value and the reference attribute value of the current attribute of the placeholder child node of the current node in the transformed tree structure;
[0492] Based on the first original attribute value, the reference attribute value, and the scaling parameter, the cross-attribute prediction residuals of the current attribute and the reference attribute are obtained; the scaling parameter is used to characterize the correlation between the current attribute and the reference attribute.
[0493] The cross-attribute prediction residuals are transformed, quantized, and encoded to obtain a bitstream.
[0494] In this embodiment, the encoding end constructs a transformation tree structure of the current attributes of the point cloud to be encoded, obtains the first original attribute value and the reference attribute value of the current attribute of the placeholder child node of the current node in the transformation tree structure, and obtains the cross-attribute prediction residual of the current specification attribute and the reference attribute based on the first original attribute value, the reference attribute value, and the scaling parameter. Then, the cross-attribute prediction residual is transformed, quantized, and encoded to obtain the bitstream. This embodiment can utilize the correlation between the attribute types of the current node to perform cross-attribute prediction, obtain the cross-attribute prediction residual, and then transform and encode the cross-attribute prediction residual to obtain the bitstream, which can help improve the compression rate of the bitstream and thus improve encoding efficiency.
[0495] When terminal 1200 executes the decoding method, processor 1210 is used to:
[0496] The bitstream is decoded and dequantized to obtain the reconstructed value of the first transform coefficient residual of the point cloud to be decoded;
[0497] Construct a transformation tree structure of the current attributes of the point cloud to be decoded;
[0498] Obtain the reference attribute value of the placeholder child node of the current node in the transformed tree structure;
[0499] Cross-attribute prediction is performed based on the reference attribute value and scaling parameters to obtain the first attribute prediction value; the scaling parameters are used to characterize the correlation between the current attribute and the reference attribute.
[0500] The attribute reconstruction values of the placeholder child nodes of the current node are obtained by performing an inverse transformation based on the first attribute prediction value and the reconstructed value of the first transformation coefficient residual.
[0501] In this embodiment, the decoding end decodes and dequantizes the bitstream to obtain the reconstructed value of the first transform coefficient residual of the point cloud to be decoded, constructs a transform tree structure of the current attribute of the point cloud to be decoded, obtains the reference attribute value of the reference attribute of the placeholder child node of the current node in the transform tree structure, and performs cross-attribute prediction based on the reference attribute value and scaling parameters to obtain the first attribute prediction value. Then, based on the first attribute prediction value and the reconstructed value of the first transform coefficient residual, an inverse transform is performed to obtain the attribute reconstruction value of the placeholder child node of the current node. This embodiment can utilize the correlation between the attribute types of the current node to perform cross-attribute prediction and obtain the reconstructed attribute value of the placeholder child node of the current node, which can help improve the compression rate of the bitstream.
[0502] It is understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description of the method embodiment and achieve the same or corresponding technical effect. To avoid repetition, it will not be described again here.
[0503] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described encoded method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0504] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described decoding method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0505] The processor mentioned above is the processor in the terminal described in the above embodiments. The readable storage medium includes computer-readable storage media, such as ROM, RAM, magnetic disk, or optical disk. In some examples, the readable storage medium may be a non-transient readable storage medium.
[0506] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described encoded method embodiment and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0507] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described decoding method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0508] It should be understood that the chips mentioned in the embodiments of this application may include system-on-a-chip (also known as system chip, chip system, or system-on-a-chip) or discrete display chips, etc.
[0509] This application also provides a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-encoded method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0510] This application also provides a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described decoding method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0511] This application also provides an encoding / decoding system, including an encoding end device and a decoding end device. The encoding end device can be used to perform the steps of the encoding method described above, and the decoding end device can be used to perform the steps of the decoding method described above.
[0512] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0513] From the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of computer software products plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes several instructions to cause the terminal or network-side device to execute the methods described in the various embodiments of this application.
[0514] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other implementations under the guidance of this application without departing from the spirit and scope of the claims. All of these implementations are within the protection scope of this application.
Claims
1. An encoding method, wherein, include: The encoding end constructs a transformation tree structure of the current attributes of the point cloud to be encoded; Obtain the first original attribute value and the reference attribute value of the current attribute of the placeholder child node of the current node in the transformed tree structure; Based on the first original attribute value, the reference attribute value, and the scaling parameter, obtain the cross-attribute prediction residual of the current attribute and the reference attribute; The scaling parameter is used to characterize the correlation between the current attribute and the reference attribute; The cross-attribute prediction residuals are transformed, quantized, and encoded to obtain a bitstream.
2. The method according to claim 1, wherein, The step of obtaining the cross-attribute prediction residuals of the current attribute and the reference attribute based on the first original attribute value, the reference attribute value, and the scaling parameter includes: If the current node makes a prediction, then obtain the first predicted value of the current attribute and the second predicted value of the reference attribute of the current node; The first prediction residual of the current attribute is obtained based on the first original attribute value and the first predicted value; and the second prediction residual of the reference attribute is obtained based on the reference attribute value and the second predicted value. The cross-attribute prediction residual is obtained based on the first prediction residual, the second prediction residual, and the scaling parameter.
3. The method according to claim 1, wherein, The step of obtaining the cross-attribute prediction residuals of the current attribute and the reference attribute based on the first original attribute value, the reference attribute value, and the scaling parameter includes: If the current node does not make a prediction, the cross-attribute prediction residual is obtained based on the first original attribute value, the reference attribute value, and the scaling parameter.
4. The method according to any one of claims 1-3, wherein, Each node in each level of the transformation tree structure corresponds to the same scaling parameter; or Each node in the transformation tree structure corresponds to the same scaling parameter.
5. The method according to any one of claims 1-3, wherein, Also includes: For the current layer of the transformed tree structure, obtain the nth node of the kth group; k and n are integers greater than or equal to 1; Obtain the second original attribute value of the current attribute and the third original attribute value of the reference attribute of the i-th node of the current layer; where i∈kn~(k+1)n-1 and are integers; The scaling parameter is determined based on the second original attribute value and the third original attribute value of the i-th node; wherein the scaling parameter is shared by the n nodes.
6. The method according to any one of claims 1-3, wherein, Also includes: For the current layer in the transformation tree structure, obtain the second original attribute value of the current attribute and the third original attribute value of the reference attribute of the i-th node of the current layer; where i is less than or equal to the number of nodes in the current layer and is a positive integer; The scaling parameter is determined based on the second original attribute value and the third original attribute value of the i-th node; wherein the scaling parameter is shared by all nodes in the current layer.
7. The method according to claim 5 or 6, wherein, Determining the scaling parameter based on the second original attribute value and the third original attribute value of the i-th node includes: Determine the product of the second original attribute value and the third original attribute value of the i-th node, and sum the products corresponding to the nodes within the range of the i-th node; Determine the sum of squares of the third original attribute values of the nodes within the range of the i-th node; The scaling parameter is obtained by comparing the sum of the products of the nodes within the range of the i-th node with the sum of squares.
8. The method according to any one of claims 1-3, wherein, Also includes: For the current layer of the transformed tree structure, obtain the first attribute reconstruction data of the current attribute and the second attribute reconstruction data of the reference attribute of the j-th node of the current layer; the j-th node belongs to the already encoded and reconstructed node; j is a positive integer; The scaling parameter is determined based on the first attribute reconstruction data and the second attribute reconstruction data of the j-th node.
9. The method according to claim 8, wherein, The step of determining the scaling parameter based on the reconstructed data of the first attribute and the reconstructed data of the second attribute of the j-th node includes: Determine the product of the first attribute reconstruction data and the second attribute reconstruction data of the j-th node, and sum the products corresponding to the nodes within the range of the j-th node; Determine the sum of squares of the second attribute reconstructed data of the nodes within the range of the j-th node; The scaling parameter is obtained by comparing the sum of the products of the nodes within the range of the j-th node with the sum of squares.
10. The method according to claim 8 or 9, wherein, The j-th node belongs to the pre-encoded and reconstructed nodes of the current layer; or The j-th node belongs to the first m child nodes whose preceding order has been reconstructed for the current node; or The j-th node belongs to the first k nodes of the current layer that have been encoded and reconstructed in the preceding sequence; m and k are both positive integers.
11. The method according to claim 10, wherein, The number of child nodes of the previously encoded and reconstructed nodes in the current layer is greater than or equal to q; or The number of child nodes of the first k nodes is greater than or equal to q; Where q∈[0,7] and is an integer.
12. The method according to claim 10, wherein, Also includes: The m or the k is passed into the bitstream.
13. The method according to claim 12, wherein, Also includes: The q is passed into the bitstream.
14. The method according to any one of claims 1-13, wherein, Also includes: The scaling parameter is passed into the bitstream.
15. The method according to any one of claims 1-14, wherein, The current node belongs to the first A layers of the transformation tree structure, where A is a positive integer.
16. The method according to claim 15, wherein, Also includes: Pass the first parameter into the bitstream; The first parameter is used to instruct the nodes of the first A layer to perform cross-attribute prediction.
17. A decoding method, wherein, include: The decoding end decodes and dequantizes the bitstream to obtain the reconstructed value of the first transform coefficient residual of the point cloud to be decoded; Construct a transformation tree structure of the current attributes of the point cloud to be decoded; Obtain the reference attribute value of the placeholder child node of the current node in the transformed tree structure; Cross-attribute prediction is performed based on the reference attribute value and scaling parameters to obtain the first attribute prediction value; the scaling parameters are used to characterize the correlation between the current attribute and the reference attribute. The attribute reconstruction values of the placeholder child nodes of the current node are obtained by performing an inverse transformation based on the first attribute prediction value and the reconstructed value of the first transformation coefficient residual.
18. The method according to claim 17, wherein, The step of obtaining the first attribute prediction value based on the reference attribute value and the scaling parameter includes: If the current node allows prediction, then obtain the second predicted value of the reference attribute of the current node; Based on the second predicted value and the reference attribute value, obtain the second predicted residual of the reference attribute; The predicted value of the first attribute is obtained based on the second prediction residual and the scaling parameter.
19. The method of claim 17, wherein, The step of obtaining the first attribute prediction value based on the reference attribute value and the scaling parameter includes: If the current node does not make a prediction, then the predicted value of the first attribute is obtained based on the reference attribute value and the scaling parameter.
20. The method according to any one of claims 17-19, wherein, Each node in the transformation tree structure corresponds to the same scaling parameter; or The nodes in the current layer share the scaling parameters; or The scaling parameter is shared by every n nodes in the current layer, where n is an integer greater than or equal to 1.
21. The method according to any one of claims 17-19, wherein, Also includes: The scaling parameters are obtained by parsing the bitstream.
22. The method according to any one of claims 17-19, wherein, Also includes: For the current layer of the transformed tree structure, obtain the first attribute reconstruction data of the current attribute and the second attribute reconstruction data of the reference attribute of the j-th node of the current layer; the j-th node belongs to the nodes that have been decoded and reconstructed; j is a positive integer; The scaling parameter is determined based on the first attribute reconstruction data and the second attribute reconstruction data of the j-th node.
23. The method according to claim 22, wherein, The step of determining the scaling parameter based on the reconstructed data of the first attribute and the reconstructed data of the second attribute of the j-th node includes: Determine the product of the first attribute reconstruction data and the second attribute reconstruction data of the j-th node, and sum the products corresponding to the nodes within the range of the j-th node; Determine the sum of squares of the second attribute reconstructed data of the nodes within the range of the j-th node; The scaling parameter is obtained by comparing the sum of the products of the nodes within the range of the j-th node with the sum of squares.
24. The method according to claim 22 or 23, wherein, The j-th node belongs to the previously decoded and reconstructed nodes of the current layer; or The j-th node belongs to the previous m child nodes that have been decoded and reconstructed before the current node; or The j-th node belongs to the previous k nodes that have been decoded and reconstructed in the current layer; m and k are positive integers.
25. The method according to any one of claims 22-24, wherein, The number of child nodes of the previously decoded and reconstructed nodes in the current layer is greater than or equal to q; or The number of child nodes of the first k nodes is greater than or equal to q; Where q∈[0,7] and is an integer.
26. The method according to claim 24, wherein, Also includes: Parse the bitstream to obtain either m or k.
27. The method according to claim 25, wherein, Also includes: The bitstream is parsed to obtain q.
28. The method according to any one of claims 17-27, wherein, The current node belongs to the first A layer of the transformation tree structure.
29. The method according to claim 28, wherein, Also includes: The bitstream is parsed to obtain a first parameter; the first parameter is used to instruct the nodes of the first A layer to perform cross-attribute prediction.
30. The method according to any one of claims 17-29, wherein, The step of performing an inverse transformation based on the reconstructed value of the first attribute prediction value and the first transformation coefficient residual to obtain the attribute reconstruction value of the occupier child node of the current node includes: The second transformation coefficient residual is obtained based on the DC coefficient value inherited by the current node and the DC coefficient value corresponding to the first attribute prediction value of the current node; The reconstructed attribute residual value is obtained by performing an inverse transformation based on the reconstructed values of the second transformation coefficient residual and the first transformation coefficient residual; Based on the reconstructed attribute residual value and the first attribute prediction value, the attribute reconstruction value of the placeholder child node of the current node is obtained.
31. The method according to claim 30, wherein, The step of obtaining the second transformation coefficient residual based on the DC coefficient value inherited from the current node and the DC coefficient value predicted by the first attribute of the current node includes: If the current node does not make a prediction, the DC coefficient is inherited and the difference between the DC coefficient value of the first attribute prediction value is calculated to obtain the second transformation coefficient residual.
32. The method according to claim 31, wherein, The step of obtaining the attribute reconstruction value of the placeholder child node of the current node based on the reconstructed attribute residual value and the first attribute prediction value includes: The attribute reconstruction value of the placeholder child node of the current node is obtained by summing the reconstructed attribute residual value and the first attribute prediction value.
33. The method according to claim 30, wherein, The step of obtaining the second transformation coefficient residual based on the DC coefficient value inherited by the current node and the DC coefficient value corresponding to the first attribute prediction value of the current node includes: If the current node is to be predicted, then the predicted value of the second attribute of the current node is obtained; The DC coefficient is inherited, and the difference between the DC coefficient value corresponding to the second attribute prediction value and the DC coefficient value corresponding to the first attribute prediction value is successively calculated to obtain the second transformation coefficient residual.
34. The method according to claim 33, wherein, The step of obtaining the attribute reconstruction value of the placeholder child node of the current node based on the reconstructed attribute residual value and the first attribute prediction value includes: The attribute reconstruction values of the placeholder child nodes of the current node are obtained by summing the reconstructed attribute residual value, the second attribute prediction value, and the first attribute prediction value.
35. An encoding device, wherein, include: The building module is used by the encoding end to construct the transformation tree structure of the current attributes of the point cloud to be encoded; The acquisition module is used to acquire the first original attribute value and the reference attribute value of the current attribute of the placeholder child node of the current node of the transformed tree structure; The prediction module is used to obtain the cross-attribute prediction residuals of the current attribute and the reference attribute based on the first original attribute value, the reference attribute value and the scaling parameter; The scaling parameter is used to characterize the correlation between the current attribute and the reference attribute; The encoding module is used to transform, quantize, and encode the cross-attribute prediction residuals to obtain a bitstream.
36. The apparatus according to claim 35, wherein, The prediction module is specifically used for: If the current node makes a prediction, then obtain the first predicted value of the current attribute and the second predicted value of the reference attribute of the current node; The first prediction residual of the current attribute is obtained based on the first original attribute value and the first prediction value; And obtain the second prediction residual of the reference attribute based on the reference attribute value and the second prediction value; The cross-attribute prediction residual is obtained based on the first prediction residual, the second prediction residual, and the scaling parameter.
37. The apparatus according to claim 35, wherein, The prediction module is specifically used for: If the current node does not make a prediction, the cross-attribute prediction residual is obtained based on the first original attribute value, the reference attribute value, and the scaling parameter.
38. The apparatus according to any one of claims 35-37, wherein, The acquisition unit is also used for: For the current layer of the transformed tree structure, obtain the n nodes of the k-th group; k and n are integers greater than or equal to 1. Obtain the second original attribute value of the current attribute and the third original attribute value of the reference attribute of the i-th node of the current layer; where i∈kn~(k+1)n-1 and are integers; The scaling parameter is determined based on the second original attribute value and the third original attribute value of the i-th node; wherein the scaling parameter is shared by the n nodes.
39. The apparatus according to any one of claims 35-37, wherein, The acquisition unit is also used for: For the current layer in the transformation tree structure, obtain the second original attribute value of the current attribute and the third original attribute value of the reference attribute of the i-th node of the current layer; where i is less than or equal to the number of nodes in the current layer and is a positive integer; The scaling parameter is determined based on the second original attribute value and the third original attribute value of the i-th node; wherein the scaling parameter is shared by all nodes in the current layer.
40. The apparatus according to any one of claims 35-37, wherein, The acquisition unit is also used for: For the current layer of the transformed tree structure, obtain the first attribute reconstruction data of the current attribute and the second attribute reconstruction data of the reference attribute of the j-th node of the current layer; the j-th node belongs to the already encoded and reconstructed node; j is a positive integer; The scaling parameter is determined based on the first attribute reconstruction data and the second attribute reconstruction data of the j-th node.
41. The apparatus according to claim 40, wherein, The j-th node belongs to the pre-encoded and reconstructed nodes of the current layer; or The j-th node belongs to the first m child nodes whose preceding order has been reconstructed for the current node; or The j-th node belongs to the first k nodes of the current layer that have been encoded and reconstructed in the preceding sequence; m and k are both positive integers.
42. A decoding apparatus, wherein, include: The parsing module is used by the decoding end to decode and dequantize the bitstream to obtain the reconstructed value of the first transform coefficient residual of the point cloud to be decoded; A construction module is used to construct the transformation tree structure of the current attributes of the point cloud to be decoded; The acquisition module is used to acquire the reference attribute values of the placeholder child nodes of the current node in the transformed tree structure. The prediction module is used to perform cross-attribute prediction based on the reference attribute value and the scaling parameter to obtain a first attribute prediction value; the scaling parameter is used to characterize the correlation between the current attribute and the reference attribute. The inverse transformation module is used to perform an inverse transformation based on the first attribute prediction value and the reconstructed value of the first transformation coefficient residual to obtain the attribute reconstruction value of the placeholder child node of the current node.
43. The apparatus according to claim 42, wherein, The prediction module is specifically used for: If the current node allows prediction, then obtain the second predicted value of the reference attribute of the current node; Based on the second predicted value and the reference attribute value, obtain the second predicted residual of the reference attribute; The predicted value of the first attribute is obtained based on the second prediction residual and the scaling parameter.
44. The apparatus according to claim 42, wherein, The prediction module is specifically used for: If the current node does not make a prediction, then the predicted value of the first attribute is obtained based on the reference attribute value and the scaling parameter.
45. The apparatus according to any one of claims 42-44, wherein, The acquisition module is also used for: For the current layer of the transformed tree structure, obtain the first attribute reconstruction data of the current attribute and the second attribute reconstruction data of the reference attribute of the j-th node of the current layer; the j-th node belongs to the nodes that have been decoded and reconstructed; j is a positive integer; The scaling parameter is determined based on the first attribute reconstruction data and the second attribute reconstruction data of the j-th node.
46. The apparatus according to any one of claims 42-45, wherein, The inverse transformation module is specifically used for: The second transformation coefficient residual is obtained based on the DC coefficient value inherited by the current node and the DC coefficient value corresponding to the first attribute prediction value of the current node; The reconstructed attribute residual value is obtained by performing an inverse transformation based on the reconstructed values of the second transformation coefficient residual and the first transformation coefficient residual; Based on the reconstructed attribute residual value and the first attribute prediction value, the attribute reconstruction value of the placeholder child node of the current node is obtained.
47. The apparatus according to claim 46, wherein, The inverse transformation module is specifically used for: If the current node does not make a prediction, the DC coefficient is inherited and the difference between the DC coefficient value of the first attribute prediction value is calculated to obtain the second transformation coefficient residual.
48. The apparatus according to claim 47, wherein, The inverse transformation module is specifically used for: The attribute reconstruction value of the placeholder child node of the current node is obtained by summing the reconstructed attribute residual value and the first attribute prediction value.
49. The apparatus according to claim 46, wherein, The inverse transformation module is specifically used for: If the current node is to be predicted, then the predicted value of the second attribute of the current node is obtained; The DC coefficient is inherited, and the difference between the DC coefficient value corresponding to the second attribute prediction value and the DC coefficient value corresponding to the first attribute prediction value is successively calculated to obtain the second transformation coefficient residual.
50. The apparatus according to claim 49, wherein, The inverse transformation module is specifically used for: The attribute reconstruction values of the placeholder child nodes of the current node are obtained by summing the reconstructed attribute residual value, the second attribute prediction value, and the first attribute prediction value.
51. An electronic device, wherein, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the encoding method as claimed in any one of claims 1 to 16, or to implement the steps of the decoding method as claimed in any one of claims 17 to 34.
52. A readable storage medium, wherein, The readable storage medium stores a program or instructions that, when executed by a processor, implement the encoding method as described in any one of claims 1 to 16, or the decoding method as described in any one of claims 17 to 34.
53. A chip, wherein, The chip includes a processor and a communication interface coupled to the processor. The processor is used to run programs or instructions to implement the steps of the method as claimed in any one of claims 1 to 16, or to implement the steps of the method as claimed in claims 17 to 34.
Citation Information
Patent Citations
Restriction on applicability of cross component mode
CN113692739A
Coding of color attribute components in geometry-based point cloud compression (G-PCC)
CN116325747A
Inter-component residual prediction of color attributes in geometric point cloud compression coding
CN116325748A
Point cloud data transmission method, point cloud data transmission device, point cloud data reception method, and point cloud data reception device
US20240029312A1
Data processing method and apparatus for immersive media, and device, medium and product
WO2024037137A1