Encoding method, decoding method and related devices
By directly copying point cloud information when the similarity of point cloud blocks is high, the problem of high resource consumption in inter-frame predictive coding of point clouds is solved, and coding efficiency is improved.
Patent Information
- Application Number
- CN202311312190.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-10
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2043-10-10
AI Technical Summary
Existing point cloud inter-frame predictive coding methods consume significant coding resources.
If the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold, the encoder copies the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block, and generates a target bitstream to instruct the decoder to perform the corresponding copy.
When the similarity of point cloud blocks is high, point cloud information can be directly copied to save the encoding process, reduce encoding resources and time consumption, and improve encoding efficiency.
Smart Images

Figure CN119815052B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of encoding and decoding technology, specifically relating to an encoding method, a decoding method, and related equipment. Background Technology
[0002] In related technologies, when performing inter-frame predictive coding of point clouds, a prediction entropy coding mode is used to encode the current point cloud frame. Specifically, based on the relationship between the node size corresponding to the current point cloud frame and the node size corresponding to the Largest Prediction Unit (LPU), a reference point cloud is used as prediction information to encode the current point cloud frame, or a motion-compensated point cloud is used as prediction information to encode the current point cloud frame. When the point cloud frame corresponds to a large number of point clouds, this coding method will consume a large amount of coding resources. Summary of the Invention
[0003] This application provides an encoding method, a decoding method, and related equipment that can solve the problem of existing point cloud inter-frame predictive coding methods consuming large amounts of encoding resources.
[0004] Firstly, an encoding method is provided, including:
[0005] If the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold, the encoder copies the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block.
[0006] The encoding end generates a target bitstream based on the first identification information, wherein the first identification information is used to instruct the decoding end to copy the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block, and the point cloud information includes at least one of the geometric encoding information and attribute encoding information of the point cloud corresponding to the second point cloud block.
[0007] Secondly, a decoding method is provided, including:
[0008] The decoding end decodes the target bitstream to obtain the first identification information;
[0009] When the first identification information instructs the encoding end to copy the point cloud information of the second point cloud block of the reference point cloud frame or the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud block of the target point cloud frame, the decoding end copies the point cloud information of the second point cloud block or the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud block. The point cloud information includes at least one of the geometric encoding information and attribute encoding information of the point cloud corresponding to the second point cloud block.
[0010] Thirdly, an encoding device is provided, comprising:
[0011] The first copying module is used to copy the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block when the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold.
[0012] The generation module is used to generate a target bitstream based on the first identification information, wherein the first identification information is used to instruct the decoding end to copy the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block to the reconstructed point cloud of the first point cloud block, and the point cloud information includes at least one of the geometric encoding information and attribute encoding information of the point cloud corresponding to the second point cloud block.
[0013] Fourthly, a decoding device is provided, comprising:
[0014] The fourth acquisition module is used to decode the target bitstream to obtain the first identification information;
[0015] The second copying module is configured to copy the point cloud information of the second point cloud block or the motion-compensated second point cloud block of the reference point cloud frame to the reconstructed point cloud of the first point cloud block of the target point cloud frame when the first identification information instructs the encoding end to copy the point cloud information of the second point cloud block or the motion-compensated second point cloud block to the reconstructed point cloud of the first point cloud block of the target point cloud frame. The point cloud information includes at least one of the geometric encoding information and attribute encoding information of the point cloud corresponding to the second point cloud block.
[0016] Fifthly, an electronic device is provided, the terminal including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first aspect, or implementing the steps of the method as described in the second aspect.
[0017] In a sixth aspect, an electronic device is provided, including a processor and a communication interface, wherein the processor is configured to, when the similarity between a first point cloud block of a target point cloud frame and a second point cloud block of a reference point cloud frame is greater than or equal to a preset threshold, copy the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block; and generate a target bitstream according to first identification information, wherein the first identification information is used to instruct the encoding end to copy the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block, the point cloud information... The information includes at least one of the geometric encoding information and attribute encoding information of the point cloud corresponding to the second point cloud block; or, the processor is used to decode the target bitstream to obtain first identification information; when the first identification information instructs the encoding end to copy the point cloud information of the second point cloud block of the reference point cloud frame or the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud block of the target point cloud frame, the point cloud information of the second point cloud block or the second point cloud block after motion compensation is copied to the reconstructed point cloud of the first point cloud block, wherein the point cloud information includes at least one of the geometric encoding information and attribute encoding information of the point cloud corresponding to the second point cloud block.
[0018] A seventh aspect provides an electronic device comprising: a memory configured to store video data, and processing circuitry configured to implement the steps of the method described in the first aspect, or the steps of the method described in the second aspect.
[0019] Eighthly, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect, or implement the steps of the method described in the second aspect.
[0020] A ninth aspect provides an encoding / decoding system, comprising: an encoding end device and a decoding end device, wherein the encoding end device is configured to perform the steps of the method described in the first aspect, and the decoding end device is configured to perform the steps of the method described in the second aspect.
[0021] In a tenth aspect, a chip is provided, the chip including a processor and a communication interface coupled to the processor, the processor being configured to run a program or instructions to implement the steps of the method described in the first aspect, or to implement the steps of the method described in the second aspect.
[0022] Eleventhly, a computer program / program product is provided, which is stored in a storage medium and is executed by at least one processor to implement the steps of the method as described in the first aspect, or to implement the steps of the method as described in the second aspect.
[0023] In this embodiment, when the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold, the encoding end copies the point cloud information of the second point cloud block or the motion-compensated point cloud information of the second point cloud block into the reconstructed point cloud of the first point cloud block. The encoding end generates a target bitstream based on first identifier information, which instructs the decoding end to copy the point cloud information of the second point cloud block or the motion-compensated point cloud information of the second point cloud block into the reconstructed point cloud of the first point cloud block. Through this scheme, when the similarity between two point cloud blocks is sufficiently high, the point cloud information of one point cloud block is directly copied into the reconstructed point cloud of the other, saving the process of encoding the point cloud of the other point cloud block. This saves encoding resources and time while ensuring that the point cloud quality does not fluctuate significantly, effectively improving encoding efficiency. Attached Figure Description
[0024] Figure 1 This diagram illustrates the encoding / decoding system provided in an embodiment of this application.
[0025] Figure 2 A flowchart illustrating the encoding process performed by an encoder based on the MPEG G-PCC encoding framework;
[0026] Figure 3 A flowchart illustrating the encoding process performed by an encoder based on the MPEG G-PCC encoding framework;
[0027] Figure 4 This is a flowchart illustrating the decoding process performed by the decoder within the AVS-PCC-based decoding framework.
[0028] Figure 5 A flowchart illustrating the decoding process performed by a decoder based on the MPEG G-PCC decoding framework;
[0029] Figure 6 A schematic diagram illustrating the use of octree-based geometric coding for inter-frame prediction information;
[0030] Figure 7 This diagram illustrates the motion prediction process based on the PU (Programmable Logic Unit).
[0031] Figure 8 A diagram illustrating the encoding of motion vectors and the use of prediction information;
[0032] Figure 9 One of the flowcharts illustrating the encoding method of an embodiment of this application;
[0033] Figure 10 A second flowchart illustrating the encoding method of an embodiment of this application;
[0034] Figure 11 The third flowchart illustrating the encoding method of this application embodiment;
[0035] Figure 12 This diagram illustrates the motion search process in an embodiment of this application.
[0036] Figure 13 A flowchart illustrating the decoding method according to an embodiment of this application;
[0037] Figure 14 A schematic diagram of the module of the encoding device according to an embodiment of this application;
[0038] Figure 15 A schematic diagram of the module of the decoding device according to an embodiment of this application;
[0039] Figure 16 A structural block diagram illustrating an embodiment of the electronic device of this application;
[0040] Figure 17 This is a structural block diagram illustrating the terminal in an embodiment of this application. Detailed Implementation
[0041] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0042] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, without limiting the number of objects; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, "A or B" covers three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0043] Before introducing the technical solutions provided in the embodiments of this application, the meanings of some terms will be explained first.
[0044] Point cloud: A point cloud is a set of discrete points in space that are randomly distributed and represent the spatial structure and surface properties of a three-dimensional object or scene. Point clouds can be classified into different categories according to different classification criteria. For example, according to the method of acquiring the point cloud, it can be divided into dense point clouds and sparse point clouds; or according to the temporal type of the point cloud, it can be divided into static point clouds and dynamic point clouds.
[0045] Point cloud data: Point cloud data is composed of the geometric coordinates and attribute information of each point. Geometric coordinate information, also known as 3D position information, refers to the spatial coordinates (x, y, z) of a point in the point cloud. This can include the coordinate values of the point along each coordinate axis of a 3D coordinate system, such as the coordinate value x along the X-axis, the coordinate value y along the Y-axis, and the coordinate value z along the Z-axis. The attribute information of a point in the point cloud can include at least one of the following: color information, material information, and laser reflection intensity information (also known as reflectivity). Typically, each point in the point cloud has the same number of attribute information. For example, each point in the point cloud can have both color information and laser reflection intensity information, or it can have color information, material information, and laser reflection intensity information.
[0046] Point cloud encoding (PCC) refers to the process of encoding the geometric coordinates and attribute information of each point in a point cloud to obtain a compressed bitstream. Point cloud encoding can include two main processes: geometric coordinate information encoding and attribute information encoding. Currently, point cloud encoding frameworks that can compress point clouds include the Geometry-Point Cloud Compression (G-PCC) codec framework provided by the Moving Picture Experts Group (MPEG) or the Video Point Cloud Compression (V-PCC) codec framework, or the AVS-PCC codec framework provided by the Audio Video Standard (AVS).
[0047] Point cloud decoding: Point cloud decoding refers to the process of decoding the compressed bitstream obtained from point cloud encoding to reconstruct the point cloud. More specifically, it refers to the process of reconstructing the geometric coordinates and attribute information of each point in the point cloud based on the geometric bitstream and attribute bitstream in the compressed bitstream. After obtaining the compressed bitstream at the decoding end, for the geometric bitstream, entropy decoding is first performed to obtain the quantized information of each point in the point cloud, and then inverse quantization is performed to reconstruct the geometric coordinates of each point in the point cloud. For the attribute bitstream, entropy decoding is first performed to obtain the quantized attribute residual information or quantized transform coefficients of each point in the point cloud; then, inverse quantization is performed on the quantized attribute residual information to obtain the reconstructed residual information, and inverse quantization is performed on the quantized transform coefficients to obtain the reconstructed transform coefficients. The reconstructed transform coefficients are then inversely transformed to obtain the reconstructed residual information. Based on the reconstructed residual information of each point in the point cloud, the attribute information of each point in the point cloud can be reconstructed. The reconstructed attribute information of each point in the point cloud is then matched one-to-one with the reconstructed geometric coordinate information in sequence to reconstruct the point cloud.
[0048] Figure 1 This is a schematic diagram of the encoding / decoding system provided in an embodiment of this application. The technical solution of this application embodiment relates to encoding / decoding (CODEC) point cloud data (including encoding or decoding).
[0049] like Figure 1 As shown, the encoding / decoding system includes a source device 100, which provides encoded point cloud data to be decoded and displayed by a destination device 110. Specifically, the source device 100 provides the point cloud data to the destination device 110 via a communication medium 120. The source device 100 and the destination device 110 may include any one or more of the following: desktop computer, laptop computer, tablet computer, set-top box, mobile phone, wearable device (e.g., smartwatch or wearable camera), television, camera, display device, in-vehicle device, virtual reality (VR) device, augmented reality (AR) device, mixed reality (MR) device, digital media player, video game console, video conferencing equipment, video streaming equipment, broadcast receiver equipment, broadcast transmitter equipment, spacecraft, aircraft, robot, satellite, etc.
[0050] exist Figure 1 In this example, source device 100 includes a data source 101, a memory 102, an encoder 200, and an output interface 104. Destination device 110 includes an input interface 111, a decoder 300, a memory 113, and a display device 114. Source device 100 represents an example of an encoding device, while destination device 110 represents an example of a decoding device. In other examples, source device 100 and destination device 110 may not include... Figure 1Some components, or may include Figure 1 Other components besides the source device 100. For example, the source device 100 can acquire point cloud data through an external capture device. Similarly, the destination device 110 can interface with an external display device, without including an integrated display device. Furthermore, the memory 102 and memory 113 can be external memories.
[0051] Although Figure 1 Source device 100 and destination device 110 are illustrated as separate devices, but in some examples, they may also be integrated into a single device. In such embodiments, the same hardware or software, or separate hardware or software, or any combination thereof, may be used to implement the functionality corresponding to source device 100 and destination device 110.
[0052] In some examples, source device 100 and destination device 110 can perform unidirectional or bidirectional data transmission. In the case of bidirectional data transmission, source device 100 and destination device 110 can operate in a substantially symmetrical manner, i.e., each of source device 100 and destination device 110 includes an encoder and a decoder.
[0053] Data source 101 represents the source of point cloud data (i.e., raw, unencoded point cloud data) and provides the point cloud data to encoder 200, which encodes the point cloud data. Source device 100 may include capture devices (e.g., camera devices, sensing devices, or scanning devices), archives containing previously captured point cloud data, or feed interfaces for receiving point cloud data from data content providers. Camera devices may include ordinary cameras, stereo cameras, and light field cameras; sensing devices may include laser devices, radar devices, etc.; and scanning devices may include 3D laser scanning devices, etc. Point cloud data can be obtained by capturing real-world visual scenes using capture devices. Alternatively, data source 101 may generate computer graphics-based data as source data, or combine real-time data, archived data, and computer-generated data. For example, the data source may generate point cloud data based on virtual objects (e.g., virtual 3D objects and virtual 3D scenes obtained through 3D modeling).
[0054] Encoder 200 encodes captured, pre-captured, or computer-generated data. Encoder 200 can rearrange point cloud data from the received order (sometimes referred to as the "display order") according to the encoded order. Encoder 200 can generate a bitstream including the encoded point cloud data. Source device 100 can then output the encoded point cloud data to communication medium 120 via output interface 104 for reception or retrieval, for example, by input interface 111 of destination device 110.
[0055] The memory 102 of the source device 100 and the memory 113 of the destination device 110 represent general-purpose memory. In some examples, memory 102 may store raw data from data source 101, and memory 113 may store decoded point cloud data from decoder 300. Additionally or alternatively, memories 102 and 113 may respectively store software instructions executable by, for example, encoder 200 and decoder 300. Although memories 102 and 113 are shown separately from encoder 200 and decoder 300 in this example, it should be understood that encoder 200 and decoder 300 may also include internal memory for functionally similar or equivalent purposes. If encoder 200 and decoder 300 are deployed on the same hardware device, memories 102 and 113 may be the same memory. Furthermore, memories 102 and 113 may store, for example, encoded point cloud data output from encoder 200 and input to decoder 300. In some examples, portions of memories 102 and 113 may be allocated as one or more point cloud buffers, for example, to store raw, decoded, or encoded point cloud data.
[0056] In some examples, source device 100 can output encoded data from output interface 104 to memory 113. Similarly, destination device 110 can access encoded data from memory 113 via input interface 111. Memory 113 or memory 102 can include any of a variety of distributed or locally accessed data storage media, such as hard drives, Blu-ray discs, digital versatile discs (DVDs), compact disc read-only memory (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded point cloud data.
[0057] Output interface 104 may include any type of medium or device capable of transmitting encoded point cloud data from source device 100 to destination device 110. For example, output interface 104 may include a transmitter or transceiver, such as an antenna, configured to transmit encoded point cloud data directly from source device 100 to destination device 110 in real time. The encoded point cloud data may be modulated according to the communication standards of a wireless communication protocol and transmitted to destination device 110.
[0058] Communication medium 120 may include transient media, such as wireless broadcasting or wired network transmission. For example, communication medium 120 may include radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). Communication medium 120 may form part of a packet-based network (such as a local area network, a wide area network, or a global network such as the Internet). Communication medium 120 may also take the form of a storage medium (e.g., a non-transitory storage medium), such as a hard disk, flash drive, compact disk, digital point cloud disk, Blu-ray disc, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded point cloud data.
[0059] In some implementations, the communication medium 120 may include a router, switch, base station, or any other device that can be used to facilitate communication from source device 100 to destination device 110. For example, a server (not shown) may receive encoded point cloud data from source device 100 and provide it to destination device 110, for example, via network transmission. The server may include (e.g., a web server for a website), a server configured to provide file transfer protocol services (such as File Transfer Protocol (FTP) or File DeliveryOver Unidirectional Transport (FLUTE) protocol), a content delivery network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Services (MBMS) or Evolved Multimedia Broadcast Multicast Service (eMBMS) server, or a Network-attached Storage (NAS) device, etc. The server can implement one or more HTTP streaming protocols, such as MPEG Media Transport (MMT), Dynamic Adaptive Streaming over HTTP (DASH), HTTP Live Streaming (HLS), or Real Time Streaming Protocol (RTSP).
[0060] Destination device 110 can access encoded point cloud data from a server, for example, via a wireless channel (e.g., Wi-Fi connection) or a wired connection (e.g., Digital subscriber line (DSL), cable modem, etc.) for accessing encoded point cloud data stored on the server.
[0061] Output interface 104 and input interface 111 can represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to the IEEE 802.11 or IEEE 802.15 standard (e.g., ZigBee™), Bluetooth standard, or other physical components. In an example where output interface 104 and input interface 111 include wireless components, output interface 104 and input interface 111 can be configured to transmit data, such as encoded point cloud data, via Wi-Fi, Ethernet, or cellular networks (such as 4G, LTE (Long Term Evolution), Advanced LTE, 5G, 6G, etc.).
[0062] The technology provided in this application can be applied to support one or more of the following application scenarios: machine-perceived point clouds, which can be used in autonomous navigation systems, real-time inspection systems, geographic information systems, visual sorting robots, disaster relief robots, and other scenarios; human-perceived point clouds, which can be used in point cloud application scenarios such as digital cultural heritage, free-viewpoint broadcasting, 3D immersive communication, and 3D immersive interaction.
[0063] The input interface 111 of the destination device 110 receives an encoded bitstream from the communication medium 120. The encoded bitstream may include high-level syntax elements and encoded data units (e.g., sequences, image groups, images, slices, blocks, etc.), where the high-level syntax elements are used to decode the encoded data units to obtain decoded point cloud data. The display device 114 displays the decoded point cloud data to the user. The display device 114 may include a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices. In some examples, the destination device 110 may not have a display device 114; for example, if the decoded point cloud data is used to determine the location of a physical object, the display device 114 may be replaced by a processor.
[0064] The encoder 200 and decoder 300 can be implemented as one or more of various processing circuits, which may include microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. When the technology is implemented wholly or partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and use one or more processors to execute the instructions in hardware to perform the technology provided in the embodiments of this application.
[0065] The basic principles of the encoder 200 and decoder 300 provided in this application embodiment are introduced below, taking the G-PCC and AVS-PCC codec frameworks as examples.
[0066] The encoding and decoding frameworks of G-PCC and AVS-PCC are largely the same. For example... Figure 2 The diagram illustrates the encoding flowchart executed by the encoder within the AVS-PCC-based encoding framework, as shown below. Figure 3 This diagram illustrates the encoding flowchart performed by an encoder based on the MPEG G-PCC encoding framework. The encoder described above can be... Figure 1 The encoder 200 shown above. The above encoding framework can be broadly divided into a geometric coordinate information encoding process and an attribute information encoding process. In the geometric information encoding process, the geometric coordinate information of each point in the point cloud is encoded to obtain a geometric bitstream; in the attribute information encoding process, the attribute information of each point in the point cloud is encoded to obtain an attribute bitstream; the geometric bitstream and the attribute bitstream together constitute the compressed bitstream of the point cloud.
[0067] For the geometric information encoding process, the encoding flow executed by encoder 200 is as follows:
[0068] 1. Pre-processing: This can include coordinate transformation and voxelization. Through scaling and translation operations, pre-processing converts the point cloud data in 3D space into integer form and moves its smallest geometric position to the origin. In some examples, encoder 200 may not perform pre-processing.
[0069] 2. Geometric Coding: For the AVS-PCC coding framework, geometric coding includes two modes: octree-based geometric coding and prediction tree-based geometric coding. For the G-PCC coding framework, geometric coding includes three modes: octree-based geometric coding, trisoup-based geometric coding, and prediction tree-based prediction coding. Among them:
[0070] Octree-based geometric encoding: First, the geometric information is transformed to ensure that the entire point cloud is contained within an octree-based system consisting of two extreme points (0, 0, 0) and (2... d ,2 d ,2 d The bounding box is determined by the data structure and then voxelized, which involves quantization, rounding, and removal of duplicate points (depending on the parameters). Next, following a breadth-first traversal, octree partitioning is performed on the non-empty sub-cubes (containing points from the point cloud) within the bounding box. At the same octree depth, a node is divided into 8 child nodes, continuing until the resulting leaf nodes are 1x1x1 unit cubes. The 8-bit binary code generated by determining whether a point in the sub-cube is occupied (1 for occupied, 0 for unoccupied) is called the occupancy code. The occupancy code of each node is encoded to generate a binary code stream.
[0071] The aforementioned octree is a tree-like data structure that, in three-dimensional spatial partitioning, uniformly divides a predefined bounding box, with each node having eight child nodes. By using "1" and "0" to indicate whether each child node of the octree is occupied, occupancy code information is obtained as the bitstream of point cloud geometric information.
[0072] Geometric encoding based on prediction trees: First, the input point cloud is sorted. Currently, sorting methods include unordered, Morton order, azimuth order, and radial distance order. At the encoding end, the prediction tree structure is built using two different methods: a high-latency slow mode (KD-Tree) and a low-latency fast mode (using LiDAR calibration information to assign each point to different LiDAR rays and build prediction structures according to the different rays). Next, based on the prediction tree structure, each node in the prediction tree is traversed. By selecting different prediction modes, the geometric position information of the node is predicted to obtain the prediction residual, and the geometric prediction residual is quantized using quantization parameters. Finally, through continuous iteration, the prediction residual of the prediction tree node position information, the prediction tree structure, and the quantization parameters are encoded to generate a binary code stream.
[0073] Geometric encoding based on triangulation: First, an octree is partitioned. Unlike geometric information encoding based on octree structures, this method does not need to partition the point cloud down to the bottom-level leaf nodes with side length 1·1·1, but instead partitions leaf nodes with specified side lengths. Then, the surface information composed of voxels within the nodes is represented by a series of triangular meshes. In GPCC, the parameter trisoup node size represents the size of the block containing the triangular facet. When trisoup node size is greater than 0, a geometric facet represents the set of voxels within the node. The maximum of twelve intersection points generated by the geometric facet and the twelve edges of the block are called vertices. The vertex coordinates of each block are encoded sequentially to generate a binary bitstream.
[0074] 3. Geometric Entropy Encoding: This method uses statistical compression encoding on the occupancy code information of the octree, the prediction residual information of the prediction tree, and the vertex information of the triangular representation, finally outputting a binary (0 or 1) compressed bitstream. Statistical coding is a lossless coding method that can effectively reduce the bit rate required to represent the same signal. A commonly used statistical coding method is Content Adaptive Binary Arithmetic Coding (CABAC).
[0075] 4. Geometric Reconstruction: Decoding and reconstructing the geometric information after geometric encoding.
[0076] For the attribute information encoding process, the encoding flow executed by encoder 200 is as follows:
[0077] 1. Color Transformation: Apply transformations to change the color information of an attribute to a different domain. For example, color information can be transformed from the RGB color space to the YCbCr color space.
[0078] 2. Attribute Recoloring: In lossy encoding, after encoding the geometric coordinate information, the encoding end needs to decode and reconstruct the geometric information, that is, restore the geometric information of each point in the point cloud. Attribute information corresponding to one or more neighboring points in the original point cloud is used as the attribute information for the reconstructed point.
[0079] In some examples, encoder 200 may not perform color transformation or attribute recoloring.
[0080] 3. Attribute information processing: In AVS-PCC, attribute information processing can include three modes: prediction coding, transformation coding, and prediction & transformation coding. These three coding modes can be used under different conditions.
[0081] Predictive coding refers to determining the neighboring points of the point to be coded as prediction points among the already coded points based on information such as distance or spatial relationships. Based on set criteria, the predicted attribute information of the point to be coded is calculated according to the attribute information of the prediction points. The difference between the actual attribute information and the predicted attribute information of the point to be coded is calculated as attribute residual information. This attribute residual information is then quantized, transformed (optional), and entropy encoded.
[0082] Transform coding refers to using transformation methods such as Discrete Cosine Transform (DCT) and Haar Transform (Haar) to group and transform attribute information, quantize the transformation coefficients, obtain attribute reconstruction information through inverse quantization and inverse transformation, calculate the difference between the real attribute information and the attribute reconstruction information to obtain attribute residual information and quantize it, and then entropy-encode the quantized transformation coefficients and attribute residuals.
[0083] Predictive transform coding refers to using the attribute residual information obtained from prediction to perform transformation, and then quantizing and entropy coding the transform coefficients.
[0084] In MPEG G-PCC, attribute information processing can include three modes: Prediction Transform coding, Lifting Transform coding, and Region Adaptive Hierarchical Transform (RAHT) coding. These three coding modes can be used under different conditions.
[0085] Predictive transform coding refers to dividing the point cloud into multiple different levels of detail (LoD) based on distance-selected subsets of points, achieving a multi-quality, hierarchical point cloud representation from coarse to fine. Bottom-up prediction is possible between adjacent layers, where neighboring points in the coarse layer predict the attribute information of points introduced in the fine layer, obtaining the corresponding attribute residual information. The points at the lowest level are encoded as reference information.
[0086] Lift transform coding refers to introducing a weight update strategy for neighboring points on the basis of LoD neighboring layer prediction, and finally obtaining the predicted attribute information of each point and the corresponding attribute residual information.
[0087] Hierarchical region adaptive transform coding refers to the process of transforming attribute information into the transform domain, which is called the transform coefficient.
[0088] 4. Attribute Quantization: The fineness of quantization is usually determined by the quantization parameters. The transformation coefficients or attribute residuals obtained from attribute information processing are quantized, and the quantized results are entropy-coded. For example, in predictive transform coding and boost transform coding, entropy coding is performed on the quantized attribute residuals; in RAHT, entropy coding is performed on the quantized transform coefficients.
[0089] 5. Entropy Coding: The quantized attribute residual information and / or transform coefficients are generally compressed using run-length coding and arithmetic coding. The corresponding coding mode, quantization parameters, and other information are also encoded using an entropy encoder.
[0090] The encoder 200 encodes the geometric coordinate information of each point in the point cloud to obtain a geometric bitstream, and encodes the attribute information of each point in the point cloud to obtain an attribute bitstream. The encoder 200 can transmit the encoded geometric bitstream and attribute bitstream together to the decoder 300.
[0091] Figure 4 The following is a flowchart illustrating the decoding process performed by the decoder in the AVS-PCC-based decoding framework: Figure 5 This diagram illustrates the decoding flowchart performed by the decoder within the MPEG G-PCC-based decoding framework. The decoder can be... Figure 1 The decoder 300 is shown. After receiving the compressed bitstream (i.e., attribute bitstream and geometric bitstream) transmitted by the encoder 200, the decoder 300 decodes the geometric bitstream to reconstruct the geometric coordinate information of each point in the point cloud, and decodes the attribute bitstream to reconstruct the attribute information of each point in the point cloud.
[0092] The decoding process performed by decoder 300 is as follows:
[0093] 1. Entropy Decoding: Perform entropy decoding on the geometric bitstream and attribute bitstream respectively to obtain geometric syntax elements and attribute syntax elements.
[0094] 2. Geometric Decoding: For the AVS-PCC coding framework, geometric decoding includes two modes: octree-based geometric decoding and prediction tree-based geometric decoding. For the G-PCC coding framework, geometric decoding includes three modes: octree-based geometric decoding, trisoup-based geometric decoding, and prediction tree-based prediction decoding.
[0095] Octree-based geometric decoding: Following the breadth-first traversal order, the placeholder code of each node is obtained by continuously parsing, and the nodes are continuously divided until a 1x1x1 unit cube is obtained. The division stops when the division is stopped. The number of points contained in each leaf node is obtained by parsing, and finally the geometric reconstruction point cloud information is recovered.
[0096] Geometric Decoding Based on Prediction Tree: The decoder continuously parses the bitstream to reconstruct the prediction tree structure. Then, it obtains the geometric position prediction residual information and quantization parameters of each prediction node through parsing. The prediction residual is then dequantized to recover the reconstructed geometric position information of each node, and finally, the geometric reconstruction of the decoder is completed.
[0097] Geometric Decoding Based on Triangular Representation: To decode the geometric coordinates of a point cloud from a node's triangular facet, it is necessary to check whether each voxel within the node cube intersects with the triangular facet. This technique is called triangulation. It uses six unit vectors (0,0,1), (0,0,1), (0,0,1), (0,0,1), (0,0,1), (0,0,1), (0,0,1) to perform an intersection check. If each unit vector intersects with the triangular facet, the intersection point is calculated and the decoded cube is output. The number of points generated in the decoder is determined by the grid distance d.
[0098] 3. Geometric Reconstruction: Perform reconstruction to obtain the geometric coordinate information of the points in the point cloud.
[0099] 4. Inverse coordinate transformation: Perform an inverse transformation on the reconstructed geometric coordinate information to convert the reconstructed coordinates (positions) of points in the point cloud from the transformation domain back to the initial domain.
[0100] 5. Dequantization: Dequantizes attribute syntax elements.
[0101] 6. Attribute Information Processing: In AVS-PCC, attribute information processing determines the color information of points in the point cloud by predicting or predicting the transformation of the inverse-quantized prediction residual or prediction residual transformation coefficients, or by transforming the transformation coefficients of the inverse-quantized transformation.
[0102] In MPEG G-PCC, attribute information processing determines the color information of points in the point cloud by using RAHT to invert the attribute information, or by using LOD and inverse boosting to determine the color information of points in the point cloud.
[0103] 7. Inverse Color Transformation: Transforms color information from the YCbCr color space to the RGB color space. In some examples, the inverse color transformation operation may not be necessary.
[0104] The following information relates to this application.
[0105] (a) Use of inter-frame prediction information;
[0106] like Figure 6 As shown, the predicted node placeholder code can be obtained directly from the reference frame point cloud or from the compensated point cloud, depending on whether the current node has undergone motion compensation. Based on the occupancy status of the predicted nodes, the inter-frame prediction information is divided into the following categories:
[0107] 1) No pred: When the predated node placeholder code is zero (bP=0), that is, none of its child nodes occupy the placeholder, the inter-frame prediction information is not used.
[0108] 2) Pred0: When the predicted child node i is empty, the predicted child node i is not to occupy bP. i =0.
[0109] Pred1: When the predicted child node i is not empty, the predicted child node i is to occupy bP. i =1; at this point, based on the number of points contained in the node, there are two further cases:
[0110] Case 1: predL=1: When the predicted child node i is not empty and the number of points in it exceeds the threshold th, the child node i is strongly predicted to be occupied.
[0111] Case 2: predL = 0: When the predicted child node i is not empty and the number of points in it does not exceed the threshold th, the child node i is not strongly predicted.
[0112] Optionally, this threshold is set to 2 in TMC13 v23 and GES.
[0113] (ii) Local motion estimation;
[0114] For non-radar dense point clouds, G-PCC only performs local motion estimation. The local motion enable flag (localMotionEnabled) of the geometric point cloud layer (GBS layer) determines whether local motion estimation is enabled for a certain layer. Local motion estimation is based on block (prediction unit) inter-frame prediction. First, the size of the LPU and the number of layers for block prediction are read from the configuration parameters, and the size of the minimum prediction unit (min LPUsize) is calculated.
[0115] a) When the current layer node size is greater than the LPU size, there are no motion vectors to perform motion compensation on the reference point cloud. Therefore, the occupancy information of the reference point cloud (without motion compensation) is directly used as the inter-frame prediction context.
[0116] b) When the current layer node size equals LPUsize, first determine if the number of points in the prediction block is greater than 50 to decide whether to enable local motion, and then write the recursive prediction unit structure (PU_tree). Each node can continue to partition downwards and use the motion vectors of its child nodes' PUs to perform motion compensation on the reference point cloud, or directly use the motion vectors of the current node that has not been partitioned to perform motion compensation on the reference point cloud. The PU_tree records the flag for whether to partition downwards (split_flag), the flag for whether to perform compensation in the current layer (isCompensated), and the set of motion vectors (MVs); if a node partitions to minLPUsize, further partitioning is terminated (split_flag == 0), and motion compensation is performed (isCompensated == 1). Finally, based on the flag for whether to compensate, it is decided whether to use the reference point cloud or the compensation point cloud occupancy information as the inter-frame prediction context.
[0117] (III) Motion prediction process based on prediction unit or prediction block (PU);
[0118] A PU contains the following parameters:
[0119] popul_flags: PU occupancy status;
[0120] split_flags: Flags for splitting downwards;
[0121] MVs: Motion Vectors;
[0122] isCompensated: If it is 1, it means that the reference point cloud has been motion compensated; if it is 0, it means that the reference point cloud has not been compensated.
[0123] hasMotion: This indicates whether the node contains motion information. It is 1 if it contains motion information, and 0 otherwise.
[0124] Please refer to the PU-based motion prediction process. Figure 7 .
[0125] (iv) Determination of motion vectors and motion compensation based on rate distortion optimization.
[0126] 1. Optimal motion vector selection and encoding;
[0127] 1.1) Motion Vector Search (MV) (How to obtain the motion vector MV);
[0128] i. Motion estimation criterion: The matching metric is log() of the sum of the absolute values of the differences between each point in the predicted block and the block to be encoded (Manhattan distance);
[0129] D(B,P)=∑ b∈B log2(1+min p∈P ||bp||1);
[0130] Where B represents the block to be encoded, P represents the predicted block, D(B,P) represents the distortion measure between the predicted block and the block to be encoded, b represents a point in the block to be encoded, p represents a point in the predicted block, and l represents the 1 norm.
[0131] ii. Search algorithm: Within the search window, starting from the location of the reference node, search for the two best motion vectors in the surrounding 18 directions. By changing the search step size, the search distance is continuously reduced until the best motion vector is finally obtained.
[0132] 1.2) Calculate the bitrate of each encoded MV:
[0133] Set the context based on the value of MV: whether the motion vector is 0 (mvIsZero), whether the motion vector is 1 (mvIsOne), the sign bit of the motion vector (mvSign), the context index of the exponentially Golomb encoded motion vector (ctxLocalMV), and calculate the entropy of the encoded MV;
[0134] The optimal MV determination is associated with whether or not the downsegmentation flag is used.
[0135] For details on encoding motion vectors and using prediction information, please refer to [link / reference]. Figure 8 .
[0136] (v) Selecting whether to partition the PU downwards based on Rate Distortion Optimization (RDO).
[0137] Whether to split down (split_flag) is determined based on the total cost of the distortion of the reference node and the current node, the encoded MV, and the encoded split_flag (if it is 0, the reference node is not compensated; otherwise, the reference node is compensated).
[0138] Cost calculation process:
[0139] The flag indicating whether to split downwards and the different PU motion vectors (MVs) are both optional coding parameters. Using a specific set of coding parameters (e.g., when split_flag is not split downwards and motion vector MV1 is selected), the bit rate and distortion under that condition can be obtained, i.e., the rate-distortion performance (R, D). In order to find the coding parameters that minimize distortion (D) while satisfying a certain bit rate limit (R), the Lagrange factor is introduced to calculate the following cost:
[0140] C=∑i[D(B,P(W,Vi))+λR(Vi)]+λR(split flags)+λR(pop flags);
[0141] Where C represents the rate-distortion cost, B represents the block to be encoded, P represents the predicted block, R represents the size of the encoded bitstream, λ represents the Lagrange factor between D and R, W represents the search window size, Vi represents the motion translation vector, and pop flags represent the occupancy flags of the sub-blocks after partitioning.
[0142] The optimal motion vector and the split_flag flag for the current PU block are determined by comparing the costs of splitting downwards or not.
[0143] The coding method provided in this application will be described in detail below with reference to the accompanying drawings, through some embodiments and application scenarios.
[0144] like Figure 9 As shown, this application provides an encoding method, including:
[0145] Step 901: If the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold, the encoding end copies the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block.
[0146] Optionally, the aforementioned reference point cloud frame is the preceding point cloud frame adjacent to the target point cloud frame.
[0147] The aforementioned point cloud information includes at least one of the geometric encoding information and attribute encoding information of the point cloud corresponding to the second point cloud block.
[0148] The method of copying the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block can also be described as an inter-frame prediction skip mode.
[0149] By copying the geometric encoding information of the second point cloud block or the second point cloud block after motion compensation into the reconstructed point cloud of the first point cloud block, it is convenient to directly use the geometric encoding information to encode the attribute information of the first point cloud block in the future, omitting the process of encoding the geometric information of the first point cloud block and saving encoding resources.
[0150] By copying the attribute encoding information of the second point cloud block or the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud block, the attribute encoding information of the first point cloud block can be obtained directly using the attribute encoding information of the second point cloud block without having to encode the attribute information of the first point cloud block again. This eliminates the need to encode the attribute information of the first point cloud block, thus saving encoding resources.
[0151] Step 902: The encoding end generates a target bitstream based on the first identification information, wherein the first identification information is used to instruct the decoding end to copy the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block to the reconstructed point cloud of the first point cloud block, and the point cloud information includes at least one of the geometric encoding information and attribute encoding information of the point cloud corresponding to the second point cloud block.
[0152] In this embodiment, the first identifier information can be represented as PU_copy_flag, for example. This first identifier information is used to identify whether the target point cloud frame uses copy mode for predictive coding. For example, when the value of the first identifier information is set to 1, it indicates that the target point cloud frame uses copy mode for predictive coding, that is, the encoder copies the point cloud information of the second point cloud block or the point cloud information of the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud. Alternatively, when the value of the first identifier information is set to 0, it indicates that the target point cloud frame does not use copy mode for predictive coding, that is, the encoder does not copy the point cloud information of the second point cloud block or the point cloud information of the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud. The purpose of motion compensation is to make the second point cloud block of the reference point cloud frame more closely matched with the first point cloud block of the target point cloud frame, thereby improving the coding quality.
[0153] In the above-described scheme of this application embodiment, when the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold, the encoding end copies the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block. The encoding end generates a target bitstream based on the first identification information, which is used to instruct the decoding end to copy the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block. Through the above scheme, when the similarity between two point cloud blocks is sufficiently high, the point cloud information of one point cloud block is directly copied into the reconstructed point cloud of the other point cloud block, saving the process of encoding the point cloud of the other point cloud block. This can save encoding resources and encoding time while ensuring that the point cloud quality does not fluctuate significantly, effectively improving encoding efficiency.
[0154] Optionally, when the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold, the encoding end copies the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block, including:
[0155] The encoding end obtains the tree data structure corresponding to the target point cloud frame;
[0156] The encoding end obtains the size of the node to be encoded in the target level of the tree data structure;
[0157] If the preset conditions are met and the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold, the encoding end copies the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block. The first point cloud block is the point cloud block corresponding to the node to be encoded in the target level. The inter-frame prediction skip mode refers to the mode of not encoding the point cloud information corresponding to the point cloud frame.
[0158] The preset conditions include one of the following:
[0159] The second identification information indicates that the target point cloud frame enables inter-frame prediction skip mode;
[0160] The second identification information indicates that the target point cloud frame has enabled the inter-frame prediction skip mode, and that the size of the node to be encoded is within a preset range.
[0161] As one implementation, the aforementioned tree-like data structure is an octree. The encoding end performs octree partitioning and placeholder code encoding on the target point cloud frame to obtain the octree corresponding to the target point cloud frame.
[0162] The target level mentioned above is a level in the tree structure mentioned above, for example, an octree level in an octree.
[0163] The size of the node to be encoded can refer to the length of the side corresponding to the node to be encoded.
[0164] The aforementioned second identifier information can be represented by gps.Skip_mode_flag on. Optionally, when the value corresponding to the second identifier information is set to 1, it indicates that the target point cloud frame has the inter-frame prediction skip mode enabled, and when the value corresponding to the second identifier information is set to 0, it indicates that the target point cloud frame does not have the inter-frame prediction skip mode enabled.
[0165] For example, the aforementioned preset range includes the size values of the first node and the second node. For instance, the first node size is represented by the maximum value of the copy mode (max_size_CopyPU), and the second node size is represented by the minimum value of the copy mode (min_size_CopyPU). The size of the node to be encoded can be represented by the current node value (CurNode_size), and the size of the node to be encoded falling within the preset range can be expressed as CurNode_size ≤ max_size_CopyPU && CurNode_size ≥ min_size_CopyPU.
[0166] Since not all image frames have nodes suitable for encoding using the inter-frame prediction skip mode, setting the second identifier information allows for the selection of suitable image frames. Furthermore, setting the preset range allows for the selection of appropriately sized nodes for encoding using the inter-frame prediction skip mode. These second identifier information and preset range enhance the flexibility in selecting nodes for encoding using the inter-frame prediction skip mode.
[0167] Optionally, the encoding end generates the target bitstream based on the first identification information, including:
[0168] The encoding end generates the target bitstream based on the first identification information and the second identification information;
[0169] Alternatively, the target bitstream can be generated based on the first identification information, the second identification information, and the preset range.
[0170] Here, when the point cloud information of the second point cloud block or the point cloud information of the second point cloud block after motion compensation is copied to the reconstructed point cloud of the first point cloud block at the encoding end, the first identification information, the second identification information (and the preset range) are encoded to obtain the encoded stream corresponding to the target point cloud frame. The process of encoding the point cloud information of the first point cloud block is skipped, saving encoding resources and encoding time.
[0171] Optionally, the encoding end obtains the size of the node to be encoded in the target level of the tree data structure, including:
[0172] When the third identification information indicates that the inter-frame prediction skip mode is enabled in the point cloud frame sequence, the encoding end obtains the size of the node to be encoded in the target level of the tree data structure, and the point cloud frame sequence includes the target point cloud frame.
[0173] In this embodiment of the application, the aforementioned third identification information can be represented by Sps.Skip_mode_flag on. Optionally, when the value of the third identification information is set to 1, it indicates that the point cloud frame sequence enables the inter-frame prediction skip mode; when the value of the third identification information is set to 0, it indicates that the point cloud frame sequence does not enable the inter-frame prediction skip mode.
[0174] Here, by setting the aforementioned third identifier information, point cloud frame sequences that use the inter-frame prediction skip mode can be filtered out. In other words, this third identifier information provides flexibility in selecting point cloud frame sequences that use the inter-frame prediction skip mode.
[0175] Optionally, the encoding end generates the target bitstream based on the first identification information, including:
[0176] The encoding end generates the target bitstream based on the first identifier information, the second identifier information, and the third identifier information, or generates the target bitstream based on the first identifier information, the second identifier information, the third identifier information, and the preset range.
[0177] Here, when the point cloud information of the second point cloud block or the point cloud information of the second point cloud block after motion compensation is copied to the reconstructed point cloud of the first point cloud block at the encoding end, the first identification information, the second identification information, the third identification information (and the preset range) are encoded to obtain the encoded stream corresponding to the target point cloud frame. The process of encoding the point cloud information of the first point cloud block is skipped, saving encoding resources and encoding time.
[0178] Optionally, the method in this application embodiment further includes:
[0179] Obtain the distortion rate of the first point cloud block relative to the second point cloud block or the second point cloud block after motion compensation;
[0180] Based on the distortion rate, the similarity between the first point cloud block and the second point cloud block, or between the first point cloud block and the second point cloud block after motion compensation, is obtained.
[0181] The aforementioned distortion rate is inversely proportional to the similarity; that is, the smaller the distortion rate, the higher the similarity.
[0182] In this embodiment, when the similarity between two point cloud blocks is high (i.e., the difference is small), the distortion-rate cost of using or not using the inter-frame prediction skip mode can be evaluated through rate-distortion cost. For example, if the distortion-rate cost of using or not using the inter-frame prediction skip mode is high, the value corresponding to the first identifier information can be set to 0, indicating that the inter-frame prediction skip mode is not used. In this case, the prediction entropy coding mode is used to encode the target point cloud frame layer by layer. If the rate-distortion cost of using the inter-frame prediction skip mode is low, the value corresponding to the first identifier information can be set to 1, indicating that the inter-frame prediction skip mode is used. In this case, no encoding is required below the PU layer node of the target point cloud frame up to the leaf node layer, and the point cloud information of the second point cloud block or the motion-compensated second point cloud block is directly copied into the corresponding reconstructed point cloud.
[0183] In addition, in this embodiment, the similarity between the first and second point cloud blocks can be obtained based on the difference between the centroid offset of the first point cloud block and the centroid offset of the second point cloud block. Since the centroid is calculated based on the distribution of points in the point cloud block, the smaller the difference between the centroid offsets of the first and second point cloud blocks, the more similar the distribution of points in the two point cloud blocks are, and thus the higher the similarity between the first and second point cloud blocks.
[0184] Of course, in addition to obtaining the similarity of two point cloud blocks based on the above-mentioned distortion rate and centroid offset, other methods can also be used to obtain the similarity of two point cloud blocks in the embodiments of this application, and no specific limitation is made here.
[0185] In addition, in scenarios with high encoding latency requirements during real-time dynamic point cloud transmission, the aforementioned inter-frame prediction skipping mode can also be adopted. The above prediction entropy coding mode is a general inter-frame prediction coding mode. When processing the nodes to be encoded in the current point cloud frame (hereinafter referred to as nodes), different processing is performed based on the size relationship between the node and the LPU: when the current point cloud frame node size > the LPU node size, since there are no available motion vectors for motion compensation of the reference frame point cloud, the reference point cloud is directly used as prediction information to encode the current frame; when the current frame node size ≤ the LPU node size, it is divided according to the PU mode. First, it is determined whether the PU needs to be further divided. If the PU stops dividing downwards, the motion-compensated point cloud is used as prediction information to encode the current frame. Conversely, if the PU continues to divide downwards, the reference point cloud is directly used as prediction information to encode the current frame. Finally, the constructed inter-frame context is merged with the intra-frame context information of the current frame node, and the merged context is input into the entropy encoder to encode the point cloud of the current point cloud frame.
[0186] Optionally, the method in this application embodiment further includes:
[0187] When the size of the node to be encoded in the target level of the tree data structure corresponding to the target point cloud frame is the same as the size of the node to be encoded corresponding to the maximum prediction unit (LPU) and the reference point cloud frame satisfies the local motion estimation enabling condition, motion estimation is performed on the first point cloud block to determine the second point cloud block in the reference point cloud frame that matches the first point cloud block and to determine the motion vector of the first point cloud block relative to the second point cloud block.
[0188] Motion compensation is performed on the second point cloud block based on the motion vector to obtain the motion-compensated second point cloud block.
[0189] In this embodiment of the application, motion compensation is performed on the second point cloud block to make the second point cloud block and the first point cloud block more closely matched, thereby making the point cloud information of the second point cloud block more closely matched with that of the first point cloud block, and thus effectively improving the coding quality of the point cloud.
[0190] The following is combined Figure 10 and Figure 11 The encoding method of the embodiments of this application will be described as follows.
[0191] In this embodiment, the target point cloud frame is first divided into an octree and encoded with placeholder codes to obtain the octree corresponding to the target point cloud frame. When performing PU-based inter-frame prediction coding based on this octree, the following steps can be followed:
[0192] Step 1: Determine whether Sps.Skip_mode_flag on is 1. If it is 1, it means that the point cloud frame sequence enables the inter-frame prediction skip mode. If it is 0, it means that the point cloud frame sequence disables the inter-frame prediction skip mode.
[0193] Step 2: When Sps.Skip_mode_flag on is 1, partition the current point cloud frame into slices.
[0194] These Step 1 and Step 2 are optional steps.
[0195] Step 3: When Sps.Skip_mode_flag on is 1, based on each slice of the point cloud, read all the nodes of each level (0 ≤ depth < maxDepth) of the current slice from the First Input First Output (FIFO) queue of the octree data structure. Here, maxDepth is the maximum depth (the number of layers from the root node to the leaf node) under the octree partition of the current frame of the point cloud.
[0196] Step 4: Determine whether gps.Skip_mode_flag on is 1. If it is, it means that the target point cloud frame (such as the current point cloud frame) enables the inter-frame prediction skip mode. If it is 0, it means that the target point cloud frame disables the inter-frame prediction skip mode.
[0197] If gps.Skip_mode_flag on is 1, determine whether the size of the current node to be encoded meets the Skip_size qualification condition. That is, only the nodes that satisfy CurNode_size ≤ max_size_CopyPU && CurNode_size ≥ min_size_CopyPU are allowed to use the inter-frame prediction skip (SKip) mode. When encoding such nodes, the encoder will evaluate the distortion-rate cost brought by using or not using this mode through rate-distortion cost. If the rate-distortion cost of using the copy mode for the current node is large, mark that the current node does not use the copy mode (PU_copy_flag = 0), and at this time, use the predictive entropy coding mode to encode the current frame node layer; if the rate-distortion cost of using the copy mode for the current node is small, mark that the current node uses the inter-frame prediction skip mode (PU_copy_flag = 1), and at this time, the current node does not need to be encoded, and the reference block or the compensated reference block is directly copied to the reconstructed point cloud.
[0198] If gps.Skip_mode_flag on is 0, at this time, use the predictive entropy coding mode to encode the current frame node layer.
[0199] In addition, as Figure 11As shown, for nodes in the depth layer, the following judgment is also performed: the node size reaches the LPU size and the reference point cloud frame meets the local motion estimation enabling condition (e.g., the number of points in the reference point cloud frame is greater than 50). If the condition is met, then for each PU, we can try to partition it or not, and perform motion estimation for each PU to determine the matching position of the current PU in the reference frame and the corresponding motion vector (e.g., ...). Figure 12 As shown, the optimal motion vector is found under different PU partitioning modes for motion compensation. Each PU includes the current layer PU and all sub-PUs that may be iteratively partitioned. Inter-frame prediction is performed for each PU. Based on the best-matching motion vector, various prediction modes are tried, such as inter-frame skip mode and prediction entropy coding mode, to predict the current PU. Rate-distortion optimization techniques are used to select the optimal PU partitioning mode, motion vector, and prediction coding mode. The optimization goal is to minimize coding distortion while maintaining an appropriate bit rate. For each PU, different partitioning modes, motion vectors, and prediction coding modes are tried, and the resulting distortion and bit rate are calculated. Then, by comparing the distortion-bit rate trade-offs of different options, the PU partitioning mode, motion vector, and prediction coding mode with the best performance are selected. Finally, based on the selected PU partitioning mode, motion vector, and prediction coding mode, the optimal PU partitioning, motion compensation, and prediction coding are performed.
[0200] Step 5: Finally, determine whether the size of the nodes in the current depth layer has reached the size of the nodes for trisoup encoding. If so, perform trisoup encoding; otherwise, return to step 1 and loop through the nodes until all nodes are encoded.
[0201] Using the current ges v3.0_rc1 as a reference or anchor, the geometric compression gain of the scheme in this application is shown in Table 1. Here, Cat2-A, Cat2-B, and Cat2-C are sequence names; C2 represents the point cloud encoding test condition; Luma, Chroma Cb, and Chroma Cr represent the three color components; and D1 and D2 represent two different point cloud geometric quality distortion evaluation parameters. Negative values for D1 and D2 indicate that, under test condition C2, encoding the point cloud using the scheme in this application produces a higher overall point cloud quality than existing technologies.
[0202] Table 1
[0203]
[0204] In the above-described scheme of this application embodiment, when the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold, the encoding end copies the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block. The encoding end generates a target bitstream based on the first identification information, which is used to instruct the decoding end to copy the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block. Through the above scheme, when the similarity between two point cloud blocks is sufficiently high, the point cloud information of one point cloud block is directly copied into the reconstructed point cloud of the other point cloud block, saving the process of encoding the point cloud of the other point cloud block. This can save encoding resources and encoding time while ensuring that the point cloud quality does not fluctuate significantly, effectively improving encoding efficiency.
[0205] like Figure 13 As shown in the embodiments of this application, a decoding method is also provided, including:
[0206] Step 1301: The decoding end decodes the target bitstream to obtain the first identification information.
[0207] The first identification information has been described in detail in the method embodiment at the encoding end, and will not be repeated here.
[0208] Step 1302: When the first identification information instructs the encoding end to copy the point cloud information of the second point cloud block of the reference point cloud frame or the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud block of the target point cloud frame, the decoding end copies the point cloud information of the second point cloud block or the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud block. The point cloud information includes at least one of the geometric coding information and attribute coding information of the point cloud corresponding to the second point cloud block.
[0209] Step 1302: When the first identification information instructs the encoding end to copy the point cloud information of the second point cloud block of the reference point cloud frame or the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud block of the target point cloud frame, the decoding end copies the point cloud information of the second point cloud block or the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud block. The point cloud information includes at least one of the geometric coding information and attribute coding information of the point cloud corresponding to the second point cloud block.
[0210] In this embodiment, the decoding end decodes the target bitstream to obtain first identification information. When the first identification information instructs the encoding end to copy the point cloud information of the second point cloud block of the reference point cloud frame or the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud block of the target point cloud frame, the decoding end copies the point cloud information of the second point cloud block or the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud block. That is, the point cloud information of the first point cloud block and the second point cloud block are the same. When decoding, the decoding end only needs to decode the point cloud information of the second point cloud block to obtain the decoding information of the first point cloud block and the second point cloud block, which saves decoding resources and decoding time and improves decoding efficiency.
[0211] Optionally, the decoding end decodes the target bitstream to obtain first identification information, including:
[0212] The decoding end decodes the first bitstream in the target bitstream to obtain the second identification information;
[0213] The decoding end obtains the size of the node to be encoded in the target level of the tree data structure corresponding to the target point cloud frame;
[0214] When the second identification information indicates that the target point cloud frame has enabled the inter-frame prediction skip mode, and the size of the node to be encoded is within the preset range, the second bitstream in the target bitstream is decoded to obtain the first identification information. The inter-frame prediction skip mode refers to a mode in which the point cloud information corresponding to the point cloud frame is not encoded.
[0215] Here, the decoding end can obtain the preset range from the first bitstream mentioned above, or the decoding end can directly obtain the preset range that has been pre-configured or pre-set.
[0216] Optionally, the first bitstream may further include third identification information;
[0217] Obtaining the size of the node to be encoded in the target level of the tree data structure includes:
[0218] When the third identification information indicates that the inter-frame prediction skip mode is enabled in the point cloud frame sequence, the size of the node to be encoded in the target level of the tree data structure is obtained, and the point cloud frame sequence includes the target point cloud frame.
[0219] Optionally, the method further includes:
[0220] When the size of the node to be encoded in the target level is the same as the size of the node corresponding to the maximum prediction unit (LPU) and the reference point cloud frame meets the local motion estimation start condition, obtain the second point cloud block in the reference point cloud frame that matches the first point cloud block according to the target bitstream, and obtain the motion vector of the first point cloud block relative to the second point cloud block;
[0221] Perform motion compensation on the second point cloud block according to the motion vector to obtain the motion-compensated second point cloud block.
[0222] In the embodiments of the present application, when performing inter-frame prediction decoding based on a prediction unit (PU) according to octree partitioning, the decoding end can operate according to the following steps:
[0223] Step 1: Parse whether the current Sps.Skip_mode_flag on in the frame-level syntax element is 1; if it is 1, it means that the inter-frame prediction skip mode is enabled for the point cloud frame sequence, and if it is 0, it means that the inter-frame prediction skip mode is disabled for the point cloud frame sequence.
[0224] Step 2: Based on the slice point cloud, parse the occupancy information of all nodes at each level (0 ≤ depth < maxDepth) of the current slice. Here, maxDepth is the maximum depth (from the root node to the leaf node layer) under the octree partitioning of the current frame point cloud.
[0225] Step 3: Parse whether the current gps.Skip_mode_flag on is 1; if it is 1, it means that the target point cloud frame (such as the current point cloud frame) enables the inter-frame prediction skip mode, and if it is 0, it means that the target point cloud frame disables the inter-frame prediction skip mode.
[0226] If gps.Skip_mode_flag on is 1, determine whether the current node size meets the Skip_size eligibility condition. That is, only nodes that satisfy CurNode_size ≤ max_size_CopyPU && CurNode_size ≥ min_size_CopyPU are allowed to use the inter-frame prediction skip (SKip) mode. For nodes that meet this condition during decoding, parse whether the PU_split_flag of the current node is 0. If it is 0, at this time, the current frame nodes are encoded layer by layer using the prediction entropy coding mode; if it is 1, at this time, the current node does not need to be decoded, and the reference block or the compensated reference block is directly copied to the reconstructed point cloud.
[0227] If gps.Skip_mode_flag on is 0, at this time, the current frame node layer is decoded using the prediction entropy decoding mode.
[0228] Additionally, for nodes in the depth layer, the following checks are performed: the node size reaches the LPU size and the reference point cloud frame meets the local motion estimation enabling condition. If the condition is met, operations are performed for each PU. If the PU is less than or equal to the minimum PU size (sps_minPU_size), then PU_split_flag is inferred to be 0; otherwise, the PU_split_flag partitioning flag of the decoded PU layer is determined. If PU_split_flag is false, motion compensation is performed on the current layer PU; if PU_split_flag is true, motion compensation is performed on the nodes iteratively partitioned into PUs where PU_split_flag is false and the PU is greater than the minimum PU size. For PUs requiring motion compensation, the three direction values of the motion vector are decoded and used to perform motion compensation on the PU. The PU_copy_flag of the decoded PU layer is then used to determine the prediction mode. If PU_copy_flag is true, it indicates that the decoder has selected the copy mode, directly copying the reference point into the reconstructed point cloud, in which case no further decoding operations are required. If PU_copy_flag is false, then the current frame node needs to be decoded based on the inter-frame information decoded by PU and the intra-frame context.
[0229] Step 4: Finally, determine whether the size of the current depth layer node has reached the size of the node to be decoded by trisoup. If so, perform trisoup decoding; otherwise, return to step 1 and loop through the nodes until all nodes have been decoded.
[0230] It should be noted that the decoding method executed by the decoding end is the same as the decoding method executed by the encoding end mentioned above, which will not be elaborated here.
[0231] The encoding method provided in this application can be executed by an encoding device. This application uses an encoding device executing the encoding method as an example to illustrate the encoding device provided in this application.
[0232] like Figure 14 As shown, this application embodiment provides an encoding device 1400, including:
[0233] The first copying module 1401 is used to copy the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block when the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold.
[0234] The generation module 1402 is used to generate a target bitstream based on the first identification information, wherein the first identification information is used to instruct the decoding end to copy the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block to the reconstructed point cloud of the first point cloud block, and the point cloud information includes at least one of the geometric encoding information and attribute encoding information of the point cloud corresponding to the second point cloud block.
[0235] Optionally, the first replication module includes:
[0236] The first acquisition submodule is used to acquire the tree data structure corresponding to the target point cloud frame;
[0237] The second acquisition submodule is used to acquire the size of the node to be encoded in the target level of the tree data structure;
[0238] The copying submodule is used to copy the point cloud information of the second point cloud block or the motion-compensated point cloud information of the second point cloud block into the reconstructed point cloud of the first point cloud block when the preset conditions are met and the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold. The first point cloud block is the point cloud block corresponding to the node to be encoded in the target level. The inter-frame prediction skip mode refers to the mode of not encoding the point cloud information corresponding to the point cloud frame.
[0239] The preset conditions include one of the following:
[0240] The second identification information indicates that the target point cloud frame enables inter-frame prediction skip mode;
[0241] The second identification information indicates that the target point cloud frame has enabled the inter-frame prediction skip mode, and that the size of the node to be encoded is within a preset range.
[0242] Optionally, the generation module is used to generate the target bitstream based on the first identification information and the second identification information;
[0243] Alternatively, the target bitstream can be generated based on the first identification information, the second identification information, and the preset range.
[0244] Optionally, the second acquisition submodule is used to acquire the size of the node to be encoded in the target level of the tree data structure when the third identification information indicates that the inter-frame prediction skip mode of the point cloud frame sequence is enabled, wherein the point cloud frame sequence includes the target point cloud frame.
[0245] Optionally, the generation module is configured to generate the target bitstream based on the first identification information, the second identification information, and the third identification information, or to generate the target bitstream based on the first identification information, the second identification information, the third identification information, and the preset range.
[0246] Optionally, the apparatus in this application embodiment further includes:
[0247] The first acquisition module is used to acquire the distortion rate of the first point cloud block relative to the second point cloud block or the second point cloud block after motion compensation;
[0248] The second acquisition module is used to acquire the similarity between the first point cloud block and the second point cloud block or with the second point cloud block after motion compensation, based on the distortion rate.
[0249] Optionally, the apparatus in this application embodiment further includes:
[0250] The determination module is used to perform motion estimation on the first point cloud block when the size of the node to be encoded in the target level of the tree data structure corresponding to the target point cloud frame is the same as the size of the node to be encoded corresponding to the maximum prediction unit (LPU) and the reference point cloud frame satisfies the local motion estimation enabling condition, and to determine the second point cloud block in the reference point cloud frame that matches the first point cloud block and to determine the motion vector of the first point cloud block relative to the second point cloud block.
[0251] The third acquisition module is used to perform motion compensation on the second point cloud block according to the motion vector to obtain the motion-compensated second point cloud block.
[0252] In this embodiment, when the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold, the encoding end copies the point cloud information of the second point cloud block or the motion-compensated point cloud information of the second point cloud block into the reconstructed point cloud of the first point cloud block. The encoding end generates a target bitstream based on first identifier information, which instructs the decoding end to copy the point cloud information of the second point cloud block or the motion-compensated point cloud information of the second point cloud block into the reconstructed point cloud of the first point cloud block. Through this scheme, when the similarity between two point cloud blocks is sufficiently high, the point cloud information of one point cloud block is directly copied into the reconstructed point cloud of the other, saving the process of encoding the point cloud of the other point cloud block. This saves encoding resources and time while ensuring that the point cloud quality does not fluctuate significantly, effectively improving encoding efficiency.
[0253] The encoding device provided in this application embodiment can achieve... Figures 9 to 12The various processes implemented in the method embodiments achieve the same technical effect, and will not be described again here to avoid repetition.
[0254] like Figure 15 As shown, this application embodiment provides a decoding device 1500, including:
[0255] The fourth acquisition module 1501 is used to decode the target bitstream to obtain the first identification information;
[0256] The second copying module 1502 is used to copy the point cloud information of the second point cloud block or the motion-compensated second point cloud block of the reference point cloud frame to the reconstructed point cloud of the first point cloud block of the target point cloud frame when the first identification information instructs the encoding end to copy the point cloud information of the second point cloud block or the motion-compensated second point cloud block to the reconstructed point cloud of the first point cloud block of the target point cloud frame. The point cloud information includes at least one of the geometric encoding information and attribute encoding information of the point cloud corresponding to the second point cloud block.
[0257] Optionally, the fourth acquisition module includes:
[0258] The third acquisition submodule is used to decode the first bitstream in the target bitstream and obtain the second identification information;
[0259] The fourth acquisition submodule is used to acquire the size of the node to be encoded in the target level of the tree data structure corresponding to the target point cloud frame;
[0260] The fifth acquisition submodule is used to decode the second bitstream in the target bitstream to obtain the first identification information when the second identification information indicates that the target point cloud frame has enabled the inter-frame prediction skip mode and the size of the node to be encoded is within the preset range. The inter-frame prediction skip mode refers to a mode in which the point cloud information corresponding to the point cloud frame is not encoded.
[0261] Optionally, the first bitstream may further include third identification information;
[0262] The fourth acquisition submodule is used to acquire the size of the node to be encoded in the target level of the tree data structure when the third identification information indicates that the inter-frame prediction skip mode of the point cloud frame sequence is enabled. The point cloud frame sequence includes the target point cloud frame.
[0263] Optionally, the apparatus in this application embodiment further includes:
[0264] The fifth acquisition module is used to acquire, according to the target bitstream, a second point cloud block that matches the first point cloud block in the reference point cloud frame and acquire the motion vector of the first point cloud block relative to the second point cloud block, when the size of the node to be encoded in the target level is the same as the size of the node corresponding to the maximum prediction unit (LPU) and the reference point cloud frame satisfies the local motion estimation enabling condition.
[0265] The sixth acquisition module is used to perform motion compensation on the second point cloud block according to the motion vector to obtain the motion-compensated second point cloud block.
[0266] In this embodiment, the decoding end decodes the target bitstream to obtain first identification information. When the first identification information instructs the encoding end to copy the point cloud information of the second point cloud block of the reference point cloud frame or the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud block of the target point cloud frame, the decoding end copies the point cloud information of the second point cloud block or the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud block. That is, the point cloud information of the first point cloud block and the second point cloud block are the same. When decoding, the decoding end only needs to decode the point cloud information of the second point cloud block to obtain the decoding information of the first point cloud block and the second point cloud block, which saves decoding resources and decoding time and improves decoding efficiency.
[0267] like Figure 16 As shown, this application embodiment also provides an electronic device 1600, including a processor 1601 and a memory 1602. The memory 1602 stores a program or instructions that can run on the processor 1601. For example, when the electronic device 1600 is an encoding device, the program or instructions executed by the processor 1601 implement the various steps of the above-described encoding method embodiment and achieve the same technical effect. When the electronic device 1600 is a decoding device, the program or instructions executed by the processor 1601 implement the various steps of the above-described decoding method embodiment and achieve the same technical effect. To avoid repetition, this will not be described again here. Optionally, the memory 1602 may be... Figure 1 The processor 1601 can implement the memory 102 or memory 113 in the illustrated embodiment. Figure 1-3 The functions of the encoder 200 or decoder 300 in the illustrated embodiment.
[0268] This application also provides an electronic device, including: a memory configured to store video data; and a processing circuit configured to implement the steps of the encoding or decoding method embodiments described above. Optionally, the memory may be... Figure 1 The processing circuitry of memory 102 or memory 113 in the illustrated embodiment can implement... Figure 1-3 The functions of the encoder 200 or decoder 300 in the illustrated embodiment.
[0269] This application embodiment also provides an electronic device, including a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement, for example... Figure 9 or Figure 13 The steps in the method embodiment shown are illustrated. This device embodiment corresponds to the above method embodiment, and all implementation processes and methods of the above method embodiments can be applied to this terminal embodiment and achieve the same technical effect.
[0270] The aforementioned electronic devices can be terminals or other devices besides terminals, such as servers, network attached storage (NAS), etc.
[0271] The terminal can be a mobile phone, tablet computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile internet device (MID), augmented reality (AR), virtual reality (VR) device, mixed reality (MR) device, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipboard equipment, pedestrian user equipment (PUE), smart home (home devices with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), game console, personal computer (PC), ATM or self-service machine, etc. Wearable devices include: smartwatches, smart bracelets, smart earphones, smart glasses, smart jewelry (smart bracelets, smart chains, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. Among these, in-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units, etc. It should be noted that the embodiments in this application do not limit the specific type of terminal.
[0272] A server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server. A cloud server can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), or cloud computing services based on big data and artificial intelligence platforms.
[0273] For example, the aforementioned electronic devices may include, but are not limited to, those described above. Figure 1 The type of source device 100 or destination device 110 shown.
[0274] Taking electronic devices as terminals as an example, Figure 17 A schematic diagram of the hardware structure of a terminal to implement an embodiment of this application.
[0275] The terminal 1700 includes, but is not limited to, at least some of the following components: radio frequency unit 1701, network module 1702, audio output unit 1703, input unit 1704, sensor 1705, display unit 1706, user input unit 1707, interface unit 1708, memory 1709, and processor 1710.
[0276] Those skilled in the art will understand that the terminal 1700 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1710 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 17 The terminal structure shown does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0277] It should be understood that, in this embodiment, the input unit 1704 may include a graphics processing unit (GPU) 17041 and a microphone 17042. The GPU 17041 processes image data of still images or videos obtained by an image acquisition device (such as a camera) in video acquisition mode or image acquisition mode, or it may process the obtained point cloud data. The display unit 1706 may include a display panel 17061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1707 includes at least one of a touch panel 17071 and other input devices 17072. The touch panel 17071 is also called a touch screen. The touch panel 17071 may include a touch detection device and a touch controller. Other input devices 17072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, joysticks, etc., which will not be described in detail here.
[0278] In this embodiment, after receiving downlink data from the network-side device, the radio frequency unit 1701 can transmit it to the processor 1710 for processing; in addition, the radio frequency unit 1701 can send uplink data to the network-side device. Typically, the radio frequency unit 1701 includes, but is not limited to, antennas, amplifiers, transceivers, couplers, low-noise amplifiers, duplexers, etc.
[0279] The memory 1709 can be used to store software programs or instructions, as well as various data. The memory 1709 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1709 may include volatile memory or non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1709 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0280] Processor 1710 may include one or more processing units; optionally, processor 1710 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1710.
[0281] In one embodiment of this application, the processor 1710 is configured to, when the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold, copy the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block.
[0282] The encoding end generates a target bitstream based on the first identification information, wherein the first identification information is used to instruct the decoding end to copy the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block, and the point cloud information includes at least one of the geometric encoding information and attribute encoding information of the point cloud corresponding to the second point cloud block.
[0283] Optionally, the processor 1710 is also used for:
[0284] Obtain the tree data structure corresponding to the target point cloud frame;
[0285] Obtain the size of the node to be encoded in the target level of the tree data structure;
[0286] If the preset conditions are met and the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold, the encoding end copies the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block. The first point cloud block is the point cloud block corresponding to the node to be encoded in the target level. The inter-frame prediction skip mode refers to the mode of not encoding the point cloud information corresponding to the point cloud frame.
[0287] The preset conditions include one of the following:
[0288] The second identification information indicates that the target point cloud frame enables inter-frame prediction skip mode;
[0289] The second identification information indicates that the target point cloud frame has enabled the inter-frame prediction skip mode, and that the size of the node to be encoded is within a preset range.
[0290] Optionally, the processor 1710 is also used for:
[0291] The target bitstream is generated based on the first identification information and the second identification information;
[0292] Alternatively, the target bitstream can be generated based on the first identification information, the second identification information, and the preset range.
[0293] Optionally, the processor 1710 is also used for:
[0294] When the third identification information indicates that the inter-frame prediction skip mode is enabled in the point cloud frame sequence, the size of the node to be encoded in the target level of the tree data structure is obtained, and the point cloud frame sequence includes the target point cloud frame.
[0295] Optionally, the processor 1710 is also used for:
[0296] The target bitstream is generated based on the first identifier information, the second identifier information, and the third identifier information; or, the target bitstream is generated based on the first identifier information, the second identifier information, the third identifier information, and the preset range.
[0297] Optionally, the processor 1710 is also used for:
[0298] Obtain the distortion rate of the first point cloud block relative to the second point cloud block or the second point cloud block after motion compensation;
[0299] Based on the distortion rate, the similarity between the first point cloud block and the second point cloud block, or between the first point cloud block and the second point cloud block after motion compensation, is obtained.
[0300] Optionally, the processor 1710 is also used for:
[0301] In the target level of the tree data structure corresponding to the target point cloud frame, the size of the node to be encoded is the same as the size of the node to be encoded corresponding to the maximum prediction unit (LPU), and the reference point cloud frame satisfies the local motion estimation enabling condition, motion estimation is performed on the first point cloud block to determine the second point cloud block in the reference point cloud frame that matches the first point cloud block and to determine the motion vector of the first point cloud block relative to the second point cloud block.
[0302] Motion compensation is performed on the second point cloud block based on the motion vector to obtain the motion-compensated second point cloud block.
[0303] In this embodiment, when the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold, the encoding end copies the point cloud information of the second point cloud block or the motion-compensated point cloud information of the second point cloud block into the reconstructed point cloud of the first point cloud block. The encoding end generates a target bitstream based on first identifier information, which instructs the decoding end to copy the point cloud information of the second point cloud block or the motion-compensated point cloud information of the second point cloud block into the reconstructed point cloud of the first point cloud block. Through this scheme, when the similarity between two point cloud blocks is sufficiently high, the point cloud information of one point cloud block is directly copied into the reconstructed point cloud of the other, saving the process of encoding the point cloud of the other point cloud block. This saves encoding resources and time while ensuring that the point cloud quality does not fluctuate significantly, effectively improving encoding efficiency.
[0304] It is understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description of the coding method in the method embodiment and achieve the same or corresponding technical effect. To avoid repetition, it will not be described again here.
[0305] In one embodiment of this application, the processor 1710 is used for:
[0306] Decode the target bitstream to obtain the first identifier information;
[0307] When the first identification information instructs the encoding end to copy the point cloud information of the second point cloud block of the reference point cloud frame or the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud block of the target point cloud frame, the point cloud information of the second point cloud block or the second point cloud block after motion compensation is copied to the reconstructed point cloud of the first point cloud block. The point cloud information includes at least one of the geometric encoding information and attribute encoding information of the point cloud corresponding to the second point cloud block.
[0308] Optionally, the processor 1710 is also used for:
[0309] Decode the first bitstream in the target bitstream to obtain the second identifier information;
[0310] Obtain the size of the node to be encoded in the target level of the tree data structure;
[0311] When the second identification information indicates that the target point cloud frame has enabled the inter-frame prediction skip mode, and the size of the node to be encoded is within the preset range, the second bitstream in the target bitstream is decoded to obtain the first identification information. The inter-frame prediction skip mode refers to a mode in which the point cloud information corresponding to the point cloud frame is not encoded.
[0312] Optionally, the first bitstream may further include third identification information;
[0313] The processor 1710 is also used for:
[0314] When the third identification information indicates that the inter-frame prediction skip mode is enabled in the point cloud frame sequence, the size of the node to be encoded in the target level of the tree data structure is obtained, and the point cloud frame sequence includes the target point cloud frame.
[0315] Optionally, the processor 1710 is also used for:
[0316] When the size of the node to be encoded in the target level is the same as the size of the node corresponding to the maximum prediction unit (LPU) and the reference point cloud frame satisfies the local motion estimation enabling condition, the second point cloud block in the reference point cloud frame that matches the first point cloud block is obtained according to the target bitstream, and the motion vector of the first point cloud block relative to the second point cloud block is obtained.
[0317] Motion compensation is performed on the second point cloud block based on the motion vector to obtain the motion-compensated second point cloud block.
[0318] It is understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description of the decoding method in the method embodiment and achieve the same or corresponding technical effect. To avoid repetition, it will not be described again here.
[0319] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described encoding or decoding method embodiments and achieve the same technical effect. To avoid repetition, these will not be described again here.
[0320] The processor mentioned above is the processor in the terminal described in the above embodiments. The readable storage medium includes computer-readable storage media, such as ROM, RAM, magnetic disk, or optical disk. In some examples, the readable storage medium may be a non-transient readable storage medium.
[0321] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described encoding or decoding method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0322] It should be understood that the chips mentioned in the embodiments of this application may include system-on-a-chip (also known as system chip, chip system, or system-on-a-chip) or discrete display chips, etc.
[0323] This application also provides a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described encoding or decoding method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0324] This application also provides an encoding system, including an encoding device and a decoding device. The encoding device can be used to perform the steps of the encoding method described above, and the decoding device can be used to perform the steps of the decoding method described above.
[0325] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0326] From the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of computer software products plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.) and includes several instructions to cause the terminal or network-side device to execute the methods described in the various embodiments of this application.
[0327] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other implementations under the guidance of this application without departing from the spirit and scope of the claims. All of these implementations are within the protection scope of this application.
Claims
1. An encoding method, characterized in that, include: If the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold, the encoder copies the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block. The encoding end generates a target bitstream based on the first identification information, wherein the first identification information is used to instruct the decoding end to copy the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block, and the point cloud information includes at least one of the geometric encoding information and attribute encoding information of the point cloud corresponding to the second point cloud block.
2. The method according to claim 1, characterized in that, When the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold, the encoding end copies the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block, including: The encoding end obtains the tree data structure corresponding to the target point cloud frame; The encoding end obtains the size of the node to be encoded in the target level of the tree data structure; If the preset conditions are met and the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to the preset threshold, the encoding end copies the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block, wherein the first point cloud block is the point cloud block corresponding to the node to be encoded in the target level. The preset conditions include one of the following: The second identification information indicates that the target point cloud frame enables inter-frame prediction skip mode; The second identification information indicates that the target point cloud frame has enabled the inter-frame prediction skip mode, and the size of the node to be encoded is within a preset range. The inter-frame prediction skip mode refers to a mode in which the point cloud information corresponding to the point cloud frame is not encoded.
3. The method according to claim 2, characterized in that, The encoding end generates the target bitstream based on the first identifier information, including: The encoding end generates the target bitstream based on the first identification information and the second identification information; Alternatively, the encoding end generates the target bitstream based on the first identification information, the second identification information, and the preset range.
4. The method according to claim 2 or 3, characterized in that, The encoding end obtains the size of the node to be encoded in the target level of the tree data structure, including: When the third identification information indicates that the inter-frame prediction skip mode is enabled in the point cloud frame sequence, the encoding end obtains the size of the node to be encoded in the target level of the tree data structure, and the point cloud frame sequence includes the target point cloud frame.
5. The method according to claim 4, characterized in that, The encoding end generates the target bitstream based on the first identifier information, including: The encoding end generates the target bitstream based on the first identifier information, the second identifier information, and the third identifier information, or generates the target bitstream based on the first identifier information, the second identifier information, the third identifier information, and the preset range.
6. The method according to claim 1, characterized in that, Also includes: Obtain the distortion rate of the first point cloud block relative to the second point cloud block or the second point cloud block after motion compensation; Based on the distortion rate, the similarity between the first point cloud block and the second point cloud block, or between the first point cloud block and the second point cloud block after motion compensation, is obtained.
7. The method according to claim 1, characterized in that, The method further includes: In the target level of the tree data structure corresponding to the target point cloud frame, the size of the node to be encoded is the same as the size of the node to be encoded corresponding to the maximum prediction unit (LPU), and the reference point cloud frame satisfies the local motion estimation enabling condition, motion estimation is performed on the first point cloud block to determine the second point cloud block in the reference point cloud frame that matches the first point cloud block and to determine the motion vector of the first point cloud block relative to the second point cloud block. Motion compensation is performed on the second point cloud block based on the motion vector to obtain the motion-compensated second point cloud block.
8. A decoding method, characterized in that, include: The decoding end decodes the target bitstream to obtain the first identification information; When the first identification information indicates that the encoding end copies the point cloud information of the second point cloud block of the reference point cloud frame or the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud block of the target point cloud frame, the decoding end copies the point cloud information of the second point cloud block or the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud block. The point cloud information includes at least one of the geometric encoding information and attribute encoding information of the point cloud corresponding to the second point cloud block.
9. The method according to claim 8, characterized in that, The decoding end decodes the target bitstream to obtain first identification information, including: The decoding end decodes the first bitstream in the target bitstream to obtain the second identification information; The decoding end obtains the size of the node to be encoded in the target level of the tree data structure corresponding to the target point cloud frame; When the second identification information indicates that the target point cloud frame has enabled the inter-frame prediction skip mode, and the size of the node to be encoded is within a preset range, the second bitstream in the target bitstream is decoded to obtain the first identification information. The inter-frame prediction skip mode refers to a mode in which the point cloud information corresponding to the point cloud frame is not encoded.
10. The method according to claim 9, characterized in that, The first bitstream also includes third identification information; Obtaining the size of the node to be encoded in the target level of the tree data structure includes: When the third identification information indicates that the inter-frame prediction skip mode is enabled in the point cloud frame sequence, the size of the node to be encoded in the target level of the tree data structure is obtained, and the point cloud frame sequence includes the target point cloud frame.
11. The method according to claim 9, characterized in that, The method further includes: When the size of the node to be encoded in the target level is the same as the size of the node corresponding to the maximum prediction unit (LPU) and the reference point cloud frame satisfies the local motion estimation enabling condition, the second point cloud block in the reference point cloud frame that matches the first point cloud block is obtained according to the target bitstream, and the motion vector of the first point cloud block relative to the second point cloud block is obtained. Motion compensation is performed on the second point cloud block based on the motion vector to obtain the motion-compensated second point cloud block.
12. An encoding device, characterized in that, include: The first copying module is used to copy the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block when the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold. The generation module is used to generate a target bitstream based on the first identification information, wherein the first identification information is used to instruct the decoding end to copy the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block, and the point cloud information includes at least one of the geometric encoding information and attribute encoding information of the point cloud corresponding to the second point cloud block.
13. The apparatus according to claim 12, characterized in that, The first copy module includes: The first acquisition submodule is used to acquire the tree data structure corresponding to the target point cloud frame; The second acquisition submodule is used to acquire the size of the node to be encoded in the target level of the tree data structure; The copying submodule is used to copy the point cloud information of the second point cloud block or the point cloud information of the motion-compensated second point cloud block into the reconstructed point cloud of the first point cloud block when the preset conditions are met and the similarity between the first point cloud block of the target point cloud frame and the second point cloud block of the reference point cloud frame is greater than or equal to a preset threshold. The first point cloud block is the point cloud block corresponding to the node to be encoded in the target level. The preset conditions include one of the following: The second identification information indicates that the target point cloud frame enables inter-frame prediction skip mode; The second identification information indicates that the target point cloud frame has enabled the inter-frame prediction skip mode, and the size of the node to be encoded is within a preset range. The inter-frame prediction skip mode refers to a mode in which the point cloud information corresponding to the point cloud frame is not encoded.
14. The apparatus according to claim 13, characterized in that, The generation module is used to generate the target bitstream based on the first identification information and the second identification information; Alternatively, the target bitstream can be generated based on the first identification information, the second identification information, and the preset range.
15. The apparatus according to claim 13 or 14, characterized in that, The second acquisition submodule is used to acquire the size of the node to be encoded in the target level of the tree data structure when the third identification information indicates that the inter-frame prediction skip mode of the point cloud frame sequence is enabled. The point cloud frame sequence includes the target point cloud frame.
16. The apparatus according to claim 15, characterized in that, The generation module is used to generate the target bitstream based on the first identifier information, the second identifier information, and the third identifier information, or to generate the target bitstream based on the first identifier information, the second identifier information, the third identifier information, and the preset range.
17. The apparatus according to claim 12, characterized in that, Also includes: The first acquisition module is used to acquire the distortion rate of the first point cloud block relative to the second point cloud block or the second point cloud block after motion compensation; The second acquisition module is used to acquire the similarity between the first point cloud block and the second point cloud block or with the second point cloud block after motion compensation, based on the distortion rate.
18. The apparatus according to claim 12, characterized in that, Also includes: The determination module is used to perform motion estimation on the first point cloud block when the size of the node to be encoded in the target level of the tree data structure corresponding to the target point cloud frame is the same as the size of the node to be encoded corresponding to the maximum prediction unit (LPU), and the reference point cloud frame satisfies the local motion estimation enabling condition, thereby determining the second point cloud block in the reference point cloud frame that matches the first point cloud block and determining the motion vector of the first point cloud block relative to the second point cloud block. The third acquisition module is used to perform motion compensation on the second point cloud block according to the motion vector to obtain the motion-compensated second point cloud block.
19. A decoding device, characterized in that, include: The fourth acquisition module is used to decode the target bitstream to obtain the first identification information; The second copying module is used to copy the point cloud information of the second point cloud block or the second point cloud block after motion compensation from the reference point cloud frame to the reconstructed point cloud of the first point cloud block in the target point cloud frame when the first identification information indicates that the encoding end copies the point cloud information of the second point cloud block or the second point cloud block after motion compensation to the reconstructed point cloud of the first point cloud block. The point cloud information includes at least one of the geometric encoding information and attribute encoding information of the point cloud corresponding to the second point cloud block.
20. The apparatus according to claim 19, characterized in that, The fourth acquisition module includes: The third acquisition submodule is used to decode the first bitstream in the target bitstream and obtain the second identification information; The fourth acquisition submodule is used to acquire the size of the node to be encoded in the target level of the tree data structure corresponding to the target point cloud frame; The fifth acquisition submodule is used to decode the second bitstream in the target bitstream to obtain the first identification information when the second identification information indicates that the target point cloud frame has enabled the inter-frame prediction skip mode and the size of the node to be encoded is within a preset range. The inter-frame prediction skip mode refers to a mode in which the point cloud information corresponding to the point cloud frame is not encoded.
21. The apparatus according to claim 20, characterized in that, The first bitstream also includes third identification information; The fourth acquisition submodule is used to acquire the size of the node to be encoded in the target level of the tree data structure when the third identification information indicates that the inter-frame prediction skip mode of the point cloud frame sequence is enabled. The point cloud frame sequence includes the target point cloud frame.
22. The apparatus according to claim 20, characterized in that, Also includes: The fifth acquisition module is used to acquire, according to the target bitstream, a second point cloud block that matches the first point cloud block in the reference point cloud frame and acquire the motion vector of the first point cloud block relative to the second point cloud block, when the size of the node to be encoded in the target level is the same as the size of the node corresponding to the maximum prediction unit (LPU) and the reference point cloud frame satisfies the local motion estimation enabling condition. The sixth acquisition module is used to perform motion compensation on the second point cloud block according to the motion vector to obtain the motion-compensated second point cloud block.
23. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the method as claimed in any one of claims 1 to 7, or to implement the steps of the method as claimed in any one of claims 8 to 11.
24. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the method as claimed in any one of claims 1 to 7, or the steps of the method as claimed in any one of claims 8 to 11.
25. A chip, characterized in that, The chip includes a processor and a communication interface coupled to the processor. The processor is used to run programs or instructions to implement the steps of the method as described in any one of claims 1 to 7, or to implement the steps of the method as described in any one of claims 8 to 11.
Citation Information
Patent Citations
Point cloud compression
CN111133476A
Patch data unit coding and decoding for point-cloud data
CN113906757A