Multiscale Inter-Prediction for Dynamic Point Cloud Compression
Multi-scale inter-prediction using neural networks addresses reconstruction errors in point cloud compression by predicting latent features and encoding residuals, enhancing compression efficiency and reducing artifacts.
Patent Information
- Application Number
- JP2025522652
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-09-21
- Filing Date
- 2023-10-05
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2043-10-05
AI Technical Summary
Conventional point cloud compression techniques suffer from reconstruction errors and artifacts due to significant differences between current and previous frames, requiring additional bandwidth for multiple bitstreams, which is impractical for limited communication channels.
Implement multi-scale inter-prediction using neural networks to predict latent features from previous frames, determining residuals, and encoding these residuals for transmission, allowing reconstruction with reduced bitstream size and improved accuracy.
Reduces transmission throughput and minimizes reconstruction artifacts by leveraging multi-scale latent features for efficient compression and decoding of point cloud frames.
Smart Images

Figure 2025538860000001_ABST
Abstract
Description
[Technical Field]
[0001] Cross-reference to related applications / incorporation by reference This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 380,089, filed October 19, 2022, which claims priority to U.S. Patent Application No. 18 / 471,753, filed with the U.S. Patent and Trademark Office on September 21, 2023, the contents of which are incorporated herein by reference in their entireties. [Background technology]
[0002] Advances in the field of dynamic point cloud compression (PCC) have led to the development of techniques that enable efficient representation of data associated with 3D points in a point cloud. Point clouds typically contain numerous unstructured 3D points. Each 3D point can contain geometric and attribute information (e.g., color, transparency, reflectance, opacity, texture, and material) associated with the corresponding 3D point. Therefore, each 3D point in a point cloud can contain a significant amount of data. The point cloud data may require compression (i.e., encoding) using a PCC encoder for storage, processing, or transmission of the point cloud. A PCC decoder can then reconstruct the point cloud based on the encoded point cloud data received from the PCC encoder. The PCC encoder can generate encoded point cloud data based on the current point cloud frame and previously decoded point cloud frames to encode the current point cloud frame. The encoded point cloud data can be transmitted to a PCC decoder, which can reconstruct the corresponding points by decoding the encoded point cloud data. Reconstruction of the current point cloud frame based on encoded point cloud data (generated based on previously decoded point cloud frames) is prone to errors, and artifacts or irregularities may appear in the geometry of the reconstructed point cloud. Summary of the Invention
[0003] The limitations and disadvantages of conventional approaches will become apparent to those skilled in the art by comparing the described system with certain aspects of the present disclosure illustrated in the remainder of this application and with reference to the drawings.
[0004] An electronic apparatus and method for multi-scale inter-prediction for dynamic point cloud compression is provided substantially as shown and / or described in connection with at least one of the figures and more fully set forth in the claims.
[0005] These and other features and advantages of the present disclosure will become apparent from a consideration of the following detailed description of the disclosure when taken in conjunction with the accompanying drawings, in which like reference characters refer to like elements throughout. [Brief explanation of the drawings]
[0006] [Figure 1] FIG. 1 illustrates an exemplary network environment for multi-scale inter-prediction for dynamic point cloud compression, according to embodiments of the present disclosure. [Figure 2] FIG. 2 is a block diagram illustrating an exemplary first electronic device for multi-scale inter prediction for dynamic 3D point cloud frame compression, according to an embodiment of the present disclosure. [Figure 3] FIG. 2 is a block diagram illustrating an exemplary second electronic device for multi-scale inter-prediction for 3D point cloud frame reconstruction, according to an embodiment of the present disclosure. [Figure 4] FIG. 1 illustrates an exemplary architecture for multi-scale inter-prediction for dynamic 3D point cloud compression and 3D point cloud reconstruction, according to an embodiment of the present disclosure. [Figure 5] FIG. 10 is a block diagram illustrating an example operation of predicting a feature set associated with a current 3D point cloud frame based on multi-scale features associated with a reference 3D point cloud frame set, according to an embodiment of the present disclosure. [Figure 6]FIG. 10 is a block diagram illustrating an example operation of predicting a feature set associated with a current 3D point cloud frame based on multi-scale features associated with a reference 3D point cloud frame set, according to an embodiment of the present disclosure. [Figure 7A] FIG. 1 illustrates an exemplary scenario for encoding or decoding a 3D point cloud frame based on a previous 3D point cloud frame, according to an embodiment of the present disclosure. [Figure 7B] FIG. 2 illustrates an exemplary scenario for encoding or decoding a 3D point cloud frame based on two previous 3D point cloud frames, according to an embodiment of the present disclosure. [Figure 8] FIG. 1 illustrates an exemplary scenario for encoding / decoding a 3D point cloud frame based on a preceding 3D point cloud frame and a subsequent 3D point cloud frame, according to an embodiment of the present disclosure. [Figure 9] 1 is a flowchart illustrating the operation of an exemplary method for multi-scale inter-prediction for dynamic 3D point cloud frame compression, in accordance with an embodiment of the present disclosure. [Figure 10] 1 is a flowchart illustrating the operation of an exemplary method for multi-scale inter-prediction for 3D point cloud frame reconstruction, in accordance with an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0007] The disclosed first electronic device, second electronic device, and method for multi-scale inter-prediction for dynamic point cloud compression may include the following embodiments. An exemplary aspect of the present disclosure provides a first electronic device (e.g., a computer device, a game console, or virtual reality goggles) that predicts features associated with a 3D point cloud frame based on multi-scale features of a reference 3D point cloud frame and determines a difference between the actual features associated with the 3D point cloud frame and the predicted features. Specifically, the first electronic device may receive a three-dimensional (3D) point cloud sequence that may include a reference 3D point cloud frame set and a current 3D point cloud frame to be encoded. After receiving the 3D point cloud sequence, the first electronic device may generate reference frame data that may include a feature set associated with 3D points of each reference 3D point cloud frame of the reference 3D point cloud frame set. The first electronic device may further generate current frame data associated with 3D points of the current 3D point cloud frame. The current frame data may include a first feature set (i.e., an actual feature set) associated with the occupancy of the 3D points in the current 3D point cloud frame. The first electronic device can predict a second set of features associated with 3D points of the current 3D point cloud frame. The prediction of the second set of features can be based on application of a first neural network predictor to the reference frame data. The first electronic device can calculate a set of residual features based on the generated first set of features and the predicted second set of features. The first electronic device can then generate a quantized residual feature set based on application of a quantization scheme to the residual feature set. Finally, the first electronic device can generate a bitstream of encoded point cloud data for the current 3D point cloud frame based on application of an encoding scheme to the quantized residual feature set.
[0008] An exemplary embodiment of the present disclosure further provides a second electronic device (e.g., a computing device, a gaming console, virtual reality goggles, or a smart wearable device) that predicts features associated with a current 3D point cloud frame based on multi-scale features of a reference 3D point cloud frame set. The current 3D point cloud frame can be reconstructed based on the predicted features and the encoded point cloud data. Specifically, the second electronic device can receive a 3D point cloud sequence including the reference 3D point cloud frame set. After receiving the 3D point cloud sequence, the second electronic device can generate reference frame data including a feature set associated with 3D points of each reference 3D point cloud frame of the reference 3D point cloud frame set. The second electronic device can then receive (from the first electronic device) a bitstream of encoded point cloud data associated with the current 3D point cloud frame to be decoded. The second electronic device can predict a third feature set associated with 3D points of the current 3D point cloud frame based on application of a second neural network predictor to the reference frame data. The second electronic device can generate a fourth set of features associated with the 3D points of the current 3D point cloud frame based on the received bitstream of encoded point cloud data and the predicted third set of features, and can generate (i.e., reconstruct) the current 3D point cloud frame based on application of the decoding scheme to the determined fourth set of features.
[0009] Typically, 3D point clouds can be compressed using point cloud compression (PCC) techniques and then reconstructed using decoding techniques. Encoding is necessary because each 3D point in a 3D point cloud can contain a significant amount of point data, making transmission of the point cloud data over a communication channel with limited bandwidth impractical. Encoding can involve predicting features of a current point cloud frame based on features of a previously decoded point cloud frame. A bitstream of encoded point cloud data can be generated based on the predicted features. A PCC decoder can reconstruct the current point cloud frame based on the bitstream. However, if there are significant differences between the current point cloud frame and the previous point cloud frame, the reconstructed point cloud frame may contain artifacts or surface irregularities. To prevent the appearance of artifacts or irregularities, a difference between the current point cloud frame and the previously decoded point cloud frame can be determined. This difference can be used to estimate the motion of objects in the current point cloud frame relative to the previously decoded point cloud frame. Based on the motion estimation, another bitstream of encoded point cloud data can be generated. Because multiple bitstreams are generated, lossless reconstruction or decoding of each point cloud frame may require additional bandwidth resources to transmit multiple bitstreams of encoded point cloud data to the PCC decoder.
[0010] To address this issue, the first electronic device can perform dynamic point cloud compression based on predicting latent features associated with the current point cloud frame using multi-scale latent features associated with a previously decoded reference point cloud frame set that may precede the current point cloud frame. Furthermore, the second electronic device can determine multi-scale latent features associated with the reference point cloud frame set and reconstruct the current point cloud frame using the determined multi-scale latent features. Determining the multi-scale features by the first electronic device can facilitate a reduction in the size of the bitstream representing the encoded point cloud data associated with the current point cloud frame, thereby reducing transmission throughput. This reduction is achieved because a PCC encoder on the first electronic device can determine actual features associated with each point cloud frame, and residuals generated based on the actual and predicted features can be used to generate the encoded point cloud data. The residuals can represent differences between the actual and predicted features. The first electronic device can further compress the residuals using an entropy encoder and transmit the compressed residuals as encoded point cloud data to the second electronic device. The second electronic device may use a PCC encoder to determine multi-scale latent features associated with the reference point cloud frame set. Based on the determined multi-scale latent features associated with the reference point cloud frame set, latent features associated with the current point cloud frame may be predicted. The second electronic device may then reconstruct the received residuals and accumulate the reconstructed residuals with the predicted features to determine actual features associated with the current point cloud frame. The second electronic device may use a PCC decoder to reconstruct the current point cloud frame based on the determined actual features.
[0011] FIG. 1 illustrates an exemplary network environment for multi-scale inter-prediction for dynamic point cloud compression, according to an embodiment of the present disclosure. FIG. 1 shows a network environment 100. The network environment 100 includes a first electronic device 102, a second electronic device 104, and a server 106. The first electronic device 102 can communicate with the second electronic device 104 and the server 106 via one or more networks (e.g., a communication network 108). The first electronic device 102 can include a first point cloud compression (PCC) encoder 110, a first neural network predictor 112, and an octree-based encoder 114. The second electronic device 104 can include a second PCC encoder 116, an octree-based decoder 118, a second neural network predictor 120, and a PCC decoder 122. The first electronic device 102 can receive as input the reference 3D point cloud frame set 124 and the current 3D point cloud frame 126 and generate as output encoded point cloud data 128 related to the current 3D point cloud frame 126. The second electronic device 104 can receive as input the reference 3D point cloud frame set 124 and the encoded point cloud data 128 and generate as output a decoded 3D point cloud frame 130.
[0012] In FIG. 1, a first electronic device 102 is responsible for encoding point cloud data (e.g., a point cloud sequence), and a second electronic device 104 is responsible for decoding and reconstructing the point cloud data from a compressed point cloud representation shared by the first electronic device 102.
[0013] In some embodiments, the first electronic device 102 and the second electronic device 104 can be the same device. In such cases, the first electronic device 102 or the second electronic device 104 can be responsible for both encoding and reconstructing the point cloud data. The first PCC encoder 110 can be the same as the second PCC encoder 116, and the first neural network predictor 112 can be the same as the second neural network predictor 120.
[0014] The first electronic device 102 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive the reference 3D point cloud frame set 124, the current 3D point cloud frame 126, and target coordinate information associated with the current 3D point cloud frame 126. The first electronic device 102 may further determine, via the first PCC encoder 110, features associated with each reference 3D point cloud frame of the reference 3D point cloud frame set 124 and features associated with the current 3D point cloud frame 126. The first electronic device 102 may further predict, via the first neural network predictor 112, features associated with the current 3D point cloud frame 126 based on the features associated with each reference 3D point cloud frame of the reference 3D point cloud frame set 124. The first electronic device 102 may further be configured to determine a residual based on the features determined by the first PCC encoder 110 and the features predicted by the first neural network predictor 112, and transmit the residual. Examples of the first electronic device 102 may include, but are not limited to, a computing device such as a server, a video conferencing system, an augmented reality (AR) device, a virtual reality (VR) device, a mixed reality (MR) device, a gaming console, a server, a smart wearable device, a mainframe machine, a computer workstation, and / or a consumer electronics (CE) device.
[0015] The second electronic device 104 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive the reference 3D point cloud frame set 124, the encoded target coordinate information, and the compressed residual. The second electronic device 104 may determine features associated with each reference 3D point cloud frame of the reference 3D point cloud frame set 124 via the second PCC encoder 116 and predict features associated with the current 3D point cloud frame 126 via the second neural network predictor 120. Furthermore, the second electronic device 104 may reconstruct (i.e., decode) the current 3D point cloud frame 126 via the PCC decoder 122 to obtain the decoded 3D point cloud frame 130. The reconstruction may be based on the features and residual predicted by the second neural network predictor 120. Examples of the second electronic device 104 may include, but are not limited to, a computer device, a video conferencing system, an AR device, a VR device, a MR device, a gaming console, a smart wearable device, a server, a mainframe machine, a computer workstation, and / or a CE device.
[0016] The server 106 may include suitable logic, circuitry, interfaces, and / or code that may be configured to generate a reference 3D point cloud frameset 124 of a 3D object in 3D space. The server 106 may be configured to generate each 3D reference point cloud frame of the reference 3D point cloud frameset 124 using image and depth information of the object. The server 106 may be configured to store the reference 3D point cloud frameset 124 and information related to the reference 3D point cloud frameset 124. The server 106 may be further configured to receive a request from the first electronic device 102 or the second electronic device 104 for the reference 3D point cloud frameset 124. The server 106 may transmit the reference 3D point cloud frameset 124 to the first electronic device 102 or the second electronic device 104 based on the request. In some embodiments, the server 106 may include a first PCC encoder 110, a first neural network predictor 112, an octree-based encoder 114, a second PCC encoder 116, an octree-based decoder 118, a second neural network predictor 120, and a PCC decoder 122.
[0017] The server 106 may perform operations via web applications, cloud applications, HTTP requests, repository operations, file transfers, etc. Example implementations of the server 106 include, but are not limited to, a database server, a file server, a web server, an application server, a mainframe server, a cloud computing server, or a combination thereof. In at least one embodiment, the server 106 may be implemented as a plurality of distributed cloud-based resources using techniques known to those skilled in the art. Those skilled in the art will appreciate that the scope of the present disclosure may not be limited to implementing the server 106 and the first electronic device 102 as two separate entities, the server 106 and the second electronic device 104 as two separate entities, or the server 106, the first electronic device 102, and the second electronic device 104 as three separate entities. In some embodiments, the functionality of the server 106 may be incorporated, in whole or at least in part, into the first electronic device 102 or the second electronic device 104 without departing from the scope of the present disclosure.
[0018] The communication network 108 may include a communication medium that enables the first electronic device 102, the second electronic device 104, and the server 106 to communicate with each other. The communication network 108 may be a wired or wireless communication network. Examples of the communication network 108 may include, but are not limited to, the Internet, a Wireless Fidelity (Wi-Fi) network, a personal area network (PAN), a local area network (LAN), or a metropolitan area network (MAN). The first electronic device 102 and the second electronic device 104 may be configured to connect to the communication network 108 according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols include, but are not limited to, at least one of Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), ZigBee, EDGE, IEEE802.11, Light Fidelity (Li-Fi), 802.16, IEEE802.11s, IEEE802.11g, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocols, and Bluetooth (BT) communication protocols.
[0019] Each of the first PCC encoder 110 and the second PCC encoder 116 may include suitable logic, circuitry, interfaces and / or code that may be configured to encode each reference 3D point cloud frame of the reference 3D point cloud frame set 124 to generate a set of features associated with the 3D points of the corresponding reference 3D point cloud frame of the reference 3D point cloud frame set 124. The first PCC encoder 110 may be further configured to encode the current 3D point cloud frame 126 to generate a first set of features associated with the 3D points of the current 3D point cloud frame 126.
[0020] Each of the first PCC encoder 110 and the second PCC encoder 116 may be implemented as a deep neural network (including a model file and associated inference code) that can run on a graphics processing unit (GPU), a central processing unit (CPU), a tensor processing unit (TPU), a reduced instruction set computing (RISC) processor, an application specific integrated circuit (ASIC) processor, or a complex instruction set computing (CISC) processor, coprocessor, and / or combinations thereof. In some embodiments, the first PCC encoder 110 may be implemented as a deep neural network on dedicated hardware in conjunction with other computational circuitry of the first electronic device 102. Similarly, the second PCC encoder 116 may be implemented as a deep neural network on dedicated hardware in conjunction with other computational circuitry of the second electronic device 104. In such implementations, the first PCC encoder 110 and the second PCC encoder 116 may be associated with a particular form factor on a particular computational circuit. Examples of specific computing circuits include, but are not limited to, field programmable gate arrays (FPGAs), programmable logic devices (PLDs), ASICs, programmable ASICs (PL-ASICs), application specific integrated components (ASSPs), and standard microprocessor (MPU) or digital signal processor (DSP) based systems on chips (SOCs). According to an embodiment, the first PCC encoder 110 or the second PCC encoder 116 may also be interfaced with a GPU to parallelize the operation of the first PCC encoder 110 or the second PCC encoder 116, respectively.
[0021] Each of the first neural network predictor 112 and the second neural network predictor 120 can be referred to as a neural network, which is a system of computational networks or artificial neurons that can typically be arranged in multiple layers. A neural network can be defined by hyperparameters, such as activation function(s), number of weights, cost function, regularization function, input size, and number of layers. These layers can further include an input layer, one or more hidden layers, and an output layer. Each of the multiple layers can include one or more nodes (or artificial neurons). The output of every node in the input layer can be connected to at least one node in the hidden layer(s). Similarly, the input of each hidden layer can be connected to the output of at least one node in another layer of the neural network. The output of each hidden layer can be connected to the input of at least one node in another layer of the neural network. The node(s) in the final layer can receive input from at least one hidden layer and output a result. The number of layers and the number of nodes in each layer can be determined from the hyperparameters of the neural network. Such hyperparameters can be set before or after training the neural network.
[0022] Each node can correspond to a mathematical function (e.g., a sigmoid function or a rectified linear unit) with parameters that can be adjusted during training of the neural network. The parameter set can include, for example, weight parameters and regularization parameters. Each node can calculate an output using a mathematical function based on one or more inputs from nodes in other layer(s) of the neural network (e.g., previous layer(s)). All or some of the nodes in a neural network can correspond to the same or different mathematical functions. In training a neural network, one or more parameters of each node of the neural network can be updated based on whether the output of the final layer for a given input (from the training dataset) matches the correct result according to the neural network's loss function. The above process can be repeated for the same or different inputs until a minimum value of the loss function is achieved and the training error is minimized. Several training methods are known in the art, such as gradient descent, stochastic gradient descent, batch gradient descent, gradient boosting, and metaheuristic methods.
[0023] Each of the first neural network predictor 112 and the second neural network predictor 120 can be a machine learning model trained to generate multi-scale features associated with an input 3D point cloud frame. Each of the first neural network predictor 112 and the second neural network predictor 120 can receive as input a feature set associated with 3D points in each reference 3D point cloud frame of the reference 3D point cloud frame set 124 and target coordinate information associated with the current 3D point cloud frame 126. Each of the first neural network predictor 112 and the second neural network predictor 120 can generate as output a prediction in response to the input. The prediction can indicate features associated with the current 3D point cloud frame 126. The first neural network predictor 112 can predict a second feature set associated with the current 3D point cloud frame 126 based on application of the first neural network predictor 112 to the feature set (generated by the first PCC encoder 110) and the target coordinate information. The second neural network predictor 120 can predict a third feature set associated with the current 3D point cloud frame 126 based on application of the second neural network predictor 120 to the feature set (produced by the second PCC encoder 116) and the target coordinate information (decoded by the octree-based decoder 118).
[0024] In some embodiments, each of the first neural network predictor 112 and the second neural network predictor 120 may include electronic data that may be implemented as a software component of an application executable on the first electronic device 102 and the second electronic device 104. Each of the first neural network predictor 112 and the second neural network predictor 120 may rely on libraries, external scripts, or logic / instructions for execution by processing units included in the first electronic device 102 and the second electronic device 104. In one or more embodiments, each of the first neural network predictor 112 and the second neural network predictor 120 may be implemented using hardware that may include a processor, a microprocessor (e.g., that performs or controls the execution of one or more operations), an FPGA, or an ASIC. Alternatively, in some embodiments, each of the first neural network predictor 112 and the second neural network predictor 120 may be implemented using a combination of hardware and software. Examples of the first neural network predictor 112 and the second neural network predictor 120 may include, but are not limited to, a deep neural network (DNN), a convolutional neural network (CNN), an artificial neural network (ANN), a fully connected neural network, and / or a combination of such networks.
[0025] The octree-based encoder 114 may include suitable logic, circuitry, interfaces, and / or code that can be configured to encode data based on the use of a data structure (i.e., an octree). The octree-based encoder 114 may recursively divide a 3D space (such as a portion of the current 3D point cloud frame 126 that includes 3D points located at target coordinates specified in the target coordinate information) into smaller regions (e.g., blocks) known as octants. This division may continue until each block meets a stopping criterion (i.e., the points within each block represent similar density or texture). In one embodiment, the octant-based encoder 114 may encode partitioning hierarchy information and point cloud data associated with the 3D points located at the target coordinates to generate encoded target coordinate information.
[0026] In some embodiments, the octree-based encoder 114 can be a machine learning-based octree encoder that can use techniques such as deep-octree coding to encode (compress) 3D points at target coordinates in the current 3D point cloud frame 126. Examples of such machine learning-based octree encoders include, but are not limited to, G-PCC, OctSqueeze, and VoxelContext-Net.
[0027] The octree-based decoder 118 may include suitable logic, circuitry, interfaces, and / or code that may be configured to decode the encoded target coordinate information. At each level of the octree (based on which the encoded target coordinate information may be decoded), an octant may be reconstructed until the entire volume (i.e., the portion of the current 3D point cloud frame 126 that includes 3D points located at target coordinates) is reconstructed. Based on the octants, reconstructed decoded target coordinate information may be obtained.
[0028] In some embodiments, the octree-based encoder 114 can be a machine learning-based octree encoder that can use deep octree decoding techniques to decode encoded target coordinate information associated with the current 3D point cloud frame 126. Examples of such machine learning-based octree decoders include, but are not limited to, G-PCC, OctSqueeze, and VoxelContext-Net.
[0029] The PCC decoder 122 may include suitable logic, circuitry, and / or interfaces that may be configured to reconstruct the current 3D point cloud frame 126 based on the encoded point cloud data 128. The PCC decoder 122 may receive a fourth feature set as input. The fourth feature set may be generated based on the third feature set and a decompressed residual. The decoded 3D point cloud frame 130 may be generated based on application of the PCC decoder 122 to the fourth feature set. The PCC decoder 122 may be implemented as a deep neural network on a GPU, CPU, TPU, RISC processor, ASIC processor, CISC processor, coprocessor, and / or combinations thereof. In some other embodiments, the PCC decoder 122 may be implemented as a deep neural network on dedicated hardware in conjunction with other computational circuitry of the second electronic device 104. In such implementations, the PCC decoder 122 may be associated with a particular form factor on a particular computational circuit. Examples of specific computational circuits include, but are not limited to, FPGAs, PLDs, ASICs, PL-ASICs, ASSPs, and standard MPU- or DSP-based SOCs. In some embodiments, the PCC decoder 122 can be interfaced with a GPU to parallelize the operation of the PCC decoder 122.
[0030] Each of the reference 3D point cloud frames and the current 3D point cloud frame 126 of the reference 3D point cloud frame set 124 may correspond to a geometric representation of one or more 3D objects in a 3D environment (e.g., a real-world environment). Each 3D point cloud frame may comprise a set of 3D points arranged at different positions according to a 3D coordinate system. According to an embodiment, the first electronic device 102 may obtain the reference 3D point cloud frame set 124 and the current 3D point cloud frame 126 from the server 106. Similarly, the second electronic device 104 may obtain the reference 3D point cloud frame set 124 from the server 106. Each 3D point in the reference 3D point cloud frame may include geometric information (i.e., coordinates of the corresponding 3D point in the corresponding reference 3D point cloud frame) and attribute information associated with the corresponding 3D point. The attribute information may include, for example, color information, reflectance information, opacity information, normal vector information, material identifier information, or texture information.
[0031] The encoded point cloud data 128 may be generated based on encoding each reference 3D point cloud frame of the reference 3D point cloud frame set 124 and the current 3D point cloud frame 126. The encoded point cloud data 128 may also be generated based on predicting multi-scale features associated with the current 3D point cloud frame 126 and encoding target coordinate information associated with the current 3D point cloud frame 126. The encoded point cloud data 128 may include a compressed residual (generated based on the generated first feature set and the predicted second feature set) and the encoded target coordinate information. The first electronic device 102 may transmit the encoded point cloud data 128 to the second electronic device 104 as a bitstream.
[0032] The decoded 3D point cloud frame 130 can be a reconstructed point cloud frame that can correspond to the current 3D point cloud frame 126. The second electronic device 104 can reconstruct the current 3D point cloud frame 126 based on the encoded point cloud data 128 via the PCC decoder 122. The reconstruction can be based on decompressing the compressed residual and accumulating the decompressed residual (i.e., the original generated residual) with a predicted third feature set (i.e., a fourth feature set). The decoded 3D point cloud frame 130 can be generated based on applying the PCC decoder 122 to this accumulation.
[0033] In operation, the first electronic device 102 may be configured to receive a 3D point cloud sequence that may include a reference 3D point cloud frame set 124 and a current 3D point cloud frame 126 to be encoded. According to an embodiment, each reference 3D point cloud frame in the reference 3D point cloud frame set 124 may be a previously decoded 3D point cloud frame and may precede (i.e., be received earlier than) or follow the current 3D point cloud frame 126 in a timeline for receiving the 3D point cloud sequence.
[0034] The first electronic device 102 may be further configured to generate reference frame data including a set of features associated with the 3D points of each reference 3D point cloud frame of the reference 3D point cloud frame set 124. The generation of the reference frame data may be performed based on application of the first PCC encoder 110 to each reference 3D point cloud frame of the reference 3D point cloud frame set 124. The reference frame data may be generated as an output of the first PCC encoder 110. The set of features associated with the 3D points of each reference 3D point cloud frame may include reference features associated with the occupancy of the 3D points in the corresponding reference 3D point cloud frame of the reference 3D point cloud frame set 124 and reference coordinate information associated with the 3D points in the corresponding reference 3D point cloud frame. The reference coordinate information may include, for example, coordinates of the 3D points in the corresponding reference 3D point cloud frame.
[0035] The first electronic device 102 may be further configured to generate current frame data associated with the 3D points of the current 3D point cloud frame 126. The generation of the current frame data may be based on application of the first PCC encoder 110 to the current 3D point cloud frame 126. The generated current frame data may include a first set of features associated with the occupancy of the 3D points of the current 3D point cloud frame 126. The first set of features may be generated as an output of the first PCC encoder 110. The first set of features may be referred to as actual features associated with the 3D points of the current 3D point cloud frame 126. According to an embodiment, the first set of features may include features associated with the occupancy of the 3D points having coordinates included in the target coordinate information.
[0036] The first electronic device 102 may be further configured to predict a second set of features associated with 3D points of the current 3D point cloud frame 126 based on application of the first neural network predictor 112 to the reference frame data. A feature set associated with 3D points of each reference 3D point cloud frame of the reference 3D point cloud frame set 124 may be provided as input to the first neural network predictor 112. Note that each feature set may be associated with a 3D point of the corresponding reference 3D point cloud frame having coordinates included in the reference coordinate information. The feature set may include reference features related to occupancy of the 3D point in the corresponding reference 3D point cloud frame.
[0037] According to an embodiment, the first electronic device 102 can provide coordinate information associated with 3D points of the current 3D point cloud frame 126 (i.e., target coordinate information) as input to the first neural network predictor 112. The 3D points can be 3D points for which features need to be predicted, and the target coordinate information can include coordinates of these 3D points. The target coordinate information can include target coordinates included in the current 3D point cloud frame 126. Based on application of the first neural network predictor 112 to the input, a second set of features can be predicted as an output of the first neural network predictor 112. The second set of features can include features related to the occupancy of 3D points having coordinates included in the target coordinate information.
[0038] According to one embodiment, the feature set associated with each reference 3D point cloud frame can be downsampled by a scale of 2, 3, ..., K for prediction. For example, the feature set associated with a first reference 3D point cloud frame in the reference 3D point cloud frame set 124 can be downsampled by a scale of 2, 3, ..., K. Thus, (K-1) feature sets can be generated for the first reference 3D point cloud frame. Similarly, (K-1) feature sets can be generated for each of the other reference 3D point cloud frames in the reference 3D point cloud frame set 124. Then, a space-time tensor can be constructed for each downsampling scale based on the feature sets of all reference 3D point cloud frames downsampled by the same scale. Because the feature sets associated with each reference 3D point cloud frame are downsampled by a scale of 2, 3, ..., K, (K-1) space-time tensors can be constructed. For example, for a downsampling scale of "2," a first space-time tensor can be constructed using the feature sets associated with all reference 3D point cloud frames downsampled by a scale of 2. Similarly, for a downsampling scale "K," a (K-1)th spatiotemporal sensor can be constructed using feature sets associated with all reference 3D point cloud frames 124 downsampled by the "K" scale. Construction of the spatiotemporal sensor for the corresponding downsampling scale can be performed based on spatiotemporal concatenation of the feature sets of all reference 3D point cloud frames that can be downsampled by the corresponding downsampling scale. Spatiotemporal tensor analysis can then be performed based on application of sparse convolution or self-attention operations to each of the (K-1) spatiotemporal tensors.After space-time tensor analysis, each of the (K-1) downsampled space-time tensors can be further downsampled (apart from the space-time tensor constructed based on the space-time concatenation of the feature sets downsampled by K scales). The (K-1) downsampled space-time tensors can then be concatenated to generate a multi-scale feature concatenation vector.
[0039] Based on the target coordinate information and the multi-scale feature connection vector associated with the 3D points of the current 3D point cloud frame 126, a second set of features associated with the 3D points of the current 3D point cloud frame 126 can be predicted as an output of the first neural network predictor 112. The second set of features can include features associated with the 3D points having coordinates included in the target coordinate information.
[0040] The first electronic device 102 may be further configured to calculate a residual feature set based on the first feature set (i.e., current frame data associated with 3D points in the current 3D point cloud frame 126) and the predicted second feature set. According to an embodiment, the first electronic device 102 may calculate a difference between the first feature set and the predicted second feature set. The calculated difference may indicate an error associated with the second feature set (i.e., features associated with 3D points in the current 3D point cloud frame 126 predicted based on the reference frame data) relative to the first feature set (i.e., actual features associated with 3D points in the current 3D point cloud frame 126). The difference may also correspond to a residual feature set associated with 3D points in the current 3D point cloud frame 126 having coordinates included in the target coordinate information.
[0041] The first electronic device 102 may be further configured to generate a quantized residual feature set based on application of a quantization scheme to the residual feature set. The quantization scheme may be based on an entropy model and may include a quantization level set. According to an embodiment, the value of each residual feature in the residual feature set may be quantized to a certain quantization level in the quantization level set. The residual feature set may be quantized for subsequent compression and encoding of the residual feature set.
[0042] The first electronic device 102 may be further configured to generate a bitstream of encoded point cloud data 128 for the current 3D point cloud frame 126 based on application of an encoding scheme to the generated quantized residual feature set. The encoding scheme may be based on an entropy model (on which the quantization scheme may be based). Based on the encoding scheme, each quantized residual feature of the quantized residual feature set may be compressed to generate the encoded point cloud data 128. The bitstream of encoded point cloud data 128 may constitute a compressed quantized residual feature set.
[0043] According to an embodiment, the encoded point cloud data 128 may further include encoded target coordinate information. The target coordinate information may be encoded based on applying an octree-based encoder 114 to the target coordinate information. The first electronic device 102 may transmit a bitstream of the encoded point cloud data 128 to the second electronic device 104.
[0044] According to one embodiment, the second electronic device 104 can be configured to receive (from the first electronic device 102) the encoded point cloud data 128 and a 3D point cloud sequence, which may include the reference 3D point cloud frame set 124. The second electronic device 104 can be configured to extract the compressed quantized residual feature set and the encoded target coordinate information from the encoded point cloud data 128 (i.e., the bitstream generated by the first electronic device 102). The second electronic device 104 can then apply a decoding scheme to the compressed quantized residual feature set to decompress the compressed quantized residual feature set. The decoding scheme can generate a residual feature set that can correspond to the residual feature set calculated (by the first electronic device 102) based on the first feature set and the predicted second feature set. The second electronic device 104 can apply the octree-based decoder 118 to the encoded target coordinate information to generate the target coordinate information as an output of the octree-based decoder 118.
[0045] The second electronic device 104 may be further configured to generate reference frame data that may include a feature set associated with the 3D points of each reference 3D point cloud frame of the reference 3D point cloud frame set 124. The feature set associated with each reference 3D point cloud frame may be generated based on application of the second PCC encoder 116 to the corresponding reference 3D point cloud frame. The feature set may be generated as an output of the second PCC encoder 116. The features associated with each reference 3D point cloud frame may include reference features associated with the occupancy of the 3D points of the corresponding reference 3D point cloud frame and reference coordinate information (including coordinates of the 3D points of the corresponding reference 3D point cloud frame).
[0046] The second electronic device 104 may be further configured to predict a third set of features associated with 3D points of the current 3D point cloud frame 126 based on application of the second neural network predictor 120 to the reference frame data. The third set of features may be associated with 3D points having coordinates included in the target coordinate information decoded by the octree-based decoder 118. According to an embodiment, the second neural network predictor 120 may receive as input the feature sets associated with 3D points of each reference 3D point cloud frame of the reference 3D point cloud frame set 124 (i.e., the output of the second PCC encoder 116) and the target coordinate information (i.e., the output of the octree-based decoder 118). Based on application of the second neural network predictor 120 to the input, the third set of features may be predicted as the output of the second neural network predictor 120. If the first neural network predictor 112 and the second neural network predictor 120 are identical, then the prediction of the third feature set (by the second neural network predictor 120) can be identical to the prediction of the second feature set (by the first neural network predictor 112). Furthermore, if the first PCC encoder 110 and the second PCC encoder 116 are identical and the decoded target coordinate information (produced by the octree-based decoder 118) matches the target coordinate information encoded by the octree-based encoder 114, then the second feature set and the third feature set can be identical.
[0047] The second electronic device 104 may be further configured to generate a fourth feature set associated with 3D points of the current 3D point cloud frame 126 (having coordinates included in the target coordinate information) based on the received bitstream of encoded point cloud data 128 (i.e., the residual feature set that may be generated by the decoding scheme) and the predicted third feature set. According to an embodiment, the residual feature set and the predicted third feature set may be accumulated to generate the fourth feature set. The 3D points of the current 3D point cloud frame 126 may be reconstructed based on application of the PCC decoder 122 to the fourth feature set. The PCC decoder 122 may generate a decoded 3D point cloud frame 130 as an output. The decoded 3D point cloud frame 130 may be a reconstructed version of the current 3D point cloud frame 126.
[0048] FIG. 2 is a block diagram illustrating an exemplary first electronic device for multi-scale inter-prediction for dynamic 3D point cloud frame compression, according to an embodiment of the present disclosure. The description of FIG. 2 is provided with reference to the elements of FIG. 1 . FIG. 2 illustrates a block diagram 200 of a first electronic device 102. The first electronic device 102 may include a circuit 202, a memory 204, an input / output (I / O) device 206, and a network interface 208. In at least one embodiment, the memory 204 may include a first PCC encoder 110, a first neural network predictor 112, and an octree-based encoder 114. In at least one embodiment, the I / O device 206 may include a display device 210. The circuit 202 may be communicatively coupled to the memory 204, the I / O device 206, and the network interface 208 via wired or wireless communication of the first electronic device 102.
[0049] Circuitry 202 may include suitable logic, circuits, and interfaces that may be configured to execute program instructions associated with different operations performed by first electronic device 102. Circuitry 202 may include one or more processing units, which may be implemented as an integrated processor or processors that collectively perform the functions of one or more dedicated processing units. Circuitry 202 may be implemented based on multiple processor technologies known in the art. Example implementations of circuitry 202 may be an x86-based processor, a GPU, a CPU, a RISC processor, an ASIC processor, a CISC processor, a microcontroller, and / or other computing circuitry.
[0050] The memory 204 may include suitable logic, circuitry, and / or interfaces that may be configured to store instructions executable by the circuit 202. The memory 204 may be configured to store an operating system and associated applications. The memory 204 may further be configured to store the 3D point cloud sequence, the generated reference frame data, the generated current frame data, coordinate information associated with 3D points of the current 3D point cloud frame 126, the second feature set, the residual feature set, the quantized residual feature set, and a bitstream of the encoded point cloud data 128, etc. In at least one embodiment, the first PCC encoder 110, the first neural network predictor 112, and the octree-based encoder 114 included in the memory 204 may be implemented as a combination of programmable instructions stored in the memory 204 or logic units on a hardware circuit (i.e., a programmable logic unit) of the first electronic device 102. Examples of implementations of memory 204 include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), hard disk drive (HDD), solid state drive (SSD), CPU cache, and / or secure digital (SD) cards.
[0051] The I / O device 206 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive user input that may trigger the reception of a 3D point cloud sequence or the transmission of a bitstream of encoded point cloud data 128. The I / O device 206 may be further configured to provide output in response to the user input. The I / O device 206 may include a variety of input and output devices that may be configured to communicate with the circuit 202. Examples of input devices may include, but are not limited to, a touch screen, a keyboard, a mouse, a joystick, and / or a microphone. Examples of output devices may include a display device 210.
[0052] The display device 210 may include suitable logic, circuitry, interfaces, and / or code that may be configured to render each reference 3D point cloud frame of a set of reference 3D point cloud frames included in a 3D point cloud sequence on a display screen of the display device 210. According to some embodiments, the display device 210 may include a touch screen for receiving user input. The display device 210 may be implemented through a number of known technologies, including, but not limited to, liquid crystal display (LCD) displays, light emitting diode (LED) displays, plasma displays, and / or organic LED (OLED) display technologies, and / or other display technologies. According to some embodiments, the display device 210 may represent a display screen of a smart glasses device, a 3D display, a see-through display, a projection display, an electrochromic display, and / or a transparent display.
[0053] The network interface 208 may include suitable logic, circuitry, interfaces, and / or code that may be configured to establish communications between the first electronic device 102, the second electronic device 104, and the server 106 over the communications network 108. The network interface 208 may be implemented using various known technologies to support wired or wireless communications between the first electronic device 102 and the communications network 108. The network interface 208 may include, but is not limited to, an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber identity module (SIM) card, and / or a local buffer.
[0054] The network interface 208 can communicate via wireless communication with networks such as the Internet, an intranet, and / or wireless networks such as a cellular network, a wireless local area network (LAN), and / or a metropolitan area network (MAN). The wireless communication can use any of a number of communication standards, protocols, and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Long Term Evolution (LTE), Fifth Generation (5G) New Radio (NR), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wireless Fidelity (WiFi) (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Light Fidelity (Li-Fi), Wi-MAX, protocols for email, instant messaging, and / or short message service.
[0055] The functions or operations performed by first electronic device 102 described in Figure 1 may be performed by circuitry 202. The operations performed by circuitry 202 are described in detail in, for example, Figures 3, 4, 5, 6, 7A, 7B, and 8.
[0056] FIG. 3 is a block diagram illustrating an exemplary second electronic device for multi-scale inter-prediction for 3D point cloud frame reconstruction, according to an embodiment of the present disclosure. The description of FIG. 3 is provided with reference to the elements of FIG. 1 . FIG. 3 illustrates a block diagram 300 of the second electronic device 104. The second electronic device 104 may include a circuit 302, a memory 204, an I / O device 306, and a network interface 308. In one embodiment, the memory 304 may include a second PCC encoder 116, an octree-based decoder 118, a second neural network predictor 120, and a PCC decoder 122. In at least one embodiment, the I / O device 306 may include a display device 310. The circuit 302 may be communicatively coupled to the memory 304, the I / O device 306, and the network interface 308 via wired or wireless communication of the second electronic device 104.
[0057] The circuitry 302 may include suitable logic, circuits, and interfaces that may be configured to execute program instructions associated with different operations performed by the second electronic device 104. The circuitry 302 may include one or more processing units, which may be implemented as an integrated processor or a collection of processors that collectively perform the functions of one or more dedicated processing units. The circuitry 302 may be implemented based on multiple processor technologies known in the art. Example implementations of the circuitry 302 may be an x86-based processor, a GPU, a RISC processor, an ASIC processor, a CISC processor, a microcontroller, a CPU, and / or other computing circuitry.
[0058] The memory 304 may include suitable logic, circuitry, and / or interfaces that can be configured to store instructions executable by the circuit 302. The memory 304 may be configured to store an operating system and associated applications. The memory 304 may be further configured to store the 3D point cloud sequence, the reference frame data, decoded coordinate information associated with 3D points of the current 3D point cloud frame to be decoded, the predicted third feature set, the decompressed residual feature set, and the 3D decoded point cloud frame 130. In at least one embodiment, the second PCC encoder 116, the octree-based decoder 118, the second neural network predictor 120, and the PCC decoder 122 included in the memory 304 are implemented as a combination of programmable instructions stored in the memory 304 or logic units on a hardware circuit (i.e., a programmable logic unit) of the second electronic device 104. Example implementations of the memory 304 may include, but are not limited to, RAM, ROM, EEPROM, HDD, SSD, CPU cache, and / or SD card.
[0059] The I / O devices 306 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive user input that may trigger receipt of the 3D point cloud sequence and the bitstream of encoded point cloud data 128. The I / O devices 306 may be further configured to provide output in response to the user input. The I / O devices 306 may include a variety of input and output devices that may be configured to communicate with the circuitry 302. Examples of input devices may include, but are not limited to, a touch screen, a keyboard, a mouse, a joystick, and / or a microphone. Examples of output devices may include a display device 310.
[0060] The display device 310 may include suitable logic, circuitry, interfaces, and / or code that may be configured to render each reference 3D point cloud frame of the reference 3D point cloud frame set and the decoded 3D point cloud frame 130 on a display screen of the display device 310. According to some embodiments, the display device 310 may include a touch screen for receiving user input. The display device 310 may be implemented through a number of known technologies, including, but not limited to, LCD display, LED display, plasma display, and / or OLED display technology, and / or other display technologies. According to some embodiments, the display device 310 may represent a display screen of a smart glasses device, a 3D display, a see-through display, a projection display, an electrochromic display, and / or a transparent display.
[0061] The network interface 308 may include suitable logic, circuitry, interfaces, and / or code that may be configured to establish communications between the first electronic device 102, the second electronic device 104, and the server 106 over the communications network 108. The network interface 308 may be implemented using various known technologies to support wired or wireless communications between the second electronic device 104 and the communications network 108. The network interface 308 may include, but is not limited to, an antenna, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, a SIM card, and / or a local buffer.
[0062] The network interface 308 can communicate via wireless communication with networks such as the Internet, an intranet, and / or wireless networks such as a cellular network, a wireless LAN, and / or a MAN. The wireless communication can use any of a number of communication standards, protocols, and technologies such as GSM, EDGE, W-CDMA, LTE, 5G NR, CDMA, TDMA, Bluetooth, Wi-Fi (such as IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n), VoIP, Li-Fi, Wi-MAX, protocols for email, instant messaging, and / or SMS.
[0063] The functions or operations performed by second electronic device 104 described in Figure 1 may be performed by circuitry 302. The operations performed by circuitry 302 are described in detail, for example, in Figures 4, 5, 6, 7A, 7B, and 8.
[0064] Figure 4 illustrates an exemplary architecture for multi-scale inter-prediction for dynamic 3D point cloud compression and 3D point cloud reconstruction, according to an embodiment of the present disclosure. The description of Figure 4 is provided with reference to elements in Figures 1, 2, and 3. Figure 4 illustrates an exemplary architecture 400 for dynamic 3D point cloud compression and 3D point cloud reconstruction. The architecture 400 includes a first PCC encoder 110, a first neural network predictor 112, an octree-based encoder 114, a second PCC encoder 116, an octree-based decoder 118, a second neural network predictor 120, a PCC decoder 122, a subtractor 402, a quantizer 404, an autoencoder 406, an autodecoder 408, and an accumulator 410.
[0065] At any point in time, the first PCC encoder 110 may receive a reference 3D point cloud frame set and a current 3D point cloud frame to be encoded, i.e., P(t). The reference 3D point cloud frame set may include “N” 3D point cloud frames, i.e., P(t-1), P(t-2), ..., and P(tN). The reference 3D point cloud frame set, i.e., P(t-1), P(t-2), ..., or P(tN), may precede or follow the current 3D point cloud frame. In some cases, such frames may be referred to as previously decoded 3D point cloud frames (i.e., frames decoded before the reception of the current 3D point cloud frame). The first PCC encoder 110 may receive the reference 3D point cloud frame set and the current 3D point cloud frame as inputs. The first PCC encoder 110 may generate reference frame data and current frame data as respective outputs. The reference frame data may include a feature set associated with a 3D point in each reference 3D point cloud frame of the reference 3D point cloud frame set. For example, an output feature set associated with a 3D point in an input reference 3D point cloud frame P(t-1) may be F(t-1). Similarly, an output feature set associated with a 3D point in an input reference 3D point cloud frame P(tN) may be F(tN). The feature set (e.g., F(t-1)) associated with a 3D point in each reference 3D point cloud frame (e.g., P(t-1)) of the reference 3D point cloud frame set may include reference features associated with the occupancy of the 3D point in the corresponding reference 3D point cloud frame and reference coordinate information associated with the coordinates of the 3D point in the corresponding reference 3D point cloud frame. The current frame data may be associated with a 3D point in the current 3D point cloud frame (i.e., P(t)) and may include a first feature set, i.e., F(t). The first feature set may be associated with the occupancy of the 3D point in the current 3D point cloud frame and may be generated as an output of the first PCC encoder 110. According to an embodiment, the 3D points can be points of the current 3D point cloud frame that need to be encoded.
[0066] The first neural network predictor 112 may receive as input a feature set associated with 3D points of each reference 3D point cloud frame of the reference 3D point cloud frame set (i.e., F(t-1)...F(tN)). The feature set associated with 3D points of each reference 3D point cloud frame of the reference 3D point cloud frame set may be downsampled by at least one scaling factor. According to an embodiment, the feature set associated with 3D points of each reference 3D point cloud frame (having coordinates included in the reference coordinate information) may be downsampled by a scaling factor of 2, 3,..., K, etc. For example, F(t-1) may be downsampled by a scaling factor of 2, 3,..., K, etc., based on the multi-scale feature set F(t-1). 2 (t-1), F 3 (t-1), ..., F K (t-1), respectively. Similarly, F(tN) can be downsampled by a scaling factor of 2, 3, ..., K to generate the multi-scale feature set F(tN). 2 (tN), F 3 (tN),...,F K (tN).
[0067] For each scaling factor (i.e., 2, 3, ..., or K), the first neural network predictor 112 can generate a spatiotemporal tensor using feature sets associated with all reference 3D point cloud frames in the reference 3D point cloud frame set. For example, for a scaling factor of "2," F 2 (t-1), ..., and F 2 (tN) can be used to generate a space-time tensor. Similarly, for the scaling factor "K", F K (t-1), ..., F K(tN). According to one embodiment, the generation of the space-time tensor for a particular scaling factor can be based on the spatiotemporal concatenation of feature sets associated with all reference 3D point cloud frames of the reference 3D point cloud frame set downsampled by the same scaling factor. For example, the space-time tensor for a scaling factor of "2" can be generated as F 2 (t-1), ..., F 2 (tN). Similarly, the space-time tensor for the scaling factor "K" can be constructed as F K (t-1), ..., F K (tN) space-time tensors can be constructed based on the spatiotemporal concatenation of (tN). Thus, (K-1) space-time tensors can be constructed for scaling factors 2, 3, ..., and K. Each space-time tensor (of the (K-1) space-time tensors) for a particular scaling factor can represent the features of all the reference 3D point cloud frames that can be downsampled by the scaling factor.
[0068] The first neural network predictor 112 may further receive (target) coordinate information (i.e., C(t)) associated with 3D points in the current 3D point cloud frame (i.e., P(t)). C(t) may indicate coordinates in the current 3D point cloud frame that may include 3D points having features that need to be predicted (or encoded) based on feature sets associated with 3D points in each reference 3D point cloud frame (having coordinates included in the reference coordinate information). Each 3D point in the reference 3D point cloud frame having coordinates associated with the reference coordinate information may correspond to a 3D point in the current 3D point cloud frame having coordinates included in C(t). In some embodiments, a space-time tensor for each scaling factor may be constructed based on the feature sets associated with all reference 3D point cloud frames downsampled by the corresponding scaling factor and the current feature set associated with the 3D point in the current 3D point cloud frame. The coordinates of the 3D points are included in C(t), and the features in the current feature set associated with the 3D points may be padded with "1" for all scaling factors (since the features in the current feature set are to be predicted).
[0069] After constructing the space-time tensors, a space-time analysis can be performed on each of the (K-1) space-time tensors. According to an embodiment, the space-time analysis can include performing a sparse convolution operation or a self-attention operation on each of the space-time tensors (i.e., the (K-1) space-time tensors). The sparse convolution operation on the space-time tensors can be performed based on a filter. The filter can weight features represented by the space-time tensors in different spatial domains with the same weight value. The sparse convolution operation can modify each feature based on the corresponding feature, a set of neighboring features of the corresponding feature, and the weight values used to weight each of the corresponding feature and the neighboring features. The total number of neighboring features can be based on the size of the filter. Modifying each feature can generate a modified space-time tensor. Thus, performing a sparse operation on the (K-1) space-time tensors can generate (K-1) modified space-time tensors. Similarly, a self-attention operation on the space-time tensors (of the (K-1) space-time tensors) can be performed to modify the features represented by the space-time tensors. The modification can be based on a filter of a predetermined size that weights the features represented by the space-time tensors. The filter can modify each feature based on the corresponding feature, the weight of the corresponding feature, neighboring features of the corresponding feature, and the weight of each of the neighboring features. Modifying the features can generate modified space-time tensors. Thus, (K-1) modified space-time tensors can be generated based on applying a self-attention network to the (K-1) space-time tensors.
[0070] The features represented by each modified space-time tensor (generated based on sparse convolution or self-attention) can be downsampled by a specific scaling factor. The scaling factor by which the modified space-time tensor needs to be downsampled can be based on the scaling factors (such as 2, 3, ..., or K) by which the original version of the modified space-time tensor can be constructed and the highest scaling factor (i.e., K) by which the feature set of each reference 3D point cloud frame in the reference 3D point cloud frame set is downsampled. Then, an inception residual network can be applied to each downsampled modified space-time tensor to generate a final space-time tensor. Thus, (K-1) final space-time tensors can be generated as the output of the inception residual network.
[0071] The (K-1) final spatiotemporal tensors may be concatenated to generate a multi-scale feature connection vector. A sparse convolution operation or a self-attention operation may be performed on the multi-scale feature connection vector. Based on the results of the sparse convolution operation or the self-attention operation, the first neural network predictor 112 predicts a second set of features (i.e., F) associated with 3D points (having coordinates included in C(t)) of the current 3D point cloud frame (i.e., P(t)). ~ Therefore, the second feature set can be predicted further based on coordinate information (i.e., C(t)) associated with the 3D points of the current 3D point cloud frame.
[0072] The subtractor 402 subtracts the first feature set (i.e., F(t)) and the second feature set (i.e., F ~The PCC encoder 110 may receive as input the first feature set F(t) and the second feature set F(t). A residual feature set (i.e., R(t)) may be calculated based on the first feature set and the second feature set. The residual feature set is the difference between the actual feature set (i.e., F(t)) generated by the first PCC encoder 110 and the predicted feature set (i.e., F(t)) generated by the first neural network predictor 112. ~ The difference between the current 3D point cloud frame (i.e., P(t)) and the 3D points in P(t) can be calculated. This difference can be used to correct for possible errors in predicting features associated with the current 3D point cloud frame (i.e., P(t)). The predicted features can be used to reconstruct the current 3D point cloud frame (i.e., the 3D points in P(t) whose coordinates are contained in C(t)).
[0073] The quantizer 404 may receive the residual feature set (i.e., R(t)) and may quantize each residual feature in the residual feature set to a quantization level in a predetermined quantization level set. The quantizer 404 may generate a quantized residual feature set as an output based on application of a quantization scheme to the residual feature set. The quantization scheme may be based on an entropy model. The autoencoder 406 may receive the quantized residual feature set as an input and generate a bitstream of encoded point cloud data for the current 3D point cloud frame as an output. This generation may be based on application of an encoding scheme to the quantized residual feature set. The encoding scheme may be based on an entropy model and may compress each quantized residual feature in the quantized residual feature set to encode the quantized residual feature set.
[0074] According to an embodiment, the octree-based encoder 114 can receive the coordinate information (i.e., C(t)) as an input and encode the coordinate information to generate encoded coordinate information. The circuit 202 can include the encoded coordinate information in a bitstream of the generated encoded point cloud data and transmit the generated bitstream to the second electronic device 104.
[0075] The circuit 302 can receive the generated bitstream and extract the encoded coordinate information and the encoded point cloud data from the bitstream. The octree-based decoder 118 can receive the encoded coordinate information as input and decode the encoded coordinate information to recover the coordinate information (i.e., C(t)). The coordinate information can be related to a 3D point in the current 3D point cloud frame (i.e., P(t)) to be decoded. Based on the recovered coordinate information, a target coordinate in the current 3D point cloud frame containing the 3D point to be decoded can be determined.
[0076] The autodecoder 408 receives the encoded point cloud data as input and generates a residual feature set (i.e., R~ (t)) can be reproduced as an output. This reproduction can be based on applying a decoding scheme to the encoded point cloud data. The decoding scheme can be based on an entropy model and can decompress the quantized residual feature set. ~ (t)) can be used to reconstruct the 3D points (with coordinates contained in C(t)) of the current 3D point cloud frame (i.e., P(t)).
[0077] The second PCC encoder 116 may receive as input a 3D point cloud sequence that may include a reference 3D point cloud frame set, i.e., P(t-1), P(t-2), ..., and P(tN). Based on the input, reference frame data may be generated as output, including a feature set associated with a 3D point in each reference 3D point cloud frame of the reference 3D point cloud frame set. The reference frame data may be generated based on application of the second PCC encoder 116 to the reference 3D point cloud frame set. For example, the feature set associated with a 3D point in the reference 3D point cloud frame P(t-1) may be F(t-1). Similarly, the feature set associated with a 3D point in the reference 3D point cloud frame P(tN) may be F(tN). The feature set associated with a 3D point in each reference 3D point cloud frame of the reference 3D point cloud frame set may include reference features associated with the occupancy of the 3D point in the corresponding reference 3D point cloud frame of the reference 3D point cloud frame set and reference coordinate information associated with the 3D point in the corresponding reference 3D point cloud frame. These 3D points may correspond to 3D points in the current 3D point cloud frame (i.e., P(t)) whose coordinates are contained in the reconstructed coordinate information (i.e., C(t)).
[0078] The second neural network predictor 120 may receive as input a feature set associated with the 3D points of each reference 3D point cloud frame of the reference 3D point cloud frame set. The second neural network predictor 120 may then predict as output a third feature set (i.e., F'(t)) associated with the 3D points of the current 3D point cloud frame (i.e., P(t)). The 3D points of P(t) need to be decoded, and the coordinates of the 3D points may be those contained in C(t). The prediction may be based on application of the second neural network predictor 120 to the reference frame data, i.e., the feature set associated with the 3D points of each reference 3D point cloud frame. The third feature set may be predicted further based on the (re)generated coordinate information (i.e., C(t)). The generation of the third feature set may be identical to the generation of the second feature set by the first neural network predictor 112.
[0079] The accumulator 410 calculates the third feature set (i.e., F′(t)) and the recovered residual feature set (i.e., R ~ The PCC decoder 122 may receive as input the fourth feature set (i.e., F′(t)) associated with the 3D points (i.e., those having coordinates included in C(t)) of the current 3D point cloud frame (i.e., P(t)). The accumulator 410 may generate as output a fourth feature set (i.e., F′(t)) associated with the 3D points (i.e., those having coordinates included in C(t)) of the current 3D point cloud frame (i.e., P(t)). The generation may be based on accumulating the third feature set and the reconstructed residual feature set. The PCC decoder 122 may receive as input the fourth feature set and reconstruct the current 3D point cloud frame (i.e., P(t)) based on applying a decoding scheme to the fourth feature set. The PCC decoder 122 may generate as output a 3D point cloud frame (i.e., P′(t)) (which may correspond to P(t)).
[0080] FIG. 5 is a block diagram illustrating an example operation for predicting a feature set associated with a current 3D point cloud frame based on multi-scale features associated with a reference 3D point cloud frame set, according to an embodiment of the present disclosure. The description of FIG. 5 is provided with reference to elements in FIGS. 1, 2, 3, and 4. FIG. 5 illustrates an example block diagram 500 for predicting a feature set associated with a current 3D point cloud frame. The example block diagram 500 may include a series of blocks that may represent a series of operations starting at 502 and ending at 508. The series of operations may be performed by the first neural network predictor 112 or the second neural network predictor 120. The first neural network predictor 112 or the second neural network predictor 120 may receive two feature sets, namely, F(t-1) and F(t-2), as input. F(t-1) may be associated with 3D points in the reference 3D point cloud frame P(t-1). F(t-2) may be associated with 3D points in the reference 3D point cloud frame P(t-2). Both P(t-2) and P(t-1) can precede the current 3D point cloud frame P(t), which should be encoded or decoded based on predictions of features associated with the 3D points of P(t).
[0081] F(t-1) can include reference features related to the occupancy of the 3D points of P(t-1) and reference coordinate information related to the 3D points of P(t-1). For example, F(t-1) can include reference features f related to four 3D points of P(t-1). 11 , f 13 , f 14 and f 15 and the reference coordinates of the four 3D points. Meanwhile, F(t-2) can include reference features related to the occupancy of the 3D points of P(t-2) and reference coordinate information related to the 3D points of P(t-2). For example, F(t-2) can include reference features f related to the five 3D points of P(t-2). 23 , f 24 , f 25 , f 26 and f 28 and the reference coordinates of the five 3D points.
[0082] A feature set associated with the 3D points of P(t) can be constructed based on the (target) coordinate information C(t). C(t) can include the coordinates of the 3D points of P(t). The first neural network predictor 112 or the second neural network predictor 120 can receive C(t) as input. For example, C(t) can include the coordinates of five 3D points. These coordinates can be calculated as c 33 , c 34 , c 35 , c36 and c 39 The feature set can include features associated with the five 3D points of P(t). According to one embodiment, the feature set can be constructed by padding the features associated with the five 3D points with "1". The constructed feature set can be used to construct F(t), i.e., [1 33 ,1 34 ,1 35 ,1 36 ,1 39 ]) for encoding or decoding P(t). The first neural network predictor 112 or the second neural network predictor 120 calculates a feature set (i.e., F ) associated with the five 3D points of P(t) (having coordinates contained in C(t)) based on downsampling F(t-2), F(t-1), and F(t) by scaling factors of 2 and 3, respectively. ~ (t)) (e.g., [f 33 ,f 34 ,f 35 ,f 36 ,f 39 ]) can be predicted.
[0083] The reference features of F(t-2) (i.e., f 11 ,f 13 ,f 14 and f 15 ) is F 2 (t-2) and F 3 Similarly, the reference features of F(t-1) (i.e., f23 ,f 24 ,f 25 ,f 26 and f 28 ) is F 2 (t-1) and F 3 (t-1) can be downsampled by scaling factors of 2 and 3, respectively. Since F(t) is constructed by padding the 3D point features of P(t) with 1's, downsampling F(t) by scaling factors of 2 and 3 will return F(t).
[0084] In 502A, F(t), F 2 (t-1) and F 2 (t-2) can be spatiotemporally concatenated. This spatiotemporal concatenation can generate a first space-time tensor. The first space-time tensor can be constructed for a scaling factor of 2 and generated based on features of all reference 3D point cloud frames (i.e., P(t-2) and P(t-1)) downsampled by the scaling factor of 2. For example, the features represented by the first space-time tensor can include feature vectors such as [f11,0,0], [0,0,0], [f13,f23,1], [f14,f24,1], [f15,f25,1], [0,f26,1], [0,0,0], [0,f28,0], and [0,0,1]. Here, [0,0,0] indicates an empty region.
[0085] In 502B, F(t), F 3 (t-1) and F 3 (t-2) can be spatiotemporally concatenated. This spatiotemporal concatenation can generate a second spatiotemporal tensor. The second spatiotemporal tensor can be constructed for a scaling factor of 3 and generated based on features from all reference 3D point cloud frames (i.e., P(t-2) and P(t-1)) downsampled by a scaling factor of 3.
[0086] At 504A, a spatiotemporal analysis can be performed on the first space-time tensor. The spatiotemporal analysis can be based on a self-attention operation, and features represented by the first space-time tensor can be modified. The modification of each feature can be based on one or more of a size of a filter used to perform the self-attention operation, a weight assigned to the corresponding feature, and a weight assigned to each neighboring feature of the corresponding feature. A modified first space-time tensor can be generated based on the modification of each feature represented by the first space-time tensor.
[0087] At 504B, a spatiotemporal analysis can be performed on the second space-time tensor. The spatiotemporal analysis can be based on a self-attention operation, and can modify each feature represented by the second space-time tensor. A modified second space-time tensor can be generated based on the modification of each feature.
[0088] At 506, a multi-scale feature connection vector can be generated based on a spatio-temporal connection of the modified first space-time tensor and the modified second space-time tensor.
[0089] At 508, features associated with 3D points of the current 3D point cloud frame (i.e., P(t)) can be predicted. The prediction can be based on applying a sparse convolution operation or a self-attention operation to the generated multi-scale feature connection vector. The coordinates of the 3D points of the current 3D point cloud frame (for which features are predicted) can be those contained in C(t). Thus, the prediction can be further based on the coordinate information (i.e., C(t)). The predicted features can be expressed as f 33 , f 34 , f 35 , f 36 and f 39 Based on such predicted features, a 3D point (i.e., c 33 , c 34 , c 35 , c 36 and c 39 ) can be encoded or decoded.
[0090] FIG. 6 is a block diagram illustrating an example operation for predicting a feature set associated with a current 3D point cloud frame based on multi-scale features associated with a reference 3D point cloud frame set, according to an embodiment of the present disclosure. The description of FIG. 6 is provided with reference to elements in FIGS. 1, 2, 3, 4, and 5. FIG. 6 illustrates an example block diagram 600 for predicting a feature set associated with a current 3D point cloud frame (i.e., P(t)). The block diagram 600 may include a series of blocks that may represent a series of operations starting at 602 and ending at 608. The series of operations may be performed by the first neural network predictor 112 or the second neural network predictor 120. The first neural network predictor 112 or the second neural network predictor 120 may receive two feature sets, i.e., F(t-1) and F(t-2), as input. The first neural network predictor 112 or the second neural network predictor 120 predicts the 3D related feature set of P(t) (i.e., F) having coordinates contained in C(t) based on downsampling of F(t-2) and F(t-1) by scaling factors of "2" and "3", respectively. ~ (t)) can be predicted. F ~ (t) can be predicted for encoding or decoding P(t).
[0091] In 602A, F 2 (t-1) and F 2 (t-2) can be spatiotemporally concatenated. This spatiotemporal concatenation can generate a first spatiotemporal tensor. The first spatiotemporal tensor can be constructed for a scaling factor of 2 and generated based on features of all reference 3D point cloud frames (i.e., P(t-2) and P(t-1)) downsampled by a scaling factor of 2. For example, the features represented by the first spatiotemporal tensor can be [f11,0], [0,0], [f13,f 23 ], [f 14 ,f 24 ], [f15 ,f 25 ],[0,f 26 ],[0,0],[0,f 28 ], [0,0], etc., where [0,0,0] indicates an empty region.
[0092] In 602B, F 3 (t-1), F 3 (t-2) can be spatiotemporally concatenated. This spatiotemporal concatenation can generate a second spatiotemporal tensor. The second spatiotemporal tensor can be constructed for a scaling factor of 3 and can therefore be generated based on features from all reference 3D point cloud frames (i.e., P(t-2) and P(t-1)) downsampled by a scaling factor of 3.
[0093] At 604A, a space-time analysis may be performed on the first space-time tensor. In one embodiment, the space-time analysis may be based on a sparse convolution operation, and each feature represented by the first space-time tensor may be modified. The modification of each feature may be based on one or more of the size of a filter used to perform the sparse convolution operation and weights assigned to the corresponding feature and its neighboring features. A modified first space-time tensor may be generated based on the modification of each feature represented by the first space-time tensor. The modified first space-time tensor may then be downsampled to generate a downsampled modified first space-time tensor. The downsampled first space-time tensor may be passed as an input to a first Inception residual network, which generates a final space-time tensor as an output.
[0094] At 604B, a spatiotemporal analysis can be performed on the second space-time tensor. The spatiotemporal analysis can be based on a sparse convolution operation, and each feature represented by the second space-time tensor can be modified. A modified second space-time tensor can be generated based on the modification of each feature. The modified second space-time tensor can be passed as an input to a first Inception residual network, which generates a final space-time tensor as an output.
[0095] At 606, a multi-scale feature concatenation vector can be generated based on a spatiotemporal concatenation of the final spatiotemporal tensor generated by the first Inception residual network and the final spatiotemporal tensor generated by the second Inception residual network.
[0096] At 608, features associated with 3D points of the current 3D point cloud frame (i.e., P(t)) can be predicted. The prediction can be based on applying a sparse convolution operation or a self-attention operation to the generated multi-scale feature connection vector. The coordinates of the 3D points of the current 3D point cloud frame (for which features are predicted) can be those contained in C(t). Thus, the prediction can be further based on the coordinate information (i.e., C(t)). Based on the predicted features, the 3D points of P(t) can be encoded or decoded.
[0097] FIG. 7A illustrates an exemplary scenario for encoding or decoding a 3D point cloud frame based on a previous 3D point cloud frame, according to an embodiment of the present disclosure. The description of FIG. 7A is provided with reference to elements in FIGS. 1, 2, 3, 4, 5, and 6. FIG. 7A illustrates an exemplary scenario 700A. The exemplary scenario 700A illustrates an exemplary sequence of 3D point cloud frames, which may include an "I" 3D point cloud frame and a "P" 3D point cloud frame set. The "I" 3D point cloud frame can be encoded / decoded independently, while each "P" 3D point cloud frame can be encoded / decoded based on the previous 3D point cloud frame. For example, at time t0, an "I" 3D point cloud frame can be encoded / decoded. At time t1, a "P" 3D point cloud frame of the "P" 3D point cloud frame set can be encoded / decoded based on predicted features related to the occupancy of the 3D points in the "I" 3D point cloud frame. At time t2, another "P" 3D point cloud frame of the "P" 3D point cloud frame set can be encoded / decoded based on predicted features related to the occupancy of the 3D points of the "P" 3D point cloud frame encoded / decoded at time t1. Other "P" 3D point cloud frames of the "P" 3D point cloud frame set can be encoded / decoded in a similar manner.
[0098] Figure 7B illustrates an exemplary scenario for encoding or decoding a 3D point cloud frame based on two preceding 3D point cloud frames, according to an embodiment of the present disclosure. Figure 7B is described with reference to elements in Figures 1, 2, 3, 4, 5, 6, and 7A. Figure 7B illustrates an exemplary scenario 700B. The exemplary scenario 700B illustrates an exemplary sequence of 3D point cloud frames that may include two "I" 3D point cloud frames and a "P" 3D point cloud frame set. The "I" 3D point cloud frames can be encoded / decoded independently, while each "P" 3D point cloud frame can be encoded / decoded based on two preceding 3D point cloud frames.
[0099] For example, a first "I" 3D point cloud frame may be encoded or decoded at time t0, and a second "I" 3D point cloud frame may be encoded or decoded at time t1. At time t2, a "P" 3D point cloud frame of the "P" 3D point cloud frame set may be encoded / decoded based on predicted features associated with the 3D point occupancy of the first "I" 3D point cloud frame and predicted features associated with the 3D point occupancy of the second "I" 3D point cloud frame. At time t3, another "P" 3D point cloud frame of the "P" 3D point cloud frame set may be encoded / decoded based on predicted features associated with the 3D point occupancy of the second "I" 3D point cloud frame (encoded / decoded at time t1) and predicted features associated with the 3D point occupancy of the "P" 3D point cloud frame encoded / decoded at time t2. At time t4, another "P" 3D point cloud frame of the "P" 3D point cloud frame set can be encoded / decoded based on the predicted features associated with the occupancy of 3D points of the "P" 3D point cloud frame encoded / decoded at time t2 and the predicted features associated with the occupancy of 3D points of the "P" 3D point cloud frame encoded / decoded at time t2. Other "P" 3D point cloud frames of the "P" 3D point cloud frame set can be encoded / decoded in a similar manner.
[0100] FIG. 8 illustrates an exemplary scenario for encoding / decoding a 3D point cloud frame based on a preceding 3D point cloud frame and a subsequent 3D point cloud frame, according to an embodiment of the present disclosure. The description of FIG. 8 is provided with reference to elements in FIGS. 1, 2, 3, 4, 5, 6, 7A, and 7B. FIG. 8 illustrates an exemplary scenario 800. The exemplary scenario 800 illustrates an exemplary sequence of 3D point cloud frames that may include two "I" 3D point cloud frames, three "P" 3D point cloud frames, and three "B" 3D point cloud frames. The "I" 3D point cloud frames can be encoded / decoded independently. Each "P" 3D point cloud frame can be encoded / decoded based on the previous 3D point cloud frame, such as the "I" 3D point cloud frame or the "P" 3D point cloud frame. Each "B" 3D point cloud frame can be encoded / decoded based on the previous 3D point cloud frame (such as the "I" 3D point cloud frame or the "P" 3D point cloud frame) and the subsequent 3D point cloud frame (such as the "P" 3D point cloud frame).
[0101] For example, at time t0, a first "I" 3D point cloud frame (i.e., the first frame with frame index -0) can be encoded / decoded, and at time t1, a second "I" 3D point cloud frame (i.e., the second frame with frame index -1) can be encoded / decoded. At time t2, a first "P" 3D point cloud frame (i.e., the fourth frame with frame index -3) can be encoded or decoded based on predicted features related to the occupancy of 3D points in the second "I" 3D point cloud frame. Because the first "P" 3D point cloud frame is the third frame to be encoded / decoded (after the encoding / decoding of the first "I" 3D point cloud frame and the second "I" 3D point cloud frame), the encoding / decoding order of the first "P" 3D point cloud frame can be "2." At time t3, a first "B" 3D point cloud frame (i.e., the third frame with frame index -2) can be encoded or decoded based on predicted features related to the occupancy of 3D points of the second "I" 3D point cloud frame (i.e., the preceding frame) and the first "P" 3D point cloud frame (i.e., the subsequent frame). The encoding / decoding order of the first "B" 3D point cloud frame can be "3" (i.e., the fourth frame to be encoded or decoded). At time t4, a second "P" 3D point cloud frame (i.e., the sixth frame with frame index -5) can be encoded / decoded based on predicted features related to the occupancy of 3D points of the first "P" 3D point cloud frame. The encoding / decoding order of the second "P" 3D point cloud frame can be "4" (i.e., the fifth frame to be encoded or decoded). At time t5, a second "B" 3D point cloud frame (i.e., the fifth frame having frame index -4) can be encoded or decoded based on predicted features related to the occupancy of 3D points of the first "P" 3D point cloud frame (i.e., the preceding frame) and predicted features related to the occupancy of 3D points of the second "P" 3D point cloud frame (i.e., the subsequent frame). The encoding / decoding order of the second "B" 3D point cloud frame can be "5" (i.e., the sixth frame to be encoded or decoded).The third "P" 3D point cloud frame (i.e., the eighth frame with frame index -7) and the third "B" 3D point cloud frame (i.e., the seventh frame with frame index -6) can also be encoded / decoded in a similar manner.
[0102] Figure 9 is a flowchart illustrating operations of an exemplary method for multi-scale inter prediction for dynamic 3D point cloud frame compression, according to embodiments of the present disclosure. Figure 9 is described with reference to elements of Figures 1, 2, 3, 4, 5, 6, 7A, 7B, and 8. Figure 9 shows a flowchart 900. Operations 902-916 may be performed by any computer system, such as the first electronic device 102 or the circuitry 202 of the first electronic device 102. Operations may begin at 902 and proceed to 904.
[0103] At 904, a 3D point cloud sequence may be received that may include a reference 3D point cloud frame set and a current 3D point cloud frame to be encoded. In at least one embodiment, the circuit 202 may be configured to receive a 3D point cloud sequence that may include a reference 3D point cloud frame set and a current 3D point cloud frame to be encoded. Details of receiving the reference 3D point cloud frame set and the current 3D point cloud frame are described, for example, in Figures 1 and 4.
[0104] At 906, reference frame data can be generated that includes a set of features associated with the 3D points of each reference 3D point cloud frame of the set of reference 3D point cloud frames. In at least one embodiment, the circuit 202 can be configured to generate the reference frame data that includes a set of features associated with the 3D points of each reference 3D point cloud frame of the set of reference 3D point cloud frames. Details of generating the reference frame data are described, for example, in FIGS. 1 and 4.
[0105] At 908, current frame data related to the 3D points of the current 3D point cloud frame can be generated. In at least one embodiment, the circuit 202 can be configured to generate the current frame data related to the 3D points of the current 3D point cloud frame. The current frame data can include a first set of features related to the occupancy of the 3D points in the current 3D point cloud frame. Details of generating the current frame data are described, for example, in FIGS. 1 and 4.
[0106] At 910, a second set of features associated with the 3D points of the current 3D point cloud frame can be predicted based on application of the first neural network predictor to the reference frame data. In at least one embodiment, the circuit 202 can be configured to predict the second set of features associated with the 3D points of the current 3D point cloud frame based on application of the first neural network predictor to the reference frame data. Details of predicting the second set of features are described, for example, in Figures 1, 4, 5, and 6.
[0107] At 912, a residual feature set can be calculated based on the first feature set and the second feature set. In at least one embodiment, the circuit 202 can be configured to calculate the residual feature set based on the first feature set and the second feature set. Details of calculating the residual feature set are described, for example, in FIGS. 1 and 4.
[0108] At 914, a quantized residual feature set can be generated based on application of the quantization scheme to the residual feature set. In at least one embodiment, the circuit 202 can be configured to generate the quantized residual feature set based on application of the quantization scheme to the residual feature set. Details of generating the quantized residual feature set are described, for example, in FIGS. 1 and 4.
[0109] At 916, a bitstream of encoded point cloud data may be generated for the current 3D point cloud frame based on application of the encoding scheme to the quantized residual feature set. In at least one embodiment, the circuit 202 may be configured to generate a bitstream of encoded point cloud data for the current 3D point cloud frame based on application of the encoding scheme to the quantized residual feature set. Details of generating the bitstream of encoded point cloud data are described, for example, in Figures 1 and 4. Control may proceed to an end.
[0110] Although flowchart 700 depicts discrete operations such as 904, 906, 908, 910, 912, 914, and 916, the disclosure is not so limited. Thus, in some embodiments, such discrete operations may be further divided into additional operations, combined into fewer operations, or eliminated, depending on the implementation, without departing from the essence of the disclosed embodiments.
[0111] Figure 10 is a flowchart illustrating operations of an exemplary method for multi-scale inter-prediction for 3D point cloud frame reconstruction, according to embodiments of the present disclosure. Figure 10 is described with reference to elements in Figures 1, 2, 3, 4, 5, 6, 7A, 7B, 8, and 9. Figure 10 shows a flowchart 1000. Operations 1002-1014 can be performed by any computer system, such as the second electronic device 104 or the circuitry 302 of the second electronic device 104. Operations can begin at 1002 and proceed to 1004.
[0112] At 1004, a 3D point cloud sequence may be received, which may include a reference 3D point cloud frame set. In at least one embodiment, the circuit 302 may be configured to receive a 3D point cloud sequence, which may include a reference 3D point cloud frame set. Details of receiving the 3D point cloud sequence are described, for example, in FIGS. 1 and 4.
[0113] At 1006, reference frame data can be generated that includes a set of features associated with the 3D points of each reference 3D point cloud frame of the set of reference 3D point cloud frames. In at least one embodiment, the circuit 302 can be configured to generate the reference frame data that includes a set of features associated with the 3D points of each reference 3D point cloud frame of the set of reference 3D point cloud frames. Details of generating the reference frame data are described, for example, in FIGS. 1 and 4.
[0114] At 1008, a bitstream of encoded point cloud data associated with the current 3D point cloud frame to be decoded may be received. In at least one embodiment, the circuit 302 may be configured to receive a bitstream of encoded point cloud data associated with the current 3D point cloud frame to be decoded. Details of receiving the bitstream are described, for example, in Figures 1 and 4.
[0115] At 1010, a third set of features associated with the 3D points of the current 3D point cloud frame can be predicted based on application of the second neural network predictor to the reference frame data. In at least one embodiment, the circuit 302 can be configured to predict the third set of features associated with the 3D points of the current 3D point cloud frame based on application of the second neural network predictor to the reference frame data. Details of predicting the third set of features are described, for example, in Figures 1, 4, 5, and 6.
[0116] At 1012, a fourth set of features associated with the 3D points of the current 3D point cloud frame can be generated based on the received bitstream of encoded point cloud data and the predicted third set of features. In at least one embodiment, the circuit 302 can be configured to generate the fourth set of features associated with the 3D points of the current 3D point cloud frame based on the received bitstream of encoded point cloud data and the predicted third set of features. Details of generating the fourth set of features are described, for example, in Figures 1 and 4.
[0117] At 1014, the current 3D point cloud frame may be reconstructed based on application of the decoding scheme to the determined fourth feature set. In at least one embodiment, the circuit 302 may be configured to reconstruct the current 3D point cloud frame based on application of the decoding scheme to the determined fourth feature set. Details of the reconstruction of the current 3D point cloud frame are described, for example, in Figures 1 and 4. Control may proceed to an end.
[0118] Although flowchart 1000 depicts discrete operations such as 1004, 1006, 1008, 1010, 1012, and 1014, the disclosure is not so limited. Thus, in some embodiments, such discrete operations may be further divided into additional operations, combined into fewer operations, or eliminated, depending on the implementation, without departing from the essence of the disclosed embodiments.
[0119] An example embodiment of the present disclosure may include an electronic device (such as the first electronic device 102 of FIG. 1 ) that may include a circuit (such as the circuit 202 of FIG. 2 ) that may be communicatively coupled to another electronic device (such as the second electronic device 104 of FIG. 1 ). The first electronic device 102 may further include a memory (such as the memory 204 of FIG. 2 ) that may be configured to store a predictor (such as the first neural network predictor 112 of FIG. 1 ). The memory 204 may be configured to store a PCC encoder (such as the first PCC encoder 110 of FIG. 1 ). The circuit 202 may be configured to receive a 3D point cloud sequence that may include a reference 3D point cloud frame set and a current 3D point cloud frame to be encoded. The reference 3D point cloud frame set may include at least one reference 3D point cloud frame that may precede the current 3D point cloud frame or at least one reference 3D point cloud frame that may follow the current 3D point cloud frame. The circuit 202 may be further configured to generate reference frame data including a set of features associated with 3D points in each reference 3D point cloud frame of the reference 3D point cloud frame set. The reference frame data may be generated based on application of the first PCC encoder 110 to the reference 3D point cloud frame set. The set of features associated with 3D points in each reference 3D point cloud frame of the reference 3D point cloud frame set may include reference features associated with the occupancy of the 3D points in the corresponding reference 3D point cloud frame of the reference 3D point cloud frame set and reference coordinate information associated with the 3D points in the corresponding reference 3D point cloud frame. The circuit 202 may be further configured to generate current frame data associated with 3D points in a current 3D point cloud frame. The current frame data may be generated based on application of the first PCC encoder 110 to the current 3D point cloud frame. The current frame data may include a first set of features associated with the occupancy of the 3D points in the current 3D point cloud frame. The circuit 202 may be further configured to predict a second set of features associated with the 3D points of the current 3D point cloud frame based on application of the first neural network predictor 112 to the reference frame data.The second feature set can be predicted further based on coordinate information associated with 3D points of the current 3D point cloud frame. The circuit 202 can be further configured to calculate a residual feature set based on the first feature set and the second feature set. The circuit 202 can be further configured to generate a quantized residual feature set based on application of a quantization scheme to the residual feature set. The circuit 202 can be further configured to generate a bitstream of encoded point cloud data for the current 3D point cloud frame based on application of an encoding scheme to the quantized residual feature set.
[0120] According to an embodiment, the circuit 202 may be further configured to encode the coordinate information based on application of an octree-based encoder (such as the octree-based encoder 114) to the coordinate information. The bitstream of encoded point cloud data may include the encoded coordinate information.
[0121] According to an embodiment, the circuit 202 may be further configured to downsample the feature set associated with the 3D points of each reference 3D point cloud frame of the set of reference 3D point cloud frames by at least one scaling factor.
[0122] An exemplary embodiment of the present disclosure may include an electronic device (such as the second electronic device 104 of FIG. 1 ) that may include a circuit (such as the circuit 302 of FIG. 3 ) communicatively coupleable to another electronic device (such as the first electronic device 102 of FIG. 1 ). The second electronic device 104 may further include a memory (such as the memory 304 of FIG. 3 ) that may be configured to store a predictor (such as the second neural network predictor 120 of FIG. 1 ). The memory 204 may be configured to store a PCC encoder (such as the second PCC encoder 116 of FIG. 1 ) and a PCC decoder (such as the PCC decoder 122 of FIG. 1 ). The circuit 302 may be configured to receive a 3D point cloud sequence that may include a reference 3D point cloud frame set. The circuit 302 may further be configured to generate reference frame data that includes a feature set associated with 3D points in each reference 3D point cloud frame of the reference 3D point cloud frame set. The reference frame data may be generated based on application of the second PCC encoder 116 to the reference 3D point cloud frame set. The feature set associated with the 3D points of each reference 3D point cloud frame of the reference 3D point cloud frame set may include reference features associated with the occupancy of the 3D points in a corresponding reference 3D point cloud frame of the reference 3D point cloud frame set and reference coordinate information associated with the 3D points of the corresponding reference 3D point cloud frame. The circuit 302 may be further configured to receive a bitstream of encoded point cloud data associated with the current 3D point cloud frame to be decoded. The received bitstream of encoded point cloud data may further include encoded coordinate information associated with the 3D points of the current 3D point cloud frame. The circuit 302 may be further configured to predict a third feature set associated with the 3D points of the current 3D point cloud frame based on application of the second neural network predictor 120 to the reference frame data. The circuit 302 may be further configured to generate a fourth feature set associated with the 3D points of the current 3D point cloud frame based on the received bitstream of encoded point cloud data and the predicted third feature set.The circuit 302 may be further configured to reconstruct the current 3D point cloud frame based on application of the decoding scheme to the determined fourth feature set.
[0123] According to an embodiment, the circuit 302 may be further configured to generate coordinate information associated with the 3D points of the current 3D point cloud frame based on application of an octree-based decoder (such as the octree-based decoder 118) to the encoded coordinate information. The third set of features may be predicted further based on the generated coordinate information.
[0124] The present disclosure can be implemented in hardware or a combination of hardware and software. The present disclosure can be implemented in a centralized manner in at least one computer system, or in a distributed manner where different elements can be distributed across several interconnected computer systems. Any computer system or other device adapted to perform the methods described herein can be suitable. The combination of hardware and software can be a general-purpose computer system that includes a computer program that, when loaded and executed, can control the computer system to perform the methods described herein. The present disclosure can be implemented in hardware, including portions of integrated circuits that also perform other functions.
[0125] The present disclosure may also be embodied in a computer program product, which includes all features that enable the implementation of the methods described herein and which is capable of executing these methods when loaded into a computer system. A computer program in this context means any expression, in any language, code or notation, of a set of instructions intended to cause a system having information processing capabilities to perform a particular function, either directly, or after a) conversion into another language, code or notation, or b) reproduction in a different content form, or both.
[0126] While the present disclosure has been described with reference to several embodiments, those skilled in the art will recognize that various modifications may be made and equivalents may be substituted without departing from the scope of the disclosure. Additionally, many modifications may be made to adapt a particular situation or material to the teachings of the disclosure without departing from the scope of the disclosure. Therefore, it is not intended that the disclosure be limited to the particular embodiments disclosed, but rather, it is intended to include all embodiments falling within the scope of the appended claims. [Explanation of symbols]
[0127] 102 First Electronic Device 104 Second Electronic Device 110 First PCC Encoder 112 First NN predictor 114 Octree-based Encoder 116 Second PCC Encoder 118 Octree-based decoder 120 Second NN predictor 122 PCC decoder 402 Subtractor 404 Quantizer 406 Autoencoder 408 Autodecoder 410 Accumulator
Claims
1. a first electronic device, a memory configured to store a first neural network predictor; The circuit and wherein the circuit comprises: receiving a three-dimensional (3D) point cloud sequence including a reference 3D point cloud frame set and a current 3D point cloud frame to be encoded; generating reference frame data including a set of features associated with 3D points in each reference 3D point cloud frame of the reference 3D point cloud frame set; generating current frame data including a first set of features associated with 3D points of the current 3D point cloud frame and related to occupancy of the 3D points in the current 3D point cloud frame; predicting a second set of features associated with the 3D points of the current 3D point cloud frame based on application of the first neural network predictor to the reference frame data; calculating a residual feature set based on the first feature set and the second feature set; generating a quantized residual feature set based on applying a quantization scheme to the residual feature set; generating a bitstream of encoded point cloud data for the current 3D point cloud frame based on applying an encoding scheme to the quantized residual feature set. It is configured as follows: A first electronic device comprising:
2. the second set of features is predicted further based on coordinate information associated with the 3D points of the current 3D point cloud frame. The first electronic device of claim 1 .
3. the circuitry is further configured to encode the coordinate information based on application of an octree-based encoder to the coordinate information; the bit stream of encoded point cloud data includes the encoded coordinate information; The first electronic device of claim 2 .
4. the memory is further configured to store a first point cloud compression (PCC) encoder; the reference frame data is generated based on application of the first PCC encoder to the reference 3D point cloud frame set; the current frame data is generated based on application of the first PCC encoder to the current 3D point cloud frame. The first electronic device of claim 1 .
5. The set of features associated with a 3D point of each reference 3D point cloud frame of the reference 3D point cloud frame set comprises: a reference feature related to the occupancy of the 3D points in a corresponding reference 3D point cloud frame of the set of reference 3D point cloud frames; Reference coordinate information associated with the 3D points of the corresponding reference 3D point cloud frame; The first electronic device of claim 1 , comprising:
6. the set of reference 3D point cloud frames includes at least one reference 3D point cloud frame preceding the current 3D point cloud frame or at least one reference 3D point cloud frame following the current 3D point cloud frame; The first electronic device of claim 1 .
7. the circuitry is further configured to downsample the feature set associated with a 3D point in each reference 3D point cloud frame of the set of reference 3D point cloud frames by at least one scaling factor. The first electronic device of claim 1 .
8. a second electronic device, a memory configured to store a second neural network predictor; The circuit and wherein the circuit comprises: receiving a three-dimensional (3D) point cloud sequence comprising a reference 3D point cloud frame set; generating reference frame data including a set of features associated with 3D points in each reference 3D point cloud frame of the reference 3D point cloud frame set; receiving a bitstream of encoded point cloud data associated with a current 3D point cloud frame to be decoded; predicting a third set of features associated with 3D points of the current 3D point cloud frame based on application of the second neural network predictor to the reference frame data; generating a fourth set of features associated with the 3D points of the current 3D point cloud frame based on the received bitstream of encoded point cloud data and the predicted third set of features; reconstructing the current 3D point cloud frame based on application of a decoding scheme to the determined fourth feature set. It is configured as follows: A second electronic device.
9. the received bitstream of encoded point cloud data further includes encoded coordinate information associated with the 3D points of the current 3D point cloud frame. The second electronic device of claim 8 .
10. the circuitry is further configured to generate coordinate information associated with the 3D points of the current 3D point cloud frame based on application of an octree-based decoder to the encoded coordinate information. The second electronic device of claim 9 .
11. the third set of features is predicted further based on the generated coordinate information. The second electronic device of claim 10.
12. the memory is further configured to store a point cloud compression (PCC) encoder; the reference frame data is generated based on application of the PCC encoder to the reference 3D point cloud frame set; The second electronic device of claim 8 .
13. The set of features associated with a 3D point of each reference 3D point cloud frame of the reference 3D point cloud frame set comprises: a reference feature related to the occupancy of the 3D points in a corresponding reference 3D point cloud frame of the set of reference 3D point cloud frames; Reference coordinate information associated with the 3D points of the corresponding reference 3D point cloud frame; The second electronic device of claim 8 , comprising:
14. In the first electronic device, receiving a three-dimensional (3D) point cloud sequence including a reference 3D point cloud frame set and a current 3D point cloud frame to be encoded; generating reference frame data including a set of features associated with 3D points in each reference 3D point cloud frame of the reference 3D point cloud frame set; generating current frame data including a first set of features associated with 3D points of the current 3D point cloud frame and related to occupancy of the 3D points in the current 3D point cloud frame; predicting a second set of features associated with the 3D points of the current 3D point cloud frame based on application of a first neural network predictor to the reference frame data; and calculating a residual feature set based on the first feature set and the second feature set; generating a quantized residual feature set based on application of a quantization scheme to the residual feature set; generating a bitstream of encoded point cloud data for the current 3D point cloud frame based on applying an encoding scheme to the quantized residual feature set; A method comprising:
15. the second set of features is predicted further based on coordinate information associated with the 3D points of the current 3D point cloud frame.
15. The method of claim 14.
16. encoding the coordinate information based on application of an octree-based encoder to the coordinate information; the bit stream of encoded point cloud data includes the encoded coordinate information; 16. The method of claim 15.
17. the reference frame data is generated based on application of a first point cloud compression (PCC) encoder to the reference 3D point cloud frame set; the current frame data is generated based on application of the first PCC encoder to the current 3D point cloud frame.
15. The method of claim 14.
18. The set of features associated with a 3D point of each reference 3D point cloud frame of the reference 3D point cloud frame set comprises: a reference feature related to the occupancy of the 3D points in a corresponding reference 3D point cloud frame of the set of reference 3D point cloud frames; Reference coordinate information associated with the 3D points of the corresponding reference 3D point cloud frame; 15. The method of claim 14, comprising:
19. the set of reference 3D point cloud frames includes at least one reference 3D point cloud frame preceding the current 3D point cloud frame or at least one reference 3D point cloud frame following the current 3D point cloud frame; 15. The method of claim 14.
20. downsampling the feature set associated with the 3D points of each reference 3D point cloud frame of the reference 3D point cloud frame set by at least one scaling factor.
15. The method of claim 14.
Citation Information
Patent Citations
Point cloud geometric compression method based on deep convolutional network
CN110691243A