Point cloud encoding and decoding method, device and electronic device
Through the hybrid codec order, combined with depth priority and breadth priority codec order, the problem of low traversal efficiency of node traversal in the octree structure is solved, and efficient parallel processing and rapid reconstruction of point cloud codec is realized.
Patent Information
- Application Number
- CN202080038882.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-05
- Filing Date
- 2020-10-07
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-10-07
AI Technical Summary
In the existing point cloud encoding and decoding technology, the node traversal efficiency of the octree structure is low, resulting in low encoding and decoding efficiency.
Using a hybrid codec order, combining depth-first and breadth-first codec order, the nodes in the octree structure are reconstructed and point clouds are reconstructed by decoding the nodes in the octree structure in parallel, especially allowing the child nodes of at least one node to decode without waiting for the decoding of nodes of the same size.
It improves the efficiency of point cloud encoding and decoding, realizes parallel processing of nodes and faster reconstruction processes.
Smart Images

Figure CN113892128B_ABST
Abstract
Description
[0001] Reference added
[0002] This application claims priority to U.S. Patent Application No. 17 / 063,411, "METHOD AND APPARATUS FOR POINT CLOUD CODING," filed on October 5, 2020, which claims priority to U.S. Provisional Patent Application No. 62 / 942,549, "HYBRID CODING ORDER FOR POINT CLOUD CODING," filed on December 2, 2019. The entire contents of the two prior applications are incorporated herein by reference. Technical Field
[0003] This application relates to point cloud coding and decoding, and in particular to a method, apparatus, and electronic device for point cloud coding and decoding. Background Art
[0004] The background art description provided herein is for the purpose of generally presenting the context of the present application. To the extent described in this background art section, the work of the currently named inventors, as well as aspects that may not qualify as prior art at the time of filing of the present application, are neither expressly nor implicitly considered prior art to the present application.
[0005] Various technologies have been developed to capture the world and represent it in three-dimensional (3D) space, e.g., objects in the world, environments in the world, etc. A 3D representation of the world can enable more immersive forms of interaction and communication. A point cloud can be used as a 3D representation of the world. A point cloud is a set of points in 3D space, each point having associated attributes such as color, material properties, texture information, intensity attributes, reflectivity attributes, motion-related attributes, morphological attributes, and various other attributes. Such point clouds may include a large amount of data, and their storage and transmission can be costly and time-consuming.
[0006] In point cloud coding and decoding (PCC) technology, an octree structure is used. However, existing schemes for traversing the nodes of the octree structure for coding and decoding still need to be improved in terms of coding and decoding efficiency. Summary of the Invention
[0007] Aspects of the present application provide a point cloud encoding and decoding method and apparatus. In some examples, the point cloud encoding and decoding apparatus includes a processing circuit. In some embodiments, the processing circuit receives encoded data of a node in an octree structure of a point cloud from an encoded bitstream of the point cloud. A node in the octree structure corresponds to a three-dimensional (3D) partition of the space of the point cloud, and the size of the node is associated with the size of the corresponding 3D partition. In addition, the processing circuit decodes an occupancy code of the node from the encoded data. The decoding of the first occupancy code of at least the children nodes of the first node does not need to wait for the decoding of the second occupancy code of the second node. The second node has the same node size as the first node. Then, the processing circuit reconstructs the octree structure based on the decoded occupancy code of the node; and reconstructs the point cloud based on the octree structure.
[0008] In some embodiments, the processing circuit decodes a first set of occupancy codes of a first set of nodes in a first sub-octree, where the first node is the root of the first sub-octree; and decodes a second set of occupancy codes of a second set of nodes in a second sub-octree, where the second node is the root of the second sub-octree. In one embodiment, the processing circuit decodes the first set of occupancy codes of the first set of nodes in the first sub-octree in parallel with decoding the second set of occupancy codes of the second set of nodes in the second sub-octree.
[0009] In another embodiment, the processing circuit decodes the first set of occupancy codes of the first set of nodes in the first sub-octree using a first encoding and decoding mode; and decodes the second set of occupancy codes of the second set of nodes in the second sub-octree using a second encoding and decoding mode. In one example, the processing circuit decodes a first index indicating the first encoding and decoding mode of the first sub-octree from the encoded bitstream; and decodes a second index indicating the second encoding and decoding mode of the second sub-octree from the encoded bitstream.
[0010] In some embodiments, for a larger node among the nodes, the processing circuit uses a first encoding and decoding order to decode a first part of the occupancy code. The larger node is larger than a specific node size for changing the encoding and decoding order; and for a smaller node among the nodes, the processing circuit uses a second encoding and decoding order different from the first encoding and decoding order to decode a second part of the occupancy code. The smaller node is equal to or smaller than the specific node size for changing the encoding and decoding order. In one example, the first encoding and decoding order is a breadth-first encoding and decoding order, and the second encoding and decoding order is a depth-first encoding and decoding order. In another example, the first encoding and decoding order is a depth-first encoding and decoding order, and the second encoding and decoding order is a breadth-first encoding and decoding order.
[0011] In some examples, the processing circuit determines a specific node size for changing the encoding / decoding order based on signals in the encoded bitstream of the point cloud. In one example, the processing circuit decodes a control signal from the encoded bitstream of the point cloud, and the control signal indicates a change in the encoding / decoding order. Then, the processing circuit decodes the signal and determines a specific node size for changing the encoding / decoding order.
[0012] Aspects of the present application also provide a non - volatile computer - readable storage medium storing instructions that, when executed by a point cloud encoding / decoding computer, cause the computer to execute any one or a combination of the point cloud encoding / decoding methods.
[0013] Aspects of the present application provide a point cloud encoding / decoding method, including: receiving encoded data of a node in an octree structure of a point cloud from an encoded bitstream of the point cloud, where the node in the octree structure corresponds to a three - dimensional (3D) partition of the space of the point cloud, and the size of the node is associated with the size of the corresponding 3D partition; decoding an occupancy code of the node from the encoded data, where at least the decoding of the first occupancy code of the child nodes of the first node does not need to wait for the decoding of the second occupancy code of the second node, and the second node has the same node size as the first node; reconstructing the octree structure based on the decoded occupancy code of the node; and reconstructing the point cloud based on the octree structure.
[0014] Aspects of the present application also provide an electronic device, including a memory for storing computer - readable instructions; and a processor for reading the computer - readable instructions and executing any one or a combination of the point cloud encoding / decoding methods according to the instructions of the computer - readable instructions.
[0015] Aspects of the present application provide a point cloud encoding / decoding apparatus, including: a receiving unit for receiving an encoded occupancy code of a node in an octree structure of the point cloud from an encoded bitstream of the point cloud, where the node in the octree structure corresponds to a three - dimensional (3D) partition of the space of the point cloud, and the size of the node is associated with the size of the corresponding 3D partition; a decoding unit for decoding the occupancy code of the node from the encoded occupancy code, where at least the decoding of the first occupancy code of the child nodes of the first node does not need to wait for the decoding of the second occupancy code of the second node, and the second node has the same node size as the first node; a reconstruction unit for reconstructing the octree structure based on the decoded occupancy code of the node and reconstructing the point cloud based on the octree structure.
[0016] The point cloud encoding / decoding method, apparatus, electronic device, and storage medium provided by aspects of the present application use a technology of a hybrid encoding / decoding order, which can fully realize the parallel encoding / decoding of nodes in the octree structure and improve the encoding / decoding efficiency. Description of the Drawings
[0017] In conjunction with the following detailed description and the accompanying drawings, other features, natures, and various advantages of the subject matter of the present application will become more apparent, where:
[0018] Figure 1 is a simplified block diagram schematic of a communication system according to an embodiment;
[0019] Figure 2 is a simplified block diagram schematic of a streaming system according to an embodiment;
[0020] Figure 3 shows an encoder block diagram for encoding a point cloud frame according to some embodiments;
[0021] Figure 4 shows a decoder block diagram for decoding a compressed bitstream corresponding to a point cloud frame according to some embodiments;
[0022] Figure 5 is a simplified block diagram schematic of a video decoder according to an embodiment;
[0023] Figure 6 is a simplified block diagram schematic of a video encoder according to an embodiment;
[0024] Figure 7 shows an encoder block diagram for encoding a point cloud frame according to some embodiments;
[0025] Figure 8 shows a block diagram of a decoder for decoding a compressed bitstream corresponding to a point cloud frame according to some embodiments;
[0026] Figure 9 shows an illustration of cube partitioning based on octree partitioning technology according to some embodiments of the present application.
[0027] Figure 10 shows an example of an octree partitioning and an octree structure corresponding to the octree partitioning according to some embodiments of the present application.
[0028] Figure 11 shows a diagram of an octree structure illustrating a breadth-first encoding / decoding order.
[0029] Figure 12 shows a diagram of an octree structure illustrating a depth-first encoding / decoding order.
[0030] Figure 13 shows a syntax example of a geometric parameter set according to some embodiments of the present application.
[0031] Figure 14 shows another syntax example of a geometric parameter set according to some embodiments of the present application.
[0032] Figure 15 Shows a pseudo-code example for octree encoding and decoding according to some embodiments of the present application.
[0033] Figure 16 Shows a pseudo-code example for depth-first encoding and decoding order according to some embodiments of the present application.
[0034] Figure 17 Shows a flowchart outlining an example method according to some embodiments.
[0035] Figure 18 Is a schematic diagram of a computer system according to an embodiment. Detailed Description
[0036] Aspects of the present application provide point cloud coding and decoding (PCC) techniques. PCC can be performed according to various schemes, for example, a geometry-based scheme (referred to as G-PCC), a video coding and decoding-based scheme (referred to as V-PCC), etc. According to some aspects of the present application, G-PCC directly encodes 3D geometric structures and is a purely geometry-based method with little shared content with video coding and decoding, while V-PCC is mainly based on video coding and decoding. For example, V-PCC can map the points of a 3D cloud to the pixels of a 2D grid (image). The V-PCC scheme can utilize a general video codec for point cloud compression. The Moving Picture Experts Group (MPEG) is researching G-PCC standards and V-PCC standards that use G-PCC schemes and V-PCC schemes respectively.
[0037] Aspects of the present application provide techniques for a hybrid coding and decoding order that can be used in PCC (e.g., G-PCC schemes and V-PCC schemes). The hybrid coding encoder can include a depth-first traversal scheme and a breadth-first traversal scheme in the coding and decoding order. The present application also provides techniques for signaling the coding and decoding order.
[0038] Point clouds can be widely used in many applications. For example, point clouds can be used for object detection and positioning in autonomous driving vehicles; point clouds can be used for mapping in a Geographic Information System (GIS), and can be used for visualizing and archiving cultural heritage objects and collections in cultural heritage, etc.
[0039] In the following, a point cloud generally refers to a set of points in 3D space, where each point has associated attributes such as color, material, texture information, intensity attributes, reflectivity attributes, motion-related attributes, modality attributes, and various other attributes. The point cloud can be used to reconstruct an object or a scene as a combination of these points. These points can be acquired using multiple cameras, depth sensors, or lidar in various settings, and can consist of thousands to billions of points in order to realistically represent the reconstructed scene. A patch generally refers to an adjacent subset of the surface described by the point cloud. In one example, a patch includes points whose deviation of surface normal vectors from each other is less than a threshold amount.
[0040] Compression techniques can reduce the amount of data required to represent a point cloud, so as to enable faster transmission or reduced storage. Thus, techniques for lossy compression of point clouds are needed for real-time communication and six degrees of freedom (6DoF) virtual reality. Additionally, techniques for lossless point cloud compression in dynamic mapping contexts such as autonomous driving and cultural heritage applications are also being sought.
[0041] According to one aspect of the present application, the main idea behind V-PCC is to use existing video codecs to compress the geometry, occupancy, and texture of a dynamic point cloud into three separate video sequences. The additional metadata required to explain these three video sequences is compressed separately. A small portion of the entire bitstream is metadata, which can be efficiently encoded / decoded using a software implementation. Most of the information is processed by the video codec.
[0042] Figure 1 A simplified block diagram of a communication system (100) according to an embodiment of the present application is shown. The communication system (100) includes a plurality of terminal devices that can communicate with each other via, for example, a network (150). For example, the communication system (100) includes a pair of terminal devices (110) and (120) interconnected via a network (150). In Figure 1 an example, the first pair of terminal devices (110) and (120) can perform unidirectional transmission of point cloud data. For example, the terminal device (110) can compress the point cloud (e.g., points representing a structure) acquired by a sensor (105) connected to the terminal device (110). The compressed point cloud can be sent via the network (150), for example, in the form of a bitstream, to another terminal device (120). The terminal device (120) can receive the compressed point cloud from the network (150), decompress the bitstream to reconstruct the point cloud, and appropriately display the reconstructed point cloud. Unidirectional data transmission can be common in media service applications and the like.
[0043] In Figure 1In the example, the terminal device (110) and the terminal device (120) may be illustrated as a server and a personal computer, but the principles of this application are not limited thereto. The embodiments of this application are applicable to laptop computers, tablet computers, smart phones, game terminals, media players, and / or dedicated three-dimensional (3D) devices. The network (150) represents any number of networks capable of transmitting compressed point clouds between the terminal device (110) and the terminal device (120). The network (150) may include, for example, wired and / or wireless communication networks. The network (150) may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of the discussion of this application, the architecture and topology of the network (150) may not be important for the operation of this application, unless otherwise explained below.
[0044] Figure 2 A simplified block diagram of a streaming system (200) according to an embodiment is illustrated. Figure 2 The example is an application of the subject matter disclosed in this application to point clouds. The subject matter disclosed in this application is equally applicable to other applications implemented with point clouds, such as, for example, 3D telepresence applications, virtual reality applications, and the like.
[0045] The streaming system (200) may include an acquisition subsystem (213). The acquisition subsystem (213) may include a point cloud source (201), such as a light detection and ranging (LIDAR) system, a 3D camera, a 3D scanner, a graphics generation component that generates an uncompressed point cloud in software form, etc., to generate components such as an uncompressed point cloud (202). In one example, the point cloud (202) includes points collected by a 3D camera. The point cloud (202) is depicted as a thick line to emphasize that it has a higher data volume compared to the compressed point cloud (204) (the bitstream of the compressed point cloud). The compressed point cloud (204) may be generated by an electronic device (220) that includes an encoder (203) coupled to the point cloud source (201). The encoder (203) may include hardware, software, or a combination thereof to implement or carry out aspects of the subject matter disclosed in this application described in more detail below. The compressed point cloud (204) (or the bitstream (204) of the compressed point cloud) is depicted as a thin line to emphasize that it has a lower data volume compared to the point cloud stream (202), and it may be stored on the streaming server (205) for future use. One or more streaming client subsystems, such as Figure 2The client subsystems (206) and client subsystems (208) therein can access the streaming server (205) to retrieve copies (207) and copies (209) of the compressed point cloud (204). The client subsystem (206) can include, for example, a decoder (210) in an electronic device (230). The decoder (210) decodes the incoming copy (207) of the compressed point cloud and generates an output stream (211) of the reconstructed point cloud that can be presented on a rendering device (212).
[0046] Note that the electronic devices (220) and (230) can include other components (not shown). For example, the electronic device (220) can include a decoder (not shown), and the electronic device (230) can also include an encoder (not shown).
[0047] In some streaming systems, the compressed point cloud (204), compressed point cloud (207), and compressed point cloud (209) (e.g., the bitstream of the compressed point cloud) can be compressed according to specific standards. In some examples, video coding standards are used for the compression of point clouds. Examples of these standards include High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), and so on.
[0048] Figure 3 A block diagram of a V-PCC encoder (300) for encoding point cloud frames according to some embodiments is shown. In some embodiments, the V-PCC encoder (300) can be used in a communication system (100) and a streaming system (200). For example, the encoder (203) can be configured and operated in a manner similar to the V-PCC encoder (300).
[0049] The V-PCC encoder (300) receives a point cloud frame as an uncompressed input and generates a bitstream corresponding to the compressed point cloud frame. In some embodiments, the V-PCC encoder (300) can receive a point cloud frame from a point cloud source such as a point cloud source (201).
[0050] In Figure 3 the example, the V-PCC encoder (300) includes a patch generation module (306), a patch packing module (308), a geometric image generation module (310), a texture image generation module (312), a patch information module (304), an occupancy map module (314), a smoothing module (336), image padding modules (316) and (318), a group dilation module (320), video compression modules (322), (323) and (332), an auxiliary patch information compression module (338), an entropy compression module (334), and a multiplexer (324).
[0051] According to one aspect of the present application, a V-PCC encoder (300) converts a 3D point cloud frame together with some metadata (e.g., occupancy map and patch information) for converting the compressed point cloud back to the decompressed point cloud into an image-based representation. In some examples, the V-PCC encoder (300) may convert the 3D point cloud frame into a geometry image, a texture image, and an occupancy map, and then use video coding techniques to encode the geometry image, the texture image, and the occupancy map into a bitstream. Generally, a geometry image is a 2D image whose pixels are filled with geometric values associated with the points projected onto the pixels, and the pixels filled with geometric values can be called geometric samples. A texture image is a 2D image whose pixels are filled with texture values associated with the points projected onto the pixels, and the pixels filled with texture values can be called texture samples. An occupancy map is a 2D image whose pixels are filled with values indicating whether a patch occupies or does not occupy.
[0052] A patch generation module (306) divides the point cloud into a set of patches (e.g., a patch is defined as an adjacent subset of the surface described by the point cloud), and the set of patches may or may not overlap, such that each patch can be described by a depth field relative to a plane in 2D space. In some embodiments, the patch generation module (306) aims to decompose the point cloud into the minimum number of patches with smooth boundaries while also minimizing the reconstruction error.
[0053] A patch information module (304) may collect patch information indicating the size and shape of the patches. In some examples, the patch information can be packed into an image frame and then encoded by an auxiliary patch information compression module (338) to generate compressed auxiliary patch information.
[0054] A patch packing module (308) is configured to map the extracted patches onto a two-dimensional (2D) grid while minimizing the unused space and ensuring that each M×M (e.g., 16×16) block of the grid is associated with a unique patch. Effective patch packing can directly affect the compression efficiency by minimizing the unused space or ensuring temporal consistency.
[0055] The geometric image generation module (310) can generate a 2D geometric image associated with the geometric structure of the point cloud at a given patch location. The texture image generation module (312) can generate a 2D texture image associated with the texture of the point cloud at a given patch location. The geometric image generation module (310) and the texture image generation module (312) use the 3D-to-2D mapping calculated during the packing process to store the geometric structure and texture of the point cloud as images. To better handle the case where multiple points are projected onto the same sample, each patch is projected onto two images, called layers. In one example, the geometric image is represented by a monochrome frame of W×H in YUV420-8-bit format. To generate the texture image, the texture generation step uses the reconstructed / smoothed geometric structure to calculate the color to be associated with the resampled points.
[0056] The occupancy map module (314) can generate an occupancy map that describes the occupancy information at each cell. For example, the occupancy image includes a binary map that indicates, for each cell of the grid, whether the cell belongs to the empty space or to the point cloud. In one example, the occupancy map uses binary information to describe each pixel, indicating whether the pixel is filled. In another example, the occupancy map uses binary information to describe each pixel block, indicating whether the pixel block is filled.
[0057] The occupancy map generated by the occupancy map module (314) can be compressed using lossless coding or lossy coding. When using lossless coding, the entropy compression module (334) is used to compress the occupancy map. When using lossy coding, the video compression module (332) is used to compress the occupancy map.
[0058] Note that the patch packing module (308) can leave some blank space between the 2D patches packed in the image frame. The image filling modules (316) and (318) can fill the blank space (referred to as padding) to generate an image frame that can be applied to 2D video and image codecs. Image filling is also known as background filling, which can fill the unused space with redundant information. In some examples, good background filling minimally increases the bit rate without introducing significant coding distortion around the patch boundaries.
[0059] The video compression modules (322), (323), and (332) can encode 2D images, such as the filled geometric image, the filled texture image, and the occupancy map, based on an appropriate video coding standard, e.g., HEVC, VVC, etc. In one example, the video compression modules (322), (323), and (332) are separate components that operate separately. Note that in another example, the video compression modules (322), (323), and (332) can be implemented as a single component.
[0060] In some examples, the smoothing module (336) is configured to generate a smoothed image of the reconstructed geometry image. The smoothed image can be provided to texture image generation (312). Then, the texture image generation (312) can adjust the generation of the texture image based on the reconstructed geometry image. For example, when the patch shape (e.g., geometry) has slight distortion during encoding and decoding, the distortion can be considered during texture image generation to correct the distortion of the patch shape.
[0061] In some embodiments, group dilation (320) is configured to fill pixels around the object boundary with redundant low-frequency content to improve the coding and decoding gain and the visual quality of the reconstructed point cloud.
[0062] The multiplexer (324) can multiplex the compressed geometry image, the compressed texture image, the compressed occupancy map, and the compressed auxiliary patch information into the compressed bitstream.
[0063] Figure 4 A block diagram of a V-PCC decoder (400) for decoding a compressed bitstream corresponding to a point cloud frame according to some embodiments is shown. In some embodiments, the V-PCC decoder (400) can be used in a communication system (100) and a streaming system (200). For example, the decoder (210) can be configured to operate in a manner similar to the V-PCC decoder (400). The V-PCC decoder (400) receives the compressed bitstream and generates a reconstructed point cloud based on the compressed bitstream.
[0064] In Figure 4 the example of, the V-PCC decoder (400) includes a demultiplexer (432), a video decompression module (434) and (436), an occupancy map decompression module (438), an auxiliary patch information decompression module (442), a geometry reconstruction module (444), a smoothing module (446), a texture reconstruction module (448), and a color smoothing module (452).
[0065] The demultiplexer (432) can receive the compressed bitstream and separate it into a compressed texture image, a compressed geometry image, a compressed occupancy map, and compressed auxiliary patch information.
[0066] The video decompression module (434) and (436) can decode the compressed image according to appropriate standards (e.g., HEVC, VVC, etc.) and output the decompressed image. For example, the video decompression module (434) decodes the compressed texture image and outputs the decompressed texture image; the video decompression module (436) decodes the compressed geometry image and outputs the decompressed geometry image.
[0067] The occupancy map decompression module (438) can decode the compressed occupancy map according to appropriate standards (e.g., HEVC, VVC, etc.) and output the decompressed occupancy map.
[0068] The auxiliary patch - information decompression module (442) can decode the compressed auxiliary patch information according to appropriate standards (e.g., HEVC, VVC, etc.) and output the decompressed auxiliary patch information.
[0069] The geometry reconstruction module (444) can receive the decompressed geometry image and generate a reconstructed point - cloud geometry structure based on the decompressed occupancy map and the decompressed auxiliary patch information.
[0070] The smoothing module (446) can smooth the inconsistencies at the patch edges. The smoothing process aims to mitigate the discontinuities that may occur at the patch boundaries due to compression artifacts. In some embodiments, a smoothing filter can be applied to the pixels located on the patch boundaries to mitigate the distortions that may be caused by compression / decompression.
[0071] The texture reconstruction module (448) can determine the texture information of the points in the point cloud based on the decompressed texture image and the smoothed geometry structure.
[0072] The color smoothing module (452) can smooth the coloring inconsistencies. Non - adjacent patches in 3D space are typically packed adjacent to each other in 2D video. In some examples, the pixel values from non - adjacent patches can be blended by a block - based video codec. The goal of color smoothing is to reduce the visible artifacts that appear at the patch boundaries.
[0073] Figure 5 A block diagram of a video decoder (510) according to an embodiment of the present application is shown. The video decoder (510) can be used in the V - PCC decoder (400). For example, the video decompression modules (434) and (436), and the occupancy map decompression module (438) can be configured in a similar manner to the video decoder (510).
[0074] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from compressed images such as an encoded video sequence. The categories of these symbols include information for managing the operations of the video decoder (510). The parser (520) may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be performed according to a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (520) may extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. The subgroups may include Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), and so on. The parser (520) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.
[0075] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory to create symbols (521).
[0076] Depending on the type of the encoded video picture or the encoded video picture part (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (521) may involve multiple different units. Which units are involved and the way they are involved may be controlled by subgroup control information parsed by the parser (520) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (520) and multiple units below are not described.
[0077] In addition to the functional blocks already mentioned, the video decoder (510) may be conceptually subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the functional units below.
[0078] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives quantized transform coefficients as symbols (521) and control information from the parser (520), including which transform mode to use, block size, quantization factor, quantization scaling matrix, and so on. The scaler / inverse transform unit (551) may output a block including sample values, and the sample values may be input into an aggregator (555).
[0079] In some cases, the output samples of the scaler / inverse transform unit (551) may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra picture prediction unit (552). In some cases, the intra picture prediction unit (552) generates a block having the same size and shape as the block being reconstructed by using surrounding reconstructed information extracted from the current picture buffer (558). For example, the current picture buffer (558) buffers a portion of the reconstructed current picture and / or the fully reconstructed current picture. In some cases, the aggregator (555) adds the prediction information generated by the intra prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a per-sample basis.
[0080] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to an inter-coded and potentially motion-compensated block. In this case, the motion compensation prediction unit (553) may access the reference picture memory (557) to extract samples for prediction. After motion compensating the extracted samples according to the signs (521) belonging to the block, these samples may be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (which is referred to as residual samples or a residual signal in this case), thereby generating output sample information. The motion compensation prediction unit (553) obtaining prediction samples from addresses within the reference picture memory (557) may be controlled by motion vectors, and the motion vectors are in the form of the signs (521) for use by the motion compensation prediction unit (553), and the signs (521) include, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values extracted from the reference picture memory (557) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and the like.
[0081] The output samples of the aggregator (555) may be employed by various loop filtering techniques in the loop filter unit (556). Video compression techniques may include in-loop filter techniques that are controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream), and the parameters are available to the loop filter unit (556) as signs (521) from the parser (520), but may also respond to meta-information obtained during decoding of a previous (in decoding order) portion of the encoded picture or the encoded video sequence, and in response to previously reconstructed and loop-filtered sample values.
[0082] The output of the loop filter unit (556) may be a sample stream, which may be output to a display device and stored in a reference picture memory (557) for subsequent inter-picture prediction.
[0083] Once fully reconstructed, some of the encoded pictures can be used as reference pictures for future prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (e.g., by a parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before starting to reconstruct subsequent encoded pictures.
[0084] The video decoder (510) can perform decoding operations according to a predetermined video compression technique in a standard such as the ITU-T H.265 recommendation. In the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard, the encoded video sequence can comply with the syntax specified by the video compression technique or standard used. Specifically, the profile can select some tools from all the tools available in the video compression technique or standard as the only tools available under the profile. For compliance, it is also required that the complexity of the encoded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, megasamples per second), the maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0085] Figure 6 A block diagram of a video encoder (603) according to an embodiment of the present application is shown. The video encoder (603) can be used to compress a point cloud in a V-PCC encoder (300). In one example, the video compression modules (322) and (323) and the video compression module (332) are configured in a similar manner to the encoder (603).
[0086] The video encoder (603) can receive images, such as filled geometric images, filled texture images, etc., and generate compressed images.
[0087] According to an embodiment, a video encoder (603) may encode and compress pictures (images) of a source video sequence into an encoded video sequence (compressed images) in real time or under any other time constraints required by an application. Implementing an appropriate encoding speed is a function of a controller (650). In some embodiments, the controller (650) controls other functional units as described below and is functionally coupled to these units. For simplicity, couplings are not labeled in the figures. Parameters set by the controller (650) may include rate control related parameters (picture skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) may be used for other suitable functions that relate to optimizing the video encoder (603) for a certain system design.
[0088] In some embodiments, the video encoder (603) operates in an encoding loop. As a simple description, in one example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, e.g., a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data (since in the video compression techniques contemplated by the subject matter disclosed in this application, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input into a reference picture memory (634). Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory (634) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples “seen” by the prediction part of the encoder are exactly the same as the sample values that the decoder will “see” when using the prediction during decoding. This reference picture synchronization principle (and the drift that occurs in cases where synchronization cannot be maintained, e.g., due to channel errors) is also used in some related technologies.
[0089] The operation of the “local” decoder (633) may be the same as that of the “remote” decoder that has been described in detail above in connection with Figure 5 the video decoder (510). However, briefly referring additionally to Figure 5 , when symbols are available and the entropy encoder (645) and the parser (520) can encode / decode the symbols losslessly into the encoded video sequence, the entropy decoding part of the video decoder (510), including the parser (520), may not be fully implemented in the local decoder (633).
[0090] At this time, it can be observed that any decoder technology other than parsing / entropy decoding existing in the decoder must also exist in the corresponding encoder in substantially the same functional form. For this reason, the present application focuses on decoder operations. The description of encoder technology can be simplified because encoder technology is reciprocal to the decoder technology described comprehensively. More detailed descriptions are only needed in certain areas and are provided below.
[0091] During operation, in some examples, the source encoder (630) may perform motion-compensated predictive coding, predicting and encoding an input picture with reference to one or more previously encoded pictures designated as "reference pictures" in the video sequence. In this way, the encoding engine (632) encodes the difference between the pixel blocks of the input picture and the pixel blocks of the reference picture, and the reference picture can be selected as the prediction reference for the input picture.
[0092] The local video decoder (633) may decode the encoded video data of the picture that can be designated as a reference picture based on the symbols created by the source encoder (630). The operation of the encoding engine (632) may be a lossy process. When the encoded video data can be decoded at a video decoder (not shown), the reconstructed video sequence is usually a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that can be performed by the video decoder on the reference picture and may store the reconstructed reference picture in the reference picture cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference picture, which has the same content (without transmission errors) as the reconstructed reference picture to be obtained by the remote video decoder. Figure 6 The predictor (635) may perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search in the reference picture memory (634) for sample data (as candidate reference pixel blocks) or some metadata that can be used as an appropriate prediction reference for the new picture, such as reference picture motion vectors, block shapes, etc. The predictor (635) may operate on a per-pixel block basis of the sample blocks to find a suitable prediction reference. In some cases, according to the search results obtained by the predictor (635), it can be determined that the input picture may have a prediction reference obtained from multiple reference pictures stored in the reference picture memory (634).
[0093] The controller (650) may manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.
[0094] The controller (650) may manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.
[0095] The outputs of all the above functional units can be entropy encoded in an entropy encoder (645). The entropy encoder (645) losslessly compresses the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.
[0096] A controller (650) can manage the operation of the video encoder (603). During encoding, the controller (650) can assign a certain encoded picture type to each encoded picture, which may affect the encoding techniques applicable to the corresponding picture. For example, pictures can generally be assigned to any of the following picture types:
[0097] An intra picture (I picture), which can be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variants of I pictures and their corresponding applications and characteristics.
[0098] A predictive picture (P picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.
[0099] A bi - predictive picture (B picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.
[0100] Source pictures can generally be spatially subdivided into multiple sample blocks (e.g., blocks each having 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be prediction - encoded with reference to other (encoded) blocks, which are determined according to the encoding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non - prediction - encoded, or the blocks can be prediction - encoded with reference to already - encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be prediction - encoded with reference to a previously encoded reference picture either through spatial prediction or through temporal prediction. Blocks of a B picture can be prediction - encoded with reference to one or two previously encoded reference pictures either through spatial prediction or through temporal prediction.
[0101] The video encoder (603) may perform an encoding operation according to a predetermined video encoding technique or standard such as the ITU-T H.265 recommendation. In operation, the video encoder (603) may perform various compression operations, including predictive encoding operations that utilize the temporal and spatial redundancies in the input video sequence. Accordingly, the encoded video data may conform to the syntax specified by the video encoding technique or standard being used.
[0102] Video may be in the form of a plurality of source pictures (images) in a time series. Intra-picture prediction (often simplified to intra prediction) utilizes the spatial correlation within a given picture, while inter-picture prediction utilizes the (temporal or other) correlation between pictures. In an embodiment, a particular picture being encoded / decoded is partitioned into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector may have a third dimension identifying the reference picture.
[0103] In some embodiments, bidirectional prediction techniques may be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, e.g., a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be past and future respectively in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.
[0104] In addition, merge mode techniques may be used in inter-picture prediction to improve encoding efficiency.
[0105] According to some embodiments disclosed in the present application, the execution of predictions such as inter-picture prediction and intra-picture prediction is performed in units of blocks. For example, according to the HEVC standard, pictures in a video picture sequence are segmented into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, for example, 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Further, each CTU can be split into one or more coding units (CUs) by a quadtree. For example, a 64×64 pixel CTU can be split into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In an embodiment, each CU is analyzed to determine a prediction type for the CU, for example, an inter-prediction type or an intra-prediction type. In addition, depending on the temporal and / or spatial predictability, the CU is split into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Taking the luminance prediction block as the prediction block as an example, the prediction block includes a matrix of pixel values (for example, luminance values), for example, 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, and so on.
[0106] Figure 7 FIG. shows a block diagram of a G-PPC encoder (700) according to an embodiment. The encoder (700) can be configured to receive point cloud data and compress the point cloud data to generate a bitstream carrying the compressed point cloud data. In an embodiment, the encoder (700) can include a position quantization module (710), a duplicate point removal module (712), an octree coding module (730), an attribute transfer module (720), a level of detail (LOD) generation module (740), an attribute prediction module (750), a residual quantization module (760), an arithmetic coding module (770), an inverse residual quantization module (780), an addition module (781), and a memory (790) for storing reconstructed attribute values.
[0107] As Figure 7As shown, an input point cloud (701) can be received at an encoder (700). The positions (e.g., 3D coordinates) of the point cloud (701) are provided to a quantization module (710). The quantization module (710) is configured to quantize the coordinates to generate quantized positions. A duplicate point removal module (712) is configured to receive the quantized positions and perform a filtering process to identify and remove duplicate points. An octree encoding module (730) is configured to receive the filtered positions from the duplicate point removal module (712) and perform an octree-based encoding process to generate an occupancy code sequence for a 3D grid describing voxels. The occupancy codes are provided to an arithmetic coding module (770).
[0108] An attribute transfer module (720) is configured to receive the attributes of the input point cloud and, when multiple attribute values are associated with respective voxels, perform an attribute transfer process to determine the attribute values for each voxel. The attribute transfer process can be performed on the re-ordered points output from the octree encoding module (730). The attributes after the transfer operation are provided to an attribute prediction module (750). A LOD generation module (740) is configured to operate on the re-ordered points output from the octree encoding module (730) and re-organize these points into different LODs. The LOD information is provided to the attribute prediction module (750).
[0109] The attribute prediction module (750) processes the points according to the LOD-based order indicated by the LOD information from the LOD generation module (740). The attribute prediction module (750) generates an attribute prediction value for a current point based on the reconstructed attributes of a set of adjacent points of the current point stored in a memory (790). Subsequently, a prediction residual can be obtained based on the original attribute values received from the attribute transfer module (720) and the locally generated attribute prediction values. When candidate indices are used in respective attribute prediction processes, the indices corresponding to the selected prediction candidates are provided to the arithmetic coding module (770).
[0110] A residual quantization module (760) is configured to receive the prediction residuals from the attribute prediction module (750) and perform quantization to generate quantized residuals. The quantized residuals are provided to the arithmetic coding module (770).
[0111] An inverse residual quantization module (780) is configured to receive the quantized residuals from the residual quantization module (760) and generate reconstructed prediction residuals through an inverse operation of the quantization operation performed at the residual quantization module (760). An addition module (781) is configured to receive the reconstructed prediction residuals from the inverse residual quantization module (780) and the respective attribute prediction values from the attribute prediction module (750). By combining the reconstructed prediction residuals and the attribute prediction values, reconstructed attribute values are generated and stored in the memory (790).
[0112] The arithmetic coding module (770) is configured to receive occupancy codes, candidate indices (if used), quantized residuals (if generated), and other information, and perform entropy coding to further compress the received values or information. As a result, a compressed bitstream (702) carrying the compressed information can be generated. The bitstream (702) can be transmitted or otherwise provided to a decoder that decodes the compressed bitstream, or can be stored in a storage device.
[0113] Figure 8 FIG. shows a block diagram of a G-PCC decoder (800) according to an embodiment. The decoder (800) can be configured to receive a compressed bitstream and perform point cloud data decompression to decompress the bitstream to generate decoded point cloud data. In an embodiment, the decoder (800) can include an arithmetic decoding module (810), an inverse residual quantization module (820), an octree decoding module (830), a LOD generation module (840), an attribute prediction module (850), and a memory (860) for storing the reconstructed attribute values.
[0114] As Figure 8 shown, the compressed bitstream (801) can be received at the arithmetic decoding module (810). The arithmetic decoding module (810) is configured to decode the compressed bitstream (801) to obtain the quantized residuals (if generated) and occupancy codes of the point cloud. The octree decoding module (830) is configured to determine the reconstructed positions of the points in the point cloud based on the occupancy codes. The LOD generation module (840) is configured to reorganize the points into different LODs based on the reconstructed positions and determine the order based on the LOD. The inverse residual quantization module (820) is configured to generate reconstructed residuals based on the quantized residuals received from the arithmetic decoding module (810).
[0115] The attribute prediction module (850) is configured to perform an attribute prediction process to determine the attribute prediction of a point according to the order based on the LOD. For example, the attribute prediction of the current point can be determined based on the reconstructed attribute values of the adjacent points of the current point stored in the memory (860). The attribute prediction module (850) can combine the attribute prediction with the respective reconstructed residuals to generate the reconstructed attribute of the current point.
[0116] In one example, the sequence of reconstructed attributes generated from the attribute prediction module (850) together with the reconstructed positions generated from the octree decoding module (830) corresponds to the decoded point cloud (802) output from the decoder (800). Additionally, the reconstructed attributes are also stored in the memory (860) and can subsequently be used to derive the attribute prediction of subsequent points.
[0117] In various embodiments, the encoder (300), decoder (400), encoder (700), and / or decoder (800) can be implemented in hardware, software, or a combination thereof. For example, the encoder (300), decoder (400), encoder (700), and / or decoder (800) can be implemented using processing circuitry such as one or more integrated circuits (ICs) that operate with or without software such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc. In another example, the encoder (300), decoder (400), encoder (700), and / or decoder (800) can be implemented as software or firmware, including instructions stored on a non-volatile (or non-transitory) computer-readable storage medium. These instructions, when executed by processing circuitry such as one or more processors, cause the processing circuitry to perform the functions of the encoder (300), decoder (400), encoder (700), and / or decoder (800).
[0118] Note that an attribute prediction module (750) or (850) configured to implement the attribute prediction techniques disclosed herein can be included in other decoders or encoders that are similar to or different from the structures shown in Figure 7 and Figure 8 In addition, the encoder (700) and decoder (800) can be included in the same device or, in various examples, in separate devices.
[0119] According to some aspects of the present application, a geometric octree structure can be used in PCC. In some related examples, the geometric octree structure is traversed in breadth-first order. In breadth-first order, the octree nodes at the current level can be accessed after the octree nodes at the previous level have been accessed. According to one aspect of the present application, the breadth-first order scheme is not suitable for parallel processing because the current level must wait for the previous level to be encoded and decoded. The present application provides techniques for adding a depth-first encoding / decoding order to the encoding / decoding order techniques for the geometric octree structure. In some embodiments, the depth-first encoding / decoding order can be combined with the breadth-first order or, in some embodiments, can be used alone. The encoding / decoding order (e.g., a combination of depth-first encoding / decoding order, depth-first encoding / decoding order and breadth-first encoding / decoding order, etc.) can be referred to as the hybrid encoding / decoding order of PCC in the present application.
[0120] The proposed methods can be used alone or in any order combination. In addition, each of the methods (or embodiments), encoders, and decoders can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-volatile computer-readable storage medium.
[0121] According to some aspects of the present application, the geometric information of a point cloud and associated attributes, such as color, reflectivity, etc., can be compressed separately (e.g., in the Test Model 13 (TMC13) model). The geometric information of the point cloud, including the 3D coordinates of the points in the point cloud, can be encoded by an octree partitioning with occupancy information. Attributes can be compressed based on the reconstructed geometry using, for example, prediction, lifting, and region-adaptive hierarchical transform techniques.
[0122] According to some aspects of the present application, an octree partitioning can be used to partition a three-dimensional space. An octree is a three-dimensional analogue of a quadtree in two-dimensional space. Octree partitioning techniques refer to partitioning techniques that recursively subdivide a three-dimensional space into eight octants, and an octree structure refers to a tree structure representing these partitions. In one example, each node in the octree structure corresponds to a three-dimensional space, and the node can be a terminal node (no longer divided, also called a leaf node in some examples) or a node with further divisions. The division at a node can divide the three-dimensional space represented by the node into eight octants. In some examples, the nodes corresponding to the partitions of a particular node can be called the children of that particular node.
[0123] Figure 9 A partitioning illustration of a 3D cube (900) (corresponding to a node) based on octree partitioning techniques according to some embodiments of the present application is shown. This partitioning can divide the 3D cube (900) into 8 smaller cubes of equal size, 0 - 7, as Figure 9 shown.
[0124] Octree partitioning techniques (e.g., in TMC13) can recursively divide the original 3D space into smaller units and can encode the occupancy information of each subspace to represent geometric positions.
[0125] In some embodiments (e.g., in TMC13), an octree geometry codec is used. The octree geometry codec can perform geometric encoding. In some examples, geometric encoding is performed on a cube box. For example, the cube box can be an axis-aligned bounding box B defined by two points (0, 0, 0) and (2 M-1 , 2 M-1 , 2 M-1 ), where 2 M-1 defines the size of the bounding box B and M can be specified in the bitstream.
[0126] Then, an octree structure is constructed by recursively subdividing the cube box. For example, the cube box defined by two points (0, 0, 0) and (2 M-1 , 2 M-1 , 2 M-1)The defined cubic box is divided into 8 sub-cubic boxes to generate an 8-bit code, called the occupancy code. Each bit in the occupancy code is associated with a sub-cubic box, and the value of the bit is used to indicate whether the associated sub-cubic box contains any points of the point cloud. For example, a value of 1 for a bit indicates that the sub-cubic box associated with that bit contains one or more points of the point cloud; a value of 0 for a bit indicates that the sub-cubic box associated with that bit does not contain points of the point cloud.
[0127] In addition, for an empty sub-cubic box (e.g., the value of the bit associated with the sub-cubic box is 0), the sub-cubic box is no longer divided. When a sub-cubic box has one or more points of the point cloud (e.g., the value of the bit associated with the sub-cubic box is 1), the sub-cubic box is further divided into 8 smaller sub-cubic boxes, and an occupancy code can be generated for the sub-cubic box to indicate the occupancy of the smaller sub-cubic boxes. In some examples, the subdivision operation can be repeatedly performed on non-empty sub-cubic boxes until the size of the sub-cubic box is equal to a predetermined threshold, such as a size of 1. In some examples, a sub-cubic box with a size of 1 is called a voxel, and a sub-cubic box with a size larger than a voxel can be called a non-voxel.
[0128] Figure 10 An example of an octree partition (1010) and an octree structure (1020) corresponding to the octree partition (1010) according to some embodiments of the present application is shown. Figure 10 A two-level partition in the octree partition (1010) is shown. The octree structure (1020) includes a node (N0) corresponding to the cubic box of the octree partition (1010). At the first level, according to Figure 9 the numbering technique shown, the cubic box is divided into 8 sub-cubic boxes numbered from 0 to 7. The occupancy code of the partition of node N0 is binary "10000001", which indicates that the first sub-cubic box represented by node N0-0 and the eighth sub-cubic box represented by node N0-7 contain points in the point cloud, and the other sub-cubic boxes are empty.
[0129] Then, in the second-level partition, the first sub-cubic box (represented by node N0-0) and the eighth sub-cubic box (represented by node N0-7) are further divided into eight octants respectively. For example, according to Figure 9The numbered technique shown divides the first sub-cube box (represented by node N0-0) into 8 smaller sub-cube boxes numbered from 0 to 7. The occupancy code for the division of node N0-0 is binary "00011000", which indicates that the fourth smaller sub-cube box (represented by node N0-0-3) and the fifth smaller sub-cube box (represented by node N0-0-4) contain points in the point cloud, and the other smaller sub-cube boxes are empty. At the second level, the seventh sub-cube box (represented by node N0-7) is divided into 8 smaller sub-cube boxes in a similar manner, as Figure 10 shown.
[0130] In Figure 10 the example of, nodes corresponding to non-empty cube spaces (such as cube boxes, sub-cube boxes, smaller sub-cube boxes, etc.) are colored gray and are called shaded nodes.
[0131] According to some aspects of the present application, appropriate coding techniques can be used to appropriately compress the occupancy code. In some embodiments, an arithmetic encoder is used to compress the occupancy code of the current node in the octree structure. This occupancy code can be represented as an 8-bit integer S, and each bit in S indicates the occupancy status of the child nodes of the current node. In one embodiment, bit-by-bit coding is used to encode the occupancy code. In another embodiment, byte-by-byte coding is used to encode the occupancy code. In some examples (such as TMC13), bit-by-bit coding is enabled by default. Both bit-by-bit coding and byte-by-byte coding can perform arithmetic coding with context modeling to encode the occupancy code. The context state can be initialized at the start of the entire occupancy code encoding process and updated during the occupancy code encoding process.
[0132] In the embodiment of bit-by-bit coding for encoding the occupancy code of the current node, the eight binary numbers (bins) in the occupancy code S of the current node are encoded in a specific order. Each binary number in S is encoded by referring to the occupancy status of the adjacent nodes of the current code and / or the child nodes of the adjacent nodes. Adjacent nodes are at the same level as the current node and can be called the sibling nodes of the current node.
[0133] In the embodiment of byte-by-byte coding for encoding the occupancy code of the current node, the occupancy code S (one byte) can be encoded by referring to: (1) an adaptive lookup table (A-LUT) that tracks P (such as 32) of the most frequently used occupancy codes; and (2) a cache that tracks the last Q (such as 16) different occupancy codes observed.
[0134] In some examples of byte-by-byte encoding, a binary flag indicating whether S is in the A-LUT is encoded. If S is in the A-LUT, its index in the A-LUT is encoded by using a binary arithmetic coder. If S is not in the A-LUT, a binary flag indicating whether S is in the cache is encoded. If S is in the cache, the binary representation of its index in the cache is encoded by using a binary arithmetic coder. Otherwise, if S is not in the cache, the binary representation of S is encoded by using a binary arithmetic coder.
[0135] In some embodiments, on the decoder side, the decoding process can start by parsing the dimensions of the bounding box from the bitstream. The bounding box represents a cubic box corresponding to the root node in an octree structure, and is used to partition the cubic box according to the geometric information of the point cloud (e.g., the occupancy information of the points in the point cloud). Then, the cubic box is subdivided according to the decoded occupancy code to construct the octree structure.
[0136] In some related examples, (e.g., in the version of TMC13), to encode the occupancy code, the octree structure is traversed in breadth-first order. In breadth-first order, the octree nodes in one level (the nodes in the octree structure) can be accessed after all the octree nodes in the previous level have been accessed. In one implementation example, a first-in-first-out (FIFO) data structure can be used.
[0137] Figure 11 A figure showing an octree structure (1100) illustrating the breadth-first encoding and decoding order is presented. The shaded nodes in the octree structure (1100) are the nodes corresponding to non-empty cubic spaces. The occupancy codes of the shaded nodes can be encoded and decoded in the breadth-first encoding and decoding order Figure 11 shown from 0 to 8. In the breadth-first encoding and decoding order, the octree nodes are accessed level by level. The breadth-first encoding and decoding order itself is not suitable for parallel processing because the current level must wait for the previous level to be encoded and decoded.
[0138] Some aspects of the present application provide a hybrid encoding and decoding order that includes using a depth-first encoding and decoding order instead of a breadth-first encoding and decoding order to encode and decode at least one level. Thus, in some embodiments, the nodes and their descendant nodes in the level using the depth-first encoding and decoding order can form a sub-octree structure of the octree structure. When the level using the depth-first encoding and decoding order includes multiple nodes respectively corresponding to non-empty cubic spaces, the multiple nodes and their corresponding descendant nodes can form multiple sub-octree structures. In some embodiments, the multiple sub-octree structures can be encoded and decoded in parallel.
[0139] Figure 12Shows an octree structure (1200) diagram illustrating a depth - first encoding / decoding order. The shaded nodes in the octree structure (1200) are nodes corresponding to non - empty cubic spaces. The octree structure (1200) can have the same occupancy geometry as the corresponding point cloud of the octree structure (1100). The occupancy codes of the shaded nodes can be encoded and decoded in the depth - first encoding / decoding order from 0 to 8 as shown Figure 12 and decoded.
[0140] In Figure 12 the example, the node "0" can be at any suitable division depth, such as PD0, the children of the node "0" are at division depth PD0 + 1, and the grandchildren of the node "0" are at division depth PD0 + 2. In Figure 12 the example, the nodes at division depth PD0 + 1 can be encoded and decoded in the depth - first encoding / decoding order. The nodes at division depth PD0 + 1 include two nodes corresponding to non - empty spaces. These two nodes and their respective descendant nodes can form a first sub - octree structure (1210) and a second sub - octree structure (1220), and these two nodes can be respectively called the root nodes of the two sub - octree structures.
[0141] Figure 12 The depth - first encoding / decoding order in
[0142] is called the pre - order version of the depth - first encoding / decoding order. In this pre - order version of the depth - first encoding / decoding order, for each sub - octree structure, the root node of the sub - octree structure is visited before visiting the children nodes of the sub - octree structure. Further, the deepest nodes are visited first, and then backtracked to the sibling nodes of the parent node.
[0142] In Figure 12 the example, in some implementations, the first sub - octree structure (1210) and the second sub - octree structure (1220) can be encoded and decoded in parallel processing. For example, nodes 1 and 5 can be accessed simultaneously. In some examples, recursive programming or a stack data structure can be used to implement the depth - first encoding / decoding order.
[0143] In some embodiments, the hybrid encoding / decoding order starts with a breadth - first traversal (encoding / decoding), and after several levels of breadth - first traversal, depth - first traversal (encoding / decoding) can be enabled.
[0144] It should be noted that the hybrid encoding / decoding order can be used in any suitable PCC system, such as a PCC system based on TMC 13, a PCC system based on MPEG - PCC, etc.
[0145] According to one aspect of the present application, the hybrid encoding and decoding order may include both a breadth-first encoding and decoding order and a depth-first encoding and decoding order for encoding and decoding geometric information of a point cloud. In one embodiment, a node size at which a node in an octree structure changes from a breadth-first encoding and decoding order to a depth-first encoding and decoding order may be specified. In one example, during PCC, the encoding and decoding of the octree structure starts with a breadth-first encoding and decoding order, and at a level where the node size is equal to the specified node size for changing the encoding and decoding order, the encoding and decoding order may change to a depth-first encoding and decoding order at that level. Note that in some examples, the node size is associated with the division depth.
[0146] In another embodiment, a node size at which a node in an octree structure changes from a depth-first encoding and decoding order to a breadth-first encoding and decoding order may be specified. In one example, during PCC, the encoding and decoding of the octree structure starts with a depth-first encoding and decoding order, and at a level where the node size is equal to the specified node size for changing the encoding and decoding order, the encoding and decoding order may change to a breadth-first encoding and decoding order. Note that in some examples, the node size is associated with the division depth.
[0147] More specifically, in an embodiment starting with a breadth-first encoding and decoding order, the node size may be represented on a log2 scale and denoted by d = 0, 1,..., M - 1, where M - 1 is the node size of the root node and M is the maximum number of octree division depths (also referred to as levels in some examples). Additionally, a parameter d referred to as the encoding and decoding order change size may be defined. t . In one example, the parameter d t (1 ≤ d t ≤ M - 1) is used to specify that the breadth-first order is applied to nodes with sizes from M - 1 to d t , and the depth-first order is applied to nodes with sizes from d t - 1 to 0. When d t = M - 1, the depth-first scheme is applied to all octree nodes starting from the root node. When d t = 1, only the breadth-first encoding and decoding order is used to encode and decode the octree structure.
[0148] In some embodiments, the encoding and decoding order of the octree structure can start with a breadth-first encoding and decoding order, and then at a specific depth (corresponding to a specific node size), each node at that specific depth and the descendant nodes of that node form a separate sub-octree structure of the point cloud. Thus, at that specific depth, multiple sub-octree structures are formed. The sub-octree structures can be encoded and decoded separately using any suitable encoding and decoding mode. In one example, a depth-first encoding and decoding order can be used to encode and decode the sub-octree structures. In another example, a breadth-first encoding and decoding order can be used to encode and decode the sub-octree structures. In another example, a hybrid encoding and decoding order can be used to encode and decode the sub-octree structures. In another example, a bit-by-bit encoding and decoding scheme can be used to encode and decode the occupancy codes in the sub-octree structures. In another example, a byte-by-byte encoding and decoding scheme can be used to encode and decode the occupancy codes in the sub-octree structures. In another example, a predictive geometry encoding and decoding technique can be used to encode and decode the sub-octree structures, and the predictive geometry encoding and decoding technique is an alternative encoding and decoding mode to the depth-first octree encoding and decoding mode. In some examples, the predictive geometry encoding and decoding technique can predict points based on the encoded correction vectors and the previously encoded adjacent points.
[0149] In some embodiments, on the encoder side, for each sub-octree structure, the encoder can select an encoding and decoding mode from multiple encoding and decoding modes based on the encoding and decoding efficiency. For example, the selected encoding and decoding mode for the sub-octree structure can achieve the best encoding and decoding efficiency for the sub-octree structure. Then, the encoder can use the encoding and decoding modes separately selected for the sub-octree structures to encode and decode the sub-octree structures respectively. In some embodiments, the encoder can signal in the bitstream an index for the sub-octree structure, and the index indicates the encoding and decoding mode selected for the sub-octree structure. On the decoder side, the decoder can determine the encoding and decoding mode for the sub-octree structure based on the index in the bitstream, and then decode the sub-octree structure according to the encoding and decoding mode.
[0150] Aspects of the present application also provide signal representation techniques for a hybrid encoding and decoding order. According to one aspect of the present application, the control parameters to be used in the hybrid encoding and decoding order can be signaled in a high-level syntax, for example, the sequence parameter set (SPS) of the bitstream, the slice header, the geometry parameter set, etc. Note that specific examples are provided in the following description. The techniques disclosed in the present application illustrated by these specific examples are not limited to these specific examples and can be appropriately adjusted for use in other examples.
[0151] In one embodiment, the parameter d t (encoding and decoding order change size) is specified in the high-level syntax.
[0152] Figure 13Illustrate a syntax example (1300) of a set of geometric parameters according to some embodiments of the present application. As shown in (1310), gps_depth_first_node_size_log2_minus_1 is specified in the set of geometric parameters. Parameter d can be determined based on gps_depth_first_node_size_log2_minus_1 t , for example, according to (Equation 1).
[0153] d t = gps_depth_first_node_size_log2_minus1 + 1 (Equation 1)
[0154] Note that when gps_depth_first_node_size_log2_minus_1 is equal to 0, the depth-first coding order is disabled.
[0155] In another embodiment, a control flag is explicitly signaled to indicate whether to use a hybrid coding order.
[0156] Figure 14 Illustrate another syntax example (1400) of a set of geometric parameters according to some embodiments of the present application. As shown in (1410), a control flag represented by gps_hybrid_coding_order_flag is used. When the control flag gps_hybrid_coding_order_flag is "true" (e.g., value is 1), the hybrid coding order scheme is enabled; when gps_hybrid_coding_order_flag is "false" (e.g., value is 0), the hybrid coding order scheme is disabled. When gps_hybrid_coding_order_flag is "true" (e.g., value is 1), parameter d can be determined based on gps_depth_first_node_size_log2_minus_2 t , for example, according to (Equation 2):
[0157] d t = gps_depth_first_node_size_log2_minus2 + 2 (Equation 2)
[0158] When gps_hybrid_coding_order_flag is "false" (e.g., value is 0), d t is set to 1 by default to indicate that the depth-first coding order is disabled and only the breadth-first coding order is applied in the example.
[0159] In one embodiment, when the hybrid encoding / decoding order is enabled, a breadth-first order is applied to nodes of size M-1 to d t and a depth-first order is applied to nodes of size d t -1 to 0.
[0160] Figure 15 Shows a pseudo-code example (1500) for octree encoding / decoding according to some embodiments of the present application. As shown in (1510), when depth >= MaxGeometryOctreeDepth - d t , a depth-first encoding / decoding order can be used. In the Figure 15 example, the pseudo-code "geometry_node_depth_first" can be applied to the depth-first encoding / decoding order.
[0161] Figure 16 Shows a pseudo-code example (1600) for the depth-first encoding / decoding order according to some embodiments of the present application. The pseudo-code "geometry_node_depth_first" is a recursive function. In the recursive function, first the "geometry_node" function is called to obtain the occupancy code of the current octree node, and then the pseudo-code "geometry_node_depth_first" is called to encode / decode each child node until reaching a leaf node, for example, when depth >= MaxGeometryOctreeDepth - 1.
[0162] Figure 17 Shows a flowchart outlining a method (1700) according to embodiments of the present application. The method (1700) can be used during the encoding / decoding process for a point cloud. In various embodiments, the method (1700) is executed by a processing circuit, such as the processing circuit in a terminal device (110), the processing circuit performing the functions of an encoder (203) and / or a decoder (201), the processing circuit performing the functions of an encoder (300), a decoder (400), an encoder (700) and / or a decoder (800), etc. In some embodiments, the method (1700) is implemented by software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the method (1700). The process starts at (S1701) and proceeds to (S1710).
[0163] At (S1710), an encoded bitstream of the point cloud is received. The encoded bitstream includes geometric information in the form of encoded occupancy codes of nodes in the octree structure of the point cloud. The nodes in the octree structure correspond to 3D partitions of the point cloud space. The size of a node is associated with the size of the corresponding 3D partition.
[0164] At (S1720), the occupancy code of the node is decoded from the encoded occupancy code. The decoding of the first occupancy code of at least the child nodes of the first node does not need to wait for the decoding of the second occupancy code of the second node, and the second node has the same node size as the first node.
[0165] In one embodiment, the child nodes are among the first set of nodes (first descendant nodes) in the first sub-octree, and the first node is the root of the first sub-octree. The first node and the second node are sibling nodes with the same node size. The second node is the root node of the second sub-octree including the second set of nodes (second descendant nodes). Then, in some examples, the first set of occupancy codes of the first set of nodes and the second set of occupancy codes of the second set of nodes can be decoded separately. In one example, the first set of occupancy codes of the first set of nodes and the second set of occupancy codes for the second set of nodes can be decoded in parallel. In another example, the first set of occupancy codes of the first set of nodes is decoded using the first encoding / decoding mode, and the second set of occupancy codes of the second set of nodes is decoded using the second encoding / decoding mode.
[0166] The first encoding / decoding mode and the second encoding / decoding mode can use any one of depth-first encoding / decoding order, breadth-first encoding / decoding order, predictive geometric encoding / decoding technology, etc. In some examples, the encoded bitstream includes a first index indicating the first encoding / decoding mode of the first sub-octree and a second index indicating the second encoding / decoding mode of the second sub-octree.
[0167] In another embodiment, the first node and the second node have a specific node size for changing the encoding / decoding order. In some examples, the larger node among the nodes is encoded / decoded using the first encoding / decoding order, and the smaller node among the nodes is encoded / decoded using the second encoding / decoding order. The node size of the larger node is greater than the specific node size for changing the encoding / decoding order. The node size of the smaller node is equal to or less than the specific node size for the encoding / decoding order. In one example, the first encoding / decoding order is breadth-first encoding / decoding order, and the second encoding / decoding order is depth-first encoding / decoding order. In another example, the first encoding / decoding order is depth-first encoding / decoding order, and the second encoding / decoding order is breadth-first encoding / decoding order.
[0168] In some examples, the specific node size for changing the encoding / decoding order is determined based on the signal in the encoded bitstream of the point cloud. In some examples, when the control signal indicates a change in the encoding / decoding order, the signal is provided.
[0169] At (S1730), the octree structure can be reconstructed based on the decoded occupancy code of the node.
[0170] At (S1740), the point cloud is reconstructed based on the octree structure. Then, the method proceeds to (S1799) and ends.
[0171] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable storage media. For example, Figure 18 illustrates a computer system (1800) suitable for implementing certain embodiments of the disclosed subject matter.
[0172] The computer software may be encoded using any suitable machine code or computer language that may be subject to assembly, compilation, linking, or similar mechanisms to create code including instructions that may be executed directly or through interpretation, microcode execution, etc. by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.
[0173] The instructions may be executed on various types of computers or computer components, including, for example, personal computers, tablets, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0174] Figure 18 The components shown in are exemplary in nature and are not intended to imply any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present application. Nor should the configuration of the components be construed as having any dependence on or requirement for any one or combination of components shown in the exemplary embodiments of the computer system (1800).
[0175] The computer system (1800) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input by one or more human users through, for example, tactile input (e.g., key presses, swipes, data glove movements), audio input (e.g., speech, taps), visual input (e.g., gestures), olfactory input (not depicted). The human-machine interface devices may also be used to capture certain media not directly related to conscious input by a human, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0176] The input human-machine interface devices may include one or more of the following (each depicted only one): keyboard (1801), mouse (1802), trackpad (1803), touch screen (1810), data glove (not shown), joystick (1805), microphone (1806), scanner (1807), camera (1808).
[0177] The computer system (1800) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., the tactile feedback of a touch screen (1810), a data glove (not shown), or a joystick (1805), but there may also be tactile feedback devices that do not act as input devices), audio output devices (e.g., speakers (1809), headphones (not depicted)), visual output devices (e.g., a screen (1810), including a cathode ray tube (CRT) screen, a liquid crystal display (LCD) screen, a plasma screen, an organic light-emitting diode (OLED) screen, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which are capable of outputting two-dimensional visual output or output greater than three-dimensional through, for example, stereoscopic flat painting output; virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted)), and printers (not depicted).
[0178] The computer system (1800) may also include human-accessible storage devices and the associated media of the storage devices, such as optical media, including CD / DVD ROM / RW (1820) with media such as CD / DVD (1821), thumb drives (1822), removable hard disk drives, or solid state drives (1823), legacy magnetic media such as tapes and floppy disks (not depicted), ROM / special application specific integrated circuit (ASIC) / programmable logic device (PLD)-based special devices, such as security protection devices (not depicted), and so on.
[0179] Those skilled in the art should also understand that the term "computer-readable storage medium" used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0180] The computer system (1800) may also include an interface to one or more communication networks. The network may be, for example, wireless, wired, optical. The network may also be local, wide area, metropolitan area, vehicular and industrial, real-time, latency-tolerant, etc. Examples of networks include, for example, Ethernet, local area networks of wireless LAN, cellular networks including Global System for Mobile Communications (GSM), third generation (3G), fourth generation (4G), fifth generation (5G), Long Term Evolution (LTE), etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular networks and industrial networks including Controller Area Network Bus (CANBus), etc. Certain networks typically require an external network interface adapter attached to certain general-purpose data ports or peripheral buses (1849) (e.g., Universal Serial Bus (USB) ports of the computer system (1800)); other networks are typically integrated into the core of the computer system (1800) by attaching to the system bus as described below (e.g., integrated into a PC computer system through an Ethernet interface, or integrated into a smart phone computer system through a cellular network interface). By using any of these networks, the computer system (1800) can communicate with other entities. Such communication can be only one-way reception (e.g., broadcast TV), only one-way transmission (e.g., CANBus connected to certain CANBus devices), or two-way, for example, using a local digital network or a wide area digital network to connect to other computer systems. Certain protocols and protocol stacks can be used on each of the networks and network interfaces as described above.
[0181] The above-mentioned human-machine interface device, human-accessible storage device, and network interface may be attached to the core (1840) of the computer system (1800).
[0182] The core (1840) may include one or more central processing units (CPUs) (1841), a graphics processing unit (GPU) (1842), a dedicated programmable processing unit in the form of field programmable gate areas (FPGAs) (1843), a hardware accelerator for certain tasks (1844), and so on. These devices, together with a read-only memory (ROM) (1845), a random access memory (1846), and internal mass storage devices such as internal hard disk drives, solid state drives (SSDs) that are not accessible by users (1847), may be connected via a system bus (1848). In some computer systems, the system bus (1848) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached directly or via a peripheral bus (1849) to the system bus (1848) of the core. Architectures for peripheral buses include Peripheral Component Interconnect (PCI), USB, and so on.
[0183] The CPU (1841), GPU (1842), FPGA (1843), and accelerator (1844) may execute certain instructions that, when combined, may constitute the aforementioned computer code. The computer code may be stored in the ROM (1845) or RAM (1846). Transitional data may also be stored in the RAM (1846), while permanent data may be stored, for example, in the internal mass storage device (1847). Fast storage and retrieval of any memory device may be achieved by using a cache memory that may be closely associated with one or more CPUs (1841), GPUs (1842), mass storage devices (1847), ROM (1845), RAM (1846), etc.
[0184] Computer-readable storage media may have computer code for performing various computer-implemented operations. The media and the computer code may be media and computer code designed and constructed specifically for the purposes of this application, or may be of the kind well-known and available to those skilled in the field of computer software.
[0185] By way of example and not limitation, a computer system having an architecture (1800) and in particular a core (1840) can provide functionality resulting from software executed by processors (including CPUs, GPUs, FPGAs, accelerators, etc.) embodied on one or more tangible computer-readable storage media. Such computer-readable storage media can be media associated with the user-accessible mass storage devices and certain non-transitory storage devices of the core (1840) introduced above (e.g., the on-core mass storage device (1847) or ROM (1845)). The software implementing various embodiments of the present application can be stored in such devices and executed by the core (1840). Depending on specific requirements, the computer-readable storage media can include one or more memory devices or chips. The software can cause the core (1840) and in particular the processors therein (including CPUs, GPUs, FPGAs, etc.) to execute specific processes or specific parts of specific processes described herein, including defining data structures stored in RAM (1846) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system can provide functionality resulting from logic hard-wired or otherwise embodied in circuitry (e.g., accelerator (1844)), which can operate in place of or in conjunction with the software to execute specific processes or specific parts of specific processes described herein. Where appropriate, references to software can encompass logic and vice versa. Where appropriate, references to computer-readable storage media can encompass circuitry (e.g., integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both types of circuitry. The present application encompasses any suitable combination of hardware and software.
[0186] Although the present application describes several exemplary embodiments, within the scope of the present application, there can be various modifications, permutations, and various alternative equivalents. Therefore, it should be understood that within the spirit and scope of the application, those skilled in the art can design various systems and methods that, although not explicitly shown or described herein, can embody the principles of the present application.
Claims
1. A point cloud decoding method, characterized in that, Comprising: Receiving encoded data of nodes in an octree structure of the point cloud from an encoded bitstream of the point cloud, wherein the nodes in the octree structure correspond to three-dimensional (3D) partitions of the space of the point cloud, and the size of the nodes is associated with the size of the corresponding 3D partitions; Decoding occupancy codes of the nodes from the encoded data, wherein decoding of the first occupancy codes of the child nodes of at least a first node does not need to wait for decoding of the second occupancy codes of a second node, and the second node has the same node size as the first node; Reconstructing the octree structure based on the decoded occupancy codes of the nodes; and Reconstructing the point cloud based on the octree structure; The method further comprises: Determining a specified node size for changing the encoding / decoding order based on a signal in the encoded bitstream of the point cloud; For larger nodes among the nodes, using a first encoding / decoding order to decode a first part of the occupancy codes, where the larger nodes are larger than the specified node size for changing the encoding / decoding order; and For smaller nodes among the nodes, using a second encoding / decoding order different from the first encoding / decoding order to decode a second part of the occupancy codes, where the smaller nodes are equal to or smaller than the specified node size for changing the encoding / decoding order.
2. The method according to claim 1, wherein Further comprising: Decoding a first set of occupancy codes of a first set of nodes in a first sub-octree, where the first node is the root of the first sub-octree; And Decoding a second set of occupancy codes of a second set of nodes in a second sub-octree, where the second node is the root of the second sub-octree.
3. The method according to claim 2, wherein Further comprising: Decoding the first set of occupancy codes of the first set of nodes in the first sub-octree is parallel to decoding the second set of occupancy codes of the second set of nodes in the second sub-octree.
4. The method according to claim 2, wherein Further comprising: Using a first encoding / decoding mode to decode the first set of occupancy codes of the first set of nodes in the first sub-octree; And Using a second encoding / decoding mode to decode the second set of occupancy codes of the second set of nodes in the second sub-octree.
5. The method according to claim 4, wherein Further comprising: Decoding a first index indicating the first encoding / decoding mode of the first sub-octree from the encoded bitstream of the point cloud; And Decoding a second index indicating the second encoding / decoding mode of the second sub-octree from the encoded bitstream of the point cloud.
6. The method according to claim 1, characterized in that, The first encoding / decoding order is a breadth-first encoding / decoding order, and the second encoding / decoding order is a depth-first encoding / decoding order.
7. The method according to claim 1, wherein The first encoding / decoding order is a depth-first encoding / decoding order, and the second encoding / decoding order is a breadth-first encoding / decoding order.
8. The method according to claim 1, wherein Further comprising: Decoding a control signal from the encoded bitstream of the point cloud, where the control signal indicates a change in the encoding / decoding order.
9. A point cloud decoding device, characterized in that, Comprising: A processing circuit configured to execute the method according to any one of claims 1 to 8.
10. A point cloud decoding device, characterized in that, Comprising: A receiving unit for receiving encoded occupancy codes of nodes in an octree structure of the point cloud from an encoded bitstream of the point cloud, wherein the nodes in the octree structure correspond to three-dimensional (3D) partitions of the space of the point cloud, and the size of the nodes is associated with the size of the corresponding 3D partitions; A decoding unit for decoding the occupancy code of the node from the encoded occupancy code, wherein the decoding of the first occupancy code of at least the child node of the first node does not need to wait for the decoding of the second occupancy code of the second node, and the second node has the same node size as the first node; A reconstruction unit for reconstructing the octree structure based on the decoded occupancy code of the node and reconstructing the point cloud based on the octree structure; The decoding unit is further configured to: determine a specified node size for changing the encoding / decoding order based on the signal in the encoded bitstream of the point cloud; for larger nodes among the nodes, use a first encoding / decoding order to decode a first part of the occupancy code, where the larger nodes are larger than the specified node size for changing the encoding / decoding order; and for smaller nodes among the nodes, use a second encoding / decoding order different from the first encoding / decoding order to decode a second part of the occupancy code, where the smaller nodes are equal to or smaller than the specified node size for changing the encoding / decoding order.
11. An electronic device, characterized in that, Comprising a memory for storing computer-readable instructions; a processor for reading the computer-readable instructions and executing the method according to any one of claims 1 to 8 as indicated by the computer-readable instructions.
12. A computer-readable storage medium having instructions stored thereon, which when executed by a point cloud decoding computer, cause the computer to execute the method according to any one of claims 1 to 8.
13. A method for storing a bitstream, characterized in that, The bitstream is decoded based on the decoding method according to any one of claims 1-8.
Citation Information
Patent Citations
Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
CA3096452A1
Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
CA3098586A1
Methods and devices for entropy coding point clouds
WO2019140510A1