Method and apparatus for point cloud decoding
By receiving signaling information in the point cloud coded bitstream, determining the segmentation mode and reconstructing the point cloud, the problem of inflexible and inefficient point cloud encoding and decompression in the prior art is solved, and more efficient point cloud compression and decompression is achieved.
Patent Information
- Application Number
- CN202180003429.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-16
- Filing Date
- 2021-03-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-03-25
AI Technical Summary
In the prior art, point cloud encoding and decoding flexibility and efficiency are not high, resulting in insufficient compression and decompression performance of point clouds.
By receiving signaling information in the encoded bitstream of the point cloud, the segmentation mode of the point cloud, including multiple segmentation levels, is determined, and the point cloud is reconstructed based on this mode. Specific methods include using predefined quad-tree and binary tree segmentation, determining the size and segmentation direction of 3D space, and including octree segmentation in the segmentation mode.
Improve the flexibility and efficiency of point cloud collections, and improve the performance of point cloud compression and decompression.
Smart Images

Figure CN113892235B_ABST
Abstract
Description
[0001] Incorporation by reference
[0002] This application claims the benefit of priority of U.S. Patent Application No. 17 / 203,155, "METHOD AND APPARATUS FOR POINT CLOUD CODING", filed on Mar. 16, 2021, which claims the benefit of priority of U.S. Provisional Application No. 63 / 004,304, "METHOD AND APPARATUS FOR FLEXIBLE QUAD-TREE AND BINARY-TREE PARTITIONING FOR GEOMETRY CODING", filed on Apr. 2, 2020. The entire disclosure of the prior application is hereby incorporated by reference in its entirety. Technical field
[0003] The present invention relates to the field of information processing, and specifically, to a method, an apparatus, a computer device, and a storage medium for performing point cloud geometry decoding. Background art
[0004] The purpose of the background art description provided herein is to present the background of the present disclosure generally. The work of the currently named inventors, to some extent, the work described in this background art section, and aspects that may not otherwise be regarded as prior art as of the filing date are neither expressly nor implicitly admitted to be prior art with respect to the present disclosure.
[0005] Various technologies have been developed to capture and represent the world in a three-dimensional (3D) space, such as objects in the world, environments in the world, etc. The 3D representation of the world can enable more immersive forms of interaction and communication. A point cloud can be used as a 3D representation of the world. A point cloud is a set of points in 3D space, each point having associated attributes such as color, material properties, texture information, intensity attributes, reflectivity attributes, motion-related attributes, morphological attributes, and / or various other attributes. Such a point cloud may include a large amount of data, and storage and transmission may be both expensive and time-consuming. In the prior art, the flexibility and efficiency of point cloud encoding and decoding in point cloud compression technologies are not high. Summary of the invention
[0006] Aspects of the present disclosure provide methods and apparatuses for point cloud compression and decompression. According to one aspect of the present disclosure, a method for performing point cloud geometry decoding is provided. In this method, first signaling information may be received from an encoded bitstream of a point cloud including a set of points in a three-dimensional (3D) space. The first signaling information may indicate segmentation information of the point cloud. Second signaling information may be determined based on the first signaling information indicating a first value. The second signaling information may indicate a segmentation pattern of the set of points in the 3D space. In addition, the segmentation pattern of the set of points in the 3D space may be determined based on the second signaling information. Subsequently, the point cloud may be reconstructed based on the segmentation pattern.
[0007] In some embodiments, the segmentation pattern may be determined as a predefined Quad-tree and Binary-tree (QtBt) segmentation based on the second signaling information indicating a second value.
[0008] In this method, third signaling information indicating that the 3D space is an asymmetric cuboid may be received. Signaled dimensions of the 3D space along the x, y, and z directions may be determined based on the third signaling information indicating a first value.
[0009] In some embodiments, 3-bit signaling information may be determined for each of a plurality of segmentation levels in the segmentation pattern based on the second signaling information indicating a first value. The 3-bit signaling information for each of the plurality of segmentation levels may indicate the segmentation directions of the corresponding segmentation level in the segmentation pattern along the x, y, and z directions.
[0010] In some embodiments, the 3-bit signaling information may be determined based on the dimensions of the 3D space.
[0011] In this method, a segmentation pattern may be determined based on the first signaling information indicating a second value, wherein the segmentation pattern may include corresponding octree segmentation at each of a plurality of segmentation levels in the segmentation pattern.
[0012] According to one aspect of the present disclosure, a method for performing point cloud geometry decoding is provided. In this method, first signaling information may be received from an encoded bitstream of a point cloud including a set of points in a three-dimensional (3D) space. The first signaling information may indicate segmentation information of the point cloud. A segmentation pattern of the set of points in the 3D space may be determined based on the first signaling information, wherein the segmentation pattern may include a plurality of segmentation levels. Subsequently, the point cloud may be reconstructed based on the segmentation pattern.
[0013] In some embodiments, 3-bit signaling information for each of a plurality of segmentation levels in a segmentation mode may be determined based on a first signaling information indicating a first value, wherein the 3-bit signaling information for each of the plurality of segmentation levels may indicate a segmentation direction in the corresponding segmentation level of the segmentation mode in the x, y, and z directions.
[0014] In some embodiments, 3-bit signaling information may be determined based on the size of a 3D space.
[0015] In some embodiments, based on a first signaling information indicating a second value, the segmentation mode may be determined to include a corresponding octree segmentation in each of a plurality of segmentation levels in the segmentation mode.
[0016] In this method, second signaling information may also be received from an encoded bitstream of a point cloud. When the second signaling information is the first value, the second signaling information may indicate that the 3D space is an asymmetric rectangular parallelepiped, and when the second signaling information is the second value, the second signaling information may indicate that the 3D space is a symmetric rectangular parallelepiped.
[0017] In some embodiments, based on the first signal information indicating the second value and the second signal information indicating the first value, the segmentation mode may be determined to include a corresponding octree segmentation in each of the first segmentation levels among the plurality of segmentation levels of the segmentation mode. The segmentation type and segmentation direction of the last segmentation level among the plurality of segmentation levels of the segmentation mode may be determined according to the following conditions: , where d x , d y and d z are the log2 sizes of the 3D space in the x, y, and z directions, respectively.
[0018] In this method, second signaling information may be determined based on the first signaling information indicating the first value. When the second signaling information indicates the first value, the second signaling information may indicate that the 3D space is an asymmetric rectangular parallelepiped, and when the second signaling information indicates the second value, the second signaling information may indicate that the 3D space is a symmetric rectangular parallelepiped. Further, the signal-represented sizes of the 3D space in the x, y, and z directions may be determined based on the second signaling information indicating the first value.
[0019] According to one aspect of the present disclosure, there is provided an apparatus for performing point cloud geometry decoding. The apparatus includes: a receiving module, configured to receive first signaling information from an encoded bitstream of a point cloud including a set of points in a three-dimensional (3D) space, the first signaling information indicating segmentation information of the point cloud; a first determination module, configured to determine second signaling information from the encoded bitstream of the point cloud based on the first signaling information indicating a first value, the second signaling information indicating a segmentation pattern of the set of points in the 3D space; a second determination module, configured to determine the segmentation pattern of the set of points in the 3D space based on the second signaling information; and a reconstruction module, configured to reconstruct the point cloud based on the segmentation pattern.
[0020] According to one aspect of the present disclosure, there is provided an apparatus for performing point cloud geometry decoding. The apparatus includes: a receiving module, configured to receive first signaling information from an encoded bitstream of a point cloud including a set of points in a three-dimensional (3D) space, the first signaling information indicating segmentation information of the point cloud; and a determination module, configured to determine a segmentation pattern of the set of points in the 3D space based on the first signaling information, the segmentation pattern including a plurality of segmentation levels; and a reconstruction module, configured to reconstruct the point cloud based on the segmentation pattern.
[0021] According to one aspect of the present disclosure, there is provided a computer device. The computer device includes: one or more non-transitory computer-readable storage media configured to store computer program code; and one or more computer processors configured to access the computer program code and operate in accordance with what is indicated by the computer program code to execute the method according to any one of the above embodiments.
[0022] According to one aspect of the present disclosure, there is provided a non-transitory computer-readable medium. A computer program is stored on the non-transitory computer-readable medium and configured to cause one or more computer processors to execute the method according to any one of the above embodiments.
[0023] In the method and apparatus for point cloud geometry decoding provided by the present invention, by using the first signaling information and the second signaling information to determine the point cloud segmentation pattern, the node decomposition type at each level is clearly represented using corresponding signals, realizing a more flexible segmentation method. Thus, the method provided by the present invention improves the flexibility and efficiency of point cloud set encoding and decoding, thereby improving the performance of point cloud compression and decompression. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0025] Figure 1 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment;
[0026] Figure 2 is a schematic diagram of a simplified block diagram of a streaming system according to an embodiment;
[0027] Figure 3 shows a block diagram of an encoder for encoding a point cloud frame according to some embodiments;
[0028] Figure 4 shows a block diagram of a decoder for decoding a compressed bitstream corresponding to a point cloud frame according to some embodiments;
[0029] Figure 5 is a schematic diagram of a simplified block diagram of a video decoder according to an embodiment;
[0030] Figure 6 is a schematic diagram of a simplified block diagram of a video encoder according to an embodiment;
[0031] Figure 7 shows a block diagram of a decoder for decoding a compressed bitstream corresponding to a point cloud frame according to some embodiments;
[0032] Figure 8 shows a block diagram of an encoder for encoding a point cloud frame according to some embodiments;
[0033] Figure 9 shows a diagram illustrating the segmentation of a cube based on octree segmentation technology according to some embodiments of the present disclosure;
[0034] Figure 10 shows an example of octree segmentation and an octree structure corresponding to the octree segmentation according to some embodiments of the present disclosure;
[0035] Figure 11 shows a point cloud having a shorter bounding box in the z direction according to some embodiments of the present disclosure;
[0036] Figure 12 shows a diagram illustrating the segmentation of a cube along the x-y, x-z, and y-z axes based on octree segmentation technology according to some embodiments of the present disclosure;
[0037] Figure 13 shows a diagram illustrating the segmentation of a cube along the x, y, and z axes based on binary tree segmentation technology according to some embodiments of the present disclosure;
[0038] Figure 14 FIG. 1 shows a first flowchart outlining a first processing example according to some embodiments.
[0039] Figure 15 FIG. 2 shows a second flowchart outlining a second processing example according to some embodiments.
[0040] Figure 16 FIG. 3 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION
[0041] Advanced 3D representations of the world enable more immersive forms of interaction and communication and also enable machines to understand, interpret, and navigate our world. 3D point clouds have emerged as an effective representation of such information. Many application cases associated with point cloud data have been identified, and corresponding requirements for point cloud representation and compression have been established. For example, 3D point clouds can be used for object detection and localization in autonomous driving. 3D point clouds can also be used for map building in a geographic information system (GIS) and for visualizing and archiving cultural heritage objects and collections in cultural heritage.
[0042] A point cloud generally can refer to a set of points in 3D space, each point having an associated attribute. The attributes can include color, material properties, texture information, intensity attributes, reflectance attributes, motion-related attributes, morphological attributes, and / or various other attributes. Point clouds can be used to reconstruct an object or scene as a combination of such points. These points can be captured using multiple cameras, depth sensors, and / or lidars with various settings and can consist of thousands to billions of points to realistically represent the reconstructed scene.
[0043] Compression techniques can reduce the amount of data required to represent a point cloud for faster transmission or reduced storage. Thus, there is a need for lossy compression of point clouds for real-time communication and six Degrees of Freedom (6DoF) virtual reality technologies. Additionally, techniques for lossless point cloud compression are sought in the context of dynamic mapping for applications such as autonomous driving and cultural heritage. Therefore, ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) has started working on standards for compressing geometry and attributes such as color and reflectance, scalable / progressive coding, encoding of point cloud sequences captured over time, and random access to subsets of point clouds.
[0044] Figure 1A simplified block diagram of a communication system (100) according to an embodiment of the present disclosure is shown. The communication system (100) includes a plurality of terminal devices that can communicate with each other via, for example, a network (150). For example, the communication system (100) includes a pair of terminal devices (110) and (120) interconnected via a network (150). In Figure 1 In an example, the first pair of terminal devices (110) and (120) can perform unidirectional transmission of point cloud data. For example, the terminal device (110) can compress the point cloud (e.g., points representing a structure) captured by a sensor (105) connected to the terminal device (110). The compressed point cloud can be transmitted, for example, in the form of a bitstream via the network (150) to another terminal device (120). The terminal device (120) can receive the compressed point cloud from the network (150), decompress the bitstream to reconstruct the point cloud, and appropriately display the reconstructed point cloud. Unidirectional data transmission may be common in media service applications and the like.
[0045] In Figure 1 In an example, the terminal devices (110) and (120) can be shown as a server and a personal computer, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure are applicable to laptop computers, tablet computers, smart phones, game terminals, media players, and / or dedicated three-dimensional (3D) equipment. The network (150) represents any number of networks that transmit the compressed point cloud between the terminal devices (110) and (120). The network (150) can include, for example, cables (wired) and / or wireless communication networks. The network (150) can exchange data in a circuit-switched channel and / or a packet-switched channel. Representative networks include telecommunication networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless otherwise explained below, the architecture and topology of the network (150) may be unimportant for the operation of the present disclosure.
[0046] Figure 2 A simplified block diagram of a streaming system (200) according to an embodiment is shown. Figure 2 An example of the disclosed subject is the application to point clouds. The disclosed subject can be equivalently applicable to other point cloud-supported applications, such as 3D telepresence applications, virtual reality applications, and the like.
[0047] A streaming system (200) may include a capture subsystem (213). The capture subsystem (213) may include a point cloud source (201), such as a light detection and ranging (LIDAR) system, a 3D camera, a 3D scanner, a graphics generation component that generates an uncompressed point cloud in software, and similar graphics generation components that generate, for example, an uncompressed point cloud (202). In an example, the point cloud (202) includes points captured by a 3D camera. The point cloud (202) is depicted as a thick line to emphasize the high data volume when compared to the compressed point cloud (204) (the bitstream of the compressed point cloud). The compressed point cloud (204) may be generated by an electronic device (220) that includes an encoder (203) coupled to the point cloud source (201). The encoder (203) may include hardware, software, or a combination thereof to implement or enforce aspects of the disclosed subject matter described in more detail below. The compressed point cloud (204) (or the bitstream of the compressed point cloud (204)), depicted as a thin line to emphasize the lower data volume when compared to the stream of the point cloud (202), may be stored on a streaming server (205) for future use. One or more streaming client subsystems, such as Figure 2 the client subsystems (206) and (208) in
[0048] Note that the electronic device (220) and the electronic device (230) may include other components (not shown). For example, the electronic device (220) may include a decoder (not shown), and the electronic device (230) may also include an encoder (not shown).
[0049] In some streaming systems, the compressed point clouds (204), (207), and (209) (e.g., the bitstreams of the compressed point clouds) may be compressed according to certain criteria. In some examples, video coding standards are used in the compression of the point cloud. Examples of these standards include High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), etc.
[0050] Figure 3A block diagram of a V-PCC encoder (300) for encoding a point cloud frame according to some embodiments is shown. In some embodiments, the V-PCC encoder (300) can be used in a communication system (100) and a streaming system (200). For example, the encoder (203) can be configured and operated in a similar manner as the V-PCC encoder (300).
[0051] The V-PCC encoder (300) receives a point cloud frame as an uncompressed input and generates a bitstream corresponding to the compressed point cloud frame. In some embodiments, the V-PCC encoder (300) can receive the point cloud frame from a point cloud source such as the point cloud source (201).
[0052] In Figure 3 the example of, the V-PCC encoder (300) includes a block generation module (306), a block encapsulation module (308), a geometry image generation module (310), a texture image generation module (312), a block information module (304), an occupancy map module (314), a smoothing module (336), an image padding module (316) and (318), a group extension module (320), a video compression module (322), (323) and (332), an auxiliary block information compression module (338), an entropy compression module (334), and a multiplexer (324).
[0053] According to one aspect of the present disclosure, the V-PCC encoder (300) converts a 3D point cloud frame into an image-based representation and some metadata (e.g., an occupancy map and block information), which are used to convert the compressed point cloud back to the decompressed point cloud. In some examples, the V-PCC encoder (300) can convert the 3D point cloud frame into a geometry image, a texture image, and an occupancy map, and then encode the geometry image, the texture image, and the occupancy map into a bitstream using video coding techniques. Generally, a geometry image is a 2D image having pixels filled with geometric values associated with points projected onto the pixels, and the pixels filled with geometric values can be referred to as geometric samples. A texture image is a 2D image having pixels filled with texture values associated with points projected onto the pixels, and the pixels filled with texture values can be referred to as texture samples. An occupancy map is a 2D image having pixels filled with values indicating whether a block is occupied or unoccupied.
[0054] A block generally may refer to a continuous subset of a surface described by a point cloud. In one example, a block includes points having surface normal vectors that deviate from each other by less than a threshold amount. A block generation module (306) divides the point cloud into a set of blocks that may or may not overlap, such that each block can be described by a depth field with respect to a plane in 2D space. In some embodiments, the block generation module (306) aims to decompose the point cloud into the minimum number of blocks with smooth boundaries while also minimizing the reconstruction error.
[0055] A block information module (304) may collect block information indicating the size and shape of a block. In some examples, the block information may be encapsulated into an image frame and then encoded by an auxiliary block information compression module (338) to generate compressed auxiliary block information.
[0056] A block encapsulation module (308) is configured to map the extracted blocks to a two-dimensional (2D) grid while minimizing unused space and ensuring that each M×M (e.g., 16x16) block of the grid is associated with a unique block. Effective block encapsulation can directly affect compression efficiency by minimizing unused space or ensuring temporal consistency.
[0057] A geometric image generation module (310) may generate a 2D geometric image associated with the geometry of the point cloud at a given block location. A texture image generation module (312) may generate a 2D texture image associated with the texture of the point cloud at a given block location. The geometric image generation module (310) and the texture image generation module (312) use the 3D-to-2D mapping computed during the encapsulation process to store the geometry and texture of the point cloud as images. To better handle the case of projecting multiple points to the same sample, each block is projected to two images called layers. In an example, the geometric image is represented by a monochrome frame of WxH in YUV420-8-bit format. To generate the texture image, the texture generation process uses the reconstructed / smoothed geometry to compute the color to be associated with the resampled points.
[0058] An occupancy map module (314) may generate an occupancy map that describes the occupancy information at each cell. For example, the occupancy map includes a binary map that indicates for each cell of the grid whether the cell belongs to empty space or to the point cloud. In an example, the occupancy map uses binary information that describes for each pixel whether the pixel is filled. In another example, the occupancy map uses binary information that describes for each pixel block whether the pixel block is filled.
[0059] The occupancy map generated by the occupancy map module (314) can be compressed using lossless coding or lossy coding. When using lossless coding, the entropy compression module (334) is used to compress the occupancy map. When using lossy coding, the video compression module (332) is used to compress the occupancy map.
[0060] Note that the block encapsulation module (308) can leave some blank space between the 2D blocks encapsulated in the image frame. The image padding modules (316) and (318) can fill the blank space (referred to as padding) to generate an image frame that can be adapted to 2D video and image codecs. Image padding is also referred to as background padding, which can fill the unused space with redundant information. In some examples, good background padding minimally increases the bit rate and does not introduce significant coding distortion around the block boundaries.
[0061] The video compression modules (322), (323), and (332) can encode 2D images such as the padded geometry image, the padded texture image, and the occupancy map based on a suitable video coding standard such as HEVC, VVC, etc. In an example, the video compression modules (322), (323), and (332) are separate components that operate separately. Note that in another example, the video compression modules (322), (323), and (332) can be implemented as a single component.
[0062] In some examples, the smoothing module (336) is configured to generate a smoothed image of the reconstructed geometry image. The smoothed image can be provided to the texture image generation module (312). Then, the texture image generation module (312) can adjust the generation of the texture image based on the reconstructed geometry image. For example, in the case where the block shape (e.g., geometric shape) is slightly distorted during encoding and decoding, the distortion can be considered when generating the texture image to correct the distortion of the block shape.
[0063] In some embodiments, the group extension (320) is configured to fill the pixels around the object boundary with redundant low-frequency content to improve the coding gain and visual quality of the reconstructed point cloud.
[0064] The multiplexer (324) can multiplex the compressed geometry image, the compressed texture image, the compressed occupancy map, and / or the compressed auxiliary block information into a compressed bitstream.
[0065] Figure 4A block diagram of a V-PCC decoder (400) for decoding a compressed bitstream corresponding to a point cloud frame is shown. In some embodiments, the V-PCC decoder (400) can be used in a communication system (100) and a streaming system (200). For example, the decoder (210) can be configured to operate in a similar manner to the V-PCC decoder (400). The V-PCC decoder (400) receives the compressed bitstream and generates a reconstructed point cloud based on the compressed bitstream.
[0066] In Figure 4 the example of, the V-PCC decoder (400) includes a demultiplexer (432), video decompression modules (434) and (436), an occupancy map decompression module (438), an auxiliary block information decompression module (442), a geometry reconstruction module (444), a smoothing module (446), a texture reconstruction module (448), and a color smoothing module (452).
[0067] The demultiplexer (432) can receive the compressed bitstream and separate it into a compressed texture image, a compressed geometry image, a compressed occupancy map, and compressed auxiliary block information.
[0068] The video decompression modules (434) and (436) can decode the compressed images according to suitable standards (e.g., HEVC, VVC, etc.) and output the decompressed images. For example, the video decompression module (434) decodes the compressed texture image and outputs the decompressed texture image; and the video decompression module (436) decodes the compressed geometry image and outputs the decompressed geometry image.
[0069] The occupancy map decompression module (438) can decode the compressed occupancy map according to suitable standards (e.g., HEVC, VVC, etc.) and output the decompressed occupancy map.
[0070] The auxiliary block information decompression module (442) can decode the compressed auxiliary block information according to suitable standards (e.g., HEVC, VVC, etc.) and output the decompressed auxiliary block information.
[0071] The geometry reconstruction module (444) can receive the decompressed geometry image and generate a reconstructed point cloud geometry based on the decompressed occupancy map and the decompressed auxiliary block information.
[0072] The smoothing module (446) can smooth the inconsistencies at the edges of the blocks. The smoothing process aims to mitigate the potential discontinuities that may occur at the block boundaries due to compression artifacts. In some embodiments, a smoothing filter can be applied to the pixels located on the block boundaries to mitigate the distortion that may be caused by compression / decompression.
[0073] The texture reconstruction module (448) can determine the texture information of the points in the point cloud based on the decompressed texture image and the smoothed geometry.
[0074] The color smoothing module (452) can smooth the coloring inconsistencies. Non-adjacent blocks in 3D space are typically encapsulated adjacent to each other in the 2D video. In some examples, the pixel values from non-adjacent blocks may be blended by a block-based video codec. The purpose of color smoothing is to reduce the visible artifacts that appear at the block boundaries.
[0075] Figure 5 A block diagram of a video decoder (510) according to an embodiment of the present disclosure is shown. The video decoder (510) can be used in the V-PCC decoder (400). For example, the video decompression modules (434) and (436), and the occupancy map decompression module (438) can be similarly configured as the video decoder (510).
[0076] The video decoder (510) can include a parser (520) to reconstruct symbols (521) based on a compressed image such as an encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (510). The parser (520) can perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence can be performed according to video coding techniques or standards and can follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) can extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroups can include Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. The parser (520) can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0077] The parser (520) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory to create symbols (521).
[0078] Depending on the type of the coded video picture or a part thereof (e.g., inter picture and intra picture, inter block and intra block) and other factors, the reconstruction of the symbol (521) may involve multiple different units. Which units are involved and the way of involvement can be controlled by subgroup control information parsed by the parser (520) from the coded video sequence. For the sake of brevity, the flow of such subgroup control information between the parser (520) and the multiple units hereinafter is not described.
[0079] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually divided into multiple functional units as described hereinafter. In a practical implementation running under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually divide into the following functional units.
[0080] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives the quantized transform coefficients as the symbol (521) and control information from the parser (520), including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (551) may output a block including sample values that can be input to the aggregator (555).
[0081] In some cases, the output samples of the scaler / inverse transform unit (551) may belong to an intra-coded block; that is, a block that does not use the predictive information from the previously reconstructed picture but may use the predictive information from the previously reconstructed part of the current picture. Such predictive information can be provided by the intra picture prediction unit (552). In some cases, the intra picture prediction unit (552) uses the surrounding reconstructed information extracted from the current picture buffer (558) to generate a block having the same size and shape as the block being reconstructed. For example, the current picture buffer (558) buffers the partially reconstructed current picture and / or the fully reconstructed current picture. In some cases, the aggregator (555) adds the prediction information already generated by the intra prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) based on each sample.
[0082] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to a block that is inter-frame coded and potentially motion compensated. In such cases, the motion compensation prediction unit (553) may access the reference picture memory (557) to extract samples for prediction. After motion compensating the extracted samples according to the symbols (521) belonging to the block, these samples may be added to the output of the scaler / inverse transform unit (551) (referred to as residual samples or residual signals in this case) by an aggregator (555) to generate output sample information. The address within the reference picture memory (557) from which the motion compensation prediction unit (553) extracts prediction samples may be controlled by motion vectors and made available to the motion compensation prediction unit (553) in the form of symbols (521) that may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values extracted from the reference picture memory (557) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and the like.
[0083] The output samples of the aggregator (555) may be subject to various loop filtering techniques in the loop filter unit (556). Video compression techniques may include in-loop filter techniques that are controlled by parameters included in the coded video sequence (also referred to as the coded video bitstream) and that are available to the loop filter unit (556) as symbols (521) from a parser (520), but video compression techniques may also respond to meta-information obtained during decoding of previous (in decoding order) portions of the coded picture or coded video sequence and to previously reconstructed and loop-filtered sample values.
[0084] The output of the loop filter unit (556) may be a sample stream that may be output to a rendering device and stored in the reference picture memory (557) for future inter-frame picture prediction.
[0085] Once fully reconstructed, some coded pictures may be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture has been fully reconstructed and the coded picture has been identified as a reference picture (by, for example, a parser (520)), the current picture buffer (558) may become part of the reference picture memory (557), and a new current picture buffer may be reallocated before starting to reconstruct subsequent coded pictures.
[0086] A video decoder (510) may perform decoding operations according to a predetermined video compression technique in a standard such as the ITU-T H.265 recommendation. In the sense that an encoded video sequence conforms to the syntax specified by the video compression technique or standard, the encoded video sequence may conform to the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard. Specifically, a profile may select certain tools from all the tools available in the video compression technique or standard as the only tools available under the profile. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, mega samples per second), maximum reference picture size, etc. In some cases, the limits set by the level may be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled in the encoded video sequence.
[0087] Figure 6 A block diagram of a video encoder (603) according to an embodiment of the present disclosure is shown. The video encoder (603) may be used in a V-PCC encoder (300) for compressing point clouds. In an example, the video compression modules (322) and (323) and the video compression module (332) are configured similarly to the encoder (603).
[0088] The video encoder (603) may receive images such as a padded geometry image, a padded texture image, etc., and may generate a compressed image.
[0089] According to an embodiment, the video encoder (603) may encode and compress pictures (images) of a source video sequence into an encoded video sequence (compressed image) in real time or under any other time constraints required by an application. Implementing an appropriate encoding speed is a function of the controller (650). In some embodiments, the controller (650) controls other functional units as described below and is functionally coupled to other functional units. For simplicity, such couplings are not depicted. The parameters set by the controller (650) may include rate control related parameters (picture skip, quantizer, λ value of rate distortion optimization technique,...), picture size, group of picture (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured to have other suitable functions related to the video encoder (603) optimized for a specific system design.
[0090] In some embodiments, the video encoder (603) is configured to operate in an encoding loop. As a simple description, in an example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols in a manner similar to how a (remote) decoder would also create sample data (since in the video compression or techniques considered in the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to the reference picture memory (634). Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content in the reference picture memory (634) is also bit-exact corresponding between the local encoder and the remote encoder. That is, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This reference picture synchronization principle (and the drift that occurs, for example, when the synchronization cannot be maintained due to channel errors) is also used in some related techniques.
[0091] The operation of the "local" decoder (633) may be the same as that of the "remote" decoder of the video decoder (510) described above in conjunction with Figure 5 However, briefly referring additionally to Figure 5 , when the symbols are available and the entropy encoder (645) and the parser (520) can encode / decode the symbols losslessly into an encoded video sequence, the entropy decoding part of the video decoder (510), including the buffer memory and the parser (520), may not be fully implemented in the local decoder (633).
[0092] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on decoder operations. The description of the encoder technology can be simplified because the encoder technology is reciprocal to the decoder technology described comprehensively. More detailed descriptions are only needed in certain areas and are provided below.
[0093] During operation, in some examples, the source encoder (630) may perform motion-compensated predictive coding that predictively encodes an input picture by referring to one or more previously encoded pictures designated as "reference pictures" from a video sequence. In this way, the encoding engine (632) encodes the difference between a pixel block of the input picture and a pixel block of the reference picture, and the reference picture can be selected as the prediction reference for the input picture.
[0094] The local video decoder (633) can decode the encoded video data of pictures that can be designated as reference pictures based on the symbols created by the source encoder (630). The operation of the encoding engine (632) can advantageously be lossy processing. When the encoded video data can be decoded at the video decoder ( Figure 6 not shown), the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local video decoder (633) repeats the decoding process that can be performed by the video decoder for the reference pictures, and can store the reconstructed reference pictures in the reference picture cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference pictures, which has the same content (no transmission errors) as the reconstructed reference pictures to be obtained by the remote video decoder.
[0095] The predictor (635) can perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) can search in the reference picture memory (634) for sample data (as a candidate reference pixel block) or some metadata that can be used as an appropriate prediction reference for the new picture, such as reference picture motion vectors, block shapes, etc. The predictor (635) can operate on a per-pixel block basis of the sample blocks to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture can have prediction references taken from multiple reference pictures stored in the reference picture memory (634).
[0096] The controller (650) can manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding the video data.
[0097] The outputs of all the above functional units can be entropy encoded in the entropy encoder (645). The entropy encoder (645) converts the symbols into an encoded video sequence by losslessly compressing the symbols generated by various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc. to generate the compressed image 643.
[0098] The controller (650) can manage the operations of the video encoder (603). During encoding, the controller (650) can specify a specific encoding picture type for each encoded picture, which may affect the encoding techniques that can be applied to the corresponding picture. For example, pictures can generally be designated as one of the following picture types:
[0099] An Intra picture (I picture) is a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of Intra pictures, including for example Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of these variations of I pictures and their corresponding applications and characteristics.
[0100] A Predictive picture (P picture) is a picture that can be encoded and decoded using Intra prediction or Inter prediction, where the Intra prediction or Inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.
[0101] A Bi - predictive picture (B picture) is a picture that can be encoded and decoded using Intra prediction or Inter prediction, where the Intra prediction or Inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.
[0102] Generally, a source picture can be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples respectively), and encoded on a block - by - block basis. A block can be predictively encoded with reference to other (already encoded) blocks as determined by the coding assignment applied to the block of the corresponding picture. For example, blocks of an I picture can be non - predictively encoded, or blocks of an I picture can be predictively encoded (spatial prediction or Intra prediction) with reference to already encoded blocks of the same picture. Pixel blocks of a P picture can be predictively encoded with reference to one previously encoded reference picture via spatial prediction or via temporal prediction. Blocks of a B picture can be predictively encoded with reference to one or two previously encoded reference pictures via spatial prediction or via temporal prediction.
[0103] The video encoder (603) can perform encoding operations according to a predetermined video coding technique or standard such as the ITU - T H.265 recommendation. In its operation, the video encoder (603) can perform various compression operations, including predictive coding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0104] The video can be in the form of multiple source pictures (images) in a time series. Intra-picture prediction (usually abbreviated as intra prediction) utilizes the spatial correlation within a given picture, while inter-picture prediction utilizes the (temporal or other) correlation between pictures. In an example, a particular picture being encoded / decoded is segmented into blocks, and the particular picture being encoded / decoded is referred to as the current picture. In the case where a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.
[0105] In some embodiments, bidirectional prediction techniques can be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be past and future respectively in display order). The block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.
[0106] In addition, merge mode techniques can be used in inter-picture prediction to improve encoding efficiency.
[0107] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, pictures in a video picture sequence are segmented into coding tree units (CTUs) for compression, and CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Each CTU can also be recursively split into one or more coding units (CUs) in a quadtree. For example, a 64×64 pixel CTU can be split into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In an example, each CU is analyzed to determine a prediction type for the CU, such as an inter-prediction type or an intra-prediction type. Depending on the temporal and / or spatial predictability, the CU is split into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, prediction operations in decoding (encoding / decoding) are performed on a prediction block basis. Using a luminance prediction block as an example of a prediction block, a prediction block includes a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0108] The G-PCC model can compress geometric information and associated attributes, such as color or reflectance, separately. Geometric information, which is the 3D coordinates of a point cloud, can be encoded by octree decomposition of its occupancy information. On the other hand, prediction and lifting techniques can be used to compress attributes based on the reconstructed geometry. For example, octree segmentation processing is discussed in Figures 7 to 13 the context.
[0109] Figure 7 FIG. shows a block diagram of a G-PCC decoder (800) applied during G-PCC decomposition processing according to an embodiment. The decoder (800) can be configured to receive a compressed bitstream and perform point cloud data decompression to decompress the bitstream to generate decoded point cloud data. In an embodiment, the decoder (800) can include an arithmetic decoding module (810), an inverse quantization module (820), an octree decoding module (830), a LOD generation module (840), an inverse quantization module (850), and an inverse interpolation-based prediction module (860).
[0110] As shown in the figure, a compressed bitstream (801) can be received at an arithmetic decoding module (810). The arithmetic decoding module (810) is configured to decode the compressed bitstream (801) to obtain the quantized prediction residuals (if generated) and occupancy codes (or symbols) of the point cloud. An octree decoding module (830) is configured to generate the quantized positions of the points in the point cloud based on the occupancy codes. An inverse quantization module (850) is configured to generate the reconstructed positions of the points in the point cloud based on the quantized positions provided by the octree decoding module (830).
[0111] A LOD generation module (840) is configured to reorganize the points into different LODs based on the reconstructed positions and determine the order based on the LOD. An inverse quantization module (820) is configured to generate the reconstructed prediction residuals based on the quantized prediction residuals received from the arithmetic decoding module (810). A prediction module (860) based on inverse interpolation is configured to perform an attribute prediction process to generate the reconstructed attributes of the points in the point cloud based on the reconstructed prediction residuals received from the inverse quantization module (820) and the order based on the LOD received from the LOD generation module (840).
[0112] In addition, in one example, the reconstructed attributes generated from the prediction module (860) based on inverse interpolation together with the reconstructed positions generated from the inverse quantization module (850) correspond to the decoded point cloud (or reconstructed point cloud) (802) output from the decoder (800).
[0113] Figure 8 A block diagram of a G-PPC encoder (700) according to an embodiment is shown. The encoder (700) can be configured to receive point cloud data and compress the point cloud data to generate a bitstream carrying the compressed point cloud data. In an embodiment, the encoder (700) can include a position quantization module (710), a duplicate point removal module (712), an octree encoding module (730), an attribute conversion module (720), a level of detail (LOD) generation module (740), a prediction module (750) based on interpolation, a residual quantization module (760), and an arithmetic encoding module (770).
[0114] As shown in the figure, an input point cloud (701) can be received at an encoder (700). The positions (e.g., 3D coordinates) of the point cloud (701) are provided to a quantization module (710). The quantization module (710) is configured to quantize the coordinates to generate quantized positions. A duplicate point removal module (712) is configured to receive the quantized positions and perform a filtering process to identify and remove duplicate points. An octree encoding module (730) is configured to receive the filtered positions from the duplicate point removal module (712) and perform an octree-based encoding process to generate a sequence of occupancy codes (or symbols) that describe the occupancy of a 3D grid of voxels. The occupancy codes are provided to an arithmetic coding module (770).
[0115] An attribute conversion module (720) is configured to receive the attributes of the input point cloud and perform an attribute conversion process to determine the attribute values of each voxel when multiple attribute values are associated with corresponding voxels. The attribute conversion process can be performed on the re-ordered points output from the octree encoding module (730). The attributes after the conversion operation are provided to an interpolation-based prediction module (750). A LOD generation module (740) is configured to operate on the re-ordered points output from the octree encoding module (730) and re-organize these points into different LODs. The LOD information is provided to the interpolation-based prediction module (750).
[0116] The interpolation-based prediction module (750) processes the points and generates prediction residuals based on the LOD-based order indicated by the LOD information from the LOD generation module (740) and the converted attributes received from the attribute conversion module (720). A residual quantization module (760) is configured to receive the prediction residuals from the interpolation-based prediction module (750) and perform quantization to generate quantized prediction residuals. The quantized prediction residuals are provided to the arithmetic coding module (770). The arithmetic coding module (770) is configured to receive the occupancy codes from the octree encoding module (730), candidate indices (if used), the quantized prediction residuals from the residual quantization module (760), and other information and perform entropy coding to further compress the received values or information. As a result, a compressed bitstream (702) carrying the compressed information can be generated. The bitstream (702) can be transmitted or otherwise provided to a decoder that decodes the compressed bitstream, or can be stored in a storage device.
[0117] Note that the interpolation-based prediction module (750) and the inverse interpolation-based prediction module (860) configured to implement the attribute prediction techniques disclosed herein can be included in other decoders or encoders, which can have the same Figure 7 and Figure 8Structures that are similar or different from those shown. Additionally, in various examples, the encoder (700) and the decoder (800) can be included in the same device or separate devices.
[0118] In various embodiments, the encoder (300), the decoder (400), the encoder (700), and / or the decoder (800) can be implemented in hardware, software, or a combination thereof. For example, the encoder (300), the decoder (400), the encoder (700), and / or the decoder (800) can be implemented using processing circuitry such as one or more integrated circuits (ICs) that operate with or without software, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc. In another example, the encoder (300), the decoder (400), the encoder (700), and / or the decoder (800) can be implemented as software or firmware that includes instructions stored in a non-volatile (or non-transitory) computer-readable storage medium. When executed by processing circuitry such as one or more processors, the instructions cause the processing circuitry to perform the functions of the encoder (300), the decoder (400), the encoder (700), and / or the decoder (800).
[0119] Segmenting a point cloud defined by a 3D cube in a symmetric manner along all axes (e.g., the x, y, and z axes) can produce eight sub-cubes, which is referred to as octree (OT) segmentation in point cloud compression (PCC). OT segmentation is similar to binary-tree (BT) segmentation in one-dimensional space and quadtree (QT) segmentation in two-dimensional space. The idea of OT segmentation can be found in Figure 9 as shown, where the 3D cube (900) represented by the solid line can be segmented into eight smaller equal-sized cubes represented by the dashed lines. As Figure 9 shown, the octree segmentation technique can segment the 3D cube (900) into eight smaller equal-sized cubes 0 to 7.
[0120] In the octree segmentation technique (e.g., in TMC13), if an octree geometry codec is used, the geometry coding is performed as follows. First, a cube axis-aligned bounding box B can be defined by two extreme points (0,0,0) and (2 d ,2 d ,2 d ), where 2 dThe size of the bounding box B is defined and d can be encoded as a bitstream. Thus, all points within the defined bounding box B can be compressed.
[0121] An octree structure can then be constructed by recursively subdividing the bounding box B. At each level of subdivision, the cube can be divided into 8 sub-cubes. After iteratively subdividing k (k ≤ d) times, the size of the sub-cubes can be (2 d-k , 2 d-k , 2 d -k ). Then, an 8-bit code such as an occupancy code can be generated by associating a 1-bit value with each sub-cube to indicate whether the corresponding sub-cube contains points (i.e., full and having a value of 1) or does not contain points (i.e., empty and having a value of 0). Only full sub-cubes with a size greater than 1 (i.e., non-voxels) can be further subdivided. Then, the occupancy code of each cube can be compressed by an arithmetic encoder.
[0122] The decoding process can start by reading the dimensions of the bounding box B from the bitstream. Then, the same octree structure can be constructed by subdividing the bounding box B according to the decoded occupancy code. Examples of two-level OT partitioning and corresponding occupancy codes can be found in Figure 10 where the shaded cubes and nodes indicate that the cubes and nodes are occupied by points.
[0123] Figure 10 An example of an octree partitioning (1010) and an octree structure (1020) corresponding to the octree partitioning (1010) according to some embodiments of the present disclosure is shown. Figure 10 A two-level partitioning in the octree partitioning (1010) is shown. The octree structure (1020) includes nodes (N0) corresponding to the cube boxes used for the octree partitioning (1010). At the first level, the cube box is divided into 8 sub-cube boxes numbered from 0 to 7 according to the numbering technique shown in Figure 9 The occupancy code for the partitioning of node N0 is "10000001" in binary form, which indicates that the first sub-cube box represented by node N0-0 and the eighth sub-cube box represented by node N0-7 contain points in the point cloud, while the other sub-cube boxes are empty.
[0124] Then, at the second level of partitioning, the first sub-cube box (represented by node N0-0) and the eighth sub-cube box (represented by node N0-7) are each further subdivided into eight octants. For example, according to Figure 9The numbered technique shown in the figure divides the first sub-cube box (represented by node N0-0) into 8 smaller sub-cube boxes numbered from 0 to 7. The occupancy code for the division of node N0-0 is "00011000" in binary form, which indicates that the fourth smaller sub-cube box (represented by node N0-0-3) and the fifth smaller sub-cube box (represented by node N0-0-4) contain points in the point cloud, while the other smaller sub-cube boxes are empty. At the second level, similarly, the seventh sub-cube box (represented by node N0-7) is divided into 8 smaller sub-cube boxes, as Figure 10 shown in
[0125] In Figure 10 the example, nodes corresponding to non-empty cube spaces (e.g., cube boxes, sub-cube boxes, smaller sub-cube boxes, etc.) are shaded in gray and are called shaded nodes.
[0126] In the original TMC13 design, for example as described above, the bounding box B can be restricted to a cube with the same size for all dimensions, and thus OT division can be performed for all sub-cubes at each node, where at each node, the size of the sub-cubes for all dimensions is halved. The OT division can be performed recursively until the size of the sub-cubes reaches 1. However, the division performed in this way may not be effective for all cases, especially in the case where points are non-uniformly distributed in a 3D scene (or 3D space).
[0127] An extreme case can be a 2D plane in 3D space, where all points can lie on the x-y plane in 3D space and the variation on the z-axis can be zero. In this case, performing OT division on the cube B as the starting point may waste a large number of bits to represent the occupancy information in the z direction, which is redundant and useless. In practical applications, the worst-case scenario may not occur frequently. However, point clouds usually have less variation in one direction than in other directions. As Figure 11 shown in
[0128] In the quadtree and binary-tree (QtBt) division, the bounding box B is not limited to a cube, but the bounding box B can be a rectangular cuboid of any size to better fit the shape of the 3D scene or object. In implementation, the size of the bounding box B can be represented as a power of 2, for example
[0129] Since the bounding box B may not be a perfect cube, in some cases, a node may not (or cannot) be split along all directions. If the split is performed in all three directions, the split is a typical OT split. If the split is performed in two of the three directions, the split is thus a QT split in 3D. If the split is performed in only one direction, the split is a BT split in 3D. In Figure 12 and Figure 13 examples of QT and BT in 3D are shown respectively.
[0130] As Figure 12 shown, the 3D cube 1201 can be split into 4 sub-cubes 0, 2, 4, and 6 along the x-y axis. The 3D cube 1202 can be split into 4 sub-cubes 0, 1, 4, and 5 along the x-z axis. The 3D cube 1203 can be split into 4 sub-cubes 0, 1, 2, and 3 along the y-z axis. In Figure 13 the 3D cube 1301 can be split into 2 sub-cubes 0 and 4 along the x axis. The 3D cube 1302 can be split into 2 sub-cubes 0 and 2. The 3D cube 1303 can be split into 2 sub-cubes 0 and 1.
[0131] To define the conditions for implicit QT and BT splits in TMC13, two parameters (i.e., K and M) can be applied. The first parameter K (0 ≤ K ≤ max(d x , d y , d z ) - min(d x , d y , d z )) can define the maximum number of implicit QT and BT splits that can be performed before an OT split. The second parameter M (0 ≤ M ≤ min(d x , d y , d z )) can define the minimum size of implicit QT and BT splits, which indicates that implicit QT and BT splits are only allowed if all dimensions are greater than M.
[0132] More specifically, the first K splits can follow the rules in Table I, and the splits after the first K splits can follow the rules in Table II. If none of the conditions listed in the tables are met, an OT split can be performed.
[0133] Table I: Conditions for performing implicit QT or BT splits for the first K splits.
[0134] Along the x-y axis QT Along the x-z axis QT Along the y-z axis QT Condition <![CDATA[d z <d x =d y > <![CDATA[d y <d x =d z > <![CDATA[d x <d y =d z > Along the x axis BT Along the y axis BT Along the z axis BT Condition <![CDATA[d y <d x and d z <d x > <![CDATA[d x <d y and d z <d y > <![CDATA[d x <d z and d y <d z >
[0135] Table II: Conditions for performing implicit QT or BT splits after the first K splits.
[0136]
[0137] In an embodiment, the bounding box B may have the size of. Without loss of generality, the condition 0 < d x ≤ d y ≤ d z can be applied to the bounding box B. Based on these conditions, at the first K (K ≤ d z - d x ) depths, according to Table I, an implicit BT split can be performed along the z-axis, and then an implicit QT split can be performed along the y-z axis. The size of the child nodes can then become where δ y and δ z values (δ z ≥ δ y ≥ 0) can depend on the value of K. Additionally, an OT split can be performed d x - M times such that the remaining child nodes can have the size of. Next, according to Table II, an implicit BT split can be performed along the z-axis δ z - δ y times, and then an implicit QT split can be performed along the y-z axis δ y times. The remaining nodes can thus have a size of 2 (M,M,M) . Thus, the OT split can be performed M times to reach the minimum unit.
[0138] In the QtBt split, implicit rules are provided on how to apply the split of a given cuboid by switching between octree, quadtree, and binary tree at each level of node decomposition. After an initial decomposition of K levels via QtBt split according to a rule (e.g., Table I), another round of QtBt split can be performed according to another rule (e.g., Table II). If any condition in the rule is not satisfied in the above process, an octree decomposition (or octree split) can be applied.
[0139] The implicit rules can affect the effectiveness of QtBt as follows: (1) For point cloud data with a cuboid bounding box that is almost symmetric along the x, y, and z dimensions, the QtBt split does not show an encoding gain compared to related methods (e.g., implicit QtBt split) that perform Ot (octree) decomposition at all levels; and (2) For point cloud data with a cuboid bounding box that is highly asymmetric along the x, y, and z dimensions, the QtBt split has shown an encoding gain by skipping sending unnecessary occupancy information during decomposition.
[0140] In the current QtBt partitioning, certain restrictions can be set as follows. First, the QtBt partitioning can always enforce the use of an asymmetric bounding box, which may be useless or even counterproductive in cases where the point cloud has an almost symmetric bounding box. Second, by implementing Qt / Bt partitioning instead of Ot partitioning, Table I together with the parameter K can reduce the larger dimension according to a rule. However, in the case of a symmetric bounding box, Table I together with the parameter K may not allow Qt or Bt partitioning at the beginning. Third, in the case where the minimum size of the sub-box reaches M, Table II can be applied after the above-mentioned K splits and kicks in. Thus, Table II can reduce the larger dimension according to a rule until all dimensions become equal to M. Fourth, the current implicit rule (or implicit QtBt partitioning) can always enforce octree decomposition after the first (at most) K levels until the current QtBt partitioning reaches level M. That is, the current QtBt partitioning can not allow an arbitrary choice of Qt / Bt / Ot partitioning between points at these two levels.
[0141] In the present disclosure, a variety of methods are provided. These methods, for example, provide a simplification of the QtBt design (e.g., implicit QtBt partitioning) in TMC 13 for typical use cases based on the above discussion. These methods also allow for a more flexible partitioning manner, for example, by explicitly signaling the type of node decomposition at each level.
[0142] In one embodiment, a first partitioning method (or simplified QtBt partitioning) can be provided. The first partitioning method can be a special case of implicit QtBt partitioning and can be applied to a dataset with a highly asymmetric bounding box by setting K = 0 & M = 0. The first partitioning method can simplify the QtBt design (e.g., QtBt partitioning) and still bring coding benefits for the above-mentioned typical cases.
[0143] Compared with the QtBt splitting in TMC 13, the first splitting method may include the following features: (1) The implicit enable flag (e.g., implicit_qtbt_enabled_flag) in the QtBt splitting in TMC13 can be removed. (2) An asymmetric bounding box flag (e.g., asymmetric_bbox_enabled_flag) can be introduced to enable the use of an asymmetric bounding box. In an example, in the case where the asymmetric bounding box flag is set to a value such as 0 (also referred to as the second value) for symmetric or nearly symmetric bounding box data and the asymmetric bounding box flag is set to a value such as 1 (also referred to as the first value) for highly asymmetric bounding box data. (3) If the asymmetric bounding box flag is the first value, when the node decomposition level reaches 0 (or the last level), the implicit QtBt rule where K = 0 & M = 0 (e.g., Tables I and II) can be applied. Otherwise, if the asymmetric bounding box flag is the second value, the first splitting method can perform octree decomposition (or octree splitting).
[0144] According to the first splitting method, M = 0 can be utilized to apply the implicit QtBt rule shown in Table II to skip sending unnecessary occupancy information along certain dimensions, as shown in Table III for example.
[0145] Table III: Conditions for performing implicit QtBt splitting at level 0
[0146]
[0147] In one embodiment, a second splitting method (or explicit QtBt splitting) can be provided to send explicit signaling of the splitting decision. Instead of using fixed implicit rules in the current QtBt splitting, explicit signaling can be provided.
[0148] The second splitting method may include the following features: (1) An explicit QtBt enable flag (e.g., explicit_qtbt_enabled_flag) may be introduced to enable / disable explicit splitting decision signaling, while still introducing the asymmetric bounding box flag from the first splitting method. (2) In the case where the explicit QtBt enable flag is set to a value such as 0 (or a second value), the second splitting method falls back to (or may be equivalent to) the first splitting method described above. Thus, if the asymmetric bounding box flag (e.g., asymmetric_bbox_enabled_flag) is a value such as 1 (or a first value), the implicit QtBt rule where K = 0 & M = 0 (e.g., Tables I and II) may be applied when the node decomposition level reaches 0 (or the last level). If the asymmetric bounding box flag is the second value, the second splitting method may perform octree decomposition (or octree splitting). In one embodiment, in the case where the asymmetric bounding box flag is not used and the explicit QtBt enable flag is set to the second value (e.g., 0), the second splitting method may apply octree decomposition (or octree splitting) for all levels. (4) Instead of always performing octree splitting until the level reaches 0 (or the last level) as mentioned in the first splitting method, in the case where the explicit QtBt enable flag is set to the first value (e.g., 1), a 3-bit signal may be sent at each octree level in the octree levels to indicate whether to split along each of the x, y, and z axes. Thus, this 3-bit signal may indicate whether Bt splitting, Qt splitting, or Ot splitting may be applied at each level in the octree levels. In some embodiments, the implicit QtBt rules in TMC 13 (e.g., Tables I and II) may be applied to determine the 3-bit signal for each octree level in the octree levels.
[0149] It should be noted that since in the second splitting method, Ot / Qt / Bt splitting is allowed in an arbitrary manner along the way when the explicit QtBt enable flag is set to the first value, the maximum possible total number of splits may be three times the difference between the maximum node depth and the minimum node depth.
[0150] In an embodiment of the present disclosure, a third splitting method (or explicit QtBt type 2 splitting) can be provided to send explicit signaling of the splitting decision. Instead of using fixed implicit rules (e.g., Table I and Table II) in the current QtBt splitting, explicit signaling can be provided. Compared with the current QtBt splitting, the third splitting method can include the following features: (1) An explicit QtBt enable flag (e.g., explicit_qtbt_enabled_flag) can replace the implicit QtBt enable flag (e.g., implicit_qtbt_enabled_flag) in the QtBt splitting in TMC 13 to enable / disable explicit splitting decision signaling. (2) The asymmetric bounding box flag (e.g., asymmetric_bbox_enabled_flag) can be additionally signaled only when the explicit QtBt enable identification is a value such as 1 (or a first value) to enable / disable the use of the asymmetric bounding box. In the case where the asymmetric bounding box flag is the first value, the dimensions (i.e., sizes) of the asymmetric bounding box along the x, y, and z can also be signaled, rather than the maximum value among the three. Therefore, the asymmetric bounding box flag can be set to a value such as 0 (e.g., a second value) for symmetric or nearly symmetric bounding box data, and the asymmetric bounding box flag can be set to the first value (e.g., 1) for highly asymmetric bounding box data. (3) In the case where the explicit QtBt enable flag is set to the first value, a 3-bit signal can be sent at each octree level (or octree splitting level) in the octree level to indicate whether to split along each of the x, y, and z axes. In one embodiment, the implicit QtBt rules in TMC 13 (e.g., Table I and Table II) can be applied to determine the 3-bit signal for each octree level in the octree level. In another embodiment, other splitting rules can be applied to determine the 3-bit signal for each octree level in the octree level. The other splitting rules can be beneficial for the encoding of octree occupancy information and can also consider the characteristics of the data (e.g., octree occupancy information) or the acquisition mechanism of the date. (4) In the case where the explicit QtBt enable flag is set to the second value (e.g., 0), the third splitting method can apply octree decomposition (octree splitting) for all levels.
[0151] It should be noted that since in the third splitting method, when the explicit QtBt enable flag is set to the first value, Ot / Qt / Bt splitting is allowed in an arbitrary manner along the way, the maximum possible total number of splits can be three times the difference between the maximum node depth and the minimum node depth.
[0152] In an embodiment of the present disclosure, a fourth splitting method (or flexible QtBt splitting) can be provided to provide greater flexibility in the use of QtBt splitting by taking the type of QtBt splitting as explicit or implicit additional signaling. Instead of using fixed implicit rules (e.g., Table I and Table II) in the current QtBt splitting, additional signaling can be provided.
[0153] The fourth splitting method can include: (1) A QtBt enable flag (e.g., qtbt_enabled_flag) can be applied to replace the implicit QtBt enable flag in the QtBt splitting in TMC 13 to indicate the use of QtBt splitting more flexibly. (2) In the case where the QtBt enable flag is set to a value such as 1, a QtBt type flag (e.g., qtbt_type_flag) can be additionally signaled. (3) If the QtBt type flag is set to a value such as 0, the current implicit QtBt scheme (e.g., Table I and Table II) can be applied. Additionally, an asymmetric bounding box flag (e.g., asymmetric_bbox_enabled_flag) can be additionally signaled to selectively enable / disable the use of the asymmetric bounding box. In one embodiment, in the case where the asymmetric bounding box flag is set to a value such as 1, the dimensions (i.e., sizes) of the asymmetric bounding box along the x, y, and z can be signaled instead of the maximum of the three. In another embodiment, in the case where the asymmetric bounding box flag is not signaled, the asymmetric bounding box can be used all the time.
[0154] The fourth splitting method may further include: (4) If the QtBt type flag has a value such as 1, a 3-bit signal may be sent to each octree level in the octree levels to indicate whether to split along each of the x, y, and z axes. In one embodiment, the implicit QtBt rules in TMC 13 (e.g., Table I and Table II) may be applied to determine the 3-bit signal for each octree level in the octree levels. In another embodiment, other splitting rules may be applied to determine the 3-bit signal for each octree level in the octree levels. The other splitting rules may facilitate the encoding of octree occupancy information and may also consider the characteristics of the data (e.g., octree occupancy information) or the acquisition mechanism of the date. In one embodiment, the asymmetric bounding box flag may additionally be signaled to selectively enable / disable the use of the asymmetric bounding box. In the case where the asymmetric bounding box flag has a value such as 1, the dimensions (i.e., sizes) of the asymmetric bounding box along the x, y, and z may be signaled instead of the maximum value among the three. In another embodiment, the asymmetric bounding box flag may not be signaled but the asymmetric bounding box may be used all the time. (5) In the case where the QtBt enable flag is set to a value such as 0, the fourth splitting method may apply octree decomposition (or octree splitting) to all levels.
[0155] It should be noted that since in the fourth splitting method, in the case where the explicit QtBt enable flag is set to a value such as 1, the Ot / Qt / Bt splitting is allowed to be performed in any way along the way, the maximum possible total number of splits may be three times the difference between the maximum node depth and the minimum node depth.
[0156] The above techniques may be implemented in a video encoder or decoder applicable to point cloud compression / decompression. The encoder / decoder may be implemented in hardware, software, or any combination thereof, and the software (if any) may be stored in one or more non-transitory computer-readable media. For example, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable media.
[0157] Figure 14 and Figure 15FIG. shows a flowchart outlining processes (1400) and (1500) according to embodiments of the present disclosure. Processes (1400) and (1500) can be used during the decoding process of a point cloud. In various embodiments, processes (1400) and (1500) can be performed by a processing circuit (e.g., a processing circuit in the terminal device (110), a processing circuit that performs the functions of the encoder (203) and / or decoder (201), a processing circuit that performs the functions of the encoder (300), decoder (400), encoder (700), and / or decoder (800), etc.). In some embodiments, processes (1400) and (1500) can be implemented as software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit performs processes (1400) and (1500) respectively.
[0158] As Figure 14 shown, process (1400) starts at (S1401) and proceeds to (S1410).
[0159] At (S1410), first signaling information can be received from an encoded bitstream of a point cloud including a set of points in a three-dimensional (3D) space. The first signaling information can indicate the segmentation information of the point cloud.
[0160] At (S1420), second signaling information can be determined based on the first signaling information indicating a first value. The second signaling information can indicate the segmentation pattern of the set of points in the 3D space.
[0161] At (S1430), the segmentation pattern of the set of points in the 3D space can be determined based on the second signaling information. Process (1400) can then proceed to (S1440), where the point cloud can subsequently be reconstructed based on the segmentation pattern.
[0162] In some embodiments, the segmentation pattern can be determined as a predefined Quad-tree and Binary-tree (QtBt) segmentation based on the second signaling information being a second value.
[0163] In process (1400), third signaling information indicating that the 3D space is an asymmetric cuboid can be received. The dimension represented by signals in the x, y, and z directions of the 3D space can be determined based on the third signaling information being a first value.
[0164] In some embodiments, 3-bit signaling information may be determined for each of a plurality of segmentation levels in a segmentation mode based on the second signaling information being a first value. The 3-bit signaling information for each of the plurality of segmentation levels may indicate a segmentation direction along the x, y, and z directions for the corresponding segmentation level in the segmentation mode.
[0165] In some embodiments, the 3-bit signaling information may be determined based on the dimensions of the 3D space.
[0166] In processing (1400), a segmentation mode may be determined based on the first signaling information being a second value, where the segmentation mode may include a corresponding octree segmentation at each of a plurality of segmentation levels in the segmentation mode.
[0167] As Figure 15 shown, processing (1500) starts at (S1501) and proceeds to (S1510).
[0168] At (S1510), first signaling information may be received from an encoded bitstream of a point cloud including a set of points in a three-dimensional (3D) space. The first signaling information may indicate segmentation information of the point cloud.
[0169] At (S1520), a segmentation mode of the set of points in the 3D space may be determined based on the first signaling information, where the segmentation mode may include a plurality of segmentation levels.
[0170] At (S1530), the point cloud may then be reconstructed based on the segmentation mode.
[0171] In some embodiments, 3-bit signaling information may be determined for each of a plurality of segmentation levels in a segmentation mode based on the first signaling information being a first value, where the 3-bit signaling information for each of the plurality of segmentation levels may indicate a segmentation direction along the x, y, and z directions for the corresponding segmentation level in the segmentation mode.
[0172] In some embodiments, the 3-bit signaling information may be determined based on the dimensions of the 3D space.
[0173] In some embodiments, the segmentation mode may be determined based on the first signaling information being a second value to include a corresponding octree segmentation at each of a plurality of segmentation levels in the segmentation mode.
[0174] In process (1500), second signaling information may also be received from the encoded bitstream of the point cloud. When the second signaling information is a first value, the second signaling information may indicate that the 3D space is an asymmetric cuboid, and when the second signaling information is a second value, the second signaling information may indicate that the 3D space is a symmetric cuboid.
[0175] In some embodiments, based on the first signaling information indicating the second value and the second signaling information indicating the first value, the segmentation mode may be determined to include corresponding octree segmentation in each of the first segmentation levels among the multiple segmentation levels of the segmentation mode. The segmentation type and segmentation direction of the last segmentation level among the multiple segmentation levels of the segmentation mode may be determined according to the following table: , where d x , d y and d z are the log2 sizes of the 3D space in the x, y, and z directions, respectively.
[0176] In process (1500), the second signaling information may be determined based on the first signaling information indicating the first value. When the second signaling information indicates the first value, the second signaling information may indicate that the 3D space is an asymmetric cuboid, and when the second signaling information indicates the second value, the second signaling information may indicate that the 3D space is a symmetric cuboid. Additionally, the signal-represented dimensions of the 3D space in the x, y, and z directions may be determined based on the second signaling information indicating the first value.
[0177] As described above, the techniques described above may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 16 illustrates a computer system (1800) suitable for implementing certain embodiments of the disclosed subject matter.
[0178] The computer software may be encoded using any suitable machine code or computer language that may be subject to mechanisms such as assembly, compilation, linking, etc. to create code including instructions that may be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc. or executed through interpretation, microcode execution, etc.
[0179] The instructions may be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0180] Figure 16The components shown for the computer system (1800) are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. Nor should the configuration of the components be construed as having any dependency or requirement related to any one component or combination of components shown in the exemplary embodiments of the computer system (1800).
[0181] The computer system (1800) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to inputs made by one or more human users through, for example, tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, taps), visual inputs (e.g., gestures), and olfactory inputs (not shown). The human-machine interface devices may also be used to capture certain media not necessarily directly related to conscious human inputs, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0182] The input human-machine interface devices may include one or more of the following (only one of each described): keyboard (1801), mouse (1802), touchpad (1803), touch screen (1810), data glove (not shown), joystick (1805), microphone (1806), scanner (1807), camera (1808).
[0183] The computer system (1800) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile outputs, sounds, lights, and odors / tastes. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback through the touch screen (1810), data glove (not shown), or joystick (1805), but there may also be tactile feedback devices that do not serve as input devices), audio output devices (e.g., speakers (1809), headphones (not shown)), visual output devices (e.g., screens (1810), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which may be capable of outputting two-dimensional visual outputs or more than three-dimensional outputs through means such as stereoscopic image output, virtual reality glasses (not shown), holographic displays, and fog machines (not shown)), and printers (not shown).
[0184] The computer system (1800) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1820) with media such as CD / DVD (1821), thumb drives (1822), removable hard disk drives or solid state drives (1823), traditional magnetic media (such as tapes and floppy disks (not shown)), dedicated ROM / ASIC / PLD-based devices (such as security dongles (not shown)), and the like.
[0185] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transient signals.
[0186] The computer system (1800) may also include an interface to one or more communication networks. The network may be, for example, a wireless network, a wired network, an optical network. The network may also be a local area network, a wide area network, a metropolitan area network, in-vehicle and industrial networks, real-time networks, delay-tolerant networks, etc. Examples of networks include: local area networks (such as Ethernet, wireless LAN), cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired connections or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, in-vehicle and industrial networks including CAN bus, etc. Some networks typically require an external network interface adapter attached to certain common data ports or peripheral buses (1849) (e.g., the USB port of the computer system (1800)); others are typically integrated into the core of the computer system (1800) by attaching to the system bus (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smart phone computer system) as described below. The computer system (1800) may communicate with other entities using any of these networks. Such communication may be only one-way receiving (e.g., broadcast television), only one-way transmitting (e.g., CAN bus to certain CAN bus devices), or two-way (e.g., using a local area digital network or a wide area digital network to other computer systems). Certain protocols and protocol stacks may be used on each of these networks and network interfaces as described above.
[0187] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces may be attached to the core (1840) of the computer system (1800).
[0188] The core (1840) may include one or more central processing units (CPUs) (1841), graphics processing units (GPUs) (1842), field programmable gate areas (FPGAs) (1843) in the form of dedicated programmable processing units, hardware accelerators (1844) for certain tasks, etc. These devices, together with read-only memory (ROM) (1845), random access memory (1846), and internal mass storage devices (e.g., internal non-user-accessible hard disk drives, SSDs, etc.) (1847), may be connected via a system bus (1848). In some computer systems, the system bus (1848) may be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be attached directly or via a peripheral bus (1849) to the system bus (1848) of the core. The architecture of the peripheral bus includes PCI, USB, etc.
[0189] The CPU (1841), GPU (1842), FPGA (1843), and accelerator (1844) may execute certain instructions that, when combined, may constitute the computer code mentioned above. The computer code may be stored in the ROM (1845) or the RAM (1846). Transitional data may also be stored in the RAM (1846), while permanent data may be stored in, for example, the internal mass storage device (1847). Fast storage and retrieval of any of the storage devices in the storage may be achieved by using a cache memory that may be closely associated with one or more CPUs (1841), GPUs (1842), mass storage devices (1847), ROM (1845), RAM (1846), etc.
[0190] Computer-readable media may have computer code for performing various computer-implemented operations. The media and the computer code may be media and computer code that are specially designed and constructed for the purposes of this disclosure, or the media and the computer code may be of the type well known and available to those skilled in the field of computer software.
[0191] By way of example and not limitation, a computer system (1800) having an architecture and in particular a core (1840) can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage device as introduced above, as well as certain storage devices of the core (1840) having a non-transitory nature, such as, for example, a mass storage device (1847) inside the core or a ROM (1845). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (1840). Depending on specific needs, the computer-readable media can include one or more storage devices or chips. The software can cause the core (1840) and in particular the processors therein (including a CPU, GPU, FPGA, etc.) to perform a specific process or a specific part of a specific process described herein, including defining data structures stored in a RAM (1846) and modifying such data structures according to a process defined by the software. Additionally or alternatively, the computer system can provide functionality provided by logical hardwiring or otherwise embodied in circuitry (e.g., an accelerator (1844)), which can operate in place of or in conjunction with the software to perform a specific process or a specific part of a specific process described herein. In appropriate cases, reference to software can include logic, and vice versa. In appropriate cases, reference to computer-readable media can include circuitry (e.g., an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both of the above. The present disclosure encompasses any suitable combination of hardware and software.
[0192] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it will be recognized that, although not explicitly shown or described herein, those skilled in the art will be able to envision many systems and methods that embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.
Claims
1. A method for performing point cloud geometric decoding, characterized in that, The method includes: Receiving first signaling information from an encoded bitstream of a point cloud including a set of points in a three-dimensional (3D) space, the first signaling information indicating segmentation information of the point cloud; Determining second signaling information from the encoded bitstream of the point cloud based on a first value of the first signaling information, the second signaling information indicating a segmentation pattern of the set of points in the 3D space; Determining the segmentation pattern of the set of points in the 3D space based on the second signaling information; Receiving third signaling information, the third signaling information indicating that the 3D space is an asymmetric rectangular parallelepiped; and Determining signaled dimensions of the 3D space along the x, y, and z directions based on a first value of the third signaling information; Reconstructing the point cloud based on the segmentation pattern and the signaled dimensions of the 3D space along the x, y, and z directions.
2. The method according to claim 1, characterized in that, Determining the segmentation pattern further includes: Determining the segmentation pattern as a predefined quadtree and binary tree (QtBt) segmentation based on the second signaling information indicating a second value.
3. The method according to claim 1, characterized in that, Determining the segmentation pattern further includes: Based on the second signaling information indicating the first value, receiving 3-bit signaling information for each of a plurality of segmentation levels in the segmentation pattern, the 3-bit signaling information for each of the plurality of segmentation levels indicating segmentation directions along the x, y, and z directions for a corresponding segmentation level in the segmentation pattern.
4. The method according to claim 3, wherein The 3-bit signaling information is determined based on the dimensions of the 3D space.
5. The method according to claim 1, wherein The method further includes: Determining the segmentation pattern based on the first signaling information indicating a second value, the segmentation pattern including a corresponding octree segmentation at each of a plurality of segmentation levels in the segmentation pattern.
6. A method for performing point cloud geometric decoding, characterized in that, The method includes: Receiving first signaling information from an encoded bitstream of a point cloud including a set of points in a three-dimensional (3D) space, the first signaling information indicating segmentation information of the point cloud; and Determining a segmentation pattern of the set of points in the 3D space based on the first signaling information, the segmentation pattern including a plurality of segmentation levels; Receiving second signaling information from the encoded bitstream of the point cloud, the second signaling information indicating that the 3D space is an asymmetric rectangular parallelepiped when the second signaling information is a first value; and Determining signaled dimensions of the 3D space along the x, y, and z directions based on the second signaling information indicating the first value; Reconstructing the point cloud based on the segmentation pattern and the signaled dimensions of the 3D space along the x, y, and z directions.
7. The method according to claim 6, characterized in that, Determining the segmentation pattern further includes: Based on the first signaling information indicating a first value, receiving 3-bit signaling information for each of a plurality of segmentation levels in the segmentation pattern, the 3-bit signaling information for each of the plurality of segmentation levels indicating segmentation directions along the x, y, and z directions for a corresponding segmentation level in the segmentation pattern.
8. The method according to claim 7, characterized in that, The 3-bit signaling information is determined based on the dimensions of the 3D space.
9. The method according to claim 6, wherein Determining the segmentation pattern further includes: Determine the segmentation mode based on the first signaling information indicating a second value, where the segmentation mode includes corresponding octree segmentations at each of a plurality of segmentation levels in the segmentation mode.
10. The method according to claim 6, characterized in that The method further includes: When the second signaling information is the second value, the second signaling information indicates that the 3D space is a symmetric cuboid.
11. The method according to claim 10, wherein Determining the segmentation mode further includes: Based on the first signaling information indicating the second value and the second signaling information indicating the first value, determine that the segmentation mode includes corresponding octree segmentations at each of the first segmentation levels in the plurality of segmentation levels of the segmentation mode; and Determine the segmentation type and segmentation direction at the last segmentation level among the plurality of segmentation levels of the segmentation mode according to the following conditions: , where the d x , d y and d z are the log2 sizes of the 3D space in the x, y, and z directions, respectively.
12. The method according to claim 6, wherein The method further includes: Determine second signaling information based on the first signaling information indicating a first value. When the second signaling information indicates the first value, the second signaling information indicates that the 3D space is an asymmetric cuboid, and when the second signaling information indicates the second value, the second signaling information indicates that the 3D space is a symmetric cuboid.
13. An apparatus for performing point cloud geometric decoding, characterized in that, The apparatus includes: A memory that stores instructions; and A processor that communicates with the memory, wherein when the processor executes the instructions, the processor is configured to cause the apparatus to perform the method according to any one of claims 1 to 5.
14. An apparatus for performing point cloud geometric decoding, characterized in that, The apparatus includes: A memory that stores instructions; and A processor that communicates with the memory, wherein when the processor executes the instructions, the processor is configured to cause the apparatus to perform the method according to any one of claims 6 to 12.
15. A computer device, characterized in that, The computer device includes: One or more computer-readable non-transitory storage media configured to store computer program code; and One or more computer processors configured to access the computer program code and operate in accordance with what is indicated by the computer program code to perform the method according to any one of claims 1 to 12.
16. A non-transitory computer-readable medium, characterized in that, A computer program is stored on the non-transitory computer-readable medium, and the computer program is configured to cause one or more computer processors to perform the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Scalable point cloud compression with transform, and corresponding decompression
US20170347122A1
Coding method, device, system
WO2020057530A1