Method, device and medium for point cloud coding and decoding
By combining multi-task learning with different point cloud feature extractors and feature mapping modules, the problem of semantic loss in existing point cloud encoding and decoding technologies is solved, and the quality of semantic reconstruction is improved while maintaining the geometric structure, which is suitable for machine vision tasks.
Patent Information
- Application Number
- CN202480010685.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-03
- Filing Date
- 2024-02-02
- Publication Date
- 2025-09-12
AI Technical Summary
Existing point cloud geometry encoding and decoding technologies fail to effectively consider semantic loss in machine vision tasks, resulting in a decrease in the accuracy of point cloud semantic segmentation and classification. Existing methods mainly use the sum of squared errors as a distortion measure for rate-distortion optimization, without considering the feature semantic loss in the encoding and decoding process.
Different point cloud feature extractors (point-based and voxel-based extractors) are combined with feature mapping modules. A multi-task learning mechanism is used to preserve semantic information while maintaining the geometric structure. Multi-objective loss constraints and feature space similarity calculation are used to design the encoding and decoding process to optimize rate distortion and semantic distortion.
It improves the accuracy of point cloud semantic segmentation and classification, optimizes the quality of semantic reconstruction during the encoding and decoding process, and meets the needs of machine vision tasks.
Smart Images

Figure CN120641944A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to point cloud encoding and decoding technology, and more particularly, to semantic enhancement-based point cloud compression. Background Art
[0002] A point cloud is a collection of individual data points in a three-dimensional (3D) plane, with each point having defined coordinates on the X, Y, and Z axes. Therefore, point clouds can be used to represent the physical contents of a three-dimensional space. Point clouds have proven to be a promising way to represent 3D visual data for a wide range of immersive applications, from augmented reality to self-driving cars.
[0003] Point cloud codec standards have evolved primarily through the development of the well-known MPEG organization. MPEG, short for Moving Picture Experts Group, is one of the main standardization organizations dealing with multimedia. In 2017, the MPEG 3D Graphics Codec Group (3DG) published a Request for Proposal (CFP) document to begin developing point cloud codec standards. The final standard will consist of two categories of solutions. Video-based point cloud compression (V-PCC or VPCC) is suitable for point sets with a relatively uniform distribution of points. Geometry-based point cloud compression (G-PCC or GPCC) is suitable for more sparse distributions. However, the codec efficiency of conventional point cloud codec techniques is generally expected to be further improved. Summary of the Invention
[0004] The embodiments of the present disclosure provide a solution for point cloud encoding and decoding.
[0005] In a first aspect, a method for point cloud encoding and decoding is proposed. The method comprises: determining the type of the current point cloud for conversion between a current point cloud in a point cloud sequence and a bitstream of the point cloud sequence; determining a codec module for the current point cloud based on the type of the current point cloud, the codec module comprising at least one of the following: a point cloud feature extractor or a point cloud geometry reconstruction module; and performing the conversion based on the codec module. The method according to the first aspect of the present disclosure determines a point cloud feature extractor or a point cloud geometry reconstruction module based on the type of the point cloud. In this way, point cloud geometry compression (PCGC) can be improved.
[0006] In a second aspect, an apparatus for processing a point cloud sequence is provided. The apparatus for processing a point cloud sequence includes a processor and a non-volatile memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.
[0007] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions for causing a processor to execute the method according to the first aspect of the present disclosure.
[0008] In a fourth aspect, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a point cloud sequence generated by a method performed by a point cloud processing apparatus. The method includes: determining the type of a current point cloud in the point cloud sequence; determining a codec module for the current point cloud based on the type of the current point cloud, the codec module comprising at least one of the following: a point cloud feature extractor or a point cloud geometry reconstruction module; and generating a bitstream based on the codec module.
[0009] In a fifth aspect, a method for storing a bitstream of a point cloud sequence is provided. The method includes: determining the type of a current point cloud in the point cloud sequence; determining a codec module for the current point cloud based on the type of the current point cloud, the codec module including at least one of the following: a point cloud feature extractor or a point cloud geometry reconstruction module; generating a bitstream based on the codec module; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other objects, features and advantages of example embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings, in which like reference numerals generally refer to like components throughout the example embodiments of the present disclosure.
[0012] Figure 1 A block diagram of an example point cloud encoding and decoding system according to some embodiments of the present disclosure is shown;
[0013] Figure 2 Another block diagram of an example point cloud encoding and decoding system according to some embodiments of the present disclosure is shown;
[0014] Figure 3 A block diagram illustrating an example of a point cloud compression (PCC) encoder according to some embodiments of the present disclosure is shown;
[0015] Figure 4 A block diagram illustrating an example of a PCC decoder according to some embodiments of the present disclosure;
[0016] Figure 5 An example of a flow of a compression method according to some embodiments of the present disclosure is shown;
[0017] Figure 6 shows an example of a feature encoder module according to some embodiments of the present disclosure;
[0018] Figure 7 A flowchart showing a method for point cloud encoding and decoding according to some embodiments of the present disclosure is shown; and
[0019] Figure 8 A block diagram is shown of a computing device in which various embodiments of the present disclosure may be implemented.
[0020] Throughout the drawings, same or similar reference numbers generally refer to same or similar elements. DETAILED DESCRIPTION
[0021] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.
[0022] In the following description and claims, unless defined otherwise, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0023] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment is required to include that particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is intended that such feature, structure, or characteristic, whether or not explicitly described, be applicable to other embodiments and that it is within the knowledge of those skilled in the art to apply such feature, structure, or characteristic.
[0024] It should be understood that although the terms "first" and "second" and the like may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0025] The terms used herein are used only for the purpose of describing specific embodiments and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "comprise," "including," "having," "including," and / or "comprising" when used herein indicate the presence of the features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Sample Environment
[0026] Figure 1 is a block diagram illustrating an example point cloud codec system 100 that may utilize the techniques of the present disclosure. As shown, the point cloud codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a point cloud encoding device, and the destination device 120 may also be referred to as a point cloud decoding device. In operation, the source device 110 may be configured to generate encoded point cloud data, and the destination device 120 may be configured to decode the encoded point cloud data generated by the source device 110. The techniques of the present disclosure are generally directed to encoding and decoding (encoding and / or decoding) point cloud data, i.e., to support point cloud compression. The codec may efficiently compress and / or decompress point cloud data.
[0027] Source device 100 and destination device 120 may include any of a wide variety of devices, including desktop computers, notebook (i.e., portable) computers, tablet computers, set-top boxes, telephone handsets (such as smartphones and mobile phones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, vehicles (e.g., land or sea vehicles, spacecraft, aircraft, etc.), robots, LIDAR devices, satellites, extended reality devices, etc. In some cases, source device 100 and destination device 120 may be equipped for wireless communication.
[0028] The source device 100 may include a data source 112, a memory 114, a PCC encoder 116, and an input / output (I / O) interface 118. The destination device 120 may include an input / output (I / O) interface 128, a PCC decoder 126, a memory 124, and a data consumer 122. According to the present disclosure, the PCC encoder 116 of the source device 100 and the PCC decoder 126 of the destination device 120 may be configured to apply the techniques of the present disclosure related to point cloud encoding and decoding. Therefore, the source device 100 represents an example of an encoding device, while the destination device 120 represents an example of a decoding device. In other examples, the source device 100 and the destination device 120 may include other components or arrangements. For example, the source device 100 may receive data (e.g., point cloud data) from an internal source or an external source. Similarly, the destination device 120 may be connected to an external data consumer interface instead of including the data consumer in the same device.
[0029] Generally speaking, data source 112 represents a source of point cloud data (i.e., raw, unencoded point cloud data) and can provide a sequential series of "frames" of point cloud data to PCC encoder 116, which encodes the frame's point cloud data. In some examples, data source 112 generates the point cloud data. Data source 112 of source device 100 can include a point cloud capture device, such as any of a variety of cameras or sensors, for example, one or more video cameras, an archive containing previously captured point cloud data, a 3D scanner, or a light detection and ranging (LIDAR) device, and / or a data feed interface that receives point cloud data from a data content provider. Thus, in some examples, data source 112 can generate point cloud data based on signals from a LIDAR device. Alternatively or additionally, point cloud data can be computer-generated from a scanner, camera, sensor, or other data. For example, data source 112 can generate point cloud data, or produce a combination of real-time point cloud data, archived point cloud data, and computer-generated point cloud data. In each case, the PCC encoder 116 encodes the captured, pre-captured, or computer-generated point cloud data. The PCC encoder 116 can rearrange frames of the point cloud data from a received order (sometimes referred to as "display order") to a codec order for encoding and decoding. The PCC encoder 116 can generate one or more bitstreams comprising the encoded point cloud data. The source device 100 can then output the encoded point cloud data via the I / O interface 118 for receipt and / or retrieval by, for example, the I / O interface 128 of the destination device 120. The encoded point cloud data can be transmitted directly to the destination device 120 via the network 130A via the I / O interface 118. The encoded point cloud data can also be stored on the storage medium / server 130B for access by the destination device 120.
[0030] The memory 114 of the source device 100 and the memory 124 of the destination device 120 can represent general purpose memory. In some examples, the memory 114 and the memory 124 can store raw point cloud data, such as the raw point cloud data from the data source 112 and the raw, decoded point cloud data from the PCC decoder 126. Additionally or alternatively, the memory 114 and the memory 124 can store software instructions, such as those executable by the PCC encoder 116 and the PCC decoder 126, respectively. Although the memory 114 and the memory 124 are shown as separate from the PCC encoder 116 and the PCC decoder 126 in this example, it should be understood that the PCC encoder 116 and the PCC decoder 126 can also include internal memory for functionally similar or equivalent purposes. Furthermore, the memory 114 and the memory 124 can store encoded point cloud data, such as the output from the PCC encoder 116 and the input to the PCC decoder 126. In some examples, portions of memory 114 and memory 124 may be allocated as one or more buffers, eg, to store raw, decoded, and / or encoded point cloud data. For example, memory 114 and memory 124 may store point cloud data.
[0031] I / O interface 118 and I / O interface 128 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., a network card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where I / O interface 118 and I / O interface 128 include wireless components, I / O interface 118 and I / O interface 128 may be configured to transmit data (such as encoded point cloud data) according to a cellular communication standard (such as 4G, 4G-LTE (Long Term Evolution), LTE Advanced, 5G, etc.). In some examples where I / O interface 118 includes a wireless transmitter, I / O interface 118 and I / O interface 128 may be configured to transmit data (such as encoded point cloud data) according to other wireless standards (such as the IEEE 802.11 specification). In some examples, source device 100 and / or destination device 120 may include corresponding system-on-chip (SoC) devices. For example, source device 100 may include a SoC device to perform the functions attributed to PCC encoder 116 and / or I / O interface 118 , and destination device 120 may include a SoC device to perform the functions attributed to PCC decoder 126 and / or I / O interface 128 .
[0032] The techniques disclosed herein may be applied to support encoding and decoding of any of a variety of applications, such as communications between autonomous vehicles, communications between scanners, cameras, sensors and processing devices such as local or remote servers, geographic mapping, or other applications.
[0033] I / O interface 128 of destination device 120 receives the encoded bitstream from source device 110. The encoded bitstream may include signaling information defined by PCC encoder 116 and used by PCC decoder 126, such as syntax elements with values representing point clouds. Data consumer 122 uses the decoded data. For example, data consumer 122 may use the decoded point cloud data to determine the position of a physical object. In some examples, data consumer 122 may include a display to present an image based on the point cloud data.
[0034] The PCC encoder 116 and the PCC decoder 126 can be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is partially implemented in software, the device can store the instructions of the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the technology of the present disclosure. Each of the PCC encoder 116 and the PCC decoder 126 can be included in one or more encoders or decoders, and any one of the PCC encoder 116 and the PCC decoder 126 can be integrated as part of a combined encoder / decoder (codec) in the corresponding device. The device including the PCC encoder 116 and / or the PCC decoder 126 can include one or more integrated circuits, microprocessors, and / or other types of devices.
[0035] The PCC encoder 116 and the GPCC decoder 126 may operate in accordance with a codec standard, such as the Video Point Cloud Compression (VPCC) standard or the Geometric Point Cloud Compression (GPCC) standard. The present disclosure may generally refer to the encoding and decoding (e.g., encoding and decoding) of a frame to include the process of encoding or decoding data. The encoded bitstream typically includes a series of values for syntax elements that may represent codec decisions (e.g., codec mode).
[0036] A point cloud can contain a collection of points in 3D space and can have attributes associated with the points. The attributes can be color information (such as R, G, B or Y, Cb, Cr) or reflectance information or other attributes. Point clouds can be captured by various cameras or sensors (such as LIDAR sensors and 3D scanners) and can also be computer-generated. Point cloud data is used in various applications, including but not limited to construction (modeling), graphics (3D models for visualization and animation), and the automotive industry (LIDAR sensors to help navigation).
[0037] In some example embodiments, the PCC encoder 116 in the system 100 may include a GPCC encoder, and the PCC decoder 126 may include a GPCC decoder. The GPCC encoder and the GPCC decoder may be collectively referred to as a GPCC codec or GPCC. In some example embodiments, the GPCC may be applied in combination with another codec.
[0038] Figure 2 Another block diagram of an example point cloud encoding and decoding system 200 according to some embodiments of the present disclosure is shown. As shown, the point cloud system 200 includes a geometry encoder 210, a geometry decoder 230, a GPCC 220, and a machine vision task 240. In some example embodiments, the GPCC 220 may be implemented as a PCC encoder (such as, Figure 1 PCC encoder 116 in ) and PCC decoder (such as, Figure 1 100). System 100 and system 200 may be used in combination or separately. For example, system 100 may be a part of system 200. Geometry encoder 210, geometry decoder 230, and / or machine vision task 240 may be added to system 100 as additional modules.
[0039] In some example embodiments, the geometry encoder 210 may be based on machine learning (ML) or artificial intelligence (AI), and the geometry decoder 230 may also be based on ML or AI. ML / AI-based geometry encoders / decoders may be applied in combination with the GPCC 220. The geometry encoder 210 and / or the geometry decoder 230 may be pre-trained or fine-tuned. The information or data output from the geometry decoder 230 may be used for a machine vision task 240. As used herein, the system 200 may be referred to as an AI-based point cloud compression (AI-PCC) system. It should be understood that in some example embodiments, the system 200 may include additional modules, such as a feature extractor module, etc. The scope of the present disclosure is not limited to this.
[0040] In some example embodiments, GPCC 220 may include a GPCC encoder and a GPCC decoder. Figure 3 is a block diagram illustrating an example of a GPCC encoder 300 according to some embodiments of the present disclosure, which may be Figure 2 An example of a GPCC encoder of GPCC 220 is shown in FIG. Figure 4 is a block diagram illustrating an example of a GPCC decoder 400 according to some embodiments of the present disclosure. The GPCC decoder 400 may be Figure 2 An example of a GPCC decoder of GPCC 220 is shown in FIG.
[0041] In both the GPCC encoder 300 and the GPCC decoder 400, the point cloud position is first encoded and decoded. The attribute encoding and decoding depends on the decoded geometry. Figure 3 and Figure 4 , the region adaptive hierarchical transform (RAHT) unit 318, the surface approximation analysis unit 312, the RAHT unit 414, and the surface approximation synthesis unit 410 are options that are commonly used for category 1 data. The level of detail (LOD) generation unit 320, the lifting unit 322, the LOD generation unit 416, and the inverse lifting unit 418 are options that are commonly used for category 3 data. All other units are common between category 1 and category 3.
[0042] For category 3 data, the compressed geometry is typically represented as an octree from the root down to the leaf level for each voxel. For category 1 data, the compressed geometry is typically represented by a pruned octree (i.e., an octree from the root down to the leaf level for blocks larger than a voxel) plus a model that approximates the surface within each leaf of the pruned octree. In this way, category 1 and category 3 data share the octree codec mechanism, while category 1 data can additionally utilize a surface model to approximate the voxels within each leaf. The surface model used is a triangulation that includes 1 to 10 triangles per block, resulting in a set of triangles. Therefore, category 1 geometry codecs are called trisoup geometry codecs, while category 3 geometry codecs are called octree geometry codecs.
[0043] exist Figure 3 In the example, the GPCC encoder 300 may include a coordinate transformation unit 302, a color transformation unit 304, a voxelization unit 306, an attribute transfer unit 308, an octree analysis unit 310, a surface approximation analysis unit 312, an arithmetic coding unit 314, a geometric reconstruction unit 316, a RAHT unit 318, an LOD generation unit 320, a lifting unit 322, a coefficient quantization unit 324 and an arithmetic coding unit 326.
[0044] As in Figure 3 As shown in the example of , the GPCC encoder 300 can receive a set of locations and a set of attributes. The locations can include the coordinates of a point in the point cloud. The attributes can include information about the point in the point cloud, such as a color associated with the point in the point cloud.
[0045] The coordinate transformation unit 302 can apply a transformation to the coordinates of the point to transform the coordinates from the original domain to the transformed domain. This disclosure may refer to the transformed coordinates as transformed coordinates. The color transformation unit 304 can apply a transformation to convert the color information of the attribute to a different domain. For example, the color transformation unit 304 can convert the color information from the RGB color space to the YCbCr color space.
[0046] In addition, Figure 3 In the example of , the voxelization unit 306 can voxelize the transformed coordinates. Voxelizing the transformed coordinates can include quantizing and removing some points of the point cloud. In other words, multiple points of the point cloud can be contained in a single "voxel", which can then be treated as a point in some aspects. In addition, the octree analysis unit 310 can generate an octree based on the voxelized transformed coordinates. In addition, in Figure 3 In the example of FIG, the surface approximation analysis unit 312 can analyze the points to potentially determine a surface representation of the set of points. The arithmetic coding unit 314 can perform arithmetic coding on syntax elements representing information about the octree and / or surface determined by the surface approximation analysis unit 312. The GPCC encoder 300 can output these syntax elements in a geometry bitstream.
[0047] The geometric reconstruction unit 316 can reconstruct the transformed coordinates of the points in the point cloud based on the octree, the data indicating the surface determined by the surface approximation analysis unit 312, and / or other information. Due to voxelization and surface approximation, the number of transformed coordinates reconstructed by the geometric reconstruction unit 316 may differ from the number of original points in the point cloud. This disclosure may refer to the generated points as reconstructed points. The attribute transfer unit 308 can transfer attributes of the original points in the point cloud to the reconstructed points in the point cloud data.
[0048] Furthermore, the RAHT unit 318 may apply RAHT coding to the attributes of the reconstructed points. Alternatively or additionally, the LOD generation unit 320 and the lifting unit 322 may apply LOD processing and lifting, respectively, to the attributes of the reconstructed points. The RAHT unit 318 and the lifting unit 322 may generate coefficients based on the attributes. The coefficient quantization unit 324 may quantize the coefficients generated by the RAHT unit 318 or the lifting unit 322. The arithmetic coding unit 326 may apply arithmetic coding to syntax elements representing the quantized coefficients. The GPCC encoder 300 may output these syntax elements in the attribute bitstream.
[0049] exist Figure 4 In the example, the GPCC decoder 400 may include a geometric arithmetic decoding unit 402, an attribute arithmetic decoding unit 404, an octree synthesis unit 406, an inverse quantization unit 408, a surface approximation synthesis unit 410, a geometric reconstruction unit 412, a RAHT unit 414, an LOD generation unit 416, an inverse lifting unit 418, a coordinate inverse transformation unit 420 and a color inverse transformation unit 422.
[0050] The GPCC decoder 400 can obtain a geometry bitstream and an attribute bitstream. The geometry arithmetic decoding unit 402 of the decoder 400 can apply arithmetic decoding (e.g., CABAC or other types of arithmetic decoding) to the syntax elements in the geometry bitstream. Similarly, the attribute arithmetic decoding unit 404 can apply arithmetic decoding to the syntax elements in the attribute bitstream.
[0051] The octree synthesis unit 406 may synthesize the octree based on syntax elements parsed from the geometry bitstream. In the case where surface approximation is used in the geometry bitstream, the surface approximation synthesis unit 410 may determine the surface model based on the syntax elements parsed from the geometry bitstream and based on the octree.
[0052] Furthermore, the geometric reconstruction unit 412 may perform reconstruction to determine the coordinates of the points in the point cloud. The coordinate inverse transformation unit 420 may apply an inverse transformation to the reconstructed coordinates to convert the reconstructed coordinates (positions) of the points in the point cloud from the transformed domain back to the original domain.
[0053] In addition, Figure 4 In the example of , the inverse quantization unit 408 may inverse quantize the property value. The property value may be based on syntax elements obtained from the property bitstream (eg, including syntax elements decoded by the property arithmetic decoding unit 404).
[0054] Depending on how the attribute values are encoded, the RAHT unit 414 may perform RAHT encoding and decoding to determine color values for points of the point cloud based on the dequantized attribute values. Alternatively, the LOD generation unit 416 and the de-lifting unit 418 may use a level of detail based technique to determine color values for points of the point cloud.
[0055] In addition, Figure 4 In the example of , the inverse color transform unit 422 can apply an inverse color transform to the color values. The inverse color transform can be the inverse of the color transform applied by the color transform unit 304 of the encoder 300. For example, the color transform unit 304 can transform the color information from the RGB color space to the YCbCr color space. Correspondingly, the inverse color transform unit 422 can transform the color information from the YCbCr color space to the RGB color space.
[0056] Figure 3 and Figure 4The various units are shown to aid in understanding the operations performed by the encoder 300 and the decoder 400. The units may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. A fixed-function circuit refers to a circuit that provides a specific function and is preset in the operations that can be performed. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provide flexible functionality in the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. A fixed-function circuit can execute software instructions (e.g., to receive parameters or output parameters), but the type of operation performed by the fixed-function circuit is generally immutable. In some examples, one or more units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more units may be integrated circuits.
[0057] Some exemplary embodiments of the present disclosure are described in detail below. It should be understood that the section titles used in this document are for ease of understanding and do not limit the embodiments disclosed in the section to that section. In addition, although specific embodiments are described with reference to PCC, AI-PCC, GPCC or other specific point cloud codecs, the disclosed technology is also applicable to other point cloud codec technologies. In addition, although some embodiments describe the point cloud encoding and decoding steps in detail, it is understood that the decoding steps of the corresponding inverse codec will be implemented by the decoder. 1. Brief Overview This disclosure relates to point cloud codec technology. Specifically, it relates to semantically enhanced point cloud codecs for machine vision. The concepts can be applied, alone or in various combinations, to any standard or non-standard point cloud codec, such as the currently under-development artificial intelligence-based point cloud compression (AI-PCC). 2. Abbreviation G-PCC geometry-based point cloud compression AI-PCC point cloud compression based on artificial intelligence MPEG Moving Picture Experts Group UAV drones PCGC point cloud geometry compression KNN K Nearest Neighbor 3. Introduction Current point cloud geometry codecs use PSNR as the primary distortion metric for rate-distortion optimization, aiming to ensure the objective quality of the reconstructed point cloud at the decoder at a certain bitrate, thereby providing viewers with high-quality point cloud content. However, with the development of 5G and AI technologies, more and more point cloud data is being directly used by machine vision and intelligent algorithms to complete various 3D computer vision tasks. For example, in scenarios such as autonomous driving, unmanned aerial vehicle (UAV) navigation, and smart cities, the receiver's task is no longer simply to reconstruct the complete raw point cloud, but to perform intelligent analysis tasks such as semantic segmentation, point cloud classification, and object detection on the point cloud data. The point clouds used in different scenarios (such as basic solid point clouds, LiDAR point clouds, and digital human point clouds) have different structural characteristics. Basic solid point clouds can have a simple structure and uniform distribution. LiDAR point clouds can have a relatively complex structure and an extremely sparse distribution. While there has been extensive research on image and video codec algorithms for machine vision, there is limited research on intelligent point cloud codec algorithms for machine vision. 3.1 Point Cloud Compression Network Based on Deep Learning In recent years, 3D deep learning has made great progress in vision tasks. Researchers have introduced deep learning-based methods to point cloud compression. These methods can be roughly divided into two categories: voxel-based point cloud compression and point-based point cloud compression. For voxel-based methods, the point cloud is voxelized into multiple blocks, which are processed by 3D convolutional neural networks (such as sparse convolution). For point-based methods, such as FoldingNet, graph-based convolutional networks and transformer-based networks are used to directly process the raw point cloud. 3.2 Point Cloud Semantic Enhancement For intelligent point cloud analysis in mixed human-machine visual scenes, compression codecs often introduce distortion to the point cloud, which leads to distortion in the intelligent analysis of the point cloud. Existing codec methods mainly consider geometric distortion, but geometric distortion does not represent the semantic distortion of the point cloud. To minimize the semantic distortion in intelligent point cloud analysis at the decoder and ensure the objective reconstruction quality of the point cloud, optimizing rate distortion and enhancing the semantics of the point cloud during the point cloud encoding and decoding process is a valuable problem. 4. Question The existing designs of point cloud geometry compression (PCGC) methods have the following problems: 1. In existing point cloud geometry encoding and decoding processes, the loss of objective geometric quality is primarily considered to provide viewers with higher-quality point clouds. However, the semantic loss of point clouds caused by encoding and decoding in machine vision tasks is not considered. These semantic losses often reduce the accuracy of point cloud semantic segmentation and classification. 2. Point cloud geometry encoding and decoding reconstruction and point cloud intelligent analysis are two completely different tasks. The extracted point cloud geometric features and point cloud semantic features are different in feature space. However, few methods take into account the differences between these two different tasks in the feature space during the encoding and decoding process. 3. In the training process of existing AI-PCC methods, the sum of squared errors is mainly used as the distortion measurement for rate-distortion optimization. This evaluation method mainly considers the error in objective reconstruction quality during the encoding and decoding process, but does not consider the loss of feature semantics during the encoding and decoding process. 5. Detailed solution In order to solve the above problems and some other problems not mentioned, the following methods are disclosed. The embodiments should be considered as examples to explain the general concept and should not be interpreted in a narrow sense. In addition, these embodiments can be applied alone or in combination in any way. 1) In the following discussion, the term "encoder" refers to a model that encodes and decodes information to be transmitted through a signal. The term "decoder" refers to a model that decodes compressed bits to obtain the information transmitted through the signal. 2) Different compact feature representation extractors are proposed using different point clouds. a. In one example, different point cloud feature extractors (such as a point-based extractor and a voxel-based extractor, etc.) can be used for different types of point clouds. i. In one example, for a basic solid point cloud with finite points (such as a basic object), a point-based extractor can be used to extract features of the point cloud. 1. In one example, a graph-based approach may be used. 2. In one example, graph convolution can be used to model the topological structure of a point cloud. This point cloud feature representation can be extracted based on the effective modeling of the point cloud topology. a. In one example, graph convolution can use the K-nearest neighbor (KNN) algorithm to obtain local point cloud neighbors and then extract features of these local neighboring points. 3. In one example, a transformer based approach may be used. 4. In one example, a transformer-based approach can be used to model the topological structure of a point cloud. The transformer’s attention mechanism can be used to model the spatial structure of the point cloud and extract feature representations of the point cloud. a. In one example, a transformer-based approach can achieve progressive geometric feature extraction from point clouds by stacking multiple layers of attention mechanisms. ii. In one example, for large-scale point clouds with complex structures and a large number of points (such as LIDAR point clouds), a voxel-based extractor can be used to extract features of the point cloud. 1. In one example, a sparse convolution based approach can be used. 2. In one example, sparse convolution can be used to extract geometric structure features of point clouds. a. In one example, sparse convolution can voxelize a point cloud and calculate geometric properties of the voxelized point cloud. i. In one example, voxelization can be used to make the point cloud data more regular for subsequent convolution processing. b. In one example, different pre-trained point cloud feature extractors can be used to extract compact feature representations for point clouds with different spatial structures and point sizes. c. In one example, which point cloud feature extractor is used can be signaled from the encoder to the decoder. d. In one example, which point cloud feature extractor to use can be inferred at the decoder. 3) It is proposed to encode and decode the indication of the final sampled point cloud and transmit it to the decoder through a signal. a. In one example, the indicator of the final sampled point cloud can be encoded and decoded by a point cloud codec. i. In one example, the point cloud codec can be G-PCC, V-PCC, Draco, etc. 4) It is proposed to encode and decode the indication of the feature and transmit it to the decoder through a signal. a. In one example, the indication of the feature (eg, index) can be encoded using a fixed-length codec, a unary codec, a truncated unary codec, etc. b. In one example, the indication of the feature may be encoded in a predictive manner. 5) It is proposed to use a point cloud feature mapping module to explore the similarity of features in different tasks. a. In one example, for the point cloud geometry reconstruction task, the extracted features can contain more point cloud spatial structure information. i. In one example, the spatial domain structure information may include high-frequency texture information and low-frequency structure information in the point cloud data. b. In one example, for point cloud intelligent analysis tasks (point cloud segmentation, point cloud classification), the extracted features can contain more point cloud semantic information. i. In one example, the point cloud semantic information may include semantic information of each point. c. In one example, a transformer-based point cloud feature mapping module can be used to minimize the feature similarity between the feature spaces of point cloud geometry encoding and decoding reconstruction and point cloud machine vision intelligent analysis. i. In one example, the design of the feature mapping module can refer to the transformer structure of the multi-head attention mechanism. ii. In one example, a self-attention mechanism can be used to calculate the similarity in the feature map space, and then the features can be weighted summed. d. In one example, the point cloud feature space mapping module can be designed and preserved symmetrically at the encoder and decoder. 6) A multi-task learning mechanism is proposed to preserve as much semantic information of the point cloud as possible while maintaining the geometric structure of the point cloud. a. In one example, whether the multi-task model is used can be signaled from the encoder to the decoder. b. In one example, it can be inferred at the decoder whether the multi-task model is used. c. In one example, during the training phase, multi-objective loss constraints can be used. i. In one example, geometric constraints on point cloud reconstruction can be used to ensure that a point cloud of basic quality can be reconstructed. 1. In one example, geometric constraints can be calculated by using chamfer distance for supervised learning. 2. In one example, the chamfer distance can be calculated as follows: Where S1 and S2 are two point clouds. x and y are the coordinates of the points in S1 and S2 respectively. ii. In one example, semantic constraints for machine vision can be used. 1. In one example, semantic constraints can be used to constrain the feature distribution distance between the encoded and decoded features and the original features. 2. In one example, semantic constraints can be calculated by using Kullback-Leibler (KL) divergence. 3. In one example, the KL divergence distance can be calculated as follows: L semantic =D KL (F ori ||F rec ) =∑ip ori (v i )log p ori (v i )-p ori (v i )log p rec (v i ) Among them F ori and Frec are the original features and the encoded and decoded features, p ori and p rec are the probability distributions of original features and encoded and decoded features respectively. iii. In one example, the final loss function can be calculated as follows: L multi =L cd +L semantic Among them L cd and L semantic They are geometric constraints and semantic constraints respectively. 7) A method based on decoded point cloud feature representation is proposed to obtain reconstructed point cloud. a. In one example, different point cloud geometry reconstruction decoders may be used for different types of point clouds. i. In one example, for small object point clouds with limited points (such as primitive objects), a folding-based approach can be used to reconstruct the point cloud from the decoded point cloud feature representation. 1. In one example, FoldingNet can be used as the base reconstruction network in folding-based methods. 2. In one example, a folding method can be used to reconstruct the 3D structure of a point cloud from the feature space. ii. In one example, for large-scale point clouds with a large number of points (such as LIDAR point clouds), sparse convolution-based methods and point-based methods can be used to reconstruct the point clouds. 1. In one example, sparse convolution upscaling can be used as the basic reconstruction network in sparse convolution-based methods. 2. In one example, a large-scale point cloud can be divided into blocks based on the number of points. 3. In one example, a point-based approach can be used to encode, decode, and reconstruct the segmented point cloud. 8) Whether and / or how the above disclosed methods are applied may be signaled in the bitstream / frame / slice / slice / octree / etc. from the encoder to the decoder. 9) Whether and / or how to apply the above disclosed methods may depend on the coded information, such as dimension, color format, color component, slice / picture type. Overview 10) The syntax elements disclosed above can be binarized as flags, fixed length codes, EG(x) codes, unary codes, truncated unary codes, truncated binary codes, etc. It can be signed or unsigned. 11) The syntax elements disclosed above can be encoded or decoded using at least one context model, or they can be bypassed. 12) The syntax elements disclosed above may be signaled in a conditional manner. a. SE is signaled only when the corresponding function is applicable. b. SE is signaled only if the dimensions (width and / or height) of the block meet the conditions. 13) The syntax elements disclosed above can be transmitted by signals at the block level / sequence level / picture group level / frame group level / picture level / frame level / slice level / slice level / slice group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / GPS / APS / slice header / slice group header / slice header. 14) Whether and / or how to apply the method disclosed above can be transmitted by signal at block level / sequence level / picture group level / frame group level / picture level / frame level / slice level / slice level / slice group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in sequence header / picture header / SPS / VPS / DPS / DCI / PPS / GPS / APS / slice header / slice group header / slice header. 15) Whether and / or how to apply the above disclosed methods may depend on coded information such as block size, frame size, color format, attribute format, single / dual tree partitioning, color component, slice / slice / picture / frame type. 16) The proposed method disclosed in this document can be used in other codecs that require chroma fusion. 6. Examples
[0058] In the present disclosure, an improved semantic enhancement-based point cloud compression method is proposed. First, a pre-trained point cloud geometric feature extraction model is reused to extract a compact feature representation of the point cloud. In this way, point cloud data with complex spatial structures is represented as high-dimensional feature information for further processing. Secondly, a point cloud feature mapping module is introduced. Because point cloud geometric encoding and decoding reconstruction and point cloud machine vision intelligent analysis are two completely different tasks, their representation feature spaces should also be different. The point cloud feature mapping module can explore the similarities of these two different tasks in the feature space as much as possible. Third, during the encoding and decoding model training process, a multi-task learning mechanism is introduced to retain as much point cloud geometric structure and semantic information as possible.
[0059] Figure 5 An example of an encoding and decoding flow 500 of a semantically enhanced point cloud compression method is depicted in FIG. Figure 5 A flow of a compression method according to an embodiment of the present disclosure is depicted. Figure 5 Can be compared to Figure 2 As shown, the geometry encoder 220 may include a graph-based extractor 510 and a depth PCC (DPCC)-based extractor 515 . The geometry decoder 230 may include a folding-based decoder 570 and a DPCC-based decoder 575 .
[0060] GPCC 220 may include multiple modules, including an octree encoder 530, an octree decoder 535, and an entropy model 560. GPCC 220 may include additional modules, such as a quantization (Q) module, an AE module, an AD module, and the like.
[0061] Features such as X from the geometry encoder 210 can be input to the feature encoder module 520. The feature encoder module 520 can process the input features and obtain output features such as Y. The output features can be input to the GPCC 220, such as Figure 5 The output features from the GPCC 220 may be input to a feature encoder module 525. The feature encoder module 520 may process the features and output the processed features to the geometry decoder 230. The geometry decoder 230 may transmit the point cloud data to the machine vision task 240, which may include a point cloud classification task 580 and a point cloud segmentation task 585.
[0062] Figure 6 An example of the feature encoder module 520 is depicted in FIG. Figure 6 An example of the flow of the feature encoder module 520 is depicted. The feature encoder module 520 may include an input embedding module 610. The input features may be input to the input embedding module 610. The feature encoder module 520 may include additional modules such as a multi-head attention module, an addition and normalization module, a feedforward module, an addition and normalization module, etc. It should be understood that Figure 5 The feature encoder module 525 in can be similar to or the same as the feature encoder module 520 and will not be repeated here.
[0063] More details will be discussed further below. Embodiments of the present disclosure relate to point cloud encoding and compression. As used herein, the term "point cloud sequence" may refer to a sequence of one or more point clouds. The term "frame" may refer to a point cloud in a point cloud sequence. The term "point cloud" may refer to a frame in a point cloud sequence. The term "codec unit" may refer to a block, box, cube, slice, piece, frame, or any other unit referring to a group of points in a PCC.
[0064] Figure 7FIG2 is a flowchart of a method 700 for point cloud encoding and decoding according to an embodiment of the present disclosure. The method 700 can be implemented for conversion between a current point cloud of a point cloud sequence and a bitstream of the point cloud sequence.
[0065] like Figure 7 As shown, method 700 begins at block 710, where the type of the current point cloud is determined. At block 720, a codec module for the current point cloud is determined based on the type of the current point cloud. For example, the codec module may include a point cloud feature extractor. As another example, the codec module may include a point cloud geometry reconstruction module. At block 730, conversion is performed based on the codec module.
[0066] Method 700 can determine the codec module for conversion based on the type of point cloud. For example, an appropriate feature extractor can be selected based on the type of point cloud. As another example, an appropriate reconstruction module can be selected, such as a reconstruction decoder. As a result, codec efficiency can be improved.
[0067] In some embodiments, the conversion includes encoding the current point cloud into a bitstream. For example, the conversion can be performed by an encoder that encodes and decodes information to be included in the bitstream.
[0068] In some embodiments, determining the codec module includes determining a point cloud feature extractor based on the type of the current point cloud. For example, the point cloud feature extractor includes at least one of: a point-based extractor or a voxel-based extractor. In some embodiments, the point cloud feature extractor may be a compact feature representation extractor. That is, different compact feature representation extractors may be used for different point clouds.
[0069] In some embodiments, the type of the current point cloud comprises a basic object or a basic solid point cloud having a number of points less than a threshold number, and the encoding and decoding module comprises a point-based extractor.
[0070] In some embodiments, the point-based extractor is based on at least one of: a graph-based approach or a transformer-based approach. For example, a graph-based approach may include a graph-based extractor such as, Figure 5 Graph-based extractor 510 in .
[0071] In some embodiments, the point-based extractor is based on a graph-based scheme, and graph convolution is used to model the topological structure of the current point cloud, and the feature representation of the current point cloud is extracted based on the modeling of the topological structure of the current point cloud.
[0072] In some embodiments, a K-nearest neighbor (KNN) algorithm is used by graph convolution to obtain local point cloud neighbors and extract features of local neighboring points.
[0073] In some embodiments, the point-based extractor is based on a transformer-based scheme for modeling the topological structure of the current point cloud, the spatial structure of the current point cloud is modeled, and the feature representation of the current point cloud is extracted based on the transformer's attention mechanism.
[0074] In some embodiments, the transformer-based approach performs progressive geometric feature extraction on the current point cloud by stacking multiple layers of attention mechanisms.
[0075] In some embodiments, the type of the current point cloud includes a Lidar point cloud or a large-scale point cloud having a complex structure and a number of points greater than a threshold number, and the encoding and decoding module includes a voxel-based extractor.
[0076] In some embodiments, a sparse convolution based scheme is used.
[0077] In some embodiments, a sparse convolution-based method is used to extract geometric structure features of the current point cloud.
[0078] In some embodiments, the sparse convolution voxelizes the current point cloud and determines geometric properties of the voxelized current point cloud.
[0079] In some embodiments, voxelization is used to regularize point cloud data for subsequent convolution processing. For example, voxelization can be used to regularize point cloud data for subsequent convolution processing.
[0080] In some embodiments, the point cloud feature extractor includes a pre-trained point cloud feature extractor that extracts compact feature representations for point clouds with different spatial structures and point sizes. That is, different pre-trained point cloud feature extractors can be used to extract compact feature representations for point clouds with different spatial structures and point sizes.
[0081] In some embodiments, an indication of the point cloud feature extractor to be used is included in the bitstream. For example, which point cloud feature extractor to use can be signaled from the encoder to the decoder.
[0082] In some embodiments, the point cloud feature extractor to be used is determined by the decoder used to decode the current point cloud from the bitstream.
[0083] In some embodiments, the conversion includes decoding the current point cloud from the bitstream. For example, the conversion can be performed by a decoder. The decoder encodes and decodes the compressed bits to determine the information in the bitstream.
[0084] In some embodiments, determining the codec module includes determining a point cloud geometry reconstruction module (also referred to as a point cloud geometry reconstruction decoder) based on the type of the current point cloud. The reconstructed point cloud can be determined by the point cloud geometry reconstruction module based on the decoded point cloud feature representation.
[0085] In some embodiments, the current point cloud includes a primitive or a small object point cloud having a number of points less than a threshold number, and the point cloud geometry reconstruction decoder uses a folding-based scheme to reconstruct the current point cloud from the decoded point cloud feature representation. For example, the folding-based method may use a folding-based decoder such as, Figure 5 The folding-based decoder 570 in .
[0086] In some embodiments, FoldingNet is used as the base reconstruction network for folding-based schemes.
[0087] In some embodiments, a folding-based scheme is used to reconstruct the 3D structure of the current point cloud from the feature space.
[0088] In some embodiments, the current point cloud includes a Lidar point cloud or a large-scale point cloud having a number of point clouds greater than a threshold number, and at least one of a sparse convolution-based scheme or a point-based scheme is used to reconstruct the current point cloud.
[0089] In some embodiments, sparse convolutional upscaling is used as the base reconstruction network in sparse convolution-based methods.
[0090] In some embodiments, the large-scale point cloud is divided into blocks based on the number of points in the large-scale point cloud.
[0091] In some embodiments, a point-based scheme is used to encode, decode, and reconstruct the tiled point cloud.
[0092] In some embodiments, the method 700 further includes applying a point cloud feature mapping module to determine the similarity of features in multiple tasks. For example, the point cloud feature mapping module can be implemented as Figure 5 The feature encoder module 520 and / or feature encoder module 525 in, or as Figure 6 An input embedding module 610 is shown.
[0093] In some embodiments, the multiple tasks include a point cloud geometry reconstruction task, and the features extracted from the current point cloud for the point cloud geometry reconstruction task include more point cloud spatial structure information than for another task. That is, for the point cloud geometry reconstruction task, the extracted features may include more point cloud spatial structure information.
[0094] In some embodiments, the spatial structure information includes at least one of the following: high-frequency texture information or low-frequency structure information in the point cloud data.
[0095] In some embodiments, the plurality of tasks includes a point cloud intelligent analysis task, and the extracted features of the current point cloud include more point cloud semantic information than for another task. For example, the point cloud intelligent analysis task may include at least one of the following: point cloud segmentation (such as, Figure 5 Point cloud segmentation 585 in) or point cloud classification (such as, Figure 5 That is, for the point cloud intelligent analysis task, the extracted features can contain more point cloud semantic information.
[0096] In some embodiments, the point cloud semantic information includes semantic information of each point in the current point cloud.
[0097] In some embodiments, the point cloud feature mapping module includes a transformer-based point cloud feature mapping module, which is used to minimize the feature similarity between the feature spaces of point cloud geometric encoding and decoding reconstruction and point cloud machine vision intelligent analysis.
[0098] In some embodiments, the design of the point cloud feature mapping module is a transformer structure based on a multi-head attention mechanism.
[0099] In some embodiments, a self-attention mechanism is used to determine feature similarity in the feature map space, and the features are weighted to be summed.
[0100] In some embodiments, the point cloud feature mapping module is designed and maintained symmetrically at at least one of: an encoder for conversion, or a decoder for conversion.
[0101] In some embodiments, method 700 further includes: preserving the semantic information of the point cloud while maintaining the geometric structure of the current point cloud based on a multi-task learning mechanism. For example, the multi-task learning mechanism can be used to preserve as much semantic information as possible while maintaining the geometric structure of the point cloud. In this way, semantic loss can be reduced, and the accuracy of point cloud semantic segmentation and classification can be improved.
[0102] In some embodiments, an indication of the use of the multi-task model is included in the bitstream. For example, whether the multi-task model is used can be signaled from the encoder to the decoder.
[0103] Alternatively or additionally, in some embodiments, whether to use the multi-task model is determined at the decoder for conversion.
[0104] In some embodiments, during the training phase for the multi-task learning mechanism, multi-objective loss constraints are used.
[0105] In some embodiments, geometric constraints for point cloud reconstruction are applied to point clouds of basic quality.
[0106] In some embodiments, the geometric constraints are determined based on chamfer distances for supervised learning.
[0107] In some embodiments, the chamber distance is determined by the following formula: Where S1 and S2 are two point clouds, x and y are the coordinates of the points in S1 and S2 respectively, and L CD (S1, S2) represents the cavity distance between two point clouds.
[0108] In some embodiments, semantic constraints for machine vision are applied to point clouds.
[0109] In some embodiments, semantic constraints are used to constrain the feature distribution distance between the encoded and decoded features of the point cloud and the original features.
[0110] In some embodiments, the semantic constraints are determined based on Kullback-Leibler (KL) divergence.
[0111] In some embodiments, the KL divergence distance L semantic is determined by the following formula: L semantic =D KL (F ori ||F rec )=∑ip ori (v i )log p ori (v i )-p ori (v i )log p rec (v i ), Among them F ori and F rec Represents the original features and the encoded and decoded features, p ori and p rec Represent the probability distribution of original features and encoded and decoded features respectively.
[0112] In some embodiments, the final loss constraint is determined by the following formula: multi =L cd +L semantic , where L cd Represents geometric constraints, L semantic represents a semantic constraint, and L multirepresents the final loss constraint. In this way, both the error in objective reconstruction quality during encoding and decoding and the loss of feature semantics during encoding and decoding can be considered. Through such evaluation, the training process of AI-PCC can be improved.
[0113] In some embodiments, an indication of the final sampled point cloud is included in the bitstream.
[0114] In some embodiments, the indication of the final sampled point cloud is encoded and decoded by a point cloud codec, such as a point cloud codec including at least one of: geometry-based point cloud compression (G-PCC), video-based point cloud compression (V-PCC), or Draco.
[0115] In some embodiments, an indication of at least one feature of the current point cloud is included in the bitstream.
[0116] In some embodiments, the indication of the at least one characteristic is encoded using at least one of: a fixed length codec, a unary codec, or a truncated unary codec.
[0117] In some embodiments, the indication of at least one characteristic is encoded in a predictive manner.
[0118] In some embodiments, information about whether to apply the method and / or how to apply the method is included in at least one of the following: a bitstream, a frame, a slice, a slice, or an octree.
[0119] In some embodiments, whether and / or how to apply the method is based on coded information including at least one of: dimension, color format, color component, slice type, or picture type.
[0120] In some embodiments, the syntax element or indication is binarized as at least one of: a flag, a fixed length code, an Exponential Golomb (x) (EG(x)) code, a unary code, a truncated unary code, or a truncated binary code.
[0121] In some embodiments, syntax elements or indications are signed or unsigned.
[0122] In some embodiments, syntax elements or indications are encoded using at least one context model.
[0123] In some embodiments, syntax elements or indications are bypassed for encoding.
[0124] In some embodiments, the syntax element or indication is included in the bitstream based on at least one condition, the at least one condition comprising at least one of: a first condition that the function corresponding to the syntax element or indication is applicable, or a second condition that the dimensions of the block of the current point cloud satisfy the condition.
[0125] In some embodiments, the syntax element or indication is included at one of: block level, sequence level, group of picture level, frame group level, picture level, frame level, slice level, slice level, or slice group level.
[0126] In some embodiments, the syntax element or indication is included in one of the following: codec tree unit (CTU), codec unit (CU), transform unit (TU), prediction unit (PU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptation parameter set (APS), slice header, slice group header or slice header.
[0127] In some embodiments, whether to apply the method and / or how to apply the method is included at one of the following: block level, sequence level, picture group level, frame group level, picture level, frame level, slice level, slice level, or slice group level.
[0128] In some embodiments, whether to apply the method and / or how to apply the method is included in one of the following: codec tree unit (CTU), codec unit (CU), transform unit (TU), prediction unit (PU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptation parameter set (APS), slice header, slice group header or slice header.
[0129] In some embodiments, whether to apply a method and / or how to apply a method is based on the encoded information.
[0130] In some embodiments, the encoded information includes at least one of: block size, color format, attribute format, single-tree partitioning or dual-tree partitioning, color component, slice type, slice type, picture type, or frame type.
[0131] In some embodiments, the method is used in a codec that requires chroma fusion.
[0132] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. A bitstream of a point cloud sequence is stored in the non-transitory computer-readable recording medium. The bitstream of the point cloud sequence is generated by a method executed by a point cloud sequence processing device. According to the method, the type of the current point cloud in the point cloud sequence is determined. The codec module for the current point cloud is determined based on the type of the current point cloud. The codec module includes at least one of the following: a point cloud feature extractor or a point cloud geometry reconstruction module. The bitstream is generated based on the codec module.
[0133] According to further embodiments of the present disclosure, a method for storing a bitstream of a point cloud sequence is provided. In this method, the type of a current point cloud in the point cloud sequence is determined. A codec module for the current point cloud is determined based on the type of the current point cloud. The codec module includes at least one of the following: a point cloud feature extractor or a point cloud geometry reconstruction module. A bitstream is generated based on the codec module. The bitstream is stored in a non-transitory computer-readable recording medium.
[0134] The embodiments of the present disclosure may be described according to the following items, the features of which may be combined in any reasonable way.
[0135] Item 1. A method for point cloud encoding and decoding, comprising: determining the type of the current point cloud for conversion between a current point cloud of a point cloud sequence and a bit stream of the point cloud sequence; determining an encoding and decoding module for the current point cloud based on the type of the current point cloud, the encoding and decoding module comprising at least one of the following: a point cloud feature extractor or a point cloud geometry reconstruction module; and performing conversion based on the encoding and decoding module.
[0136] Item 2. The method of Item 1, wherein converting comprises encoding the current point cloud into a bitstream.
[0137] Clause 3. The method of clause 2, wherein the conversion is performed by an encoder that encodes and decodes the information to be included in the bitstream.
[0138] Item 4. The method according to Item 2 or 3, wherein determining the codec module comprises: determining a point cloud feature extractor based on the type of the current point cloud.
[0139] Item 5. The method of Item 4, wherein the point cloud feature extractor comprises at least one of: a point-based extractor or a voxel-based extractor.
[0140] Item 6. The method according to Item 4 or 5, wherein the type of the current point cloud comprises a basic object or a basic solid point cloud having a number of points less than a threshold number, and the encoding and decoding module comprises a point-based extractor.
[0141] Item 7. The method of Item 6, wherein the point-based extractor is based on at least one of: a graph-based method or a transformer-based method.
[0142] Item 8. A method according to Item 7, wherein the point-based extractor is based on a graph-based scheme, and graph convolution is used to model the topological structure of the current point cloud, and the feature representation of the current point cloud is extracted based on the modeling of the topological structure of the current point cloud.
[0143] Item 9. The method of Item 8, wherein a K-nearest neighbor (KNN) algorithm is used by graph convolution to obtain local point cloud neighbors and extract features of local neighboring points.
[0144] Item 10. A method according to Item 7, wherein the point-based extractor is based on a transformer-based scheme for modeling the topological structure of the current point cloud, and the spatial structure of the current point cloud is modeled, and the feature representation of the current point cloud is extracted based on the attention mechanism of the transformer.
[0145] Item 11. The method of Item 10, wherein the transformer-based scheme performs progressive geometric feature extraction on the current point cloud by stacking multiple layers of attention mechanisms.
[0146] Item 12. The method according to Item 4 or 5, wherein the type of the current point cloud comprises a Lidar point cloud or a large-scale point cloud having a complex structure and having a number of points greater than a threshold number, and the encoding and decoding module comprises a voxel-based extractor.
[0147] Item 13. The method of Item 12, wherein a sparse convolution based scheme is used.
[0148] Item 14. The method according to Item 13, wherein a sparse convolution-based scheme is used to extract geometric structural features of the current point cloud.
[0149] Item 15. The method of Item 13 or 14, wherein the sparse convolution voxelizes the current point cloud and determines geometric properties of the voxelized current point cloud.
[0150] Item 16. The method of Item 15, wherein voxelization is used to regularize the point cloud data for subsequent convolution processing.
[0151] Item 17. A method according to any one of Items 4 to 16, wherein the point cloud feature extractor comprises a pre-trained point cloud feature extractor that extracts compact feature representations for point clouds with different spatial structures and point sizes.
[0152] Item 18. A method according to any one of Items 4 to 16, wherein an indication of the point cloud feature extractor to be used is included in the bitstream.
[0153] Item 19. A method according to any one of Items 4 to 16, wherein the point cloud feature extractor to be used is determined by a decoder used to decode the current point cloud from a bitstream.
[0154] Item 20. The method of Item 1, wherein converting comprises decoding the current point cloud from a bitstream.
[0155] Item 21. The method of Item 20, wherein the conversion is performed by a decoder that encodes and decodes the compressed bits to determine the information in the bitstream.
[0156] Item 22. The method according to Item 20 or 21, wherein determining the encoding and decoding module comprises: determining a point cloud geometry reconstruction module based on the type of the current point cloud.
[0157] Item 23. A method according to Item 22, wherein the current point cloud includes a basic object or a small object point cloud as follows: the small object point cloud has a number of points less than a threshold number, and the point cloud geometry reconstruction module uses a folding-based scheme to reconstruct the current point cloud from the decoded point cloud feature representation.
[0158] Item 24. The method of Item 23, wherein FoldingNet is used as the base reconstruction network for the folding-based scheme.
[0159] Item 25. The method according to Item 23 or 24, wherein a folding-based scheme is used to reconstruct the three-dimensional structure of the current point cloud from the feature space.
[0160] Item 26. A method according to Item 22, wherein the current point cloud comprises a Lidar point cloud or a large-scale point cloud having a number of point clouds greater than a threshold number, and at least one of a sparse convolution-based scheme or a point-based scheme is used to reconstruct the current point cloud.
[0161] Item 27. The method of Item 26, wherein sparse convolutional upscaling is used as a base reconstruction network in a sparse convolution-based scheme.
[0162] Item 28. The method of Item 26 or 27, wherein the large-scale point cloud is divided into blocks based on the number of points in the large-scale point cloud.
[0163] Item 29. The method of Item 28, wherein a point-based scheme is used to encode, decode, and reconstruct the partitioned point cloud.
[0164] Item 30. The method according to any one of Items 1 to 29, further comprising: applying a point cloud feature mapping module to determine similarities of features in multiple tasks.
[0165] Item 31. The method according to Item 30, wherein the multiple tasks include a point cloud geometry reconstruction task, and for the point cloud geometry reconstruction task, the extracted features of the current point cloud include more point cloud spatial structure information than for another task.
[0166] Item 32. The method according to Item 31, wherein the spatial structure information includes at least one of the following: high-frequency texture information or low-frequency structure information in the point cloud data.
[0167] Item 33. A method according to Item 30, wherein the multiple tasks include a point cloud intelligent analysis task, and the extracted features of the current point cloud include more point cloud semantic information than for another task, and wherein the point cloud intelligent analysis task includes at least one of the following: point cloud segmentation or point cloud classification.
[0168] Item 34. A method according to Item 33, wherein the point cloud semantic information includes semantic information of each point in the current point cloud.
[0169] Item 35. The method according to Item 30, wherein the point cloud feature mapping module includes a transformer-based point cloud feature mapping module, and the point cloud feature mapping module is used to minimize the feature similarity between the feature spaces of point cloud geometric encoding and decoding reconstruction and point cloud machine vision intelligent analysis.
[0170] Item 36. The method according to Item 35, wherein the design of the point cloud feature mapping module is a transformer structure based on a multi-head attention mechanism.
[0171] Item 37. The method of Item 35, wherein a self-attention mechanism is used to determine feature similarity in a feature map space, and the features are weighted to be summed.
[0172] Item 38. A method according to any one of Items 30 to 37, wherein the point cloud feature mapping module is symmetrically designed and retained at at least one of: an encoder for conversion, or a decoder for conversion.
[0173] Item 39. The method according to any one of Items 1 to 38 further includes: preserving the semantic information of the point cloud while maintaining the geometric structure of the current point cloud based on a multi-task learning mechanism.
[0174] Clause 40. The method of clause 39, wherein an indication of use of the multi-tasking model is included in the bitstream.
[0175] Item 41. The method of Item 39, wherein whether to use a multi-task model is determined at a decoder for conversion.
[0176] Item 42. A method according to any one of Items 39 to 41, wherein during a training phase for a multi-task learning mechanism, loss constraints for multiple objectives are used.
[0177] Item 43. The method of Item 42, wherein geometric constraints for point cloud reconstruction are applied to point clouds of basic quality.
[0178] Item 44. The method of Item 43, wherein the geometric constraints are determined based on chamfer distances for supervised learning.
[0179] Item 45. The method of Item 44, wherein the chamber distance is determined by the following formula: Where S1 and S2 are two point clouds, x and y are the coordinates of the points in S1 and S2 respectively, and L CD (S1, S2) represents the cavity distance between two point clouds.
[0180] Item 46. A method according to any one of Items 42 to 45, wherein semantic constraints for machine vision are applied to the point cloud.
[0181] Item 47. The method of Item 46, wherein semantic constraints are used to constrain the feature distribution distance between the encoded features and the original features of the point cloud.
[0182] Item 48. The method of Item 47, wherein the semantic constraint is determined based on Kullback-Leibler (KL) divergence.
[0183] Item 49. The method of Item 48, wherein the KL divergence distance L semantic is determined by the following formula: L semantic =D KL (F ori ||F rec )=∑ip ori (v i )log p ori (v i )-p ori (v i )log p rec (v i ), where F ori and F rec Represents the original features and the encoded and decoded features, p ori and p rec Represent the probability distribution of original features and encoded and decoded features respectively.
[0184] Item 50. The method of Item 42, wherein the final loss constraint is determined by the following formula: L multi =L cd +L semantic , where L cd Represents geometric constraints, L semantic represents a semantic constraint, and L multi represents the final loss constraint.
[0185] Item 51. A method according to any of Items 1 to 50, wherein an indication of the final sampled point cloud is included in the bitstream.
[0186] Item 52. The method of Item 51, wherein the indication of the final sampled point cloud is encoded and decoded by a point cloud codec.
[0187] Item 53. The method of Item 52, wherein the point cloud codec comprises at least one of: geometry-based point cloud compression (G-PCC), video-based point cloud compression (V-PCC), or Draco.
[0188] Item 54. A method according to any of Items 1 to 53, wherein an indication of at least one feature of the current point cloud is included in the bitstream.
[0189] Item 55. The method of Item 54, wherein the indication of at least one characteristic is encoded using at least one of: a fixed length codec, a unary codec, or a truncated unary codec.
[0190] Item 56. A method according to Item 54 or 55, wherein the indication of at least one feature is encoded in a predictive manner.
[0191] Item 57. A method according to any one of items 1 to 56, wherein information about whether to apply the method and / or how to apply the method is included in at least one of the following: a bitstream, a frame, a slice, a slice or an octree.
[0192] Item 58. A method according to any one of items 1 to 56, wherein whether and / or how to apply the method is based on encoded information, the encoded information comprising at least one of: dimension, color format, color component, slice type or picture type.
[0193] Item 59. A method according to any one of items 1 to 58, wherein the syntax element or indication is binarized into at least one of the following: a flag, a fixed length code, an Exponential Golomb (x) (EG(x)) code, a unary code, a truncated unary code or a truncated binary code.
[0194] Item 60. The method of Item 59, wherein the syntax element or indication is signed or unsigned.
[0195] Item 61. A method according to any one of items 1 to 60, wherein the syntax element or indication is encoded and decoded using at least one context model.
[0196] Item 62. A method according to any one of items 1 to 60, wherein the syntax element or indication is bypassed.
[0197] Item 63. A method according to any one of items 1 to 62, wherein the syntax element or indication is included in the bitstream based on at least one condition, the at least one condition comprising at least one of the following: a first condition that the function corresponding to the syntax element or indication is applicable, or a second condition that the dimensions of the block of the current point cloud satisfy the condition.
[0198] Item 64. A method according to any one of items 1 to 63, wherein the syntax element or indication is included at one of the following: block level, sequence level, picture group level, frame group level, picture level, frame level, slice level, slice level or slice group level.
[0199] Item 65. A method according to item 64, wherein the syntax element or indication is included in one of the following: a codec tree unit (CTU), a codec unit (CU), a transform unit (TU), a prediction unit (PU), a codec tree block (CTB), a codec block (CB), a transform block (TB), a prediction block (PB), a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a slice group header or a slice header.
[0200] Item 66. A method according to any one of items 1 to 65, wherein whether to apply the method and / or how to apply the method is included at one of the following: block level, sequence level, picture group level, frame group level, picture level, frame level, slice level, slice level or slice group level.
[0201] Item 67. A method according to item 66, wherein whether to apply the method and / or how to apply the method is included in one of the following: codec tree unit (CTU), codec unit (CU), transform unit (TU), prediction unit (PU), codec tree block (CTB), codec block (CB), transform block (TB), prediction block (PB), sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), decoding parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptation parameter set (APS), slice header, slice group header or slice header.
[0202] Item 68. A method according to any one of items 1 to 65, wherein whether to apply the method and / or how to apply the method is based on encoded information.
[0203] Item 69. The method of Item 68, wherein the encoded information comprises at least one of: block size, color format, attribute format, single tree partitioning or dual tree partitioning, color component, slice type, slice type, picture type, or frame type.
[0204] Item 70. A method according to any one of Items 1 to 69, wherein the method is used in a codec tool requiring chroma fusion.
[0205] Item 71. An apparatus for point cloud encoding and decoding, comprising a processor and a non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1 to 70.
[0206] Item 72. A non-transitory computer-readable storage medium storing instructions for causing a processor to perform the method of any one of Items 1 to 70.
[0207] Item 73. A non-transitory computer-readable recording medium storing a bit stream of a point cloud sequence, the bit stream being generated by a method performed by a point cloud processing device, wherein the method comprises: determining the type of a current point cloud of the point cloud sequence; determining a codec module for the current point cloud based on the type of the current point cloud, the codec module comprising at least one of the following: a point cloud feature extractor or a point cloud geometry reconstruction module; and generating a bit stream based on the codec module.
[0208] Item 74. A method for storing a bitstream of a point cloud sequence, comprising: determining the type of a current point cloud of the point cloud sequence; determining a codec module for the current point cloud based on the type of the current point cloud, the codec module comprising at least one of the following: a point cloud feature extractor or a point cloud geometry reconstruction module; generating a bitstream based on the codec module; and storing the bitstream in a non-transitory computer-readable recording medium. Example device
[0209] Figure 8A block diagram of a computing device 800 in which various embodiments of the present disclosure may be implemented is shown. The computing device 800 may be implemented as the source device 110 (or the PCC encoder 116 or the GPCC encoder 300) or the destination device 120 (or the PCC decoder 126 or the GPCC decoder 400), or may be included in the source device 110 (or the PCC encoder 116 or the GPCC encoder 300) or the destination device 120 (or the PCC decoder 126 or the GPCC decoder 400).
[0210] It should be understood that Figure 8 The computing device 800 shown in FIG. 8 is for illustration purposes only and is not intended to in any way imply any limitation on the functionality and scope of the disclosed embodiments.
[0211] like Figure 8 As shown, computing device 800 comprises a general computing device 800. Computing device 800 may include at least one or more processors or processing units 810, memory 820, storage unit 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860.
[0212] In some embodiments, the computing device 800 can be implemented as any user terminal or server terminal with computing power. The server terminal can be a server, a large computing device, etc. provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a gaming device, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 800 can support any type of interface to the user (such as a "wearable" circuit device, etc.).
[0213] The processing unit 810 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capability of the computing device 800. The processing unit 810 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0214] The computing device 800 typically includes various computer storage media. Such media can be any media accessible by the computing device 800, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 820 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 830 can be any removable or non-removable medium and can include machine-readable media, such as memory, a flash drive, a disk or other media that can be used to store information and / or data and can be accessed in the computing device 800.
[0215] The computing device 800 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 8 Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.
[0216] The communication unit 840 communicates with another computing device via a communication medium. In addition, the functionality of the components in the computing device 800 can be implemented by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 800 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0217] The input device 850 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. The output device 860 may be one or more of various output devices, such as a display, a speaker, a printer, etc. With the help of the communication unit 840, the computing device 800 may also communicate with one or more external devices (not shown), such as storage devices and display devices, and may also communicate with one or more devices that enable a user to interact with the computing device 800, or, if desired, may also communicate with any device that enables the computing device 800 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0218] In some embodiments, some or all components of the computing device 800 may also be arranged in a cloud computing architecture rather than being integrated into a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides an application via a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on servers in a remote location. Computing resources in a cloud computing environment can be consolidated or distributed across locations in remote data centers. Cloud computing infrastructure can provide services through shared data centers, although to users, they appear to be a single access point. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider in a remote location. Alternatively, the components and functionality described herein can be provided by a conventional server or installed directly or otherwise on a client device.
[0219] In embodiments of the present disclosure, computing device 800 may be used to implement point cloud encoding / decoding. Memory 820 may include one or more point cloud encoding / decoding modules 825 having one or more program instructions. These modules are accessible and executable by processing unit 810 to perform the functions of various embodiments described herein.
[0220] In an example embodiment of performing point cloud encoding, an input device 850 may receive point cloud data as input to be encoded 870. The point cloud data may be processed by, for example, a point cloud encoding / decoding module 825 to generate an encoded bitstream. The encoded bitstream may be provided as output 880 via an output device 860.
[0221] In an example embodiment of performing point cloud decoding, an input device 850 may receive an encoded bitstream as input 870. The encoded bitstream may be processed, for example, by a point cloud encoding / decoding module 825 to generate decoded point cloud data. The decoded point cloud data may be provided as output 880 via an output device 860.
[0222] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such changes are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for point cloud encoding and decoding, comprising: determining a type of the current point cloud for conversion between a current point cloud of a point cloud sequence and a bitstream of the point cloud sequence; Determining a codec module for the current point cloud based on the type of the current point cloud, the codec module comprising at least one of the following: a point cloud feature extractor or a point cloud geometry reconstruction module; as well as The conversion is performed based on the codec module. 2 . The method of claim 1 , wherein the converting comprises encoding the current point cloud into the bitstream. 3 . The method of claim 2 , wherein the conversion is performed by an encoder that encodes and decodes information to be included in the bitstream.
4. The method according to claim 2 or 3, wherein determining the codec module comprises: The point cloud feature extractor is determined based on the type of the current point cloud.
5. The method according to claim 4, wherein the point cloud feature extractor comprises at least one of the following: a point-based extractor, or Voxel-based extractor.
6. The method according to claim 4 or 5, wherein the type of the current point cloud comprises a basic object or a basic solid point cloud having a number of points less than a threshold number, and the encoding and decoding module comprises a point-based extractor.
7. The method of claim 6, wherein the point-based extractor is based on at least one of: Graph-based solutions, or Converter-based solution.
8. The method according to claim 7, wherein the point-based extractor is based on a graph-based scheme, and graph convolution is used to model the topological structure of the current point cloud, and the feature representation of the current point cloud is extracted based on the modeling of the topological structure of the current point cloud.
9. The method according to claim 8, wherein a K-nearest neighbor (KNN) algorithm is used by the graph convolution to obtain local point cloud neighbors and extract features of local neighboring points.
10. The method of claim 7, wherein the point-based extractor is based on a transformer-based scheme for modeling the topological structure of the current point cloud, and the spatial structure of the current point cloud is modeled, and the feature representation of the current point cloud is extracted based on the attention mechanism of the transformer.
11. The method according to claim 10, wherein the transformer-based scheme performs progressive geometric feature extraction on the current point cloud by stacking multiple layers of attention mechanisms.
12. The method according to claim 4 or 5, wherein the type of the current point cloud comprises a Lidar point cloud or a large-scale point cloud having a complex structure and a number of points greater than a threshold number, and the encoding and decoding module comprises a voxel-based extractor. The method according to claim 12 , wherein a sparse convolution based scheme is used. The method according to claim 13 , wherein the sparse convolution-based scheme is used to extract geometric structure features of the current point cloud. 15 . The method according to claim 13 , wherein the sparse convolution voxelizes the current point cloud and determines geometric properties of the voxelized current point cloud. The method of claim 15 , wherein voxelization is used to regularize the point cloud data for subsequent convolution processing.
17. The method according to any one of claims 4 to 16, wherein the point cloud feature extractor comprises a pre-trained point cloud feature extractor that extracts compact feature representations for point clouds with different spatial structures and point sizes.
18. A method according to any one of claims 4 to 16, wherein an indication of the point cloud feature extractor to be used is included in the bitstream.
19. The method according to any one of claims 4 to 16, wherein the point cloud feature extractor to be used is determined by a decoder used to decode the current point cloud from the bitstream.
20. The method of claim 1, wherein the converting comprises decoding the current point cloud from the bitstream.
21. The method of claim 20, wherein the converting is performed by a decoder that encodes and decodes compressed bits to determine information in the bitstream.
22. The method according to claim 20 or 21, wherein determining the codec module comprises: The point cloud geometry reconstruction module is determined based on the type of the current point cloud.
23. The method of claim 22, wherein the current point cloud comprises a primitive object or a small object point cloud having a number of points less than a threshold number, and the point cloud geometry reconstruction module uses a folding-based scheme to reconstruct the current point cloud from the decoded point cloud feature representation.
24. The method according to claim 23, wherein FoldingNet is used as a basic reconstruction network for the folding-based scheme.
25. The method according to claim 23 or 24, wherein the folding-based scheme is used to reconstruct the three-dimensional structure of the current point cloud from a feature space.
26. The method of claim 22, wherein the current point cloud comprises a Lidar point cloud or a large-scale point cloud having a number of point clouds greater than a threshold number, and at least one of a sparse convolution-based scheme or a point-based scheme is used to reconstruct the current point cloud.
27. The method of claim 26, wherein sparse convolutional upscaling is used as a basic reconstruction network in the sparse convolution-based scheme.
28. The method of claim 26 or 27, wherein the large-scale point cloud is segmented based on the number of points in the large-scale point cloud.
29. The method of claim 28, wherein the point-based scheme is used to encode, decode and reconstruct the partitioned point cloud.
30. The method according to any one of claims 1 to 29, further comprising: A point cloud feature mapping module is applied to determine the similarity of features across multiple tasks. 31 . The method according to claim 30 , wherein the plurality of tasks include a point cloud geometry reconstruction task, and for the point cloud geometry reconstruction task, the extracted features of the current point cloud include more point cloud spatial structure information than for another task.
32. The method according to claim 31, wherein the spatial structural information comprises at least one of the following: high-frequency texture information or low-frequency structural information in point cloud data.
33. A method according to claim 30, wherein the multiple tasks include a point cloud intelligent analysis task, and the extracted features of the current point cloud include more point cloud semantic information than for another task, and wherein the point cloud intelligent analysis task includes at least one of the following: point cloud segmentation or point cloud classification. The method according to claim 33 , wherein the point cloud semantic information comprises semantic information of each point in the current point cloud.
35. The method according to claim 30, wherein the point cloud feature mapping module comprises a transformer-based point cloud feature mapping module, and the point cloud feature mapping module is used to minimize the feature similarity between the feature spaces of point cloud geometric encoding and decoding reconstruction and point cloud machine vision intelligent analysis.
36. The method according to claim 35, wherein the point cloud feature mapping module is designed based on a transformer structure of a multi-head attention mechanism.
37. The method of claim 35, wherein a self-attention mechanism is used to determine the feature similarity in a feature map space, and features are weighted to be summed.
38. The method according to any one of claims 30 to 37, wherein the point cloud feature mapping module is symmetrically designed and retained at at least one of the following: an encoder for the conversion, or a decoder for the conversion.
39. The method according to any one of claims 1 to 38, further comprising: Based on a multi-task learning mechanism, the geometric structure of the current point cloud is maintained while retaining the semantic information of the point cloud.
40. The method of claim 39, wherein an indication of use of a multi-tasking model is included in the bitstream.
41. The method of claim 39, wherein whether to use a multi-task model is determined at a decoder for the conversion.
42. The method according to any one of claims 39 to 41, wherein during the training phase for the multi-task learning mechanism, multi-objective loss constraints are used.
43. The method of claim 42, wherein geometric constraints for point cloud reconstruction are applied to a base quality point cloud.
44. The method of claim 43, wherein the geometric constraint is determined based on a chamfer distance for supervised learning.
45. The method of claim 44, wherein the chamber distance is determined by the following formula: Where S1 and S2 are two point clouds, x and y are the coordinates of the points in S1 and S2 respectively, and L CD (S1, S2) represents the cavity distance between the two point clouds.
46. The method according to any one of claims 42 to 45, wherein semantic constraints for machine vision are applied to the point cloud.
47. The method of claim 46, wherein the semantic constraint is used to constrain the feature distribution distance between the encoded and decoded features of the point cloud and the original features.
48. The method of claim 47, wherein the semantic constraint is determined based on Kullback-Leibler (KL) divergence.
49. The method of claim 48, wherein the KL divergence distance L semantic is determined by the following formula: L semantic =D KL (F ori ||F rec )=∑ip ori (v i )log p ori (vi)-p ori (v i )log p rec (v i ), Among them F ori and F rec Represents the original features and the encoded and decoded features, p ori and p rec Respectively represent the probability distribution of the original feature and the encoded and decoded feature.
50. The method of claim 42, wherein the final loss constraint is determined by the following formula: L multi =L cd +L semantic , Among them L cd Represents geometric constraints, L semantic represents a semantic constraint, and L multi represents the final loss constraint.
51. A method according to any one of claims 1 to 50, wherein an indication of a final sampled point cloud is included in the bitstream.
52. The method of claim 51, wherein the indication of the final sampled point cloud is encoded and decoded by a point cloud codec.
53. The method of claim 52, wherein the point cloud codec comprises at least one of: Geometry-based point cloud compression (G-PCC), Video-based point cloud compression (V-PCC), or Draco.
54. A method according to any one of claims 1 to 53, wherein an indication of at least one feature of the current point cloud is included in the bitstream.
55. The method of claim 54, wherein the indication of the at least one characteristic is encoded using at least one of: a fixed length codec, a unary codec, or a truncated unary codec.
56. A method according to claim 54 or 55, wherein the indication of the at least one characteristic is encoded in a predictive manner.
57. The method according to any one of claims 1 to 56, wherein information on whether to apply the method and / or how to apply the method is included in at least one of the following: the bitstream, the frame, the slice, the slice or the octree.
58. The method of any one of claims 1 to 56, wherein whether and / or how to apply the method is based on coded information, the coded information comprising at least one of: dimension, color format, color component, slice type, or picture type.
59. The method according to any one of claims 1 to 58, wherein the syntax element or indication is binarized into at least one of: logo, Fixed length code, Exponential Columbus (x) (EG(x)) code, Unary code, truncated unary code, or Truncated binary code.
60. The method of claim 59, wherein the syntax element or the indication is signed or unsigned.
61. A method according to any one of claims 1 to 60, wherein syntax elements or indications are encoded or decoded using at least one context model.
62. A method according to any one of claims 1 to 60, wherein syntax elements or indications are bypass coded.
63. The method of any one of claims 1 to 62, wherein a syntax element or indication is included in the bitstream based on at least one condition, the at least one condition comprising at least one of: a first condition under which the function corresponding to the syntax element or the indication applies, or The dimension of the block of the current point cloud satisfies the second condition.
64. A method according to any one of claims 1 to 63, wherein the syntax element or indication is included in one of: Block level, Sequence level, Picture group level, Frame group level, Picture level, Frame level, Slice level, Film level, or Film group level.
65. The method of claim 64, wherein the syntax element or the indication is included in one of: Codec Tree Unit (CTU), Codec Unit (CU), Transformation Unit (TU), Prediction Unit (PU), Codec Tree Block (CTB), Codec Block (CB), Transform Block (TB), Prediction Block (PB), Sequence header, Picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Decoding Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Slice head, Film group header, or Opening credits.
66. The method of any one of claims 1 to 65, wherein whether to apply the method and / or how to apply the method is included in one of: Block level, Sequence level, Picture group level, Frame group level, Picture level, Frame level, Slice level, Film level, or Film group level.
67. The method of claim 66, wherein whether to apply the method and / or how to apply the method is included in one of: Codec Tree Unit (CTU), Codec Unit (CU), Transformation Unit (TU), Prediction Unit (PU), Codec Tree Block (CTB), Codec Block (CB), Transform Block (TB), Prediction Block (PB), Sequence header, Picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Decoding Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Slice head, Film group header, or Opening credits.
68. The method according to any one of claims 1 to 65, wherein whether to apply the method and / or how to apply the method is based on coded information.
69. The method of claim 68, wherein the encoded information comprises at least one of: block size, color format, attribute format, single-tree partitioning or dual-tree partitioning, color component, slice type, slice type, picture type, or frame type.
70. The method according to any one of claims 1 to 69, wherein the method is used in a codec requiring chroma fusion.
71. An apparatus for processing point cloud data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 70.
72. A non-transitory computer-readable storage medium storing instructions, wherein the instructions cause a processor to execute the method according to any one of claims 1 to 70.
73. A non-transitory computer-readable recording medium storing a bitstream of a point cloud sequence, the bitstream being generated by a method performed by a point cloud processing apparatus, wherein the method comprises: Determining the type of a current point cloud in the point cloud sequence; Determining a codec module for the current point cloud based on the type of the current point cloud, the codec module comprising at least one of the following: a point cloud feature extractor or a point cloud geometry reconstruction module; as well as The bitstream is generated based on the encoding and decoding module.
74. A method for storing a bitstream of a point cloud sequence, comprising: Determining the type of a current point cloud in the point cloud sequence; Determining a codec module for the current point cloud based on the type of the current point cloud, the codec module comprising at least one of the following: a point cloud feature extractor or a point cloud geometry reconstruction module; generating the bitstream based on the encoding and decoding module; as well as The bitstream is stored in a non-transitory computer-readable recording medium.