Point cloud decoding method, point cloud encoding method, decoder, electronic device, and medium
By determining the packet information based on the characteristics of the target point cloud, and adopting a diversified packet method to decode and encode the point cloud, the problem of low point cloud encoding and decoding performance in the existing technology is solved, and more efficient point cloud data processing is achieved.
Patent Information
- Application Number
- PCT/CN2024/109313
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-08-01
- Publication Date
- 2025-06-05
AI Technical Summary
In the prior art, point cloud encoding and decoding performance is not high and the grouping method is single, making it difficult to improve the encoding and decoding performance of point clouds.
By determining the decoded packet information and encoded packet information based on the target characteristics of the target point cloud, a diversified packet method is adopted to group the reconstruction location points, and group the process is performed during attribute decoding and attribute encoding.
The decoding and encoding performance of point clouds is improved, and the processing efficiency of point cloud data is improved through diversified packetization methods.
Smart Images

Figure CN2024109313_05062025_PF_FP_ABST
Abstract
Description
Point cloud decoding method, point cloud encoding method, decoder, electronic device and medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 28, 2023, with application number 202311626550X and invention name “Point cloud decoding method, point cloud encoding method, decoder, electronic device and medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of point cloud encoding and decoding technology, and in particular to a point cloud decoding method, a point cloud encoding method, a decoder, an electronic device, and a computer-readable storage medium. Background Art
[0003] In related point cloud encoding and decoding technologies, after the encoder completes geometric encoding, it groups the reconstructed points according to a fixed formula, and then performs attribute encoding on each group of reconstructed points. Similarly, after the encoder completes geometric decoding, it groups the reconstructed points according to a pre-set formula, and then performs attribute encoding on each group of reconstructed points. These technologies provide a single grouping method, which is not conducive to improving point cloud encoding and decoding performance.
[0004] Summary of the Invention
[0005] The present application provides a point cloud decoding method, a point cloud encoding method, a decoder, an electronic device, and a computer-readable storage medium, which improve the encoding and decoding performance of point clouds, at least to a certain extent.
[0006] In a first aspect, the present application provides a point cloud decoding method, which is executed by a processor, and the method includes: determining decoding grouping information based on target features of the target point cloud, wherein the above-mentioned target features include one or more of the following information: point cloud characteristic information, bounding box information, decoding parameter information and target parameter information; grouping the first reconstructed position points according to the above-mentioned decoding grouping information to obtain multiple groups, wherein the above-mentioned first reconstructed position points are obtained by decoding the geometric code stream of the above-mentioned target point cloud based on the decoder; and performing attribute decoding on each group.
[0007] In a second aspect, the present application provides a point cloud encoding method, which is executed by a processor, and the method includes: determining encoding grouping information based on target features of the target point cloud, wherein the above-mentioned target features include one or more of the following information: point cloud characteristic information, bounding box information, encoding parameter information and target parameter information; grouping the second reconstructed position points according to the above-mentioned encoding grouping information to obtain multiple groups, wherein the above-mentioned second reconstructed position points are obtained by the above-mentioned encoder decoding the geometric code stream of the above-mentioned target point cloud; and performing attribute encoding on each group.
[0008] In a third aspect, the present application provides a decoder comprising: a first determination module, a decoding grouping module, and an attribute decoding module;
[0009] Among them, the above-mentioned first determination module is used to determine the decoding grouping information based on the target characteristics of the target point cloud, wherein the above-mentioned target characteristics include one or more of the following information: point cloud characteristic information, bounding box information, decoding parameter information and target parameter information; the above-mentioned decoding grouping module is used to group the first reconstructed position points through the above-mentioned decoding grouping information to obtain multiple groups, wherein the above-mentioned first reconstructed position points are obtained by decoding the geometric code stream of the above-mentioned target point cloud based on the decoder; the above-mentioned attribute decoding module is used to perform attribute decoding on each group.
[0010] In a fourth aspect, the present application provides a decoder comprising a processor and a memory; the memory is used to store a computer program; and the processor is used to execute the computer program to implement the point cloud decoding method provided in the first aspect.
[0011] In a fifth aspect, the present application provides an encoder, comprising: a second determination module, a coding grouping module, and an attribute coding module;
[0012] Among them, the above-mentioned second determination module is used to determine the coding grouping information according to the target characteristics of the target point cloud, wherein the above-mentioned target characteristics include one or more of the following information related to the encoding of the above-mentioned target point cloud or the default information of the encoding end: point cloud characteristic information, bounding box information, coding parameter information and target parameter information; the above-mentioned coding grouping module is used to group the second reconstructed position points according to the above-mentioned coding grouping information to obtain multiple groups, wherein the above-mentioned second reconstructed position points are obtained by the above-mentioned encoder decoding the geometric code stream of the above-mentioned target point cloud; the above-mentioned attribute prediction module is used to perform attribute encoding on each group.
[0013] In a sixth aspect, the present application provides an encoder comprising a processor and a memory; the memory is used to store a computer program; and the processor is used to execute the computer program to implement the point cloud encoding method provided in the second aspect.
[0014] In a seventh aspect, the present application provides an electronic device comprising a processor and a memory; the memory is used to store a computer program; and the processor is used to execute the computer program to implement the point cloud encoding method provided in the first aspect or the second aspect above.
[0015] In an eighth aspect, the present application provides a chip for implementing the method provided in the first or second aspect above. Specifically, the chip includes: a processor for calling and running a computer program from a memory, so that a device equipped with the chip executes the method provided in the first or second aspect above.
[0016] In a ninth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program enables a computer to execute the method provided in the first or second aspect above.
[0017] In a tenth aspect, the present application provides a computer program product, comprising computer program instructions, wherein the computer program instructions enable a computer to execute the method provided in the first or second aspect above.
[0018] In an eleventh aspect, the present application provides a computer program, which, when executed on a computer, enables the computer to execute the method provided in the first or second aspect above.
[0019] In summary, in the point cloud decoding solution provided in the embodiment of the present application, the decoding grouping information is determined according to the target features of the target point cloud. During the attribute decoding process, the decoder groups the first reconstructed position points according to the above-mentioned grouping information to obtain multiple groups. Among them, the above-mentioned first reconstructed position points are obtained by the decoder decoding the geometric code stream of the target point cloud. Further, the decoder performs attribute decoding on each group. In the decoding solution provided in the embodiment of the present application, the decoder determines the decoding grouping information based on one or more of the following information obtained by default or parsing the code stream of the target point cloud on the codec side: point cloud characteristic information, bounding box information, encoding parameter information and target parameter information. It can be seen that in the embodiment of the present application, the grouping information is determined according to the diversified features related to the target point cloud itself, which can improve the grouping diversity and is conducive to improving the decoding performance of the point cloud.
[0020] Similarly, in the point cloud coding scheme provided in the embodiment of the present application, the coding grouping information is determined based on the target features of the target point cloud. During the attribute coding process, the encoder groups the second reconstructed position points by coding grouping information to obtain multiple groups. The second reconstructed position points are obtained by the encoder decoding the geometric code stream of the target point cloud. Furthermore, the encoder performs attribute encoding on each group. In the coding scheme provided in the embodiment of the present application, the encoder determines the coding grouping information based on one or more of the following information related to the code stream that is defaulted on the codec side or encodes the target point cloud: point cloud characteristic information, bounding box information, coding parameter information, and target parameter information. It can be seen that the embodiment of the present application determines the grouping information based on the diversified features related to the target point cloud itself, which can improve the grouping diversity and is conducive to improving the coding performance of the point cloud. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] FIG1 is a schematic diagram of a network architecture of a point cloud encoding and decoding method provided in an embodiment of the present application;
[0023] FIG2 is a schematic block diagram of a point cloud video encoding and decoding system provided in an embodiment of the present application;
[0024] FIG3A is a schematic block diagram of a coding framework provided in an embodiment of the present application;
[0025] FIG3B is a schematic block diagram of a decoding framework provided in an embodiment of the present application;
[0026] FIG4 is a schematic diagram of a flow chart of a point cloud decoding method provided in an embodiment of the present application;
[0027] FIG5 is a schematic diagram of a bounding box of a point cloud provided in an embodiment of the present application;
[0028] FIG6 is a schematic diagram of a flow chart of a point cloud decoding method provided in another embodiment of the present application;
[0029] FIG7 is a schematic diagram of a flow chart of a point cloud encoding method provided in an embodiment of the present application;
[0030] FIG8 is a schematic diagram of a flow chart of a point cloud encoding method provided in another embodiment of the present application;
[0031] FIG9 is a schematic diagram of the structure of a decoder provided in an embodiment of the present application;
[0032] FIG10 is a schematic diagram of the structure of an encoder provided in an embodiment of the present application;
[0033] FIG11 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;
[0034] FIG12 is a schematic diagram of the structure of the encoding and decoding system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0035] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0036] It should be noted that the terms "first," "second," and the like in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In the embodiments of the present invention, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B solely based on A, but that B can also be determined based on A and / or other information. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or devices. In the description of this application, unless otherwise specified, "plurality" refers to two or more than two.
[0037] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0038] A point cloud is a set of irregularly distributed discrete points in space that express the spatial structure and surface properties of a three-dimensional object or scene.
[0039] Point Cloud Data is a specific recording form of point cloud data. Specifically, point cloud data can record the position information and attribute information of each point in the point cloud. Among them, the position information of the point can be the three-dimensional coordinate information of the point. The position information of the point can also be called the geometric information of the point. The attribute information of the point can include color information and / or reflectivity information, etc. Among them, the color information can be information on any color space. For example, the above-mentioned color information can be (RGB). Here, "R" represents red (Red, R), "G" represents green (Green, G), and "B" represents blue (Blue, B). For another example, the above-mentioned color information can be luminance and chrominance (YCbCr, YUV) information, wherein "Y" represents brightness (Luma), "Cb (U)" represents blue color difference, and "Cr (V)" represents red color difference.
[0040] For example, a point cloud collected using the laser measurement principle may include the point cloud data of the three-dimensional coordinate information and the laser reflection intensity (reflectance) of the point. For another example, a point cloud collected using the photogrammetry principle may include the point cloud data of the three-dimensional coordinate information and the color information of the point. For another example, a point cloud collected using the laser measurement and photogrammetry principles may include the point cloud data of the three-dimensional coordinate information, the laser reflection intensity (reflectance) of the point, and the color information of the point.
[0041] Among them, the acquisition methods of point cloud data may include but are not limited to at least one of the following: (1) generation by computer equipment. Computer equipment can generate point cloud data based on virtual three-dimensional objects and virtual three-dimensional scenes. (2) acquisition by 3D (3-Dimension) laser scanning. 3D laser scanning can obtain point cloud data of static real-world three-dimensional objects or three-dimensional scenes, and millions of point cloud data can be obtained per second; (3) acquisition by 3D photogrammetry. 3D photography equipment (i.e., a group of cameras or camera equipment with multiple lenses and sensors) is used to collect point cloud data of real-world visual scenes. 3D photography can obtain point cloud data of dynamic real-world three-dimensional objects or three-dimensional scenes. (4) acquisition of point cloud data of biological tissues and organs by medical equipment. In the medical field, point cloud data of biological tissues and organs can be obtained by medical equipment such as magnetic resonance imaging (MRI), computed tomography (CT), and electromagnetic positioning information.
[0042] Point clouds can be divided into static point clouds, dynamic point clouds and dynamically acquired point clouds according to the acquisition method. Among them, static point clouds refer to the situation when the object is stationary when the point cloud is acquired, and the device that acquires the point cloud is also stationary. Dynamic point clouds refer to the situation when the object is moving but the device that acquires the point cloud is stationary when the point cloud is acquired. Dynamic point clouds refer to the situation when the device that acquires the point cloud is moving when the point cloud is acquired.
[0043] In addition, point clouds are divided into two categories according to their uses: machine-perceived point clouds and human-eye-perceived point clouds. The machine-perceived point clouds can be used in autonomous navigation systems, real-time inspection systems, geographic information systems, visual sorting robots, emergency rescue robots and other scenarios. The human-eye-perceived point clouds can be used in point cloud application scenarios such as digital cultural heritage, free viewpoint broadcasting, three-dimensional immersive communication, and three-dimensional immersive interaction.
[0044] Point clouds can flexibly and conveniently represent the spatial structure and surface properties of three-dimensional objects or scenes. Furthermore, because point clouds are directly sampled from real objects, they provide a strong sense of realism while ensuring high accuracy. Consequently, they have a wide range of applications, including virtual reality gaming, computer-aided design, geographic information systems, automated navigation systems, digital cultural heritage, free-viewpoint broadcasting, 3D immersive telepresence, and 3D reconstruction of biological tissues and organs. Advances in point cloud data acquisition methods have enabled the acquisition of large amounts of point cloud data. However, as application demands grow, the processing of massive amounts of point cloud data is facing bottlenecks due to storage space and transmission bandwidth limitations.
[0045] For example, taking a point cloud video with a frame rate of 30 frames per second (FPS), each frame contains 600,000 points, each with xyz coordinate information (float) and RGB color information (uchar). Therefore, the data volume of a 15-second point cloud video is approximately 0.6 million × (4 bytes × 3 + 1 byte × 3) × 30 fps × 10 seconds = 4.05 GB. Clearly, point cloud video data is quite large. Therefore, to save storage space and reduce point cloud data transmission traffic and time, point cloud data compression is necessary.
[0046] Currently, point cloud coding frameworks that can compress point clouds can be the Geometry-based Point Cloud Compression (G-PCC) codec framework provided by the Moving Picture Experts Group (MPEG), or the Video-based Point Cloud Compression (V-PCC) codec framework, or the AVS-PCC codec framework provided by the Audio Video Coding Standard (AVS). Specifically, the G-PCC codec framework can be used to compress the first type of static point clouds and the third type of dynamically acquired point clouds, and it can be based on the Point Cloud Compression Test Platform (Test Model Compression 13, TMC13); the V-PCC codec framework can be used to compress the second type of dynamic point clouds, and it can be based on the Point Cloud Compression Test Platform (Test Model Compression 2, TMC2). Therefore, the G-PCC codec framework is also called the Point Cloud Codec TMC13, and the V-PCC codec framework is also called the Point Cloud Codec TMC2.
[0047] FIG1 is a schematic diagram of a network architecture of a point cloud encoding and decoding system provided in an embodiment of the present application, exemplarily illustrating the network architecture of a point cloud encoding and decoding system including a decoding method and an encoding method provided in an embodiment of the present application.
[0048] 1 , the network architecture includes a communication network 100 and one or more electronic devices. The electronic devices may be a server 102, a notebook 104, a desktop 106, a mobile phone 108, and a tablet computer 110, etc., as shown in FIG1 . Specifically, electronic devices can interact with each other through the communication network 100 for video transmission. During implementation, the electronic devices may be various types of devices with point cloud encoding and decoding functions. For example, the electronic devices may include smart phones, desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, vehicle-mounted computers, navigation systems, digital phones, video phones, televisions, sensor devices, and servers, etc., which are not limited in the embodiments of the present application. The decoder or encoder in the embodiment of the present application may be the above-mentioned electronic device. The electronic device in the embodiment of the present application has point cloud encoding and decoding functions, and generally includes a point cloud encoder (which may be referred to as an encoder in the embodiment of the present application) and a point cloud decoder (which may be referred to as a decoder in the embodiment of the present application).
[0049] It should be noted that Figure 1 is only an example of the network architecture of the point cloud encoding and decoding provided in the embodiment of the present application. The point cloud encoding and decoding network construction of the embodiment of the present application includes but is not limited to that shown in Figure 1.
[0050] FIG2 is a schematic block diagram of a point cloud encoding and decoding system provided in an embodiment of the present application.
[0051] Referring to Figure 2, the point cloud encoding and decoding system exemplarily includes two electronic devices: an electronic device 210 (e.g., an encoder) serving as the encoding end, and an electronic device 220 (e.g., a decoder) serving as the decoding end. Electronic device 210 is used to encode point cloud data to generate a bitstream and transmit the bitstream to electronic device 220. Electronic device 220 decodes the received bitstream to obtain decoded point cloud data.
[0052] In some embodiments, the code stream generated by the encoding by the electronic device 210 can be transmitted to the electronic device 220 via the channel 200. The channel 200 may include one or more media and / or devices capable of transmitting the encoded point cloud data from the electronic device 210 to the electronic device 220.
[0053] In one example, channel 200 includes one or more communication media that enable electronic device 210 to transmit encoded point cloud data directly to electronic device 220 in real time. In this example, electronic device 210 may modulate the encoded point cloud data according to a communication standard and transmit the modulated point cloud data to electronic device 220. The communication media includes wireless communication media, such as radio frequency spectrum. Optionally, the communication media may also include wired communication media, such as one or more physical transmission lines.
[0054] In another example, channel 200 includes a storage medium that can store the encoded point cloud data of electronic device 210. The storage medium includes various locally accessible data storage media, such as optical disks, DVDs, and flash memory. In this example, electronic device 220 can retrieve the encoded point cloud data from the storage medium.
[0055] In another example, channel 200 may include a storage server that can store the encoded point cloud data of electronic device 210. In this example, electronic device 220 can download the stored encoded point cloud data from the storage server. Alternatively, the storage server can store the encoded point cloud data and transmit the encoded point cloud data to electronic device 220, such as a web server (e.g., for a website), a File Transfer Protocol (FTP) server, etc.
[0056] In some embodiments, the electronic device 210 includes an encoder 214 and an output interface 216. The output interface 216 may include a modulator / demodulator (modem) and / or a transmitter.
[0057] In some embodiments, in addition to the encoder 214 and output interface 216, the electronic device 210 may also include a point cloud source 212. The point cloud source 212 may include at least one of a point cloud acquisition device (e.g., a scanner), a point cloud archive, a point cloud input interface, and a computer graphics system. The point cloud input interface is used to receive point cloud data from a point cloud content provider, and the computer graphics system is used to generate point cloud data. The encoder 214 encodes the point cloud data from the point cloud source 212 to generate a bitstream. The encoder 214 transmits the encoded point cloud data directly or indirectly to the electronic device 220 via the output interface 216. The encoded point cloud data may also be stored on a storage medium or storage server for subsequent reading by the electronic device 220.
[0058] In some embodiments, the electronic device 220 may further include a display device 222 in addition to the input interface 226 and the decoder 224. The input interface 226 includes a receiver and / or a modem. The input interface 226 may receive the encoded point cloud data through the channel 200. The decoder 224 is used to decode the encoded point cloud data to obtain decoded point cloud data, and transmit the decoded point cloud data to the display device 222. The display device 222 displays the decoded point cloud data. The display device 222 may be integrated with the electronic device 220 or external to the electronic device 220. The display device 222 may include a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.
[0059] It should be noted that Figure 2 is only an example of the point cloud encoding and decoding system provided in the embodiment of the present application. The point cloud encoding and decoding system in the embodiment of the present application includes but is not limited to that shown in Figure 2.
[0060] FIG3A is a schematic block diagram of a coding framework provided in an embodiment of the present application.
[0061] 3A , the encoding framework can obtain geometric information (also called position information or geometric position) and attribute information of the point cloud from the acquisition device. The encoding of the point cloud data includes position encoding and attribute encoding.
[0062] (1) About position coding
[0063] The above-mentioned position encoding process includes: preprocessing the original point cloud such as coordinate translation and coordinate quantization; among them, by performing coordinate transformation on the original point cloud, the world coordinates of the points in the point cloud can be transformed into relative coordinates; coordinate quantization can reduce the number of coordinates; after quantization, points originally at different positions may be assigned the same coordinates, and the same coordinates can be called duplicate points.
[0064] The position encoding process also includes: octree construction; octree encoding can be used to encode the quantized point position information. For example, the point cloud can be divided into an octree, so that the point positions can be mapped one-to-one to the octree positions. The point positions in the octree are counted and marked as 1.
[0065] In addition to the aforementioned octree encoding mode, geometry encoding also includes trisoup-based geometry encoding. Specifically, in trisoup-based encoding, the point cloud is divided into blocks of a certain size, and the intersection points of the point cloud surfaces at the edges of the blocks are located and triangles are constructed. By encoding the intersection points, geometric information is compressed.
[0066] The position coding process also includes quantization. The degree of quantization is usually determined by the quantization parameter (QP). A larger QP value means that coefficients with a larger value range will be quantized into the same output, which usually results in greater distortion and a lower bit rate. Conversely, a smaller QP value means that coefficients with a smaller value range will be quantized into the same output, which usually results in less distortion and a higher bit rate.
[0067] The above-mentioned position coding process also includes: entropy coding; the position information obtained by constructing the octree can be arithmetically coded using entropy coding, that is, the position information after constructing the octree is generated using arithmetic coding to generate a geometry bitstream (geometry bitstream), also known as a geometry code stream.
[0068] The position encoding process also includes: octree reconstruction; reconstructing the position information obtained from the octree construction to reconstruct the geometric position of each point in the point cloud, thereby obtaining the reconstructed geometric position of the point. Specifically, the reconstructed geometric position of the point is used in the attribute encoding process.
[0069] (2) About attribute coding
[0070] The attribute encoding process includes: space transformation; which can be used to transform the RGB color space of the point in the point cloud into YCbCr format or other formats.
[0071] The above attribute encoding process includes: attribute interpolation.
[0072] Attribute recoloring can be achieved through attribute interpolation. In the case of lossy coding, after the geometric information is encoded, the encoder needs to decode and reconstruct the geometric information, that is, restore the 3D coordinate information of each point in the point cloud. The attribute information corresponding to one or more neighboring points in the original point cloud is then searched for as the attribute information of the reconstructed point. This attribute interpolation can be used to transform the attribute information of points in the point cloud to minimize attribute distortion. Specifically, after attribute interpolation, the true value of the point attribute information can be obtained.
[0073] The attribute encoding process also includes attribute prediction. This stage can be used to predict attribute information for points in the point cloud to obtain predicted values for the point's attribute information. For example, based on the proximity relationship of geometric information or attribute information, one or more predicted attribute values for the reconstructed location are selected and weighted averaged to obtain the predicted attribute value for the current point. Based on the predicted values of the point's attribute information, a residual value for the point's attribute information is then obtained. The residual value for the point's attribute information can be the actual value of the point's attribute information minus the predicted value.
[0074] Exemplarily, attribute transform coding includes three modes, which can be used under different conditions.
[0075] (a) Predictive Transform Coding: This method selects subsets of points based on distance and divides the point cloud into multiple levels of detail (LoDs), achieving a point cloud representation from coarse to fine. Adjacent layers can achieve bottom-up prediction, where neighboring points in the coarse layer predict the attribute information of points introduced in the fine layer to obtain the corresponding residual signal. The points in the lowest layer are encoded as reference information.
[0076] (b) Lifting Transform: Based on the prediction of adjacent layers of LoD, a weight update strategy for neighborhood points is introduced to finally obtain the predicted attribute value of each point and the corresponding residual signal.
[0077] (c) Region Adaptive Hierarchical Transform (RAHT): The attribute information is transformed into the transform domain through RAHT, which is called transform coefficients.
[0078] The attribute encoding process also includes quantization and entropy coding. Specifically, the residual value of the attribute information of the point can be quantized. Furthermore, the quantized residual value can be entropy encoded using zero runlength coding to obtain an attribute bitstream, also known as an attribute codestream.
[0079] FIG3B is a schematic block diagram of a decoding framework provided in an embodiment of the present application.
[0080] Referring to Figure 3B , after the decoding framework obtains the point cloud code stream, it parses the code stream to obtain the position information and attribute information of the points in the point cloud. Point cloud decoding includes position decoding and attribute decoding.
[0081] The position decoding process includes: performing arithmetic decoding on the geometric code stream; constructing an octree and then merging it to reconstruct the point position information to obtain the reconstructed position information of the point; and performing coordinate transformation on the reconstructed position information of the point to obtain the point position information. The point position information can also be called the point's geometric information.
[0082] The attribute decoding process includes: obtaining the residual value of the attribute information of the point in the point cloud by parsing the attribute code stream; obtaining the residual value of the attribute information of the point after dequantization by dequantizing the residual value of the attribute information of the point; based on the reconstruction information of the point position information obtained in the position decoding process, selecting one of the three prediction modes to perform point cloud prediction to obtain the reconstructed value of the attribute information of the point; and performing color space deconversion on the reconstructed value of the attribute information of the point to obtain the decoded point cloud.
[0083] In the relevant point cloud encoding and decoding technology, referring to FIG3A, in the attribute prediction stage (such as the dotted box in FIG3A), the reconstructed position points are generally grouped according to a fixed formula to perform attribute encoding based on the reconstructed position points contained in each group. Referring to FIG3B, after the encoder completes the geometric decoding, the reconstructed position points are grouped according to a pre-set formula in the attribute reconstruction stage (such as the dotted box in FIG3B). It can be seen that in the attribute prediction stage of the encoding process and the attribute reconstruction stage of the decoding process, the grouping method provided by the relevant technology is single, which is not conducive to improving the encoding and decoding performance of the point cloud.
[0084] The solution provided in the embodiment of the present application can solve the problems existing in the related technology. The point cloud decoding method provided in the embodiment of the present application is first introduced in detail below.
[0085] Figure 4 is a schematic flow chart of a point cloud decoding method P400 provided in an embodiment of the present application. The execution entity of point cloud decoding method P400 may be a decoder, an electronic device performing the decoding process, or a processor. Referring to Figure 4 , point cloud decoding method P400 includes steps S410 to S430.
[0086] In S410, decoding grouping information is determined based on target features of the target point cloud, wherein the target features include one or more of the following information obtained by parsing the code stream of the target point cloud or defaulted by the codec: point cloud characteristic information, bounding box information, decoding parameter information, and target parameter information.
[0087] The target point cloud is any point cloud data, for example, point cloud data corresponding to a certain video frame.
[0088] It should be noted that before the encoder pre-processes the target point cloud, it will determine whether to divide the entire point cloud sequence into multiple point cloud slices based on the parameter configuration. In the case where the entire point cloud sequence is determined to be divided into multiple point cloud slices based on the above parameter configuration, each point cloud slice can be treated as a single independent point cloud for serial processing. Therefore, in some embodiments, the above target point cloud can be a point cloud at the overall level; in some embodiments, the above target point cloud can also be a point cloud slice obtained by dividing the point cloud sequence at the overall level, that is, the above target point cloud can be a point cloud slice level point cloud.
[0089] Similarly, the target feature can be a feature of the entire point cloud or a feature of the point cloud slice. For example, if the target feature is point cloud density, then if the target point cloud is an entire point cloud, the target feature refers to the density of the entire point cloud; if the target point cloud is a slice, the target feature refers to the density of the slices obtained by partitioning the entire point cloud.
[0090] In one embodiment of the present application, after the decoder receives the code stream of the target point cloud, it can obtain the above-mentioned target features by parsing the code stream, for example, obtaining one or more of the following information: point cloud characteristic information, bounding box information, decoding parameter information and target parameter information.
[0091] In another embodiment of the present application, the above-mentioned target features may also be information that is a default on the codec side. In addition, the decoder may also look up a preset table to determine the above-mentioned target features, wherein the preset table is generated before or during encoding of the target point cloud. The default parameters on the codec side refer to information that is mutually defaulted between the encoding side and the corresponding decoding side, and are information that does not need to be encoded or decoded during the point cloud encoding and decoding process. The above-mentioned target parameter information may include a constant value (or fixed value, default value), and may also include at least one preset identifier, wherein each preset identifier corresponds to a fixed value.
[0092] In an exemplary embodiment, the decoding parameter information refers to feature parameters used when decoding the target point cloud, and may include one or more of the following: the maximum transform order of the attribute prediction transform, the attribute output bit depth, and the attribute quantization parameter obtained by the decoder when decoding the attribute bit stream; and the geometry output bit depth and geometry quantization parameter obtained by the decoder when decoding the geometry bit stream. The decoding parameter information may also include the transformation method, prediction method, and sorting method.
[0093] In an exemplary embodiment, the point cloud characteristic information may reflect the characteristics of the target point cloud, and may include, for example, one or more of the following information: point cloud type, point cloud density, point cloud space occupancy, resolution, and point cloud point count.
[0094] As mentioned above, the above-mentioned point cloud characteristic information provided in the embodiments of the present application may specifically refer to the point cloud density, point cloud space occupancy, resolution and point cloud point count at the overall point cloud level, and the above-mentioned point cloud characteristic information may specifically refer to the point cloud density, point cloud space occupancy, resolution and point cloud point count at the point cloud slice level.
[0095] In an exemplary embodiment, the bounding box information is specifically the features of the bounding box of the target point cloud.
[0096] It is understood that the encoder performs coordinate transformation on the geometric information of the target point cloud during the preprocessing phase, specifically to enable the point cloud to be completely contained within a bounding box. Referring to Figure 5, the bounding box represents the smallest cuboid that contains all points in the input point cloud. The origin and size of the bounding box are determined as follows:
[0097] The floating point coordinates of the k-th point of the input point cloud are expressed as (x k ,y k , z k ), k = 0, ..., K-1, K is the number of points in the point cloud, coordinate information x min ,y min , z min , x max ,y max , z max They are represented as follows: min =min(x0, x1, ..., x K-1 ) y min =min(y0, y1, ..., y K-1 ) z min =min(z0, z1, ..., z K-1 ) x max =max(x0, x1, ..., x K-1 ) ymax =max(y0, y1, ..., y K-1 ) z max =max(z0, z1, ..., z K-1 )
[0098] The above function min(s0, s1, ..., s K-1 ) means taking the current input (s0, s1, ..., s K-1 ), the minimum value among max(s0, s1, ..., s K-1 ) means taking the current input (s0, s1, ..., s K-1 ) is the maximum value among .
[0099] Furthermore, the origin of the bounding box (x origin ,y origin , z origin ) is as follows: origin =int(floor(x min )) y origin =int(floor(y min )) z origin =int(floor(z min ))
[0100] The size of the bounding box in the x, y, and z directions can be calculated as follows: Bounding_Box_Size_x = int(x max -x origin )+1 Bounding_Box_Size_y=int(y max -y origin )+1 Bounding_Box_Size_z=int(z max -z origin )+1
[0101] In this case, int(s) represents the rounding function, and the floor(s) function returns the maximum integer value less than or equal to s.
[0102] In an embodiment of the present application, the bounding box information of the target point cloud may include one or more of the following information: the lengths and directions of the three sides of the bounding box, and length statistics about the three sides of the bounding box. The length statistics about the three sides of the bounding box include one or more of the following information: the mean length of the three sides, the median length, the mode length, the maximum length of the three sides, the minimum length of the three sides, a linear combination of the lengths of at least two sides, a volume-related calculation of the bounding box, and a surface-area-related calculation of the bounding box, etc.
[0103] For example, the linear combination of the lengths of the at least two sides may include: the sum of the side lengths, the difference between the side lengths, a combination of the corresponding logarithms of the side lengths, etc.
[0104] For example, the surface area-related calculations of the above-mentioned bounding box may include: the surface area corresponding to the shortest side length, the surface area corresponding to the two longer side lengths, and combinations between different types of surface areas as mentioned above, as well as combinations between the logarithms of different types of surface areas as mentioned above.
[0105] It can be understood that any statistical value determined based on the bounding box information can be considered as the above-mentioned target feature.
[0106] After the above embodiment introduces the embodiment of obtaining the target feature, the following describes how the decoder determines the decoding group information according to the above target feature.
[0107] The embodiment of the present application does not limit the specific manner in which the decoder determines the decoding grouping information based on the target features of the target point cloud.
[0108] In some embodiments, the decoder may determine at least one of the point cloud characteristic information, the bounding box information, the decoding parameter information, and the target parameter information of the target point cloud as the decoding grouping information.
[0109] In some embodiments, the decoder may determine the decoding group information through the following steps S410-1 and S410-2A or S410-2B:
[0110] S410 - 1 : Mapping at least one of the above target features through an objective function to obtain at least one mapping value.
[0111] For example, the decoder maps the target feature x through the target function f() to obtain the mapping value f(x) corresponding to the target feature.
[0112] In the embodiment of the present application, the objective function f() can be a linear function, a quadratic function, a direct proportional function, an inverse proportional function, a trigonometric function, an exponential function, a logarithmic function, etc. For example, the objective function is a logarithmic function f(x)=log2x.
[0113] In the embodiment of the present application, when the target feature is expressed as a parameter value or a characteristic value, the decoder can map the value through the target function. For example, when the target feature includes two or more parameter values, the decoder can map different parameter values through different types of target functions. For example, when the target feature includes parameter values x1 and x2, the corresponding mapping values can be: lg(x1) and e x2 .
[0114] Exemplarily, the above mapping processing can also be a linear combination related to the number of points; for example, Log2(a*num_points+ / -b), where a and b are constants; or, the above mapping processing can also be a linear combination related to the logarithmic value of the number of points; for example, ba*Log2(num_points), or b+a*Log2(num_points) where a and b are constants.
[0115] It is understandable that the above-mentioned mapping value may also be the parameter value of the target feature itself. For example, if the parameter value of the target feature is x3, after the mapping operation of this embodiment, the mapping value is still x3.
[0116] In some embodiments, in the embodiments of the present application, after the decoder performs mapping processing on the target feature according to the above-mentioned objective function, the determined mapping value is an integer. Specifically, when the above-mentioned objective function includes logarithm processing, rounding calculation is performed after the logarithm processing; when the above-mentioned objective function includes division processing, rounding calculation is performed after the division processing. It should be noted that when decimals appear in the mapping process based on other objective functions, or when the target feature itself remains after mapping with the objective function and the target feature contains decimals, rounding operation is also required.
[0117] In an exemplary embodiment, the above rounding calculation can be: rounding up, rounding down, or rounding to the nearest integer. Of course, other forms of rounding calculation methods, such as rounding to the nearest integer, are also possible, and the embodiments of this application are not limited to this. For example, if the target feature is the number of points in the target point cloud and its value is 10000, and the objective function is a logarithmic function with base 2, then the mapping value determined by the above objective function is: round(log2(10000)).
[0118] After determining the above mapping value, the following S410 - 2A or S410 - 2B may be executed to determine the decoding group information.
[0119] In S410 - 2A, decoding group information is determined according to at least one combination of mapping values or a combination of mapping values and a constant value.
[0120] The above combination can be a linear combination or a nonlinear combination.
[0121] In an exemplary embodiment, assuming that the at least one mapping value can be expressed as f(x) and g(y), the exemplary decoded grouping information can be expressed as: A×f(x)+B×g(y), A×f(x)-B×g(y), A×f(x) 2 +A×g(y) 2 、A×f(x) 2 -A×g(y)2 、A×f(x) 3 +B×g(y), etc. Where f() and g() can represent two identical or different objective functions, and A and B are constants and positive numbers.
[0122] In S410 - 2B, decoding group information is determined according to a combination of the at least one mapping value and the unmapped target feature.
[0123] As mentioned above, the target features that have not been mapped can also be considered to be the same as before after being mapped by the target function. If the mapping value corresponding to x is still x, and at least one mapping value obtained by mapping at least one target feature by the target function can be expressed as g(y), then the exemplary decoding group information can be expressed as: C×x+B×g(y), C×xB×g(y), C×x 2 +A×g(y) 2 、C×x 2 -A×g(y) 2 、C×x 3 +B×g(y), etc. Where f() and g() can represent two identical or different objective functions, and A, B, and C are constants and positive numbers.
[0124] In an exemplary embodiment, the decoded grouping information may include an initial grouping step size, an intermediate quantity used to calculate the initial grouping step size, and an upper limit and / or lower limit of the initial grouping step size. Therefore, as a specific implementation of S410, any one or more of the following steps S410-A to S410-C may be included:
[0125] S410-A: Determine the initial grouping step size in the decoded grouping information.
[0126] S410-B: Determine an intermediate quantity for calculating the initial grouping step length, and determine the initial grouping step length according to the intermediate quantity to obtain decoded grouping information.
[0127] S410-C: Determine an upper limit and / or lower limit of the initial grouping step length in the decoded grouping information, and determine an upper limit and / or lower limit of the intermediate amount of the decoded grouping information.
[0128] The initial grouping step size is the step size used by the decoder to initially group the first reconstructed points (the reconstructed points obtained by decoding the geometry stream at the decoder) after reordering the reconstructed points. In other words, it is the grouping step size used when no updates are made to the grouping step size. Alternatively, it can be said to be the step size used by the decoder to group the reconstructed points for the first time during the attribute reconstruction process of the target point cloud.
[0129] The following describes the specific implementation of S410-A, which is to determine the initial grouping step length L. b Specific implementation method:
[0130] Specific implementation method 1: L1=Log s (gsh_bounding_box_size_x)+Log s (gsh_bounding_box_size_y)+ Log s (gsh_bounding_box_size_z); L2=Log s (slice_num_points); L b =L1–L2, and L b ≥1;
[0131] Among them, Log s () is ceilLog s (), floorLog s () or roundLog s ();
[0132] L b represents the initial grouping step size, gsh_bounding_box_size is the bounding box size at the point cloud slice level, slice_num_points is the number of point cloud points at the point cloud slice level, s represents a constant greater than 1, and L1 and L2 represent intermediate values.
[0133] Specific implementation method 2: L1=Log s (bounding_box_size_x)+Log s (bounding_box_size_y)+ Log s (bounding_box_size_z); L b =max(3,3×round(L1 / 3));
[0134] Among them, Log s () is ceilLog s (), floorLog s () or roundLog s ();
[0135] L bRepresents the initial grouping step size, bounding_box_size_x, bounding_box_size_y, and bounding_box_size_z respectively represent the lengths of the three sides of the bounding box of the target point cloud, s represents a constant greater than 1, and L1 represents an intermediate value.
[0136] Specific implementation method three: L1 = 2 × Log s (gsh_bounding_box_size_x); L2=Log s (slice_num_points); L b =L1–L2+shift; and L b ≥1; can also be expressed as: L b =max(1,L1-L2+shift); or, L b =max(1,L1-L2)+shift;
[0137] Among them, Log s () is ceilLog s (), floorLog s () or roundLog s ();
[0138] L b Indicates the initial grouping step size, gsh_bounding_box_size_x indicates the maximum length of the three sides of the point cloud slice, slice_num_points indicates the number of points in the point cloud slice, shift indicates the offset value, s indicates a constant greater than 1, and L1 and L2 indicate intermediate values.
[0139] Specific implementation method 4: L1=Log s (bounding_box_size_x)+Log s (bounding_box_size_y)+ Log s (bounding_box_size_z); L b =max(3,3×round(L1 / 3));
[0140] Among them, the Log s () is ceilLog s (), floorLog s () or roundLog s ();
[0141] L bRepresents the initial grouping step size, bounding_box_size_x, bounding_box_size_y, and bounding_box_size_z respectively represent the lengths of the three sides of the bounding box of the target point cloud, s represents a constant greater than 1, and L1 represents an intermediate value.
[0142] Specific implementation method 5: In this embodiment, L is determined based on the surface area information of the bounding box. b ; L1=gsh_bounding_box_size_x_log2+gsh_bounding_box_size_z_log2; or, L1=ceilLog2(gsh_bounding_box_size_x*gsh_bounding_box_size_z)); L2=ceillog2(num_points);
[0143] Where x and z represent the side length of the smallest area among the six sides of the bounding box; or, x and z represent the side length of the largest area among the six sides of the bounding box; or, one of x and z represents the length of the smallest length among the three sides of the bounding box; or, one of x and z represents the length of the largest length among the three sides of the bounding box; L b =max(1,max(3,3×round(L1–L2 / 3))+shift).
[0144] Specific embodiment 6: In this embodiment, L is determined based on the entire or partial surface area information of the bounding box. b ; L1=ceilLog2(gsh_bounding_box_size_x*gsh_boundingbox_size_z+ gsh_bounding_box_size_y*gsh_boundingbox_size_z); L2=ceilog2(num_points); L b =max(3,3×round(L1–L2 / 3))+shift.
[0145] From the above examples, we can see that the decoder can determine the initial grouping step length L in the decoded grouping information based on the bounding box information of the target point cloud. b Specifically, the decoder determines the intermediate quantities L1 and L2 used to calculate the initial grouping step length based on the bounding box information of the target point cloud. Then, the initial grouping step length L in the decoded grouping information is calculated based on the intermediate quantities L1 and L2. b .
[0146] Specific implementation method seven:
[0147] L1 is obtained by parsing the bitstream; it can be understood that the decoder parses the bitstream to obtain a fixed value for the initial grouping step size and uses the fixed value as L1; L b =L1+shift.
[0148] Specific implementation method eight: L1=32; L2=log2(num_points); L b =max(3,3×round(L1–L2 / 3))+shift.
[0149] Specific implementation method 9: L1 = min (A, ceillog2 (xyz)); L2 = log2 (num_points); L b =max(3,3×round(L1–L2 / 3))+shift;
[0150] Wherein, A represents the fixed value in the above target parameter information, and xyz represents the volume of the bounding box.
[0151] It should be noted that, in any specific implementation, after determining the intermediate values L1 and L2, the initial grouping step length L can be determined according to the following method: b : L b =max(3,3×round(L1–L2 / 3))+shift; or, L b =max(3,3×round(L1–L2 / 3)+shift); or L b =max(1,max(3,3×round(L1–L2 / 3))+shift).
[0152] The calculation method of shift in the above specific implementation scheme can be any of the following methods:
[0153] 1) The decoder searches for shift based on a preset lookup table T. For example, the index value of the lookup table is calculated using the geometric quantization parameter, the attribute quantization parameter, and the quantization offset parameter, and the shift value is determined in the lookup table T according to the determined index value.
[0154] 2) The decoder determines the above offset value based on the attribute features obtained by parsing the target point cloud bitstream, such as the header information parameter attrQuantParam and possible QPOffset;
[0155] 3) The decoder can also directly obtain a fixed value of shift and determine the fixed value as the value of shift. The fixed value directly obtained above can be a default value of the codec end or obtained by the decoder through decoding.
[0156] 4) The decoder may also determine the fixed value corresponding to the preset identifier as an offset value (shift), where different preset identifiers correspond to different fixed values. For example, when the preset identifier is C1, the offset value is c1; when the preset identifier is C2, the offset value is c2, where both c1 and c2 are positive numbers. The preset identifier may be a default value on the codec, obtained by the decoder through decoding, or determined using a lookup table.
[0157] Specific implementation method ten:
[0158] In this embodiment, the value of the initial grouping step in the above-mentioned decoding grouping information is an integer multiple of the target fixed value. For example, in the above-mentioned "Specific Implementation Method Eight", 'L1=32', where L1 is an integer multiple of 1, the "target fixed value" in this embodiment is 1. In this embodiment of the present application, the above-mentioned target fixed value can be determined by the decoder based on the target features of the target point cloud; or, the above-mentioned target fixed value can be determined by the decoder based on a preset identifier, and different preset identifiers correspond to different fixed values, so that the above-mentioned target fixed value can be determined based on the preset identifier; or, the above-mentioned target fixed value can be determined by the decoder as the first default value in the target parameter information.
[0159] In some embodiments, as a specific implementation of S410, the following steps S1 to S2 may be included:
[0160] S1: Determine the target fixed value based on at least one of the point cloud characteristic information, bounding box information and decoding parameter information in the target feature; or, determine the target fixed value based on a preset identifier in the target parameter information, wherein different preset identifiers correspond to different fixed values; or, determine the first default value in the target parameter information as the target fixed value.
[0161] S2: Determine the integer multiple of the above target fixed value as the value of the initial grouping step.
[0162] In some embodiments, the target feature used in S1 may be a fixed parameter set by default on the codec side, or may be one or more of the above-mentioned decoding parameter information, point cloud feature information, and bounding box information.
[0163] Exemplarily, when the above-mentioned target feature adopts the point cloud type in the point cloud characteristic information, when the point cloud type is a dense type or a human vision point cloud, the above-mentioned target fixed value is determined to be m; or, when the above-mentioned point cloud type is a sparse type or a machine vision point cloud, the above-mentioned target fixed value is determined to be n, where m and n are positive integers.
[0164] For example, the value of m is 3, and the value of n is 1. For example, if the target feature is a point cloud type and the point cloud type is a dense type, the value of the initial grouping step size can be an integer multiple of 3. Of course, the embodiment of the present application does not limit the values of m and n.
[0165] Exemplarily, when the target feature adopts bounding box information, when the correlation of the three sides of the bounding box meets the preset conditions, the target fixed value is determined to be p; or, when the correlation of the three sides of the bounding box does not meet the preset conditions, the target fixed value is determined to be q, where p and q are positive integers. The preset condition may specifically be whether the correlation is greater than a preset value (such as 80%). Specifically, the correlation between the lengths of the three sides can be determined by the length variance or standard deviation of the three sides. For example, if boundingbox_size_z is much smaller or much larger than boundingbox_size_x or boundingbox_size_y, it means that the length similarity of the three sides of the bounding box is small.
[0166] For example, the value of p is 3, and the value of q is 1. For example, if the variance between boundingbox_size_z, boundingbox_size_x, and boundingbox_size_y is less than a preset value, indicating that the lengths of the three sides of the bounding box are highly similar, then the value of the initial grouping step can be an integer multiple of 3. Of course, the embodiments of the present application do not limit the values of p and q.
[0167] Exemplarily, the above-mentioned target fixed value may also be determined according to the above-mentioned decoding parameter information, for example, according to a geometric quantization parameter and / or an attribute quantization parameter.
[0168] In other embodiments, the preset identifier in S1 may be a default identifier on the codec side, may be obtained by the decoder through decoding, or may be determined based on a preset lookup table. For example, if the decoder determines that the identifier is Q1, the fixed value corresponding to the identifier Q1 is determined as the target fixed value. Alternatively, the decoder may directly obtain the first default value and determine the first default value as the target fixed value. The first default value may be a default identifier on the codec side or obtained by the decoder through decoding.
[0169] It is understandable that the specific implementation for determining the initial grouping step size is not limited to the above embodiment. For example, other combinations between target features and between target features and constants may also be used, and this application does not limit this.
[0170] The following describes a specific implementation of S410-B, which is to determine the step length L for calculating the initial grouping.b Specific implementation of the intermediate amount:
[0171] In related technologies, when the attribute is reflectivity, during the attribute reconstruction process, the initial grouping step length is expressed as follows:
[0172] Among them, it is used to determine the initial grouping step length L base The intermediate amount is the above-mentioned offset value shift, which can also be the maximum number of bits maxBits and the minimum number of bits minBits.
[0173] The embodiment of the present application can be similar to the embodiment corresponding to S410-A. The decoder can directly determine the initial grouping step length L according to the above target characteristics. b The initial grouping step length L in the related technology can also be determined based on the above target characteristics. base The decoder may also determine a second default value in the target parameter information as the intermediate value, where the second default value may be a default value set by the codec, obtained by the decoder through decoding, or determined based on a preset lookup table. The decoder may also determine a fixed value corresponding to a preset identifier in the target parameter information as the intermediate value, where different preset identifiers correspond to different fixed values.
[0174] For example, the maximum number of bits maxbits = Log s (V(bounding_box)); minimum number of bits minbits = ceilLogs(A); where maxbits represents the maximum number of bits, V(bounding_box) represents the volume of the bounding box, minbits represents the minimum number of bits, and A represents a fixed parameter.
[0175] For example, the maximum number of bits maxbits = Log s (num_points); the minimum number of bits minbits takes different parameter values according to the type of target point cloud; among them, num_points represents the number of points in the target point cloud.
[0176] Exemplarily, the offset value shift may be determined by parsing the attribute features obtained from the code stream of the target point cloud, such as the header information parameter attrQuantParam and possible QPOffset.
[0177] Exemplarily, the offset value may also be determined by looking up a fixed parameter X1 determined in a preset table, wherein the preset table is generated before or when the encoder encodes the target point cloud.
[0178] Exemplarily, the offset value shift may be determined by parsing the fixed parameter X2 obtained from the code stream of the target point cloud, for example, shift=ceil(log2(X2)).
[0179] It can be seen that the embodiment of the present application can determine various combinations between the above target features and their mapping values through S420-2A or S420-2B, thereby providing multiple ways of determining the above intermediate quantities. The embodiment of the present application does not limit the specific combination method.
[0180] As a specific implementation of S410-C, determining the intermediate quantity or initial grouping step length L b The specific implementation methods of the upper limit and lower limit include but are not limited to the following:
[0181] (1) The decoder determines the obtained third default value as the upper limit value of the intermediate quantity (such as shift, maxbits, minbits), and determines the obtained fourth default value as the lower limit value of the intermediate quantity (such as shift, maxbits, minbits); for example, the decoder uses the fixed value a as the upper limit value of the intermediate quantity shift, and the fixed value b as the upper limit value of shift. It can be understood that a and b are positive numbers, and a is greater than b; wherein, exemplarily, the above-mentioned fixed value a and the above-mentioned fixed value b can be the default of the codec end, or can be directly obtained by the decoder decoding the code stream; in another exemplary embodiment, the decoder can determine the above-mentioned fixed value a according to the first preset identifier, and determine the above-mentioned fixed value b according to the second preset identifier, wherein the first preset identifier indicates that the upper limit value of shift is set to the corresponding fixed value a, and the second preset identifier indicates that the lower limit value of shift is set to the corresponding fixed value b. The above-mentioned first preset identifier and the second preset identifier can be the default of the codec end, or can be directly obtained by the decoder decoding the code stream.
[0182] Similarly, the decoder can also use the fixed value a' as the initial grouping step length L b The upper limit value of the fixed value b' is used as the initial grouping step length L b It is understandable that a' and b' are positive numbers, and a' is greater than b'; where, for example, the fixed value a' and the fixed value b' can be the default value of the codec end, or can be directly obtained by the decoder decoding the code stream; in another exemplary embodiment, the decoder can determine the fixed value a' according to the first preset identifier, and determine the fixed value b' according to the second preset identifier, where the first preset identifier indicates that L b The upper limit value of L is set to the corresponding fixed value a', and the second preset mark indicates that L bThe lower limit value of is set to the corresponding fixed value b'. The first preset identifier and the second preset identifier can be the default ones of the codec end, or can be directly obtained by the decoder decoding the code stream.
[0183] (2) The decoder can determine the initial grouping step size L based on the point cloud type. b Or the upper and lower limits of the intermediate amount; for example, compared with the sparse type point cloud, for the dense type point cloud, the initial grouping step length L b The upper and lower limits of are both small. For example, for dense point clouds, the initial grouping step length L b The upper limit value is z1 and the lower limit value is z2; for dense point clouds, the initial grouping step size L b The upper limit is z3 and the lower limit is z4; then the value of z1 is less than z3, and the value of z2 is less than z4;
[0184] (3) The decoder can determine the initial grouping step length L based on the decoding parameter information b Or the upper and lower limits of the intermediate quantity; such as determining L according to the geometric output bit depth, attribute output bit depth, geometric quantization parameter and attribute quantization parameter b The upper limit and lower limit of L can also be determined according to the transformation method, prediction method and sorting method. b The upper and lower limits of the transform mode are as follows: if the transform mode is predictive transform coding, the initial grouping step length L is determined. b The upper limit is s1, the lower limit is s2, s1 is greater than s2; when the transformation mode is lifting transformation coding, determine the initial grouping step size L b The upper limit value is s3, the lower limit value is s4, and s3 is greater than s4; wherein, s1 and s3 can be different positive numbers, and s2 and s4 can be different positive numbers.
[0185] (4) The decoder can determine the initial grouping step length L based on the parameter information of the bounding box, such as the length of the bounding box side, the correlation of the side length, etc. b Or the upper and lower limits of an intermediate quantity.
[0186] It is understandable that determining the initial grouping step length L b The specific implementation of the upper limit value and the lower limit value is not limited to the above content, and can be other forms determined according to the above target characteristics.
[0187] Continuing with reference to FIG4 , in S420 , the first reconstructed position points are grouped using the above-mentioned decoding grouping information to obtain a plurality of groups, wherein the above-mentioned first reconstructed position points are obtained by decoding the geometric code stream of the above-mentioned target point cloud based on the decoder; and, in S430 , attribute decoding is performed on each group.
[0188] In an exemplary embodiment, the decoder determines the initial grouping step size in the decoding grouping information according to the above-described embodiment. The decoder then groups the reordered reconstructed position points according to the initial grouping step size. The specific implementation of grouping is described in detail in the embodiment corresponding to FIG6 . Furthermore, the decoder performs attribute decoding based on the grouped reconstructed position points, thereby completing the decoding process of the target point cloud stream and obtaining a reconstructed point cloud of the target point cloud.
[0189] In the P400 point cloud decoding solution provided in the embodiment of the present application, the decoder determines the decoding grouping information based on one or more of the following information obtained by parsing the code stream of the target point cloud, which is the default at the codec end or by parsing the code stream of the target point cloud: point cloud feature information, bounding box information, decoding parameter information, and target parameter information. During the attribute decoding process, the decoder groups the first reconstructed position points using the above-mentioned grouping information to obtain multiple groups. Among them, the above-mentioned first reconstructed position points are obtained by the decoder decoding the geometric code stream of the target point cloud. Further, the decoder performs attribute decoding on each group. In the decoding solution provided in the embodiment of the present application, the decoder determines the decoding grouping information based on the above-mentioned target features. The decoding grouping information includes the initial grouping step, the intermediate amount used to determine the initial grouping step, the upper and lower limits of the above-mentioned initial grouping step, etc. It can be seen that the embodiment of the present application determines the grouping information based on the diversified features related to the target point cloud itself. Compared with the method for determining the initial grouping step provided by the relevant technology, the embodiment of the present application can improve the grouping diversity, which is conducive to improving the decoding performance of the point cloud.
[0190] FIG6 is a flow chart illustrating a point cloud decoding method P500 according to another embodiment of the present application. The execution entity of the point cloud decoding method P500 may be a decoder, or an electronic device that performs the decoding process. Referring to FIG5 , the point cloud decoding method P500 includes steps S52-S58.
[0191] In S52 , the decoder decodes the geometric code stream of the target point cloud to obtain a first reconstructed position point.
[0192] For example, the reconstruction position point obtained by decoding the geometry code stream at the decoder is recorded as the first reconstruction position point. In addition, the reconstruction position point obtained by the encoder during the attribute prediction stage is recorded as the second reconstruction position point.
[0193] In the embodiment of the present application, the order of reconstructing the geometric coordinates of the point cloud includes but is not limited to the following methods:
[0194] a) After the decoder completes decoding and reconstruction of the slice geometry, it reconstructs the overall geometry of the current slice. The final reconstructed geometry is rec_xyz = xyz + gsh_bounding_box_offset. gsh_bounding_box_offset is the xyz coordinate of the origin of the slice bounding box.
[0195] b) After decoding and reconstructing the slice attribute information, the decoder reconstructs the overall geometric information of the current slice.
[0196] In S54 , the decoder reorders the first reconstruction position points.
[0197] The embodiment of the present application does not limit the manner in which the first reconstructed position points are reordered. For example, the decoder may reorder according to the order of the point cloud input; or the decoder may reorder according to the order of the point cloud space curve (including the Morton order and the Hilbert order); or the decoder may reorder according to input parameter information (such as acquisition information, LiDAR information, etc.), etc.
[0198] In S56 , the decoder determines the decoding grouping information and groups the first reconstruction position points according to the decoding grouping information.
[0199] The specific implementation of the decoder determining the decoding group information has been described in detail in the embodiment corresponding to P400 and will not be repeated here.
[0200] In the embodiment of the present application, the first reconstructed position points are reordered in the Hilbert order, and the reordered points correspond to the Hilbert code. b In the case of , the Hilbert codes corresponding to the first reconstruction position point can be grouped in the following way:
[0201] (1) The decoder determines the initial grouping step size L b , initially grouping the first reconstruction position points;
[0202] Shift the Hilbert code right by the initial grouping step size L b Then the same points are divided into the same macroblock.
[0203] (2) The decoder determines whether the Hilbert codes within the same macroblock need to be subdivided into groups;
[0204] It should be noted that for each point in a macroblock, the decoder can determine whether there are duplicate points in the current macroblock. If there are duplicate points in the current macroblock, the group is truncated at the duplicate point and the duplicate points are grouped independently, that is, the duplicate points are grouped with a point count of 1. For example, if the current macroblock contains 4 points and the 4th point is a duplicate point, the first 3 points can be grouped as one, and the 4th duplicate point can be grouped as a separate group.
[0205] In the related art, during the color reconstruction process, if the number of points in the current block is less than or equal to the maximum transformation order colorMaxTransNum, the points in the macroblock become a group; if the number of points in the current block is greater than colorMaxTransNum, the points in the block are subdivided into groups, and the rule for the subdivision group is: obtain the current right shift bit L, then the right shift bit L1 of the current subdivision group takes the value of L-1, and then the same points after the Hilbert code is right shifted by L1 bits are grouped into a new subdivision group; and continue to judge whether the number of points in the subdivision group is greater than colorMaxTransNum, until the number of points in each subdivision group is less than or equal to colorMaxTransNum.
[0206] In the related art, during the reflectivity reconstruction process, if the number of points in the current macroblock is greater than the maximum transform order reflMaxTransNum, the decoder sequentially takes the points of the maximum transform order reflMaxTransNum as a subdivision group until the number of points in each group in the block is less than or equal to reflMaxTransNum.
[0207] The embodiment of the present application may adopt the method provided by the related technology to determine whether the points within the macroblock need to be subdivided into groups.
[0208] (3) The decoder determines whether the packet step size needs to be updated;
[0209] In the related art, during the color reconstruction process, the sum of the number of points in the first three consecutive groups is counted, and the sum is shifted right by three places (i.e., divided by 8) and recorded as B. If the value of B is less than 2, the updated step length L = L+1; if the value of B is greater than 8, the updated step length L = L-1; if B is greater than or equal to 2 and less than or equal to 8, the step length L is not updated.
[0210] In the related art, during the reflectivity reconstruction process, if the maximum transform order reflMaxTransNum is greater than 2 and the offset value shift in the initial grouping step is greater than 0, the decoder updates the grouping step; otherwise, the current grouping step is maintained. Specifically, the decoder counts the total number of points in the previous N groups and calculates the average number of points in N (e.g., N=8) groups. If the average number of points is less than 2, the updated grouping step L is L+1; if the average number of points is greater than the maximum transform order reflMaxTransNum, the updated grouping step L is L-1; if the average number of points is greater than or equal to 2 and less than or equal to the maximum transform order reflMaxTransNum, the grouping step L is also maintained unchanged.
[0211] It can be seen that in the related art, the number of points in the previous multiple groups is considered when determining whether to update the grouping step size. It should be noted that in the related art, a point group formed by repeated points can also be used as the above-mentioned previous group and used to determine whether to update the grouping step size. That is, the number of points in the group formed by repeated points is considered in the related art.
[0212] In the embodiment of the present application, the method of subdividing the points in the current macroblock into groups includes but is not limited to the following:
[0213] a) The decoder counts the points corresponding to the first N groups directly grouped by the grouping step information, where the direct grouping is a non-subdivided group, that is, a group determined by the grouping step. Specifically, the decoder determines whether to update the grouping step based on the number of points.
[0214] b) The decoder counts the number of points corresponding to the first N groups that do not contain repeated points. It should be noted that this method is different from the related art method of counting the number of repeated points when determining whether to update the grouping step size. In the embodiment of the present application, the method of counting the number of points in the group that does not contain repeated points determines whether to update the grouping step size.
[0215] In the embodiment of the present application, the updating method of the grouping step size is more flexible, which is conducive to improving decoding performance.
[0216] Continuing to refer to FIG. 6 , in S58 , the decoder performs intra-group attribute decoding.
[0217] In one embodiment, after all first reconstruction position points are grouped, attribute encoding can be performed on each grouped. In another embodiment, after each macroblock is subdivided into groups, attribute encoding can be performed on the completed groupings of the current macroblock while the next macroblock is processed, thereby improving overall encoding efficiency.
[0218] In the point cloud decoding method P500 provided in the embodiment of the present application, the decoder in method P400 determines decoding grouping information based on one or more of the following information obtained by default at the codec end or by parsing the target point cloud code stream: point cloud feature information, bounding box information, decoding parameter information, and target parameter information. Furthermore, compared to related technologies, the point cloud decoding method P500 provided in this embodiment not only improves the method for determining the initial grouping step size and its intermediate quantities, thereby increasing grouping diversity, but also expands the method for determining whether to update the grouping step size, provides an optimized processing solution for the attribute decoding process, and expands the geometric reconstruction method. All of these improvements contribute to improving point cloud decoding performance.
[0219] The point cloud encoding method provided by the embodiment of the present application is described in detail above through some embodiments. The point cloud encoding method provided by the embodiment of the present application is described in detail below through some embodiments.
[0220] FIG7 is a flow chart illustrating a point cloud encoding method P600 according to an embodiment of the present application. The execution entity of the point cloud encoding method P600 may be an encoder, or an electronic device that performs the encoding process (e.g., an encoder). Referring to FIG7 , the point cloud encoding method P600 includes steps S610 to S630.
[0221] In S610, the encoding grouping information is determined based on the target features of the target point cloud, wherein the target features include one or more of the following information related to the code stream of the encoded target point cloud or the default information of the encoding end: point cloud characteristic information, bounding box information, encoding parameter information and target parameter information.
[0222] The target point cloud is any point cloud data, for example, point cloud data corresponding to a certain video frame.
[0223] It should be noted that before the encoder pre-processes the target point cloud, it will determine whether to divide the entire point cloud sequence into multiple point cloud slices based on the parameter configuration. In the case where the entire point cloud sequence is determined to be divided into multiple point cloud slices based on the above parameter configuration, each point cloud slice can be treated as a single independent point cloud for serial processing. Therefore, in some embodiments, the above target point cloud can be a point cloud at the overall level; in some embodiments, the above target point cloud can also be a point cloud slice obtained by dividing the point cloud sequence at the overall level, that is, the above target point cloud can be a point cloud slice level point cloud.
[0224] Similarly, the target feature can be a feature of the entire point cloud or a feature of the point cloud slice. For example, if the target feature is point cloud density, then if the target point cloud is an entire point cloud, the target feature refers to the density of the entire point cloud; if the target point cloud is a slice, the target feature refers to the density of the slices obtained by partitioning the entire point cloud.
[0225] In one embodiment of the present application, the encoder may determine the information required for encoding the target point cloud as the above-mentioned target features, for example, obtaining one or more of the following information: point cloud characteristic information, bounding box information, encoding parameter information, and target parameter information. In another embodiment of the present application, the above-mentioned target features may also be information that is defaulted at the codec end. In addition, the encoder may also look up a preset table to determine the above-mentioned target features, wherein the above-mentioned preset table is generated before or during encoding of the target point cloud. The default parameters at the codec end refer to information that is mutually defaulted between the encoder end and the corresponding decoder end, and is information that does not need to be encoded or decoded during the point cloud encoding and decoding process. The above-mentioned target parameter information may include constant values (or fixed values, default values), and may also include at least one preset identifier, wherein each preset identifier corresponds to a fixed value.
[0226] In an exemplary embodiment, the encoding parameter information refers to feature parameters used when encoding the target point cloud, and may include one or more of the following information: the maximum transform order, attribute output bit depth, and attribute quantization parameter required by the encoder to encode attribute information; and the geometric output bit depth and geometric quantization parameter required by the encoder to encode geometric information. The encoding parameter information may also include a transform method, a prediction method, and a sorting method.
[0227] In an exemplary embodiment, the point cloud characteristic information may reflect the characteristics of the target point cloud, and may include, for example, one or more of the following information: point cloud type, point cloud density, point cloud space occupancy, resolution, and point cloud point count.
[0228] The above-mentioned point cloud characteristic information provided in the embodiments of the present application may specifically refer to the point cloud density, point cloud space occupancy, resolution and point cloud point count at the overall point cloud level. The above-mentioned point cloud characteristic information may also specifically refer to the point cloud density, point cloud space occupancy, resolution and point cloud point count at the point cloud slice level.
[0229] In an exemplary embodiment, the bounding box information is specifically the features of the bounding box of the target point cloud.
[0230] It is understood that the encoder performs coordinate transformation on the geometric information of the target point cloud during the pre-processing stage, specifically to make the point cloud all contained in a bounding box. Referring to Figure 5, the bounding box represents the smallest cuboid that contains all points in the input point cloud.
[0231] The origin and size of the bounding box are determined as follows:
[0232] The floating point coordinates of the k-th point of the input point cloud are expressed as (x k ,y k , z k ), k = 0, ..., K-1, K is the number of points in the point cloud, coordinate information x min ,y min , z min , x max ,y max , z max They are represented as follows: min =min(x0, x1, ..., x K-1 ) y min =min(y0, y1, ..., y K-1 ) z min =min(z0, z1, ..., z K-1 ) x max =max(x0, x1, ..., x K-1 ) y max =max(y0, y1, ..., y K-1 ) z max =max(z0, z1, ..., z K-1 )
[0233] The above function min(s0, s1, ..., s K-1 ) means taking the current input (s0, s1, ..., s K-1 ), the minimum value among max(s0, s1, ..., s K-1 ) means taking the current input (s0, s1, ..., s K-1 ) is the maximum value among .
[0234] Furthermore, the origin of the bounding box (x origin ,y origin , z origin ) is as follows: origin =int(floor(x min )) y origin =int(floor(y min )) z origin =int(floor(zmin ))
[0235] The size of the bounding box in the x, y, and z directions can be calculated as follows: Bounding_Box_Size_x = int(x max -x origin )+1 Bounding_Box_Size_y=int(y max -y origin )+1 Bounding_Box_Size_z=int(z max -z origin )+1
[0236] In this case, int(s) represents the rounding function, and the floor(s) function returns the maximum integer value less than or equal to s.
[0237] In an embodiment of the present application, the bounding box information of the target point cloud may include one or more of the following information: the length and direction of the three sides of the bounding box, and the length statistics of the three sides of the bounding box. Among them, the length statistics of the three sides of the bounding box include one or more of the following information: the mean of the lengths of the three sides, the median value of the lengths, the mode value of the lengths, the maximum length of the three sides, the minimum length of the three sides, the linear combination of the lengths of at least two sides, the volume-related calculation of the bounding box, and the surface area-related calculation of the bounding box, etc. For example, the linear combination of the lengths of at least two sides may include: the sum of the side lengths, the difference between the side lengths, the combination of the logarithms of the side lengths, etc. The surface area-related calculation of the bounding box may include: the surface area corresponding to the shortest side length, the surface area corresponding to the two longer side lengths, and the combination of different types of surface areas as mentioned above, as well as the combination of the logarithms of the different types of surface areas as mentioned above. It can be understood that all statistical values determined based on the bounding box information can be considered as the above-mentioned target features.
[0238] After the above embodiment introduces the embodiment regarding target features, the following describes how to determine the coding grouping information based on the above target features.
[0239] The embodiment of the present application does not limit the specific manner in which the encoder determines the encoding grouping information based on the target features of the target point cloud.
[0240] In some embodiments, the encoder may determine at least one of the point cloud characteristic information, the bounding box information, the encoding parameter information, and the target parameter information of the target point cloud as the decoding grouping information.
[0241] In some embodiments, the encoder may determine the decoding group information through the following steps S610-1 and S610-2A or S610-2B:
[0242] S610 - 1 : Map at least one of the above target features through an objective function to obtain at least one mapping value.
[0243] For example, the encoding end maps the target feature x through the target function f() to obtain the mapping value f(x) corresponding to the target feature.
[0244] In the embodiment of the present application, the objective function f() can be a linear function, a quadratic function, a direct proportional function, an inverse proportional function, a trigonometric function, an exponential function, a logarithmic function, etc. For example, the objective function is a logarithmic function f(x)=log2x.
[0245] In the embodiment of the present application, when the target feature is expressed as a parameter value or a feature value, the encoder can map the value through the target function. For example, when the target feature includes two or more parameter values, different parameter values can be mapped by different types of target functions. For example, when the target feature includes parameter values x1 and x2, the corresponding mapping values can be: lg(x1) and e x2 .
[0246] Exemplarily, the above mapping processing can also be a linear combination related to the number of points; for example, Log2(a*num_points+ / -b), where a and b are constants; or, the above mapping processing can also be a linear combination related to the logarithmic value of the number of points; for example, ba*Log2(num_points), or b+a*Log2(num_points) where a and b are constants.
[0247] It is understandable that the above-mentioned mapping value may also be the parameter value of the target feature itself. For example, if the parameter value of the target feature is x3, after the mapping operation of this embodiment, the mapping value is still x3.
[0248] In some embodiments, in the embodiments of the present application, after the encoder performs mapping processing on the target feature according to the above-mentioned objective function, the determined mapping value is an integer. Specifically, when the above-mentioned objective function includes logarithmic processing, rounding calculation is performed after the logarithmic processing; when the above-mentioned objective function includes division processing, rounding calculation is performed after the division processing. It should be noted that when decimals appear in the mapping process based on other objective functions, or when the target feature itself remains after mapping with the objective function and the target feature contains decimals, rounding operation is also required.
[0249] In an exemplary embodiment, the above rounding calculation can be: rounding up, rounding down, or rounding to the nearest integer. Of course, other forms of rounding calculation methods, such as rounding to the nearest integer, are also possible, and the embodiments of this application are not limited to this. For example, if the target feature is the number of points in the target point cloud and its value is 10000, and the objective function is a logarithmic function with base 2, then the mapping value determined by the above objective function is: round(log2(10000)).
[0250] After determining the above mapping value, the following S610 - 2A or S610 - 2B may be executed to determine the coding group information.
[0251] In S610 - 2A, encoding grouping information is determined according to at least one combination of mapping values or a combination of a mapping value and a constant value.
[0252] The above combination can be a linear combination or a nonlinear combination.
[0253] In an exemplary embodiment, assuming that the at least one mapping value can be expressed as f(x) and g(y), the exemplary coding grouping information can be expressed as: A×f(x)+B×g(y), A×f(x)-B×g(y), A×f(x) 2 +A×g(y) 2 、A×f(x) 2 -A×g(y) 2 、A×f(x) 3 +B×g(y), etc. Where f() and g() can represent two identical or different objective functions, and A and B are constants and positive numbers.
[0254] In S610 - 2B, encoding grouping information is determined based on a combination of the at least one mapping value and the unmapped target feature.
[0255] As mentioned above, the target feature that has not been mapped can also be considered to be the same as before after being mapped by the target function. If the mapping value corresponding to x is still x, and at least one mapping value obtained by mapping at least one target feature by the target function can be expressed as g(y), then the exemplary group coding information can be expressed as: C×x+B×g(y), C×xB×g(y), C×x 2 +A×g(y) 2 、C×x 2 -A×g(y) 2 、C×x 3 +B×g(y), etc. Where f() and g() can represent two identical or different objective functions, and A, B, and C are constants and integers.
[0256] In an exemplary embodiment, the coding grouping information may include an initial grouping step size, an intermediate quantity used to calculate the initial grouping step size, and an upper limit and / or lower limit of the initial grouping step size. Therefore, as a specific implementation of S610, any one of the following steps may be included:
[0257] S610-A: Determine the initial grouping step size in the coded grouping information.
[0258] S610-B: Determine an intermediate quantity for calculating the initial grouping step length, and determine the initial grouping step length according to the intermediate quantity to obtain the coding grouping information.
[0259] S610-C: Determine an upper limit and / or lower limit of the initial grouping step in the encoded grouping information, and determine an upper limit and / or lower limit of the intermediate amount in the decoded grouping information.
[0260] The initial grouping step size is the step size used by the encoder to initially group the second reconstructed points (the reconstructed points obtained by the encoder during the attribute prediction phase) after the reordering of the reconstructed points. In other words, it is the grouping step size used when the encoder first groups the second reconstructed points during the attribute prediction process of the target point cloud.
[0261] The following describes a specific implementation of S610-A, which is to determine the initial grouping step length L according to the above target characteristics. b Specific implementation method:
[0262] Specific implementation method 1: L1=Log s (gsh_bounding_box_size_x)+Log s (gsh_bounding_box_size_y)+ Log s (gsh_bounding_box_size_z); L2=Log s (slice_num_points); L b =L1–L2, and L b ≥1;
[0263] Among them, the Log s () is ceilLog s (), floorLog s () or roundLog s ();
[0264] L brepresents the initial grouping step size, gsh_bounding_box_size is the bounding box size of the point cloud slice level, slice_num_points is the number of point cloud points at the point cloud slice level, s represents a constant greater than 1, and L1 and L2 represent intermediate values.
[0265] Specific implementation method 2: L1=Log s (bounding_box_size_x)+Log s (bounding_box_size_y)+ Log s (bounding_box_size_z); L b =max(3,3×round(L1 / 3));
[0266] Among them, the Log s () is ceilLog s (), floorLog s () or roundLog s ();
[0267] L b represents the initial grouping step size, bounding_box_size_x, bounding_box_size_y and bounding_box_size_z respectively represent the lengths of the three sides of the bounding box of the target point cloud, s represents a constant greater than 1, and L1 represents an intermediate value.
[0268] Specific implementation method three: L1 = 2 × Log s (gsh_bounding_box_size_x); L2=Log s (slice_num_points);
[0269] L b =L1–L2+shift; and L b ≥1; can also be expressed as: L b =max(1,L1-L2+shift); or, L b =max(1,L1-L2)+shift;
[0270] Among them, the Log s () is ceilLog s (), floorLog s () or roundLog s ();
[0271] L brepresents the initial grouping step size, gsh_bounding_box_size_x represents the maximum length of the three sides of the point cloud slice, slice_num_points represents the number of points in the point cloud slice, shift represents the offset value, s represents a constant greater than 1, and L1 and L2 represent intermediate values.
[0272] Specific implementation method 4: L1=Log s (bounding_box_size_x)+Log s (bounding_box_size_y)+ Log s (bounding_box_size_z); L b =max(3,3×round(L1 / 3));
[0273] Among them, the Log s () is ceilLog s (), floorLog s () or roundLog s ();
[0274] L b represents the initial grouping step size, bounding_box_size_x, bounding_box_size_y and bounding_box_size_z respectively represent the lengths of the three sides of the bounding box of the target point cloud, s represents a constant greater than 1, and L1 represents an intermediate value.
[0275] Specific implementation method 5: In this embodiment, L is determined based on the surface area information of the bounding box. b ; L1=gsh_bounding_box_size_x_log2+gsh_bounding_box_size_z_log2; or, L1=ceilLog2(gsh_bounding_box_size_x*gsh_bounding_box_size_z)); L2=ceillog2(num_points);
[0276] Where x and z represent the side length of the smallest area among the six sides of the bounding box; or, x and z represent the side length of the largest area among the six sides of the bounding box; or, one of x and z represents the length of the smallest length among the three sides of the bounding box; or, one of x and z represents the length of the largest length among the three sides of the bounding box; L b =max(1,max(3,3×round(L1–L2 / 3))+shift).
[0277] Specific implementation scheme six: In this embodiment, L is determined based on the entire or partial surface area information of the bounding box. b ; L1=ceilLog2 (gsh_bounding_box_size_x*gsh_boundingbox_size_z+gsh_bounding_box_size_y*gsh_boundin gbox_size_z); L2=ceilog2(num_points); L b =max(3,3×round(L1–L2 / 3))+shift.
[0278] From the above examples, we can see that the decoder can determine the initial grouping step length L in the decoded grouping information based on the bounding box information of the target point cloud. b Specifically, the decoder determines the intermediate quantities L1 and L2 used to calculate the initial grouping step length based on the bounding box information of the target point cloud. Then, the initial grouping step length L in the decoded grouping information is calculated based on the intermediate quantities L1 and L2. b .
[0279] Specific implementation plan seven:
[0280] L1 is obtained by parsing the code stream; it can be understood that the encoder determines a fixed value and uses the fixed value as L1; L b =L1+shift;
[0281] Specific implementation plan eight: L1=32; L2=log2(num_points); L b =max(3,3×round(L1–L2 / 3))+shift.
[0282] Specific implementation method 9: L1 = min (A, ceillog2 (xyz)); L2 = log2 (num_points); L b =max(3,3×round(L1–L2 / 3))+shift;
[0283] Wherein, A represents the fixed value in the above target parameter information, and xyz represents the volume of the bounding box.
[0284] It should be noted that, in any specific implementation, after determining the intermediate values L1 and L2, the initial grouping step length can be determined according to the following method: b =max(3,3×round(L1–L2 / 3))+shift; or, L b =max(3,3×round(L1–L2 / 3)+shift); or L b=max(1,max(3,3×round(L1–L2 / 3))+shift).
[0285] The calculation method of shift in the above specific implementation scheme can be any of the following methods:
[0286] 1) The encoder obtains the shift value based on a preset lookup table T. For example, the index value of the lookup table is calculated based on the geometric quantization parameter, the attribute quantization parameter, and the quantization offset parameter, and the shift value is determined in the lookup table T according to the determined index value.
[0287] 2) The encoder determines the above-mentioned offset value based on the attribute features obtained by parsing the code stream of the target point cloud, such as the header information parameter attrQuantParam and possible QPOffset;
[0288] 3) The encoder can also directly obtain a fixed value of shift and determine the fixed value as the shift value. The fixed value directly obtained above can be the default value of the codec end or the parameter required by the encoder during the encoding process of the target point cloud.
[0289] 4) The encoder can also determine the fixed value corresponding to the preset identifier as the offset value shift, where different preset identifiers correspond to different fixed values. For example, when the preset identifier is C1, the offset value is c1, and when the preset identifier is C2, the offset value is c2, where both c1 and c2 are positive numbers. The preset identifier can be the default of the codec, required by the encoder during encoding of the target point cloud, or determined by a lookup table.
[0290] Specific implementation method ten:
[0291] In this embodiment, the value of the initial grouping step in the above-mentioned coding grouping information is an integer multiple of the target fixed value. For example, in the above-mentioned "Specific Implementation Method Eight", 'L1=32', where L1 is an integer multiple of 1, the "target fixed value" in this embodiment is 1. In this embodiment of the present application, the above-mentioned target fixed value may be determined by the encoder based on the target features of the target point cloud; or, the above-mentioned target fixed value may be determined by the encoder based on a preset identifier, and different preset identifiers correspond to different fixed values, so that the above-mentioned target fixed value can be determined based on the preset identifier; or, the above-mentioned target fixed value may be determined by the encoder as the first default value in the target parameter information.
[0292] In some embodiments, as a specific implementation of S610, the following steps S1 to S2 may be included:
[0293] S1: Determine a target fixed value based on target features of the target point cloud; or, determine the target fixed value based on a preset identifier, wherein different preset identifiers correspond to different fixed values; or, determine a first default value in the target parameter information as the target fixed value.
[0294] S2: Determine the integer multiple of the above target fixed value as the value of the initial grouping step.
[0295] In some embodiments, the target feature used in S1 may be a fixed parameter set by default on the codec side, or may be one or more of the above-mentioned encoding parameter information, point cloud feature information, and bounding box information.
[0296] Exemplarily, when the above-mentioned target feature adopts the point cloud type in the point cloud characteristic information, when the point cloud type is a dense type or a human eye visual point cloud, the above-mentioned target fixed value is determined to be m; or, when the point cloud type is a sparse type or a dense visual point cloud, the above-mentioned target fixed value is determined to be n, where m and n are positive integers.
[0297] For example, the value of m is 3, and the value of n is 1. For example, if the target feature is a point cloud type and the point cloud type is a dense type, the value of the initial grouping step size can be an integer multiple of 3. Of course, the embodiment of the present application does not limit the values of m and n.
[0298] Exemplarily, when the target feature adopts bounding box information, when the correlation of the three sides of the bounding box meets the preset conditions, the target fixed value is determined to be p; or, when the correlation of the three sides of the bounding box does not meet the preset conditions, the target fixed value is determined to be q, where p and q are positive integers. The preset condition may specifically be whether the correlation is greater than a preset value (such as 80%). Specifically, the correlation between the lengths of the three sides can be determined by the length variance or standard deviation of the three sides. For example, if boundingbox_size_z is much smaller or much larger than boundingbox_size_x or boundingbox_size_y, it means that the length similarity of the three sides of the bounding box is small.
[0299] For example, the value of p is 3, and the value of q is 1. For example, if the variance between boundingbox_size_z, boundingbox_size_x, and boundingbox_size_y is less than a preset value, indicating that the lengths of the three sides of the bounding box are highly similar, then the value of the initial grouping step can be an integer multiple of 3. Of course, the embodiments of the present application do not limit the values of p and q.
[0300] Exemplarily, the above-mentioned target fixed value may also be determined according to the above-mentioned encoding parameter information, for example, according to a geometric quantization parameter and / or an attribute quantization parameter.
[0301] In other embodiments, the preset identifier in S1 can be a default identifier on the codec side, can be information required by the encoder to encode the target point cloud, or can be determined based on a preset lookup table. For example, when the encoder determines that the identifier is Q1 during the encoding process, the fixed value corresponding to the identifier Q1 is determined as the above-mentioned target fixed value. In addition, the encoder can directly encode the above-mentioned first default value into the bitstream and determine the first default value as the above-mentioned target fixed value. The above-mentioned first default value can be a default identifier on the codec side, or can be required for encoding the above-mentioned target point cloud.
[0302] It is understandable that the specific implementation for determining the initial grouping step size is not limited to the above embodiment. For example, other combinations between target features and between target features and constants may also be used, and this application does not limit this.
[0303] The following describes a specific implementation of S610-B, which is to determine the step length L for calculating the initial grouping. b Specific implementation of the intermediate amount:
[0304] In related technologies, when the attribute is reflectivity, during the attribute prediction process, the initial grouping step length is expressed as follows:
[0305] Among them, it is used to determine the initial grouping step length L base The intermediate amount is the above-mentioned offset value shift, which can also be the maximum number of bits maxBits and the minimum number of bits minBits.
[0306] The embodiment of the present application can be similar to the embodiment corresponding to S610-A. The encoder can directly determine the initial grouping step length L according to the above target features. b The initial grouping step length L in the related technology can also be determined based on the above target characteristics. base The encoder may also determine a second default value in the target parameter information as the intermediate value, where the second default value may be a codec default, required by the encoder for encoding the target point cloud, or determined based on a preset lookup table. The encoder may also determine a fixed value corresponding to a preset identifier in the target parameter information as the intermediate value, where different preset identifiers correspond to different fixed values.
[0307] For example, the maximum number of bits maxbits = Log s(V(bounding_box)); minimum number of bits minbits = ceilLogs(B); wherein maxbits represents the maximum number of bits, V(bounding_box) represents the volume of the bounding box, minbits represents the minimum number of bits, and B represents a fixed parameter.
[0308] For example, the maximum number of bits maxbits = Log s (num_points); the minimum number of bits minbits takes different parameter values according to the type of target point cloud; among them, num_points represents the number of points in the target point cloud.
[0309] Exemplarily, the offset value shift may be determined by parsing the attribute features obtained from the code stream of the target point cloud, such as the header information parameter attrQuantParam and possible QPOffset.
[0310] Exemplarily, the offset value may also be determined by looking up a fixed parameter X1 determined in a preset table, wherein the preset table is generated before or when the encoder encodes the target point cloud.
[0311] Exemplarily, the offset value shift may be determined by parsing the fixed parameter X2 obtained from the code stream of the target point cloud, for example, shift=ceil(log2(X2)).
[0312] It can be seen that the embodiment of the present application can determine various combinations between the above target features and their mapping values through S420-2A or S420-2B, thereby providing multiple ways of determining the above intermediate quantities. The embodiment of the present application does not limit the specific combination method.
[0313] As a specific implementation of S410-C, that is, determining the initial grouping step length L b The specific implementation methods of the upper limit and lower limit include but are not limited to the following:
[0314] (1) The encoder determines the obtained third default value as the upper limit value of the intermediate quantity (such as shift, maxbits, minbits), and determines the obtained fourth default value as the lower limit value of the intermediate quantity (such as shift, maxbits, minbits); for example, the encoder uses the fixed value a as the upper limit value of the intermediate quantity shift, and the fixed value b as the upper limit value of shift. It can be understood that a and b are positive numbers, and a is greater than b; wherein, exemplarily, the above-mentioned fixed value a and the above-mentioned fixed value b can be the default of the codec end, or can be required by the encoder to encode the target point cloud; in another exemplary embodiment, the encoder can determine the above-mentioned fixed value a according to the first preset identifier, and determine the above-mentioned fixed value b according to the second preset identifier, wherein the first preset identifier indicates that the upper limit value of shift is set to the corresponding fixed value a, and the second preset identifier indicates that the lower limit value of shift is set to the corresponding fixed value b. The above-mentioned first preset identifier and the second preset identifier can be the default of the codec end, or can be required by the encoder to encode the target point cloud.
[0315] Specifically, the encoder uses a fixed value c as the initial grouping step size L b The upper limit value of the fixed value d is used as the initial grouping step length L b It is understandable that c and d are positive numbers, and c is greater than d; where, for example, the fixed value c and the fixed value d can be the default value of the codec end, or can be the feature required by the encoder to encode the target point cloud; in another exemplary embodiment, the encoder can determine the fixed value c according to the first preset identifier and determine the fixed value d according to the second preset identifier, where the first preset identifier indicates that L b The upper limit value of L is set to the corresponding fixed value c, and the second preset mark indicates that L b The lower limit of is set to the corresponding fixed value d. The first preset identifier and the second preset identifier may be the default ones of the codec end, or may be the features required by the encoder to encode the target point cloud.
[0316] (2) The encoder can determine the initial grouping step size L based on the point cloud type. b Or the upper and lower limits of the intermediate amount; for example, compared with the sparse type point cloud, for the dense type point cloud, the initial grouping step length L b The upper and lower limits of are both small. For example, for dense point clouds, the initial grouping step length L b The upper limit value is z1 and the lower limit value is z2; for dense point clouds, the initial grouping step size L b The upper limit is z3 and the lower limit is z4; then the value of z1 is less than z3, and the value of z2 is less than z4;
[0317] (3) The encoder can determine the initial grouping step length L based on the encoding parameter information b Or the upper and lower limits of the intermediate quantity; such as determining L according to the geometric output bit depth, attribute output bit depth, geometric quantization parameter and attribute quantization parameter b The upper limit and lower limit of L can also be determined according to the transformation method, prediction method and sorting method. b The upper and lower limits of the transform mode are as follows: if the transform mode is predictive transform coding, the initial grouping step length L is determined. b The upper limit is s1, the lower limit is s2, s1 is greater than s2; when the transformation mode is lifting transformation coding, determine the initial grouping step size L b The upper limit value is s3, the lower limit value is s4, and s3 is greater than s4; wherein, s1 and s3 can be different positive numbers, and s2 and s4 can be different positive numbers.
[0318] (4) The encoder can determine the initial grouping step length L based on the parameter information of the bounding box, such as the length of the bounding box side, the correlation of the side length, etc. b Or the upper and lower limits of an intermediate quantity.
[0319] It is understandable that determining the initial grouping step length L b Or the specific implementation of the upper limit value and the lower limit value of the intermediate quantity is not limited to the above content, and can be other forms of expression determined according to the above target characteristics.
[0320] Continuing with reference to FIG7 , in S620 , the second reconstructed position points are grouped using the above-mentioned coding grouping information to obtain a plurality of groups, wherein the above-mentioned second reconstructed position points are obtained by the above-mentioned encoder decoding the geometric code stream of the target point cloud; and, in S630 , attribute encoding is performed on each group.
[0321] In an exemplary embodiment, the encoder determines the initial grouping step size in the encoding grouping information according to the above-described embodiment. The encoder groups the reordered reconstructed position points according to the initial grouping step size. The specific implementation of grouping is described in detail in the embodiment corresponding to FIG8 . Furthermore, the encoder performs attribute encoding based on the grouped reconstructed position points, thereby completing the encoding process of the target point cloud codestream and obtaining a reconstructed point cloud of the target point cloud.
[0322] In the P600 point cloud encoding scheme provided in the embodiment of the present application, the encoder determines the above-mentioned encoding grouping information based on one or more of the following information related to the code stream of the encoding target point cloud, which is defaulted at the encoding and decoding end or encoded: point cloud characteristic information, bounding box information, encoding parameter information, and target parameter information. During the attribute encoding process, the encoder groups the first reconstructed position points by the above-mentioned grouping information to obtain multiple groups. Among them, the above-mentioned second reconstructed position points are obtained by the encoder decoding the geometric code stream of the target point cloud. Further, the encoder performs attribute encoding on each group. In the encoding scheme provided in the embodiment of the present application, the encoder determines the encoding grouping information based on the above-mentioned target features, and the encoding grouping information such as the initial grouping step, the intermediate value used to determine the initial grouping step, the upper and lower limits of the above-mentioned initial grouping step, etc. It can be seen that the embodiment of the present application determines the grouping information based on the diversified features related to the target point cloud itself. Compared with the method for determining the initial grouping step provided by the relevant technology, the embodiment of the present application can improve the grouping diversity, which is conducive to improving the encoding performance of the point cloud.
[0323] FIG8 is a flow chart illustrating a point cloud encoding method P700 according to another embodiment of the present application. The execution entity of the point cloud encoding method P700 may be an encoder or an electronic device that performs the encoding process. Referring to FIG8 , the point cloud encoding method P700 includes steps S72 to S78.
[0324] In S72 , the encoder performs geometric encoding on the geometric code stream of the target point cloud to obtain a second reconstructed position point.
[0325] In the embodiment of the present application, the reconstruction position point obtained by the encoder during the attribute prediction phase is recorded as the first reconstruction position point. In addition, the reconstruction position point obtained by encoding the geometry code stream by the decoder during the attribute reconstruction phase is recorded as the first reconstruction position point.
[0326] In the embodiment of the present application, the order of reconstructing the geometric coordinates of the point cloud includes but is not limited to the following methods:
[0327] a) After the encoder completes decoding and reconstruction of the slice geometry, it reconstructs the entire slice geometry. The final reconstructed geometry is rec_xyz = xyz + gsh_bounding_box_offset. gsh_bounding_box_offset is the xyz coordinate of the slice bounding box origin.
[0328] b) After decoding and reconstructing the slice attribute information, the encoder reconstructs the overall geometric information of the current slice.
[0329] In S74 , the encoder reorders the second reconstruction position points.
[0330] The embodiment of the present application does not limit the manner in which the second reconstructed position points are reordered. For example, the encoder may reorder according to the order of point cloud input; the encoder may reorder according to the order of point cloud space curves (including Morton order and Hilbert order); or the encoder may reorder according to input parameter information (such as acquisition information, LiDAR information, etc.).
[0331] In S76 , the encoder determines encoding grouping information, and groups the second reconstruction position points according to the encoding grouping information.
[0332] The specific implementation method of the encoder determining the encoding group information has been described in detail in the embodiment corresponding to P600 and will not be repeated here.
[0333] In the embodiment of the present application, the second reconstructed position points are reordered in the Hilbert order, and the reordered points correspond to the Hilbert code. b In the case of , the Hilbert codes corresponding to the second reconstruction position point can be grouped in the following way:
[0334] (1) The encoder determines the initial grouping step size L b , initially grouping the second reconstruction position points;
[0335] Shift the Hilbert code right by the initial grouping step size L b Then the same points are divided into the same macroblock.
[0336] (2) The encoder determines whether the Hilbert codes within the same macroblock need to be subdivided into groups;
[0337] It should be noted that for each point in a macroblock, the encoder can determine whether there are duplicate points in the current macroblock. If there are duplicate points in the current macroblock, the group is truncated at the duplicate point and the duplicate points are grouped separately, that is, the duplicate points are grouped with a point count of 1. For example, if the current macroblock contains 4 points and the 4th point is a duplicate point, the first 3 points can be grouped as one, and the 4th duplicate point can be grouped separately.
[0338] In the related art, during the color reconstruction process, if the number of points in the current block is less than or equal to the maximum transformation order colorMaxTransNum, the points in the macroblock become a group; if the number of points in the current block is greater than colorMaxTransNum, the points in the block are subdivided into groups, and the rule for the subdivision group is: obtain the current right shift bit L, then the right shift bit L1 of the current subdivision group takes the value of L-1, and then the same points after the Hilbert code is right shifted by L1 bits are grouped into a new subdivision group; and continue to judge whether the number of points in the subdivision group is greater than colorMaxTransNum, until the number of points in each subdivision group is less than or equal to colorMaxTransNum.
[0339] In the related art, during the reflectivity reconstruction process, if the number of points in the current macroblock is greater than the maximum transform order reflMaxTransNum, the encoder sequentially takes the points of the maximum transform order reflMaxTransNum as a subdivision group until the number of points in each group in the block is less than or equal to reflMaxTransNum.
[0340] The embodiment of the present application may adopt the method provided by the related technology to determine whether the points within the macroblock need to be subdivided into groups.
[0341] (3) The encoder determines whether the grouping step size needs to be updated;
[0342] In the related art, during the color reconstruction process, the sum of the number of points in the first three consecutive groups is counted, and the sum is shifted right by three places (divided by 8) and recorded as B. If the value of B is less than 2, the updated step length L = L+1; if the value of B is greater than 8, the updated step length L = L-1; if B is greater than or equal to 2 and less than or equal to 8, the step length L is not updated.
[0343] In related art, during reflectivity reconstruction, if the maximum transform order reflMaxTransNum is greater than 2 and the offset value shift in the initial grouping step is greater than 0, the encoder updates the grouping step; otherwise, the current grouping step is maintained. Specifically, the total number of points in the previous N groups is counted, and the average number of points in N (N=8) groups is calculated. If the average number of points is less than 2, the updated grouping step L is L+1; if the average number of points is greater than the maximum transform order reflMaxTransNum, the new grouping step L is L-1; if the average number of points is greater than or equal to 2 and less than or equal to the maximum transform order reflMaxTransNum, the grouping step L is also maintained unchanged.
[0344] It can be seen that in the related art, the number of points in the previous multiple groups is considered when determining whether to update the grouping step size. It should be noted that in the related art, a point group formed by repeated points can also be used as the above-mentioned previous group and used to determine whether to update the grouping step size. That is, the number of points in the group formed by repeated points is considered in the related art.
[0345] In the embodiment of the present application, the method of subdividing the points in the current macroblock into groups includes but is not limited to the following:
[0346] a) The encoder counts the number of points corresponding to the first N groups directly grouped by the grouping step information. The aforementioned direct grouping is a non-subdivided group, that is, a group determined by the grouping step. Specifically, the encoder determines whether to update the grouping step based on the number of points.
[0347] b) The encoder counts the number of points corresponding to the first N groups that do not contain repeated points. It should be noted that this method is different from the related art method of counting the number of repeated points when determining whether to update the grouping step size. In the embodiment of the present application, the method of counting the number of points in the group that does not contain repeated points determines whether to update the grouping step size.
[0348] In the embodiment of the present application, the updating method of the grouping step size is more flexible, which is conducive to improving the coding performance.
[0349] Continuing to refer to FIG. 8 , in S78 , the encoder performs intra-group attribute encoding.
[0350] In one embodiment, after all second reconstruction position points are grouped, attribute encoding can be performed on each grouped. In another embodiment, after each macroblock is subdivided into groups, attribute encoding can be performed on the completed groupings of the current macroblock while the next macroblock is processed, thereby improving overall encoding efficiency.
[0351] In the point cloud encoding method P700 provided in the embodiment of the present application, the encoder in the same method P600 determines the above-mentioned encoding grouping information based on the target parameters defaulted on the codec side or one or more of the following information related to the code stream encoding the target point cloud: point cloud characteristic information, bounding box information, encoding parameter information, and target parameter information, thereby improving grouping diversity and facilitating the improvement of point cloud encoding performance. At the same time, compared with related technologies, the point cloud encoding method P700 provided in this embodiment not only improves the method for determining the initial grouping step size and its intermediate quantity, making the grouping more diverse; it also expands the method for determining whether to update the grouping step size, provides an optimized processing solution for the attribute encoding process, and expands the geometric reconstruction method. The above improvements are all conducive to improving the encoding performance of the point cloud.
[0352] The above describes in detail the point cloud decoding method and the point cloud encoding method embodiments of the present application. The following, in conjunction with Figures 9 to 12, describes in detail the decoder and encoder embodiments provided by the embodiments of the present application.
[0353] FIG9 is a schematic diagram of the structure of a decoder provided in an embodiment of the present application. The decoder 900 may be an electronic device that decodes a target point cloud bitstream. Referring to FIG9 , the decoder 900 includes: a first determination module 910 , a decoding grouping module 920 , and an attribute decoding module 930 .
[0354] Among them, the above-mentioned first determination module 910 is used to determine the decoding grouping information based on the target characteristics of the target point cloud, wherein the above-mentioned target characteristics include one or more of the following information: point cloud characteristic information, bounding box information, decoding parameter information and target parameter information; the above-mentioned decoding grouping module 920 is used to group the first reconstructed position points through the above-mentioned decoding grouping information to obtain multiple groups, wherein the above-mentioned first reconstructed position points are obtained based on the decoder decoding the geometric code stream of the above-mentioned target point cloud; and the above-mentioned attribute decoding module 930 is used to perform attribute decoding on each group.
[0355] In some embodiments, based on the above solution, the above point cloud characteristic information includes one or more of the following information: point cloud type, point cloud density, point cloud space occupancy, resolution, and point cloud point count.
[0356] In some embodiments, based on the above scheme, the above-mentioned bounding box information includes one or more of the following information: the length and direction of the three sides of the bounding box of the target point cloud, and the length statistics of the above-mentioned three sides; wherein, the length statistics of the above-mentioned three sides include one or more of the following information: the mean value of the length of the above-mentioned three sides, the median value of the length, the mode value of the length, the maximum value of the length of the above-mentioned three sides, the minimum value of the length of the above-mentioned three sides, the linear combination of the lengths of at least two sides, the linear combination between the length of at least one side and a constant, the volume of the above-mentioned bounding box and the surface area of the above-mentioned bounding box.
[0357] In some embodiments, based on the above scheme, the above decoding parameter information includes one or more of the following information: maximum transformation order of attribute prediction transformation, geometric output bit depth, attribute output bit depth, geometric quantization parameter, attribute quantization parameter, transformation method, prediction method and sorting method.
[0358] In some embodiments, based on the above solution, the target parameter information may also be determined by the decoder searching a preset table, where the preset table is determined when the encoder encodes the target point cloud.
[0359] In some embodiments, based on the above scheme, the above-mentioned first determination module 910 is specifically used to: map at least one of the above-mentioned target features through an objective function to obtain at least one mapping value; determine the decoding group information based on the combination between the above-mentioned at least one mapping value and a constant value; or, determine the decoding group information based on the combination between the above-mentioned mapping values; or, determine the decoding group information based on the combination between the above-mentioned at least one mapping value and an unmapped target feature.
[0360] In some embodiments, based on the above solution, the decoder 900 further includes: a rounding module;
[0361] The rounding module is used to: when the objective function includes logarithm processing, perform rounding calculation after the logarithm processing; or when the objective function includes division processing, perform rounding calculation after the division processing.
[0362] In some embodiments, based on the above solution, the first determination module 910 includes: a step length determination unit;
[0363] The step length determining unit is used to determine the initial grouping step length in the decoded grouping information according to the target features of the target point cloud.
[0364] In some embodiments, the initial grouping step length in the decoded grouping information is based on the above solution, and the step length determining unit includes: a fixed value determining subunit and a step length determining subunit;
[0365] The above-mentioned fixed value determination subunit is used to: determine the target fixed value based on at least one of the point cloud characteristic information, bounding box information and decoding parameter information in the above-mentioned target features; or, determine the target fixed value based on the preset identifier in the above-mentioned target parameter information, wherein different preset identifiers correspond to different fixed values; or, determine the first default value in the above-mentioned target parameter information as the above-mentioned target fixed value; the above-mentioned step size determination subunit is used to: determine an integer multiple of the above-mentioned target fixed value as the value of the above-mentioned initial grouping step size.
[0366] In some embodiments, based on the above scheme, the fixed value determination subunit is specifically configured to: when the target feature is a point cloud type, when the point cloud type is a dense type or a human vision point cloud, determine the target fixed value to be m; or, when the point cloud type is a sparse type or a machine vision point cloud, determine the target fixed value to be n, where both m and n are positive integers;
[0367] In some embodiments, based on the above scheme, the above-mentioned fixed value determination subunit is specifically used to: when the above-mentioned target feature is the bounding box information, when the above-mentioned target feature is the three-side correlation of the above-mentioned bounding box that meets the preset conditions, determine that the above-mentioned target fixed value is p; or, when the above-mentioned target feature is the three-side correlation of the above-mentioned bounding box that does not meet the preset conditions, determine that the above-mentioned target fixed value is q, where the values of p and q are both positive integers.
[0368] In some embodiments, the initial grouping step length in the decoded grouping information is based on the above solution, and the above first determination module 910 includes: an intermediate amount determination unit;
[0369] In which, the above-mentioned intermediate quantity determination unit is used to: determine the intermediate quantity used to calculate the initial grouping step length based on at least one of the point cloud characteristic information, bounding box information and decoding parameter information in the above-mentioned target features; or, determine the fixed value corresponding to the preset identifier in the above-mentioned target parameter information as the above-mentioned intermediate quantity, wherein different above-mentioned preset identifiers correspond to different fixed values; or, determine the second default value in the above-mentioned target parameter information as the above-mentioned intermediate quantity; and, determine the above-mentioned initial grouping step length based on the above-mentioned intermediate quantity to obtain the above-mentioned decoding grouping information.
[0370] In some embodiments, the initial grouping step length in the decoded grouping information is based on the above solution, and the above first determining module 910 includes: a limit value determining unit;
[0371] The limit value determination unit is used to:
[0372] Based on at least one of the point cloud characteristic information, bounding box information and decoding parameter information in the target feature, determine at least one of the upper limit value and the lower limit value of the initial grouping step in the decoding grouping information; or, determine the third default value in the target parameter information as the upper limit value of the intermediate quantity, and determine the fourth default value in the target parameter information as the lower limit value of the intermediate quantity, and the intermediate quantity is used to determine the initial grouping step; or, determine the fixed value corresponding to the first preset identifier in the target parameter information as the upper limit value of the initial grouping step or the intermediate quantity, and determine the fixed value corresponding to the second preset identifier in the target parameter information as the lower limit value of the initial grouping step or the intermediate quantity, wherein different preset identifiers correspond to different fixed values.
[0373] It should be understood that the decoder embodiment and the point cloud decoding method embodiment may correspond to each other, and similar descriptions can refer to the point cloud decoding method embodiment. To avoid repetition, they are not described here. Specifically, the decoder shown in Figure 9 can perform the above-mentioned point cloud decoding method embodiment, and the aforementioned and other operations and / or functions of each module in the device are respectively for implementing the method embodiment corresponding to the node in the master node group. For the sake of brevity, they are not described here.
[0374] FIG10 is a schematic diagram of the structure of an encoder provided in an embodiment of the present application. The encoder 1000 may be an electronic device that encodes a target point cloud bitstream. Referring to FIG10 , the encoder 1000 includes: a second determination module 1010, a coding grouping module 1020, and an attribute coding module 1030.
[0375] Among them, the above-mentioned second determination module 1010 is used to determine the coding grouping information based on the target characteristics of the target point cloud, wherein the above-mentioned target characteristics include one or more of the following information: point cloud characteristic information, bounding box information, coding parameter information and target parameter information; the above-mentioned coding grouping module 1020 is used to group the second reconstructed position points through the above-mentioned coding grouping information to obtain multiple groups, wherein the above-mentioned second reconstructed position points are obtained by the above-mentioned encoder geometrically decoding the geometric code stream of the above-mentioned target point cloud; and the above-mentioned attribute coding module 1030 is used to perform attribute encoding on each group.
[0376] In some embodiments, based on the above solution, the above point cloud characteristic information includes one or more of the following information: point cloud type, point cloud density, point cloud space occupancy, resolution, and point cloud point count.
[0377] In some embodiments, based on the above scheme, the above-mentioned bounding box information includes one or more of the following information: the length and direction of the three sides of the bounding box of the target point cloud, and the length statistics of the above-mentioned three sides; wherein, the length statistics of the above-mentioned three sides include one or more of the following information: the mean value of the length of the above-mentioned three sides, the median value of the length, the mode value of the length, the maximum value of the length of the above-mentioned three sides, the minimum value of the length of the above-mentioned three sides, the linear combination of the lengths of at least two sides, the linear combination between the length of at least one side and a constant, the volume of the above-mentioned bounding box and the surface area of the above-mentioned bounding box.
[0378] In some embodiments, based on the above scheme, the above encoding parameter information includes one or more of the following information: maximum transformation order of attribute prediction transformation, geometric output bit depth, attribute output bit depth, geometric quantization parameter, attribute quantization parameter, transformation method, prediction method and sorting method.
[0379] In some embodiments, based on the above solution, the target parameter information may also be determined by the encoder searching a preset table, where the preset table is determined when the encoder encodes the target point cloud.
[0380] In some embodiments, based on the above scheme, the above-mentioned second determination module 1010 is specifically used to: map at least one of the above-mentioned target features through an objective function to obtain at least one mapping value; determine the coding grouping information based on the combination between the above-mentioned at least one mapping value and a constant value; or, determine the coding grouping information based on the combination between the above-mentioned mapping values; or, determine the coding grouping information based on the combination between the above-mentioned at least one mapping value and an unmapped target feature.
[0381] In some embodiments, based on the above solution, the encoder 1000 further includes: a rounding module;
[0382] The rounding module is used to: when the objective function includes logarithm processing, perform rounding calculation after the logarithm processing; or when the objective function includes division processing, perform rounding calculation after the division processing.
[0383] In some embodiments, the initial grouping step length in the decoded grouping information is based on the above solution, and the second determining module 1010 includes: a step length determining unit;
[0384] The step length determining unit is configured to determine an initial grouping step length in the coding grouping information according to target features of the target point cloud.
[0385] In some embodiments, the initial grouping step size in the decoded grouping information is based on the above solution and includes: a fixed value determination subunit and a step size determination subunit;
[0386] The above-mentioned fixed value determination subunit is used to: determine the target fixed value based on at least one of the point cloud characteristic information, bounding box information and encoding parameter information in the above-mentioned target features; or, determine the target fixed value based on the preset identifier in the above-mentioned target parameter information, wherein different preset identifiers correspond to different fixed values; or, determine the default value in the above-mentioned target parameter information as the above-mentioned target fixed value; the above-mentioned step size determination subunit is used to: determine an integer multiple of the above-mentioned target fixed value as the value of the above-mentioned initial grouping step size.
[0387] In some embodiments, based on the above scheme, the fixed value determination subunit is specifically configured to: when the target feature is a point cloud type, when the point cloud type is a dense type or a human vision point cloud, determine the target fixed value to be m; or, when the point cloud type is a sparse type or a machine vision point cloud, determine the target fixed value to be n, where both m and n are positive integers;
[0388] In some embodiments, based on the above scheme, the above-mentioned fixed value determination subunit is specifically used to: when the above-mentioned target feature is bounding box information, when the correlation of the three sides of the above-mentioned bounding box meets the preset conditions, determine that the above-mentioned target fixed value is p; or, when the correlation of the three sides of the above-mentioned bounding box does not meet the preset conditions, determine that the above-mentioned target fixed value is q, where the values of p and q are both positive integers.
[0389] In some embodiments, the initial grouping step length in the decoded grouping information is based on the above solution, and the second determining module 1010 includes: an intermediate amount determining unit;
[0390] In which, the above-mentioned intermediate quantity determination unit is used to: determine the intermediate quantity used to calculate the initial grouping step length based on at least one of the point cloud characteristic information, bounding box information and encoding parameter information in the above-mentioned target features; or, determine the fixed value corresponding to the preset identifier in the above-mentioned target parameter information as the above-mentioned intermediate quantity, wherein different above-mentioned preset identifiers correspond to different fixed values; or, determine the second default value in the above-mentioned target parameter information as the above-mentioned intermediate quantity; and, determine the above-mentioned initial grouping step length based on the above-mentioned intermediate quantity to obtain the above-mentioned encoding grouping information.
[0391] In some embodiments, the initial grouping step length in the decoded grouping information is based on the above solution, and the second determining module 1010 includes: a limit value determining unit;
[0392] In which, the above-mentioned limit determination unit is used to: determine at least one of the upper limit value and the lower limit value of the initial grouping step in the decoding grouping information based on at least one of the point cloud characteristic information, bounding box information and encoding parameter information in the target feature; or, determine the third default value in the target parameter information as the upper limit value of the intermediate quantity, and determine the fourth default value in the target parameter information as the lower limit value of the intermediate quantity, and the intermediate quantity is used to determine the initial grouping step; or, determine the fixed value corresponding to the first preset identifier in the target parameter information as the upper limit value of the initial grouping step or the intermediate quantity, and determine the fixed value corresponding to the second preset identifier in the target parameter information as the lower limit value of the initial grouping step or the intermediate quantity, wherein different preset identifiers correspond to different fixed values.
[0393] It should be understood that the encoder embodiment and the point cloud encoding method embodiment can correspond to each other, and similar descriptions can refer to the point cloud encoding method embodiment. To avoid repetition, they are not described here. Specifically, the encoder shown in Figure 10 can perform the above-mentioned point cloud encoding method embodiment, and the aforementioned and other operations and / or functions of each module in the encoder device are respectively for implementing the method embodiment corresponding to the node in the master node group. For the sake of brevity, they are not described here.
[0394] FIG11 is a schematic structural diagram of an electronic device provided in another embodiment of the present application. The electronic device 1100 shown in FIG11 can be used to execute the above-mentioned point cloud encoding or point cloud decoding method.
[0395] As shown in FIG11 , the electronic device 1100 may include a memory 1110 and a processor 1120. The memory 1110 is configured to store a computer program 1130 and transmit the program code 1130 to the processor 1120. In other words, the processor 1120 may call and execute the computer program 1130 from the memory 1110 to implement the point cloud decoding method or the point cloud encoding method in the embodiments of the present application.
[0396] For example, the processor 1120 may be configured to execute the steps of the above method according to the instructions in the computer program 1130 .
[0397] In some embodiments of the present application, the processor 1120 may include but is not limited to:
[0398] General-purpose processor, Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.
[0399] In some embodiments of the present application, the memory 1130 includes but is not limited to:
[0400] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).
[0401] In some embodiments of the present application, the computer program 1130 may be divided into one or more modules, which are stored in the memory 1110 and executed by the processor 1120 to implement the method for recording a page provided by the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 1130 in the electronic device.
[0402] As shown in FIG. 11 , the electronic device 1100 may further include a transceiver 1140 , which may be connected to the processor 1120 or the memory 1110 .
[0403] The processor 1120 may control the transceiver 1140 to communicate with other devices. Specifically, the processor 1120 may send information or data to other devices or receive information or data sent by other devices. The transceiver 1140 may include a transmitter and a receiver. The transceiver 1140 may further include one or more antennas.
[0404] It should be understood that the various components in the electronic device 1100 are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.
[0405] FIG12 is a schematic diagram of the structure of a coding and decoding system provided in an embodiment of the present application. As shown in FIG12 , a coding and decoding system 1200 may include an encoder 1210 and a decoder 1220 .
[0406] In the embodiment of the present application, the encoder 1210 may be the encoder described in any one of the aforementioned embodiments, and the decoder 1220 may be the decoder described in any one of the aforementioned embodiments.
[0407] The apparatus of the embodiment of the present application is described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in hardware form, can be implemented by instructions in software form, or can be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software form instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps in the above method embodiment in conjunction with its hardware.
[0408] According to one aspect of the present application, a computer storage medium is provided, on which a computer program is stored. When the computer program is executed by a computer, the computer is enabled to perform the method of the above-mentioned method embodiment. Alternatively, the present application also provides a computer program product containing instructions. When the computer is executed by the instructions, the computer is enabled to perform the method of the above-mentioned method embodiment.
[0409] According to another aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method of the above-described method embodiment.
[0410] In other words, when implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0411] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0412] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0413] Modules described as separate components may or may not be physically separate, and components displayed as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected based on actual needs to achieve the purpose of the present embodiment. For example, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module.
[0414] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A point cloud decoding method, executed by a processor, characterized in that: The method comprises: Determine decoding grouping information according to target features of the target point cloud, wherein the target features include one or more of the following information: point cloud characteristic information, bounding box information, decoding parameter information, and target parameter information; Performing grouping processing on the first reconstructed position points according to the decoded grouping information to obtain a plurality of groups, wherein the first reconstructed position points are obtained by decoding the geometric code stream of the target point cloud based on a decoder; Decode the attributes of each group.
2. The method according to claim 1, characterized in that The point cloud characteristic information includes one or more of the following information: point cloud type, point cloud density, point cloud space occupancy, resolution, and point cloud point count.
3. The method according to claim 1 or 2, characterized in that: The bounding box information includes one or more of the following information: the lengths and directions of three sides of the bounding box of the target point cloud, and length statistics of the three sides; Among them, the length statistics of the three sides include one or more of the following information: the mean length of the three sides, the median length, the mode length, the maximum length of the three sides, the minimum length of the three sides, a linear combination of the lengths of at least two sides, a linear combination between the length of at least one side and a constant, the volume of the bounding box and the surface area of the bounding box.
4. The method according to any one of claims 1 to 3, characterized in that: The decoding parameter information includes one or more of the following information: maximum transform order of attribute prediction transform, geometric output bit depth, attribute output bit depth, geometric quantization parameter, attribute quantization parameter, transform method, prediction method and sorting method.
5. The method according to any one of claims 1 to 4, characterized in that: The step of determining the decoding grouping information according to the target features of the target point cloud includes: Mapping at least one of the target features through an objective function to obtain at least one mapping value; The decoding group information is determined based on a combination of the at least one mapping value and a constant value; or, the decoding group information is determined based on a combination between the mapping values; or, the decoding group information is determined based on a combination between the at least one mapping value and an unmapped target feature.
6. The method according to claim 5, characterized in that The method further comprises: When the objective function includes a logarithmic process, a rounding calculation is performed after the logarithmic process; or, when the objective function includes a division process, a rounding calculation is performed after the division process.
7. The method according to any one of claims 1 to 6, characterized in that: The initial grouping step in the decoded grouping information, then determining the decoded grouping information according to the target features of the target point cloud includes: Determine the target fixed value according to at least one of the point cloud characteristic information, the bounding box information and the decoding parameter information in the target feature; or determine the target fixed value according to a preset identifier in the target parameter information, wherein different preset identifiers correspond to different fixed values; or determine the first default value in the target parameter information as the target fixed value; The value of the initial grouping step length in the decoding grouping information is determined according to the integer multiples of the target fixed value.
8. The method according to claim 7, characterized in that In the case where the target feature is of a point cloud type, determining the target fixed value according to at least one of point cloud characteristic information, bounding box information, and decoding parameter information in the target feature includes: When the point cloud type is a dense type or a human vision point cloud, the target fixed value is determined to be m; or, when the point cloud type is a sparse type or a machine vision point cloud, the target fixed value is determined to be n, where both m and n are positive integers; In the case where the target feature is bounding box information, determining the target fixed value according to at least one of point cloud characteristic information, bounding box information, and decoding parameter information in the target feature includes: When the correlation of the three sides of the bounding box meets the preset conditions, the target fixed value is determined to be p; or, when the correlation of the three sides of the bounding box does not meet the preset conditions, the target fixed value is determined to be q, wherein both p and q are positive integers.
9. The method according to any one of claims 1 to 8, characterized in that The initial grouping step in the decoded grouping information, then determining the decoded grouping information according to the target features of the target point cloud includes: Determine an intermediate quantity for calculating an initial grouping step length according to at least one of the point cloud characteristic information, the bounding box information, and the decoding parameter information in the target feature; or, determine a fixed value corresponding to a preset identifier in the target parameter information as the intermediate quantity, wherein different preset identifiers correspond to different fixed values; or, determine a second default value in the target parameter information as the intermediate quantity; The initial grouping step length is determined according to the intermediate quantity to obtain the decoded grouping information.
10. The method according to any one of claims 1 to 9, characterized in that The initial grouping step in the decoded grouping information, then determining the decoded grouping information according to the target features of the target point cloud includes: Determine at least one of an upper limit value and a lower limit value of an initial grouping step length in the decoded grouping information according to at least one of point cloud characteristic information, bounding box information, and decoding parameter information in the target feature; or, Determine the third default value in the target parameter information as the upper limit of the intermediate quantity, and determine the fourth default value in the target parameter information as the lower limit of the intermediate quantity, wherein the intermediate quantity is used to determine the initial grouping step length; or, The fixed value corresponding to the first preset identifier in the target parameter information is determined as the upper limit value of the initial grouping step or the intermediate quantity, and the fixed value corresponding to the second preset identifier in the target parameter information is determined as the lower limit value of the initial grouping step or the intermediate quantity, wherein different preset identifiers correspond to different fixed values.
11. A point cloud encoding method, characterized in that: Applied to an encoder, the method comprises: Determine encoding grouping information according to target features of the target point cloud, wherein the target features include one or more of the following information related to a code stream encoding the target point cloud or defaulted by the encoding end: point cloud characteristic information, bounding box information, encoding parameter information, and target parameter information; The second reconstruction position points are grouped according to the coding grouping information to obtain a plurality of groups, wherein the second reconstruction position points are obtained by decoding the geometric code stream of the target point cloud by the encoder; Encode the attributes of each group.
12. The method according to claim 11, characterized in that The point cloud characteristic information includes one or more of the following information: point cloud type, point cloud density, point cloud space occupancy, resolution, and point cloud point count; The bounding box information includes one or more of the following information: the lengths and directions of three sides of the bounding box of the target point cloud, and length statistics of the three sides; The length statistics of the three sides include one or more of the following information: the mean value of the lengths of the three sides, the median value of the lengths, the mode value of the lengths, the maximum value of the lengths of the three sides, the minimum value of the lengths of the three sides, a linear combination of the lengths of at least two sides, a linear combination of the length of at least one side and a constant, the volume of the bounding box, and the surface area of the bounding box; The encoding parameter information includes one or more of the following information: maximum transform order of attribute prediction transform, geometric output bit depth, attribute output bit depth, geometric quantization parameter, attribute quantization parameter, transform method, prediction method and sorting method.
13. The method according to claim 11 or 12, characterized in that: The step of determining the coding grouping information according to the target features of the target point cloud includes: Mapping at least one of the target features through an objective function to obtain at least one mapping value; The coding grouping information is determined based on a combination of the at least one mapping value and a constant value; or, the coding grouping information is determined based on a combination between the mapping values; or, the coding grouping information is determined based on a combination between the at least one mapping value and an unmapped target feature.
14. The method according to any one of claims 11 to 13, characterized in that The initial grouping step in the decoded grouping information, then determining the encoded grouping information according to the target features of the target point cloud, comprises: Determine the target fixed value according to at least one of the point cloud characteristic information, bounding box information and encoding parameter information in the target feature; or, determine the target fixed value according to a preset identifier in the target parameter information, wherein different preset identifiers correspond to different fixed values; or, determine the default value in the target parameter information as the target fixed value; The value of the initial grouping step length in the coding grouping information is determined according to the integer multiples of the target fixed value.
15. The method according to claim 14, characterized in that In the case where the target feature is of a point cloud type, determining the target fixed value according to at least one of point cloud characteristic information, bounding box information, and encoding parameter information in the target feature includes: When the point cloud type is a dense type or a human visual point cloud, determining that the target fixed value is m; or, When the point cloud type is a sparse type or a machine vision point cloud, determining the target fixed value n, wherein m and n are both positive integers; In the case where the target feature is bounding box information, determining the target fixed value according to at least one of point cloud characteristic information, bounding box information, and encoding parameter information in the target feature includes: When the correlation of the three sides of the bounding box meets the preset condition, the target fixed value is determined to be p; or, When the correlation of the three sides of the bounding box does not satisfy a preset condition, the target fixed value is determined to be q, wherein both p and q are positive integers.
16. The method according to any one of claims 11 to 15, characterized in that The initial grouping step in the decoded grouping information, then determining the encoded grouping information according to the target features of the target point cloud, comprises: Determine an intermediate quantity for calculating an initial grouping step length according to at least one of the point cloud characteristic information, the bounding box information, and the encoding parameter information in the target feature; or, determine a fixed value corresponding to a preset identifier in the target parameter information as the intermediate quantity, wherein different preset identifiers correspond to different fixed values; or, determine a second default value in the target parameter information as the intermediate quantity; The initial grouping step length is determined according to the intermediate quantity to obtain the coding grouping information.
17. The method according to any one of claims 11 to 16, characterized in that The initial grouping step in the decoded grouping information, then determining the encoded grouping information according to the target features of the target point cloud, comprises: Determine at least one of an upper limit value and a lower limit value of an initial grouping step length in the decoding grouping information according to at least one of point cloud characteristic information, bounding box information, and encoding parameter information in the target feature; or, Determine the third default value in the target parameter information as the upper limit of the intermediate quantity, and determine the fourth default value in the target parameter information as the lower limit of the intermediate quantity, wherein the intermediate quantity is used to determine the initial grouping step length; or, The fixed value corresponding to the first preset identifier in the target parameter information is determined as the upper limit value of the initial grouping step or the intermediate quantity, and the fixed value corresponding to the second preset identifier in the target parameter information is determined as the lower limit value of the initial grouping step or the intermediate quantity, wherein different preset identifiers correspond to different fixed values.
18. A decoder, characterized in that: The decoder comprises: A first determination module is used to determine decoding grouping information according to target features of the target point cloud, wherein the target features include one or more of the following information: point cloud characteristic information, bounding box information, decoding parameter information, and target parameter information; A decoding grouping module, configured to group the first reconstruction position points according to the decoding grouping information to obtain a plurality of groups, wherein the first reconstruction position points are obtained by decoding the geometric code stream of the target point cloud by the decoder; The attribute decoding module is used to decode the attributes of each group.
19. An encoder, characterized in that The encoder comprises: A second determination module is used to determine the coding grouping information according to the target features of the target point cloud, wherein the target features include one or more of the following information: point cloud characteristic information, bounding box information, coding parameter information and target parameter information; A coding grouping module, configured to group the first reconstruction position points according to the coding grouping information to obtain a plurality of groups, wherein the first reconstruction position points are obtained by decoding the geometric code stream of the target point cloud by the encoder; The attribute encoding module is used to perform attribute encoding on each group.
20. An electronic device, characterized in that: The electronic device comprises: a processor and a memory; The memory is used to store computer programs; The processor is configured to execute the computer program to implement the method as described in any one of claims 1 to 10 or 11 to 17 above.
21. A computer-readable storage medium, characterized in that: For storing computer programs; The computer program enables a computer to execute a point cloud decoding method as described in any one of claims 1 to 10, or to execute a point cloud encoding method as described in any one of claims 11 to 17.
22. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the method according to any one of claims 1 to 10 or 11 to 17.
Citation Information
Patent Citations
Grouping-based point cloud code stream packaging method and system
CN111435991A
Point cloud attribute coding method, point cloud attribute decoding method and terminal
CN115714859A
Point cloud decoding method, point cloud coding method, decoder, electronic equipment and medium
CN117615136A
Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method
US20230394712A1
Cited By
Solid waste heat value prediction method and system based on laser radar and image coupling
CN121392213A