A point cloud data compression method, device, equipment and storage medium

CN122820868APending Publication Date: 2026-09-25PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611010526.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]有鉴于此,本申请提供一种点云数据的压缩方法、装置、设备及存储介质,一方面,通过引入目标对象所属对象类型的类别语义信息,能够增强点云几何压缩中特征初始化的区分度,有利于解决现有技术中因类别语义缺失而出现的点云数据压缩结果的精准度不足的缺陷;另一方面,在特征编码过程中,通过对几何增强特征矩阵进行全局物理结构依赖关系的建模处理,使得最终的点云数据压缩结果不仅保留了点云数据中必要的几何结构信息,还融入了点云整体的结构语义,使得后续熵编码处理能够在更高的压缩效率下实现精准的几何重建

Benefits of technology

本申请实施例提供的一种点云数据的压缩方法、装置、设备及存储介质,通过将点云数据的几何信息矩阵与对应对象类型的文本特征向量组成多模态输入数据,使点云几何特征在编码初始化阶段即具备类别感知能力,并且在特征编码阶段,通过对几何增强特征矩阵进行全局物理结构依赖关系的建模处理,使得目标编码器能够捕捉点云数据中分属于不同物理结构的空间位置点集合之间的全局依赖关系,从而得到融合有全局物理结构信息的全局几何特征,进而使得最终的点云数据压缩结果不仅保留了点云数据中必要的几何结构信息,还融入了点云整体的结构语义,使得后续熵编码处理能够在更高的压缩效率下实现精准的几何重建。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820868A_ABST
    Figure CN122820868A_ABST
Patent Text Reader

Abstract

The application provides a point cloud data compression method, device and equipment and a storage medium. The compression method forms multi-modal input data by combining a geometry information matrix of point cloud data and a text feature vector of a corresponding object type, so that the point cloud geometry feature has class perception ability in the encoding initialization stage. In the feature encoding stage, the modeling processing of the global physical structure dependency relationship is performed on the geometry enhanced feature matrix, so that the target encoder can capture the global dependency relationship between the spatial position point sets belonging to different physical structures in the point cloud data, thereby obtaining the global geometry feature fused with the global physical structure information, and further making the final point cloud data compression result not only retain the necessary geometry structure information in the point cloud data, but also integrate the structure semantics of the point cloud as a whole, so that the subsequent entropy encoding processing can realize accurate geometry reconstruction at a higher compression efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of three-dimensional point cloud data processing technology, and more specifically, to a method, apparatus, device, and storage medium for compressing point cloud data. Background Technology

[0002] Advanced acquisition devices such as LiDAR can accurately capture the three-dimensional geometric structure of objects or scenes, generating point cloud data composed of the three-dimensional coordinates of massive spatial locations. Because point cloud data is extremely large, it not only requires significant storage resources but also severely restricts data transmission efficiency. Therefore, when transmitting point cloud data, it is usually necessary to first compress the data, send the compressed result to the receiving end, and then restore the received compressed result at the receiving end to obtain the reconstructed point cloud data.

[0003] Currently, existing technologies primarily use neural networks to encode point cloud data into a low-dimensional latent feature vector, and then perform entropy encoding on this latent feature vector to obtain compressed point cloud data. However, in these existing technologies, the neural networks based on deep learning typically only utilize the three-dimensional geometric coordinates of spatial points in the point cloud data, employing a uniform and fixed feature initialization method to encode the input point cloud data. This ignores the semantic differences between different types of point cloud data (equivalent to two different types of objects potentially having similar feature encoding results due to similar geometric structures), resulting in insufficient accuracy in the point cloud data compression results. Summary of the Invention

[0004] In view of this, this application provides a method, apparatus, device, and storage medium for compressing point cloud data. On the one hand, by introducing category semantic information of the object type to which the target object belongs, the discriminativeness of feature initialization in point cloud geometric compression can be enhanced, which helps to solve the defect of insufficient accuracy of point cloud data compression results due to the lack of category semantics in the prior art. On the other hand, in the feature encoding process, by modeling the global physical structure dependency of the geometric enhancement feature matrix, the final point cloud data compression result not only retains the necessary geometric structure information in the point cloud data, but also incorporates the overall structural semantics of the point cloud, so that subsequent entropy encoding processing can achieve accurate geometric reconstruction with higher compression efficiency.

[0005] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings.

[0006] In a first aspect, embodiments of this application provide a method for compressing point cloud data, the compression method comprising: By preprocessing the point cloud data of the target object, a geometric information matrix corresponding to the point cloud data is obtained, and the geometric information matrix and the text feature vector corresponding to the target object are combined to form multimodal input data; wherein, the geometric information matrix is ​​a feature matrix composed of the three-dimensional coordinate information of each spatial location point contained in the point cloud data, and the text feature vector is the text feature vector corresponding to the object type to which the target object belongs. The multimodal input data is input into a pre-trained target encoder. The first encoding module in the target encoder performs feature enhancement processing on the multimodal input data to obtain the geometric enhancement feature matrix corresponding to the geometric information matrix and the target text feature vector corresponding to the text feature vector. The first encoding module includes a multi-scale downsampling module and a target residual network module. The second encoding module in the target encoder models the global physical structure dependency of the geometric enhancement feature matrix to obtain global geometric features that incorporate global physical structure information; wherein, the global physical structure information is used to represent the dependency between sets of spatial location points belonging to different physical structures in the point cloud data; The target text feature vector and the global geometric features are subjected to entropy encoding to obtain the target point cloud compression result.

[0007] Secondly, embodiments of this application provide a point cloud data compression device, the compression device comprising: The input module is used to preprocess the point cloud data of the target object to obtain the geometric information matrix corresponding to the point cloud data, and to combine the geometric information matrix with the text feature vector corresponding to the target object to form multimodal input data; wherein, the geometric information matrix is ​​a feature matrix composed of the three-dimensional coordinate information of each spatial location point contained in the point cloud data, and the text feature vector is the text feature vector corresponding to the object type to which the target object belongs. The feature encoding module is used to input the multimodal input data into a pre-trained target encoder, and through the first encoding module in the target encoder, perform feature enhancement processing on the multimodal input data to obtain the geometric enhancement feature matrix corresponding to the geometric information matrix and the target text feature vector corresponding to the text feature vector; wherein, the first encoding module includes: a multi-scale downsampling module and a target residual network module; The second encoding module in the target encoder models the global physical structure dependency of the geometric enhancement feature matrix to obtain global geometric features that incorporate global physical structure information; wherein, the global physical structure information is used to represent the dependency between sets of spatial location points belonging to different physical structures in the point cloud data; The entropy coding module is used to perform entropy coding processing on the target text feature vector and the global geometric features to obtain the target point cloud compression result.

[0008] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the point cloud data compression method described above.

[0009] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the point cloud data compression method described above.

[0010] The technical solutions provided by the embodiments of this application may include the following beneficial effects: This application provides a method, apparatus, device, and storage medium for compressing point cloud data. By combining the geometric information matrix of the point cloud data with the text feature vector of the corresponding object type to form multimodal input data, the point cloud geometric features possess category awareness capabilities during the encoding initialization stage. Furthermore, during the feature encoding stage, by modeling the global physical structure dependency relationship of the geometric enhancement feature matrix, the target encoder can capture the global dependency relationship between sets of spatial location points belonging to different physical structures in the point cloud data. This results in global geometric features that incorporate global physical structure information, thereby ensuring that the final point cloud data compression result not only retains the necessary geometric structure information in the point cloud data but also incorporates the overall structural semantics of the point cloud. This enables subsequent entropy encoding processing to achieve accurate geometric reconstruction with higher compression efficiency. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating a point cloud data compression method provided in an embodiment of this application is shown. Figure 2 This illustration shows a flowchart of a process for modeling a geometric enhancement feature matrix using a physical token attention module, as provided in an embodiment of this application. Figure 3 This illustration shows a schematic diagram of the structure of a point cloud data compression device provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0014] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0015] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0016] In one embodiment of this application, a point cloud data compression method can run on a terminal device or a server. The terminal device can be a local terminal device. When the point cloud data compression method runs on a server, it can be implemented and executed based on a cloud interaction system, which includes a server and client devices (i.e., terminal devices).

[0017] In an optional implementation, the point cloud data compression method provided in this application embodiment can also be applied to a point cloud data transmission system, wherein the point cloud data transmission system includes a point cloud data compression end (which is also equivalent to the sending end of the point cloud data compression result) and a storage end (which is also equivalent to the receiving end of the point cloud data compression result). On the compression end side, the point cloud data can be compressed to obtain a target point cloud compression result, and the compressed target point cloud compression result is sent to the storage end. On the storage end side, the received target point cloud compression result can be decoded and restored to obtain reconstructed point cloud data for storage.

[0018] To facilitate understanding of the embodiments of this application, a method, apparatus, device, and storage medium for compressing point cloud data provided in the embodiments of this application will be described in detail below.

[0019] Reference Figure 1 As shown, Figure 1 The diagram illustrates a flow chart of a point cloud data compression method provided in an embodiment of this application, wherein the compression method includes steps S101-S104; specifically: S101, by preprocessing the point cloud data of the target object, a geometric information matrix corresponding to the point cloud data is obtained, and the geometric information matrix and the text feature vector corresponding to the target object are combined to form multimodal input data.

[0020] Here, before executing step S101, point cloud data of the target object can be acquired on the compression end side of the point cloud data. Specifically, three-dimensional coordinate information of multiple spatial points located on the surface and inside of the target object can be acquired using three-dimensional acquisition devices such as LiDAR and depth cameras to obtain point cloud data composed of the three-dimensional coordinate information of these spatial points.

[0021] It should be noted that the target object can be an independent entity such as an animal, plant, or object, or it can be a scene space containing one or more entity objects. This application embodiment does not limit the specific object type to which the target object belongs.

[0022] Here, when preprocessing the point cloud data, the specific preprocessing methods may include, but are not limited to, coordinate normalization, scale unification and quantization. The purpose of preprocessing is to ensure the consistency of geometric data (i.e., the three-dimensional coordinate information of points at different spatial locations in the point cloud data).

[0023] Specifically, after preprocessing the point cloud data, a geometric information matrix corresponding to the point cloud data can be obtained; wherein, the geometric information matrix is ​​a feature matrix composed of the three-dimensional coordinate information of each spatial location point contained in the point cloud data.

[0024] For example, taking a point cloud dataset containing N spatial location points, the aforementioned geometric information matrix can be denoted as... Among them, the geometric information matrix Each row in the table corresponds to the three-dimensional spatial coordinates (x, y, z) of a spatial location point.

[0025] Here, the aforementioned text feature vector is the text feature vector corresponding to the object type to which the target object belongs; wherein, on the point cloud data compression side, the category text corresponding to the object type to which the target object belongs can be obtained first. This category of text This is used to explicitly identify the semantic category to which the target object belongs in the semantic dimension (e.g., the target object belongs to a chair, the target object belongs to a car, etc.); then, through a pre-trained text encoder, the text of the above categories is processed. By performing feature encoding, the above-mentioned text categories can be output. Corresponding text feature vector The resulting text feature vector This is the text feature vector corresponding to the target object.

[0026] It should be noted that the above geometric information matrix and the above text feature vectors They do not match in terms of feature dimensions; among them, the geometric information matrix The dimension is the number of spatial location points N×3, where 3 represents that one spatial location point corresponds to one three-dimensional coordinate information; while the above text feature vector It is a single global semantic feature vector.

[0027] Based on this, a geometric information matrix with mismatched dimensions was obtained. and text feature vectors Subsequently, on the compression side of the point cloud data, the aforementioned multimodal input data can be obtained according to the methods shown in steps a1-a2 below, specifically: Step a1: Based on the number of spatial location points contained in the point cloud data, the text feature vector is copied to obtain multiple text feature vectors.

[0028] Here, the number of copies corresponds to the number of spatial location points mentioned above, for example, using the aforementioned geometric information matrix. For example, due to the geometric information matrix The number of spatial location points contained is N, therefore the text feature vector can be... Copy N times to obtain N text feature vectors .

[0029] Step a2: Concatenate and combine the copied multiple text feature vectors with the geometric information matrix to obtain the multimodal input data.

[0030] Here, taking N spatial location points as an example, we can generate N text feature vectors. As the initial text feature matrix At this point, based on the initial text feature matrix With geometric information matrix The dimensions have been matched, therefore the initial text feature matrix can be processed. With geometric information matrix By concatenating and combining the data, we obtain multimodal input data X, which incorporates category semantic information corresponding to the point cloud data; whereby multimodal input data X can be represented as... .

[0031] It should be noted that, because the aforementioned multimodal input data X incorporates category semantic information (i.e., the initial text feature matrix), Therefore, the multimodal input data X has the ability to distinguish categories during the feature initialization stage, so that point cloud data of different categories have different initial feature representations when inputting into the encoder. This effectively overcomes the shortcomings of insufficient feature discrimination caused by relying only on three-dimensional geometric coordinates and using a uniform fixed initialization method in the existing technology, and lays a reliable foundation for subsequent feature encoding and compression processing.

[0032] S102, the multimodal input data is input into the pre-trained target encoder, and the first encoding module in the target encoder performs feature enhancement processing on the multimodal input data to obtain the geometric enhancement feature matrix corresponding to the geometric information matrix and the target text feature vector corresponding to the text feature vector.

[0033] Here, the target encoder is deployed on the compression end of the point cloud data. The target encoder consists of a first encoding module and a second encoding module. The first encoding module includes a multi-scale downsampling module and a target residual network module. The multi-scale downsampling module consists of multiple downsampling modules for extracting features at different scales. Each downsampling module is followed by a target residual network module, which can specifically be an IRN (Inception-Residual Network) module.

[0034] Specifically, in the target encoder, the aforementioned first encoding module can be used to perform feature enhancement processing on the input multimodal input data. In the first encoding module, the aforementioned multi-scale downsampling module is used to perform progressive downsampling processing on the input multimodal input data through multiple downsampling modules. In each downsampling step, the number of feature points in the aforementioned multimodal input data is gradually reduced, while the receptive field represented by each feature point is gradually expanded, enabling different levels to capture multi-scale structural information from local details to a larger neighborhood.

[0035] For example, each downsampling module can be composed of "convolutional layer → ReLU activation function → downsampling convolutional layer → ReLU activation function". By cascading multiple downsampling modules, key features can be preserved while reducing the data scale.

[0036] Specifically, after each downsampling process, the multimodal input data of the current downsampling process can be enhanced by the IRN module (i.e., the target residual network module) connected in series after the downsampling module. The multimodal input data after feature enhancement is then input into the next downsampling module to continue the downsampling process at the next scale.

[0037] It should be noted that, since the multimodal input data is the initial text feature matrix mentioned above... (i.e., N text feature vectors) ) and the geometric information matrix mentioned above The result of the concatenation and combination, therefore, when executing S102, is also equivalent to processing the initial text feature matrix through the first encoding module mentioned above. With the above geometric information matrix Feature enhancement processes are performed separately to obtain the initial text feature matrix described above. The feature enhancement processing result is used as the feature vector of the target text mentioned above (denoted as...). The above geometric information matrix is ​​obtained. The feature enhancement processing result is used as the geometric enhancement feature matrix (denoted as C) mentioned above.

[0038] S103, the second encoding module in the target encoder performs global physical structure dependency modeling on the geometric enhancement feature matrix to obtain global geometric features that incorporate global physical structure information.

[0039] Here, the second encoding module can also be called the physical token attention module. The core processing idea of ​​the physical token attention module is as follows: For the geometric enhancement feature matrix C that belongs to geometric information, the point-level features (i.e., the three-dimensional coordinate information of spatial location points) in the geometric enhancement feature matrix C are aggregated into a small number of physical tokens with physical structure representation meaning in a learnable way (a physical token is used to represent a set of spatial location points in point cloud data that belong to the same physical structure). Thus, at the token level, the self-attention mechanism is used to capture the global physical structure dependency (equivalent to the association between a spatial location point and different physical tokens). Then, the token-level features containing the above global physical structure dependency are reprojected into point-level features, thereby obtaining global geometric features that are integrated with global physical structure information.

[0040] It should be noted that the self-attention mechanism in traditional point cloud compression methods (based on Transformer methods) usually calculates the relationship between points in a local neighborhood. Due to the limitation of the local receptive field, it is difficult to effectively capture the long-range dependency between spatially distant points in the point cloud that are physically interdependent. (For example, taking a chair as the target object, in the point cloud data of the chair, the spatially distant points belonging to different chair legs are far apart in three-dimensional space, but from the perspective of physical structure or physical state, different chair legs belong to the support structure of the chair. Therefore, the spatially distant points belonging to different chair legs may belong to the same physical structure.)

[0041] In this embodiment of the application, the aforementioned global physical structure information is used to represent the dependency relationship between sets of spatial location points belonging to different physical structures in the point cloud data. That is, by executing the aforementioned step S103, this embodiment of the application can effectively solve the problem that the traditional point cloud compression method is difficult to effectively capture the aforementioned global physical structure dependency relationship, thereby obtaining global geometric features that incorporate global physical structure information, which is beneficial to improving the comprehensiveness of feature expression.

[0042] Here, when performing step S103, the above-mentioned global geometric features can be obtained according to the method shown in steps b1-b5 below: Step b1: Perform a linear transformation on the geometric enhancement feature matrix to obtain the target enhancement feature matrix.

[0043] Specifically, Figure 2 This illustration shows a flowchart of a process for modeling a geometric enhancement feature matrix using a physical token attention module, as provided in an embodiment of this application; wherein, as... Figure 2As shown, the physical token attention module (also known as the second encoding module) includes: two linear transformation modules, two matrix multiplication operation modules, one self-attention module, and one Softmax function.

[0044] Here, after inputting the geometric enhancement feature matrix C into the physical token attention module, the input geometric enhancement feature matrix C can be linearly transformed through a linear layer contained in a linear transformation module to obtain the linear transformation result as the target enhancement feature matrix C1; where the role of the linear transformation is to perform preliminary feature space mapping on the input geometric enhancement feature matrix C.

[0045] It should be noted that the above linear transformation module may contain one or more linear layers. The specific number of linear layers actually included is not limited in the embodiments of this application.

[0046] Step b2: Normalize the target enhancement feature matrix to obtain the token weight matrix.

[0047] Here, as Figure 2 As shown, the target enhancement feature matrix C1 obtained after linear transformation is normalized by the Softmax function to obtain the token weight matrix W. The token weight matrix W is used to characterize the degree of association between each spatial location point and different physical tokens, and different physical tokens correspond to different physical structures (equivalent to a physical token corresponding to a physical structure that a spatial location point may belong to in the point cloud data).

[0048] It should be noted that the physical tokens contained in the aforementioned token weight matrix W can be obtained by clustering the three-dimensional coordinate information of all spatial location points contained in the aforementioned target enhancement feature matrix C1. That is, multiple clusters can be determined through clustering, and each cluster represents a set of spatial location points in the point cloud data that belong to the same physical structure (equivalent to a physical token).

[0049] Specifically, the token weight matrix Where N is the number of spatial location points contained in the target enhancement feature matrix C1, and M is the number of physical tokens obtained by clustering the target enhancement feature matrix C1.

[0050] Step b3: Perform matrix multiplication on the token weight matrix and the target enhancement feature matrix to obtain the physical token matrix.

[0051] Here, as Figure 2As shown, in the physical token attention module (i.e. the second encoding module), the token weight matrix W and the target augmentation feature matrix C1 are input into a matrix multiplication operation module. The matrix multiplication operation module can perform matrix multiplication operation on the input token weight matrix W and the target augmentation feature matrix C1 to obtain the matrix multiplication operation result as the physical token matrix P (which is equivalent to the aggregation of the target augmentation feature matrix C1 that implements point-level features to the physical token matrix P that implements token-level features).

[0052] An exemplary illustration, the physical token matrix ;in, C1 is the transpose of the token weight matrix W, and C2 is the target enhancement feature matrix.

[0053] Step b4: Apply a self-attention mechanism to the physical token matrix to update the global association between tokens, and obtain the updated physical token matrix.

[0054] Here, when executing step b4, the physical token matrix can be processed through three independent linear layers. Three independent linear transformations are performed to obtain the linear transformation results of the three linear layers, which are respectively used as the query matrix Q, key matrix K, and value matrix V in the self-attention mechanism (equivalent to query matrix Q = key matrix K = value matrix V = physical token matrix). (The result of the linear transformation).

[0055] Specifically, after obtaining the query matrix Q, key matrix K, and value matrix V, the updated physical token matrix P1 can be calculated using the self-attention calculation method shown in Formula 1 below: Formula 1; Where P1 is the updated physical token matrix; Q is the query matrix mentioned above; It is the transpose of the aforementioned key matrix K; V is the value matrix mentioned above; It is the vector dimension of the key vector k in the aforementioned key matrix K.

[0056] Step b5: Perform matrix multiplication on the token weight matrix and the updated physical token matrix to obtain the target geometric feature matrix, and perform a linear transformation on the target geometric feature matrix to obtain the global geometric features.

[0057] Here, as Figure 2As shown, in the physical token attention module (i.e. the second encoding module), the token weight matrix W and the updated physical token matrix P1 are input into a matrix multiplication operation module. The matrix multiplication operation module can perform matrix multiplication operation on the input token weight matrix W and the updated physical token matrix P1 to obtain the matrix multiplication operation result as the target geometric feature matrix C2 (which is equivalent to reprojecting the physical token matrix P1 with token-level features into the target geometric feature matrix C2 with point-level features).

[0058] Specifically, such as Figure 2 As shown, after obtaining the target geometric feature matrix C2, it is input into another linear transformation module. The linear transformation module performs a linear transformation on the target geometric feature matrix C2 through a linear layer, yielding the linear transformation result as the global geometric feature. .

[0059] S104, the target text feature vector and the global geometric features are subjected to entropy encoding processing to obtain the target point cloud compression result.

[0060] Here, after obtaining the target text feature vector and global geometric features Then, based on the target text feature vector With global geometric features Feature information belonging to different modalities can be encoded using different methods for the aforementioned target text feature vectors. and the aforementioned global geometric features Each is entropy encoded separately.

[0061] Specifically, global geometric features can be processed using an octree encoder. Lossless encoding is performed to obtain global geometric features. The corresponding geometric feature encoding results.

[0062] Specifically, the target text feature vector can be analyzed using a total decomposition density model. Entropy encoding is performed to obtain the target text feature vector. The corresponding text feature encoding result can then be used as the target point cloud compression result on the point cloud data compression side, along with the aforementioned geometric feature encoding result. The target point cloud compression result can then be sent to the point cloud data storage side.

[0063] Here, on the storage side of the point cloud data, after receiving the above-mentioned target point cloud compression result, the above-mentioned target point cloud compression result can be decoded and restored according to the method shown in steps c1-c2 below to obtain the reconstructed point cloud data corresponding to the target object. Specifically: Step c1: Perform entropy decoding processing on the target point cloud compression result to obtain the entropy decoding result corresponding to the target point cloud compression result.

[0064] Here, on the storage side of the point cloud data, after receiving the target point cloud compression result sent by the compression end, the target point cloud compression result first needs to be entropy decoded. Since the target point cloud compression result consists of two parts, geometric feature encoding result and text feature encoding result, and the two are encoded using different encoding methods at the compression end, different decoding methods need to be used at the decoding end to perform entropy decoding respectively.

[0065] Specifically, regarding the geometric feature encoding results, since they are obtained through lossless encoding by an octree encoder at the compression end, an octree decoder can be used on the storage side to perform lossless decoding on the geometric feature encoding results in the target point cloud compression results, and decode and restore them to obtain the global geometric features.

[0066] Specifically, for the text feature encoding results, since they are obtained by entropy encoding through a full decomposition density model at the compression end, on the storage side, a corresponding full decomposition density model can be used to entropy decode the text feature encoding results in the target point cloud compression results, and decode and restore the target text feature vector.

[0067] It should be noted that the global geometric features and target text feature vectors obtained by decoding and restoration are the entropy decoding results corresponding to the target point cloud compression results.

[0068] Step c2: Input the entropy decoding result into the target decoder that matches the target encoder, and decode the entropy decoding result through the target decoder to obtain the reconstructed point cloud data corresponding to the target object.

[0069] Here, the target decoder is deployed on the storage side of the point cloud data. The target decoder and the target encoder at the compression end have a symmetrical network structure, used to gradually reconstruct the point cloud data from the features recovered after entropy decoding. The target decoder includes a first decoding module and a second decoding module, which are functionally symmetrically mapped to the first and second encoding modules in the target encoder (equivalent to the first decoding module matching the aforementioned first encoding module, and the second decoding module matching the aforementioned second encoding module).

[0070] Specifically, the second decoding module is structurally symmetrical with the second encoding module of the encoding end. It is used to further restore and refine the global physical structure information in the features, so that the reconstructed point cloud can retain the overall physical structure features of the original point cloud.

[0071] It should be noted that the global geometric features and target text feature vectors included in the entropy decoding results above can be concatenated and combined along the channel dimension before being input into the target decoder, forming multimodal decoding input data corresponding to the multimodal input data on the compression side. Then, this multimodal decoding input data is input into the target decoder, and after being processed by the second and first decoding modules, it undergoes feature classification through convolutional layers to identify the actual voxels occupied, ultimately yielding the reconstructed point cloud data corresponding to the target object.

[0072] Specifically, the first decoding module includes a multi-scale upsampling module and a target residual network module (i.e., an IRN model). The multi-scale upsampling module consists of multiple upsampling modules for extracting features at different scales. Each upsampling module is followed by a target residual network module, which can be an IRN (Inception-Residual Network) module.

[0073] It should be noted that the number of upsampling modules in the multi-scale upsampling module is the same as the number of downsampling modules in the aforementioned multi-scale downsampling module. The aforementioned multi-scale upsampling module is used to perform progressive upsampling processing on the decoded data output by the aforementioned second decoding module through multiple upsampling modules.

[0074] For example, each upsampling module can be composed of a combination of "upsampling convolutional layer → ReLU activation function → convolutional layer → ReLU activation function". By cascading multiple upsampling modules, the geometric structure of the point cloud at different levels of abstraction is gradually recovered. After each upsampling operation, the upsampled feature data is further enhanced by a target residual network module (e.g., an IRN module) to ensure that the reconstructed point cloud data has rich feature representation capabilities and detail fidelity.

[0075] In this embodiment of the application, the target encoder and the target decoder can be trained using the method shown in steps d1-d7 below, specifically: Step d1: Based on the sample point cloud data of the sample object and the object type to which the sample object belongs, obtain the sample geometric information matrix and sample text feature vector corresponding to the sample object, and combine the sample geometric information matrix and the sample text feature vector to form sample multimodal input data.

[0076] Here, the specific implementation of step d1 is the same as the specific implementation of step S101 mentioned above, and the repeated parts will not be described again.

[0077] Step d2: Input the sample multimodal input data into the target encoder to be trained. Through the first encoding module in the target encoder to be trained, perform feature enhancement processing on the sample multimodal input data to obtain the sample geometric enhancement feature matrix and the sample target text feature vector respectively.

[0078] Here, the specific implementation of step d2 is the same as the specific implementation of step S102 mentioned above, and the repeated parts will not be described again.

[0079] Step d3: The global physical structure dependency of the sample geometric enhancement feature matrix is ​​modeled by the second encoding module in the target encoder to be trained, and the global geometric features of the sample are obtained.

[0080] Here, the specific implementation of step d3 is the same as the specific implementation of step S103 mentioned above, and the repeated parts will not be described again.

[0081] Step d4: Perform entropy encoding on the target text feature vector and the global geometric features of the sample to obtain the sample point cloud compression result.

[0082] Here, the specific implementation of step d4 is the same as the specific implementation of step S104 mentioned above, and the repeated parts will not be described again here.

[0083] Step d5: Perform entropy decoding processing on the sample point cloud compression result to obtain the sample entropy decoding result corresponding to the sample point cloud compression result.

[0084] Here, the specific implementation of step d5 is the same as that of step c1 above, and the repetitions will not be repeated here.

[0085] Step d6: Input the sample entropy decoding result into the target decoder to be trained, and use the target decoder to decode the sample entropy decoding result to obtain the sample reconstruction point cloud data corresponding to the sample object.

[0086] Here, the specific implementation of step d6 is the same as that of step c2 above, and the repetitions will not be repeated here.

[0087] Step d7: Based on the loss between the reconstructed point cloud data and the sample point cloud data, adjust the model parameters of the target encoder and the target decoder to be trained until the preset convergence condition is met, and obtain the trained target encoder and target decoder.

[0088] Here, the sample point cloud data is the real point cloud data obtained by actually collecting the three-dimensional coordinate information of each spatial location point in the sample object, while the above-mentioned sample reconstructed point cloud data is the reconstructed point cloud data obtained by processing the above-mentioned target encoder and target decoder. Therefore, based on the loss between the two, the model parameters of the above-mentioned target encoder and target decoder can be adjusted until the calculated loss reaches the minimum or the number of iterations (which is also equivalent to the number of model parameter adjustments) reaches a preset threshold. Then, it can be determined that the preset convergence condition is met, and the target encoder and target decoder (i.e., the trained target encoder and target decoder) including the adjusted model parameters are obtained.

[0089] Based on the point cloud data compression method provided in the embodiments of this application, by combining the geometric information matrix of the point cloud data with the text feature vector of the corresponding object type to form multimodal input data, the point cloud geometric features have category awareness capabilities in the encoding initialization stage. Furthermore, in the feature encoding stage, by modeling the global physical structure dependency relationship of the geometric enhancement feature matrix, the target encoder can capture the global dependency relationship between sets of spatial location points belonging to different physical structures in the point cloud data, thereby obtaining global geometric features that integrate global physical structure information. As a result, the final point cloud data compression result not only retains the necessary geometric structure information in the point cloud data, but also incorporates the overall structural semantics of the point cloud, enabling subsequent entropy encoding processing to achieve accurate geometric reconstruction with higher compression efficiency.

[0090] Based on the same inventive concept, this application also provides a point cloud data compression device corresponding to the above-mentioned point cloud data compression method. Since the principle of the point cloud data compression device in the embodiments of this application is similar to that of the above-mentioned point cloud data compression method in the embodiments of this application, the implementation of the point cloud data compression device can refer to the implementation of the above-mentioned point cloud data compression method, and the repeated parts will not be described again.

[0091] Reference Figure 3 As shown, Figure 3 A schematic diagram of a point cloud data compression device provided in an embodiment of this application is shown, wherein the compression device includes: The input module 301 is used to preprocess the point cloud data of the target object to obtain the geometric information matrix corresponding to the point cloud data, and to form multimodal input data by combining the geometric information matrix with the text feature vector corresponding to the target object; wherein, the geometric information matrix is ​​a feature matrix composed of the three-dimensional coordinate information of each spatial location point contained in the point cloud data, and the text feature vector is the text feature vector corresponding to the object type to which the target object belongs. The feature encoding module 302 is used to input the multimodal input data into a pre-trained target encoder, and perform feature enhancement processing on the multimodal input data through the first encoding module in the target encoder to obtain the geometric enhancement feature matrix corresponding to the geometric information matrix and the target text feature vector corresponding to the text feature vector; wherein, the first encoding module includes: a multi-scale downsampling module and a target residual network module; The second encoding module in the target encoder models the global physical structure dependency of the geometric enhancement feature matrix to obtain global geometric features that incorporate global physical structure information; wherein, the global physical structure information is used to represent the dependency between sets of spatial location points belonging to different physical structures in the point cloud data; The entropy coding module 303 is used to perform entropy coding processing on the target text feature vector and the global geometric features to obtain the target point cloud compression result.

[0092] In an optional implementation, when the geometric information matrix and the text feature vector corresponding to the target object are combined to form multimodal input data, the input module 301 is used to: Based on the number of spatial location points contained in the point cloud data, the text feature vector is copied to obtain multiple text feature vectors; wherein the number of copies is matched with the number of spatial location points. The copied text feature vectors are concatenated and combined with the geometric information matrix to obtain the multimodal input data.

[0093] In an optional implementation, when the geometric enhancement feature matrix is ​​modeled for global physical structure dependencies by the second encoding module in the target encoder, the feature encoding module 302 is used to: A linear transformation is performed on the geometric enhancement feature matrix to obtain the target enhancement feature matrix; The target enhancement feature matrix is ​​normalized to obtain a token weight matrix; wherein, the token weight matrix is ​​used to characterize the degree of association between each spatial location point and different physical tokens, and the different physical tokens correspond to the different physical structures respectively; Perform matrix multiplication on the token weight matrix and the target enhancement feature matrix to obtain the physical token matrix; The physical token matrix is ​​updated by applying a self-attention mechanism to perform global association updates between tokens, resulting in an updated physical token matrix. Matrix multiplication is performed on the token weight matrix and the updated physical token matrix to obtain the target geometric feature matrix, and a linear transformation is performed on the target geometric feature matrix to obtain the global geometric features.

[0094] In an optional implementation, when performing entropy encoding processing on the target text feature vector and the global geometric features, the entropy encoding module 303 is used to: The global geometric features are losslessly encoded using an octree encoder to obtain the geometric feature encoding results corresponding to the global geometric features; The target text feature vector is entropy encoded by a total decomposition density model to obtain the text feature encoding result corresponding to the target text feature vector; The geometric feature encoding result and the text feature encoding result are used as the target point cloud compression result.

[0095] In an optional embodiment, the compression device further includes: an entropy decoding module and a feature decoding module, wherein: The entropy decoding module is used to perform entropy decoding processing on the target point cloud compression result to obtain the entropy decoding result corresponding to the target point cloud compression result; The feature decoding module is used to input the entropy decoding result into the target decoder that matches the target encoder, and to decode the entropy decoding result through the target decoder to obtain the reconstructed point cloud data corresponding to the target object.

[0096] In an optional embodiment, the compression device further includes a training module, wherein the training module is used to train the target encoder and the target decoder using the following method: Based on the sample point cloud data of the sample object and the object type to which the sample object belongs, the sample geometric information matrix and sample text feature vector corresponding to the sample object are obtained respectively, and the sample geometric information matrix and sample text feature vector are combined to form sample multimodal input data; The sample multimodal input data is input into the target encoder to be trained. The first encoding module in the target encoder to be trained performs feature enhancement processing on the sample multimodal input data to obtain the sample geometric enhancement feature matrix and the sample target text feature vector, respectively. The global physical structure dependency of the sample geometric enhancement feature matrix is ​​modeled by the second encoding module in the target encoder to be trained, thereby obtaining the global geometric features of the sample. Entropy encoding is performed on the target text feature vector and the global geometric features of the sample to obtain the sample point cloud compression result; The sample point cloud compression result is subjected to entropy decoding processing to obtain the sample entropy decoding result corresponding to the sample point cloud compression result; The sample entropy decoding result is input into the target decoder to be trained, and the target decoder to be trained decodes the sample entropy decoding result to obtain the sample reconstruction point cloud data corresponding to the sample object. Based on the loss between the reconstructed point cloud data and the sample point cloud data, the model parameters of the target encoder and the target decoder to be trained are adjusted until the preset convergence condition is met, and the trained target encoder and target decoder are obtained.

[0097] In one optional implementation, the target decoder includes: a first decoding module matched with the first encoding module and a second decoding module matched with the second encoding module; wherein the first decoding module includes: a multi-scale upsampling module and a target residual network module.

[0098] like Figure 4 As shown, this application provides an electronic device 400 for executing the point cloud data compression method of this application. The device includes a memory 401, a processor 402, and a computer program stored in the memory 401 and executable on the processor 402. The memory 401 and the processor 402 are connected via a bus for communication. When the processor 402 executes the computer program, it implements the steps of the point cloud data compression method described above.

[0099] Specifically, the memory 401 and processor 402 mentioned above can be general-purpose memory and processor, without any specific limitations. When the processor 402 runs the computer program stored in the memory 401, it can execute the point cloud data compression method mentioned above.

[0100] Corresponding to the point cloud data compression method in this application, this application embodiment also provides a computer-readable storage medium storing a computer program, which is executed by a processor to perform the steps of the point cloud data compression method described above.

[0101] Specifically, the storage medium can be a general-purpose storage medium, such as a portable disk or hard disk. When the computer program on the storage medium is run, it can execute the point cloud data compression method described above.

[0102] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.

[0103] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0104] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0105] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0106] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0107] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.

Claims

1. A method for compressing point cloud data, characterized in that, The compression method includes: By preprocessing the point cloud data of the target object, a geometric information matrix corresponding to the point cloud data is obtained, and the geometric information matrix and the text feature vector corresponding to the target object are combined to form multimodal input data; wherein, the geometric information matrix is ​​a feature matrix composed of the three-dimensional coordinate information of each spatial location point contained in the point cloud data, and the text feature vector is the text feature vector corresponding to the object type to which the target object belongs. The multimodal input data is input into a pre-trained target encoder. The first encoding module in the target encoder performs feature enhancement processing on the multimodal input data to obtain the geometric enhancement feature matrix corresponding to the geometric information matrix and the target text feature vector corresponding to the text feature vector. The first encoding module includes a multi-scale downsampling module and a target residual network module. The second encoding module in the target encoder models the global physical structure dependency of the geometric enhancement feature matrix to obtain global geometric features that incorporate global physical structure information; wherein, the global physical structure information is used to represent the dependency between sets of spatial location points belonging to different physical structures in the point cloud data; The target text feature vector and the global geometric features are subjected to entropy encoding to obtain the target point cloud compression result.

2. The compression method according to claim 1, characterized in that, The step of assembling the geometric information matrix and the text feature vector corresponding to the target object into multimodal input data includes: Based on the number of spatial location points contained in the point cloud data, the text feature vector is copied to obtain multiple text feature vectors; wherein the number of copies is matched with the number of spatial location points. The copied text feature vectors are concatenated and combined with the geometric information matrix to obtain the multimodal input data.

3. The compression method according to claim 1, characterized in that, The step of modeling the global physical structure dependency of the geometric enhancement feature matrix through the second encoding module in the target encoder includes: A linear transformation is performed on the geometric enhancement feature matrix to obtain the target enhancement feature matrix; The target enhancement feature matrix is ​​normalized to obtain a token weight matrix; wherein, the token weight matrix is ​​used to characterize the degree of association between each spatial location point and different physical tokens, and the different physical tokens correspond to the different physical structures respectively; Perform matrix multiplication on the token weight matrix and the target enhancement feature matrix to obtain the physical token matrix; The physical token matrix is ​​updated by applying a self-attention mechanism to perform global association updates between tokens, resulting in an updated physical token matrix. Matrix multiplication is performed on the token weight matrix and the updated physical token matrix to obtain the target geometric feature matrix, and a linear transformation is performed on the target geometric feature matrix to obtain the global geometric features.

4. The compression method according to claim 1, characterized in that, The entropy encoding process for the target text feature vector and the global geometric features includes: The global geometric features are losslessly encoded using an octree encoder to obtain the geometric feature encoding results corresponding to the global geometric features; The target text feature vector is entropy encoded by a total decomposition density model to obtain the text feature encoding result corresponding to the target text feature vector; The geometric feature encoding result and the text feature encoding result are used as the target point cloud compression result.

5. The compression method according to claim 1, characterized in that, The compression method further includes: The target point cloud compression result is subjected to entropy decoding processing to obtain the entropy decoding result corresponding to the target point cloud compression result; The entropy decoding result is input into the target decoder that matches the target encoder. The target decoder then decodes the entropy decoding result to obtain the reconstructed point cloud data corresponding to the target object.

6. The compression method according to claim 5, characterized in that, The target encoder and the target decoder are trained using the following method: Based on the sample point cloud data of the sample object and the object type to which the sample object belongs, the sample geometric information matrix and sample text feature vector corresponding to the sample object are obtained respectively, and the sample geometric information matrix and sample text feature vector are combined to form sample multimodal input data; The sample multimodal input data is input into the target encoder to be trained. The first encoding module in the target encoder to be trained performs feature enhancement processing on the sample multimodal input data to obtain the sample geometric enhancement feature matrix and the sample target text feature vector, respectively. The global physical structure dependency of the sample geometric enhancement feature matrix is ​​modeled by the second encoding module in the target encoder to be trained, thereby obtaining the global geometric features of the sample. Entropy encoding is performed on the target text feature vector and the global geometric features of the sample to obtain the sample point cloud compression result; The sample point cloud compression result is subjected to entropy decoding processing to obtain the sample entropy decoding result corresponding to the sample point cloud compression result; The sample entropy decoding result is input into the target decoder to be trained, and the target decoder to be trained decodes the sample entropy decoding result to obtain the sample reconstruction point cloud data corresponding to the sample object. Based on the loss between the reconstructed point cloud data and the sample point cloud data, the model parameters of the target encoder and the target decoder to be trained are adjusted until the preset convergence condition is met, and the trained target encoder and target decoder are obtained.

7. The compression method according to claim 5, characterized in that, The target decoder includes: a first decoding module matched with the first encoding module and a second decoding module matched with the second encoding module; wherein, the first decoding module includes: a multi-scale upsampling module and a target residual network module.

8. A point cloud data compression device, characterized in that, The compression device includes: The input module is used to preprocess the point cloud data of the target object to obtain the geometric information matrix corresponding to the point cloud data, and to combine the geometric information matrix with the text feature vector corresponding to the target object to form multimodal input data; wherein, the geometric information matrix is ​​a feature matrix composed of the three-dimensional coordinate information of each spatial location point contained in the point cloud data, and the text feature vector is the text feature vector corresponding to the object type to which the target object belongs. The feature encoding module is used to input the multimodal input data into a pre-trained target encoder, and through the first encoding module in the target encoder, perform feature enhancement processing on the multimodal input data to obtain the geometric enhancement feature matrix corresponding to the geometric information matrix and the target text feature vector corresponding to the text feature vector; wherein, the first encoding module includes: a multi-scale downsampling module and a target residual network module; The second encoding module in the target encoder models the global physical structure dependency of the geometric enhancement feature matrix to obtain global geometric features that incorporate global physical structure information; wherein, the global physical structure information is used to represent the dependency between sets of spatial location points belonging to different physical structures in the point cloud data; The entropy coding module is used to perform entropy coding processing on the target text feature vector and the global geometric features to obtain the target point cloud compression result.

9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the point cloud data compression method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the point cloud data compression method as described in any one of claims 1 to 7.