Multi-machine collaborative SLAM point cloud data compression method based on attention network
By using a hierarchical parity-based encoding and decoding architecture based on attention networks to process octree point clouds, the problems of high computational cost and long processing time in existing technologies are solved, and efficient compression and real-time transmission of point cloud data are achieved.
Patent Information
- Application Number
- CN202511011982.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-28
AI Technical Summary
In existing technologies, point cloud compression methods based on deep learning have high computational costs and processing time, which cannot meet the real-time requirements of multi-machine collaborative SLAM systems.
A hierarchical parity-based encoding and decoding architecture based on attention networks is used to process octree point clouds, and entropy coding is used for compression.
It improves the real-time performance and transmission efficiency of point cloud data transmission, while reducing computing costs and processing time.
Smart Images

Figure CN121033192A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of point cloud compression, and particularly relates to a multi-robot cooperative SLAM point cloud data compression method based on an attention network. BACKGROUND
[0002] Multi-robot clusters exhibit strong capabilities that single robots cannot achieve when performing complex tasks. Multi-robot cooperative SLAM (simultaneous localization and mapping) technology can provide multi-robot clusters with comprehensive environmental information, global perspectives, and member pose information. However, this cooperative work requires robots to exchange a large amount of point cloud data, and point cloud data itself contains a large number of point coordinates, resulting in a large data size.
[0003] In order to alleviate the pressure of point cloud data transmission, it is particularly important to encode and compress point clouds. Current point cloud compression methods are mainly divided into traditional methods and deep learning-based methods. Traditional methods usually have low bit rates, cannot effectively reduce information entropy, and cannot meet the demand for communication bandwidth, so more and more research is turning to deep learning-based point cloud compression methods. Although deep learning-based compression methods perform well in bit rate and data volume, they have high computational costs and long processing times, and still cannot fully meet the real-time requirements of SLAM systems. SUMMARY
[0004] The present application relates to the technical field of point cloud compression, and particularly relates to a multi-robot cooperative SLAM point cloud data compression method based on an attention network.
[0005] To achieve the above-mentioned purpose, the present application provides the following technical solutions:
[0006] The present application provides a multi-robot cooperative SLAM point cloud data compression method based on an attention network, comprising:
[0007] S100, preprocessing the original point cloud data to obtain quantized point cloud data P Q ;
[0008] S200, converting the quantized point cloud data P Q according to the octree encoding principle to obtain an octree point cloud;
[0009] S300, data processing on the octree point cloud through a layered and parity coding architecture, predicting a feature probability distribution of each node of the octree point cloud, and entropy encoding according to the feature probability distribution of the node to obtain compressed point cloud data.
[0010] Optionally, the step S100 comprises:
[0011] In each of the original point cloud data, the minimum coordinate and the maximum coordinate of the original point cloud in the three-dimensional coordinate are obtained.
[0012] According to the minimum coordinate and the maximum coordinate of each of the original point cloud, the standardization processing is performed on each of the original point cloud to obtain the standardized point cloud data.
[0013] The quantization step q is performed on the standardized point cloud data. s , so as to quantize the coordinate value of the standardized point cloud data, and then to round, remove duplicates and sort the coordinate value to obtain the quantized point cloud data P Q .
[0014] Optionally, the step S200 comprises:
[0015] S201, selecting point cloud data P from the quantized point cloud data P Q , and dividing the space where the point cloud data P is located into eight equal subspaces according to Cartesian coordinates.
[0016] S202, judging whether each of the subspaces exists point cloud data.
[0017] S203, if yes, marking the node corresponding to the subspace where the point cloud data exists as an occupied state, and encoding the point cloud data in the subspace to obtain the node feature f of the node.
[0018] S204, after obtaining the node feature f, judging whether the current subspace size is less than a threshold value.
[0019] S205, if the current subspace size is less than the threshold value, stopping the algorithm outputting the octree point cloud.
[0020] S206, if the current subspace size is not less than the threshold value, constructing an octree subtree with the non-empty subspace as a new three-dimensional space, repeating the steps S201-S205 until the division reaches a given depth.
[0021] Optionally, the step S300 comprises:
[0022] The octree point cloud is unfolded into a one-dimensional node sequence based on a breadth-first rule.
[0023] the node sequence in the context window and integrate the context information;
[0024] training and testing the integrated context information through the hierarchical and parity coding architecture to obtain a predicted node feature probability distribution;
[0025] According to the predicted node feature probability distribution, each node is assigned a code, and an entropy coding method is used to compress the octree point cloud to obtain compressed point cloud data.
[0026] Optionally, the context window is used to intercept the node sequence in the context window and integrate the context information, comprising:
[0027] The fixed size context window is used to expand the receptive field of each node to be coded in the octree point cloud, and each node to be coded is batch coded in a hierarchical and parity manner to obtain node information of each node to be coded, and an input vector X is obtained by integrating the node information, wherein the node information includes the octree point cloud feature and the absolute position information corresponding to the subspace.
[0028] Optionally, the training and testing of the integrated context information through the hierarchical and parity coding architecture to obtain a predicted node feature probability distribution, comprising:
[0029] The node information in the input node sequence is converted into an embedding vector h in a high-dimensional space through an embedding layer;
[0030] According to the embedding vector h and the self-attention network, the relationship between each node in the window is calculated to obtain an attention score a;
[0031] The embedding feature of each node is input into an MLP layer, and the context fusion feature is determined according to the attention score a Through the context fusion feature The occupancy code probability distribution of the even node and the occupancy code probability distribution of the odd node are output to obtain a probability estimation model;
[0032] The parity node information and the even node information in the window to be predicted are predicted through the probability estimation model to obtain a predicted node feature probability distribution.
[0033] Optionally, the embedding vector h and the self-attention network are used to calculate the relationship between each node in the window to obtain an attention score a, comprising:
[0034] The occupancy code embedding of the current node in the embedding vector h is removed to obtain a removed embedding vector h;
[0035] input the pruned embedding vector h into a two-layer multi-head self-attention network in which position encoding and local enhancement modules are embedded, and calculate the relationship between each node in the window through the self-attention network to obtain an attention score a,
[0036]
[0037] wherein CONV is a one-dimensional convolutional network, MLPi is a fully connected neural network, t is the number of heads of the self-attention mechanism, m is the current node number being calculated, n is the nth sibling node in the window, and n≤m.
[0038] Optionally, the embedding feature of each node is input into an MLP layer, and a context fusion feature is determined according to the attention score a Through the context fusion feature output the occupancy code probability distribution of even nodes and the occupancy code probability distribution of odd nodes to obtain a probability estimation model, including:
[0039] input the embedding feature of each node into an MLP layer, and determine a context fusion feature according to the attention score a The calculation formula is:
[0040]
[0041] input the feature of even nodes into a multi-layer perceptron to output the occupancy code probability distribution of the even nodes, and input the feature of odd nodes into a multi-layer perceptron to output the occupancy code probability distribution of the odd nodes, thereby obtaining the probability estimation model, wherein the calculation formula of each node is:
[0042]
[0043] wherein P m i is the probability that the node occupancy code is m+1.
[0044] Optionally, the probability estimation model is used to predict the odd node information and the even node information in the window to be predicted, and a detailed probability distribution of each node is obtained, including:
[0045] The obtained even node occupancy code is input into an embedding layer to obtain a 128-dimensional vector, and the segmented latent space vector is spliced to obtain h={h1,...,h N / 2}∈R N / 2×(K+1)×150 , the odd node information and the even node information are predicted, and are expressed by the following formula:
[0046]
[0047] wherein A i is the information of the father node, x i2 is the information of the even node, x i1 is the information of the odd node.
[0048] Optionally, the encoding is assigned to each node according to the predicted node feature probability distribution, the octree point cloud is compressed by using an entropy encoding method to obtain compressed point cloud data.
[0049] The octree point cloud is flattened in a breadth-first manner, different length encodings are assigned according to the predicted node feature probability distribution, the point cloud data is converted into a binary stream through the encodings, the compression of the octree point cloud is completed, and the compressed point cloud data is obtained.
[0050] The above technical solutions of the present application have the following advantages or beneficial effects:
[0051] The present application converts the original point cloud data into an octree point cloud, estimates the feature probability distribution of each node in the octree point cloud by using a hierarchical and even-odd coding and decoding architecture based on a parallel attention network model, and finally compresses the octree point cloud according to the node feature probability distribution by using an entropy encoding method, thereby improving the real-time performance of point cloud data transmission. BRIEF DESCRIPTION OF DRAWINGS
[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings. The drawings are as follows:
[0053] Figure 1 is the first flowchart of the embodiment of the present application;
[0054] Figure 2 is the second flowchart of the embodiment of the present application. DETAILED DESCRIPTION
[0055] In order to make the objects, technical solutions and advantages of the present application clearer, the various exemplary embodiments to be described below will be described with reference to the corresponding drawings, which constitute a part of the exemplary embodiments and in which various exemplary embodiments that can be used to implement the present application are described. The same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all the implementations consistent with the present disclosure. It should be understood that they are only examples of processes, methods and apparatuses, etc. consistent with some aspects of the present disclosure as detailed in the appended claims, and other implementations can be used or structural and functional modifications can be made to the implementations listed herein without departing from the scope and spirit of the present application.
[0056] In the description of the present application, it should be understood that the terms "center", "longitudinal", "transverse" and the like indicate the orientation or positional relationship shown in the drawings, which are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the elements referred to must have a specific orientation, be constructed and operated in a specific orientation. The terms "first", "second" and the like are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. The term "a plurality of" means two or more. The terms "connected", "connected" should be interpreted broadly, for example, it can be fixed connection, detachable connection, integral connection, mechanical connection, electrical connection, communication connection, direct connection, indirect connection through intermediate medium, internal communication of two elements or interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more related listed items. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0057] In order to illustrate the technical solutions described in the present application, the following will be described by specific embodiments, only showing the parts related to the embodiments of the present application.
[0058] Embodiment one:
[0059] As Figure 1 shown, the present application provides a multi-machine cooperative SLAM point cloud data compression method based on attention network, comprising:
[0060] S100, preprocessing the original point cloud data to obtain quantized point cloud data P Q ;
[0061] S200, converting the quantized point cloud data P Q according to the octree encoding principle to obtain an octree point cloud;
[0062] S300, the octree point cloud is processed by the layered and odd-even coding architecture, the feature probability distribution of each node of the octree point cloud is predicted, and entropy coding is performed according to the feature probability distribution of the node, and compressed point cloud data is obtained. Specifically, after the original point cloud data is preprocessed, the quantized point cloud data P Q is converted into an octree point cloud, the octree point cloud is processed by the layered and odd-even coding architecture of the probability estimation model and the position coding and local information enhancement module, and then compressed by entropy coding, and compressed point cloud data is obtained. The layered and odd-even coding architecture can improve the running speed of the algorithm, and the position coding and local information enhancement module can improve the algorithm compression performance according to the characteristics of the data and the network.
[0063] The original point cloud data is converted into an octree point cloud, then the feature probability distribution of each node in the octree point cloud is estimated by the layered and odd-even coding architecture of the model based on the parallel attention network, and finally the octree point cloud is compressed according to the node feature probability distribution by using the entropy coding method, thereby improving the real-time performance of point cloud data transmission.
[0064] As an optional implementation, step S100 can include:
[0065] In each original point cloud of the original point cloud data, the minimum coordinate and the maximum coordinate of the original point cloud in the three-dimensional coordinate are obtained;
[0066] According to the minimum coordinate and the maximum coordinate of each original point cloud, the standardization processing is performed on each original point cloud, and the standardized point cloud data is obtained;
[0067] The quantization step q s is performed on the standardized point cloud data, so that the coordinate values of the standardized point cloud data are quantized, and then the coordinate values are rounded, de-duplicated and sorted, and the quantized point cloud data P Q is obtained. Specifically, in each original point cloud of the original point cloud data, the minimum value of each dimension in the three-dimensional coordinate is obtained, and the minimum coordinate of the original point cloud is determined, that is, the minimum coordinate of the original point cloud is the minimum value of each dimension in the xyz three-dimensional coordinate. At the same time, in each original point cloud of the original point cloud data, the maximum value of each dimension in the three-dimensional coordinate is obtained, and the maximum coordinate of the original point cloud is determined, that is, the maximum coordinate of the original point cloud is the maximum value of each dimension in the xyz three-dimensional coordinate. When each original point cloud is standardized, the minimum coordinate of the original point cloud is used as the offset, and the maximum coordinate of the original point cloud is used as the denominator, the minimum coordinate of the original point cloud is moved to the vicinity of the origin, and the standardization processing of each original point cloud is completed, that is, the movement of the original point cloud data is completed, and the standardized point cloud data is obtained. The quantization step q s, thereby completing the quantization of the coordinates of the point cloud. In the process of quantizing the coordinates of the point cloud, the coordinate values are rounded (i.e., rounding off the decimal places in the coordinate values) to facilitate subsequent calculation. After removing the repeated point cloud data and reordering, the quantized point cloud data P is obtained Q .
[0068] As an optional implementation, as shown in Figure 2 , step S200 can include:
[0069] S201, selecting point cloud data P from the quantized point cloud data P Q , and dividing the space where the point cloud data P is located into eight equal subspaces according to Cartesian coordinates;
[0070] S202, determining whether each subspace contains point cloud data;
[0071] S203, if yes, marking the node corresponding to the subspace containing point cloud data as an occupied state, and encoding the point cloud data in the subspace to obtain the node feature f of the node;
[0072] S204, after obtaining the node feature f, determining whether the current subspace size is less than a threshold value;
[0073] S205, if the current subspace size is less than the threshold value, stopping the algorithm to output the octree point cloud;
[0074] S206, if the current subspace size is not less than the threshold value, constructing an octree subtree for the non-empty subspace as a new three-dimensional space, and repeating steps S201-S205 until a given depth is reached. Specifically, selecting corresponding point cloud data from the quantized point cloud data P Q according to the point cloud information request of the collaborative SLAM, and recursively dividing the cubic space along the maximum side length of the bounding box of the selected point cloud data P Q into eight equal subspaces, each subspace corresponding to an octree node, and the occupancy states of the eight subspaces forming an eight-bit binary occupancy code. Therefore, after dividing the space where the quantized point cloud data P Q is located into eight equal subspaces according to Cartesian coordinates, the occupancy states of the eight subspaces are further determined. When determining whether each subspace contains point cloud data, if the subspace contains point cloud data, the node corresponding to the subspace containing point cloud data is marked as an occupied state; if the subspace does not contain point cloud data, the node corresponding to the subspace not containing point cloud data is marked as a non-occupied state. When the subspace contains point cloud data, the point cloud data is encoded to obtain the node feature f of the node. The node feature f is an eight-bit binary code, f = [b1, … b8] Bwhere bi represents whether there is point cloud data in the i-th subspace, if there is a point, bi = 1, otherwise bi = 0. After obtaining the node feature f, it is judged whether the current subspace size is less than the threshold value, if the current subspace size is less than the threshold value, the algorithm is stopped to output the octree point cloud, if the current subspace size is not less than the threshold value, the non-empty subspace (i.e. the subspace with point cloud data) is taken as a new three-dimensional space, and steps S201-S205 are repeated to recursively construct the octree sub-tree, until the division of the octree terminates at a given depth. The threshold value can be selected as 0.01m.
[0075] In this way, the original point cloud data is effectively converted into a more compact representation, each node only stores the information of the non-empty subspace, thereby reducing the storage amount of data and improving the transmission efficiency.
[0076] As an optional implementation, step S300 can include:
[0077] The octree point cloud is unfolded into a one-dimensional node sequence based on a breadth-first rule;
[0078] The to-be-encoded node in the node sequence is intercepted through a context window and the context information is integrated;
[0079] The integrated context information is subjected to model training and testing through a hierarchical and odd-even coding architecture to obtain a predicted node feature probability distribution;
[0080] Each node is assigned a code according to the predicted node feature probability distribution, and the octree point cloud is compressed using an entropy coding method to obtain compressed point cloud data. Specifically, the octree point cloud is unfolded into a one-dimensional node sequence based on a breadth-first rule, a fixed-size context window is used to expand the receptive field of the node, then each node in the octree is batch-encoded according to a hierarchical and odd-even manner, model training and testing are performed to obtain a predicted node feature probability distribution, the octree point cloud is compressed using entropy coding according to the predicted node feature probability distribution to obtain compressed point cloud data, and the calculation cost is reduced. The algorithm processes data in batches according to the hierarchical and odd-even manner, rather than processing each node one by one, thereby significantly speeding up the feature extraction and coding process, reducing the processing time, and meeting the real-time requirements of the SLAM system. Since the hierarchical and odd-even coding architecture can cause insufficient information of nodes in the same level, the present application introduces a position coding and local enhancement module to enhance the prediction ability of the model and improve the overall performance of the compression algorithm. The hierarchical coding architecture first uses the nodes {x a1 , x a1 , x a3} ∈ R K×6 , the node itself feature x i ∈ R 1×5and the information of the sibling nodes in the window {x1, x1, …, xj, …} ∈ R(N-1) ×5 The odd nodes in the window to be predicted are predicted. Five dimensions are missing the occupancy code information of the current node and are also information to be predicted. The final prediction outputs the occupancy code feature of the odd node. Then the predicted occupancy code feature of the odd node is merged with the upper three ancestor nodes a of the even node to predict the feature distribution of the even node to obtain the occupancy code of the odd node. Finally, the features of all nodes in the window are predicted. By the above phased decoding mode, N / 2 nodes can be decoded at a time, and the information of the sibling nodes is introduced in the prediction process of the second phase, which improves the disadvantage of serial decoding, and also uses the features of the sibling nodes to improve the prediction accuracy of the nodes.
[0081] As an optional implementation, the context window intercepts the node to be encoded in the node sequence and integrates the context information, which can include:
[0082] A fixed-size context window is used to expand the receptive field of each node to be encoded in the octree point cloud, and then each node to be encoded is batch-encoded in batches according to the hierarchical and even-odd manner to obtain node information of each node to be encoded, and an input vector X is obtained by integrating the node information, wherein the node information includes octree point cloud features and absolute position information corresponding to a subspace. Specifically, the present application selects K parent nodes as the context perception range, that is, the vertical tree visual field, and sets the size of the context window as N, that is, the horizontal level visual field, and integrates the context information of the node to be encoded to form an input vector X. The context information of the node to be encoded can include six dimensions of octree features (occupancy code, level and gop limit) and absolute position information (x, y, z) corresponding to the subspace of the octree point cloud, and the final input vector X information dimension is (N, K+1, 6). By expanding the comprehensive receptive field and the absolute position information, the feature probability distribution of the node is accurately predicted, and more abundant information basis is provided for subsequent entropy coding.
[0083] As an optional implementation, the integrated context information is model trained and tested through a hierarchical and even-odd coding and decoding architecture to obtain a predicted node feature probability distribution, which can include:
[0084] The node information in the input node sequence is converted into an embedding vector h in a high-dimensional space through an embedding layer;
[0085] The relationship between each node in the window is calculated according to the embedding vector h and the self-attention network to obtain an attention score a;
[0086] The embedding feature of each node is input into an MLP layer, and the context fusion feature is determined according to the attention score a Through the context fusion feature The occupancy code probability distribution of the even nodes and the occupancy code probability distribution of the odd nodes are outputted to obtain a probability estimation model.
[0087] The odd node information and the even node information in the to-be-predicted window are predicted through the probability estimation model to obtain a predicted node feature probability distribution. Specifically, the node information in the input node sequence is converted into an embedding vector h in a high-dimensional space through an embedding layer. Each node information E = [f, level, octant, x, y, z], wherein the node feature f ∈ [0, 255], the depth level ∈ [0, 12], and the octant ∈ [0, 8], and the filled data node E = [255, 0, 0, 0, 0, 0]. The node feature f is mapped to a 128-dimensional space, the depth level is mapped to a 6-dimensional space, the octant is mapped to a 4-dimensional space, the position is mapped to a 12-dimensional space (position encoding and local enhancement module) using a linear layer and a RELU activation function, and finally spliced in order. Each node information becomes a 150-dimensional vector, and the embedding conversion layer outputs an embedding vector X of (N, K+1, 150). The parameters of the mapping network of each part will be updated with training. The attention score a, the context fusion feature and the context fusion feature The occupancy code probability distribution of the even nodes and the occupancy code probability distribution of the odd nodes are outputted to obtain a probability estimation model. The predicted node feature probability distribution is obtained through the probability estimation model, which is convenient for compressing the octree point cloud, reducing the cost and time of data transmission, and improving the real-time performance of the transmitted data.
[0088] As an optional implementation, the relationship between each node in the window is calculated according to the embedding vector h and the self-attention network to obtain the attention score a, which can include:
[0089] The occupancy code of the middle node of the embedding vector h is excluded to obtain an excluded embedding vector h;
[0090] The excluded embedding vector h is input into a two-layer multi-head self-attention network embedded with a position encoding and a local enhancement module, and the relationship between each node in the window is calculated through the self-attention network to obtain the attention score a,
[0091]
[0092] wherein CONV is a one-dimensional convolutional network, MLPi is a fully connected neural network, t is the number of heads of the self-attention mechanism, m is the current calculated node number, n is the nth brother node in the window, and n≤m. Specifically, the embedding vector h = {h1,..., hm} ∈ R N N×(K+1)×150 The occupancy code embedding of the middle nodes is removed to obtain a removed embedding vector h = {h1,..., h N}∈R N×(K+1)×150-128 The removed embedding vector h is then input into a two-layer multi-head self-attention network in which position encoding and local enhancement modules are embedded. The self-attention network is used to calculate the relationship between each node in the window, and multiple projection heads are used to independently generate different attention weights to avoid overfitting. The calculated attention values are subjected to a mask operation to obtain attention scores a. Wherein, n≤m, and m=i-N+1, …, i.
[0093] As an optional implementation, the embedding features of each node are input into an MLP layer to determine context fusion features The context fusion features The occupancy code probability distribution of even nodes and the occupancy code probability distribution of odd nodes are output to obtain a probability estimation model, which can include:
[0094] The embedding features of each node are input into an MLP layer to determine context fusion features The calculation formula is:
[0095]
[0096] The features of even nodes are input into a multi-layer perceptron to output the occupancy code probability distribution of even nodes, and the features of odd nodes are input into a multi-layer perceptron to output the occupancy code probability distribution of odd nodes, thereby obtaining a probability estimation model, wherein the calculation formula of each node is:
[0097]
[0098] Wherein, P m i is the probability that the node occupancy code is m+1. Specifically, after the embedding features of each node are input into an MLP layer, the values obtained by multiplying the other n nodes by the attention scores a are added to obtain the weighted context fusion features Subsequently, the generated context vector is divided into two parts, one part is odd nodes, and the other part is even nodes. The features of even nodes are input into a multi-layer perceptron MLP3() of multi-head attention to output the occupancy code probability distribution of even nodes, and the features of odd nodes are input into a multi-layer perceptron MLP3() of multi-head attention to output the occupancy code probability distribution of odd nodes, thereby obtaining a probability estimation model.
[0099] As an optional implementation, the odd node information and the even node information in the to-be-predicted window are predicted by the probability estimation model to obtain a detailed probability distribution of each node, which can include:
[0100] The obtained even node occupation code is subjected to an embedding layer to obtain a 128-dimensional vector, and the segmented hidden space vector is spliced to obtain h={h1,...,h N / 2}∈R N / 2×(K+1)×150 , the odd node information and the even node information are predicted, and are expressed by the following formula:
[0101]
[0102] Wherein, A i is the information of the father node, x i2 is the even node information, and x i1 is the odd node information. More specifically, when the probability estimation model is used, the loss function L of the probability estimation model needs to be calculated, and the parameters of the probability estimation model are updated according to the loss function L, and the cycle is repeated until the loss function converges. The calculation formula of the loss function L is:
[0103]
[0104] Wherein, f i is the node feature of node i, p i (f i |x,w) is the probability value calculated by the probability estimation model. is the node feature f i is the probability of the mth feature. There are 256 kinds of octree node features in the three-dimensional point cloud space, and the invention sorts the feature categories according to the size of their binary codes, wherein is the probability of f i =[00000001] b .
[0105] As an optional implementation, each node is assigned a code according to the predicted node feature probability distribution, and an entropy coding method is used to compress the octree point cloud to obtain compressed point cloud data, including:
[0106] The octree point cloud is flattened in a breadth-first manner, different lengths of codes are assigned according to the predicted node feature probability distribution, the point cloud data is converted into a binary stream through the codes, the compression of the octree point cloud is completed, and compressed point cloud data is obtained. Specifically, different lengths of codes are assigned according to the predicted node feature probability distribution, high-frequency features are assigned shorter codes, and low-frequency features are assigned longer codes, which can effectively reduce the total number of data bits and achieve efficient data compression. The octree point cloud is converted into a compact binary stream through an entropy encoding method, the compression of the octree point cloud is completed, and the required storage space and transmission bandwidth are greatly reduced. In the process of assigning codes, for a data stream with a length of N to be encoded, the feature of the node n i is f i, and the corresponding probability distribution is P (p 1, p 2..p 256 ), wherein p k is the probability of the feature f i =f k . Starting from the real interval (0, 1], we first divide the space into 256 intervals according to P 0 (p 1, p 2..p 256 ), and the size l k of each interval is related to the corresponding probability:
[0107]
[0108] According to the node feature, the corresponding interval k is selected, and then the process is repeated.
[0109] More specifically, based on the SemanticKITTI dataset, the model is trained and tested, and the method is compared with the Voxel DNN based on the voxel network model, the VoxelContex-Net, and the advanced compression method Oct-Attention algorithm which also uses an octree structure. The method can significantly reduce the storage and transmission cost of data while ensuring data integrity and accuracy. The algorithm comparison results are shown in Table 1:
[0110] Table 1
[0111]
[0112] It can be seen that the method takes less time than other methods in terms of calculation time, and the average encoding and decoding time of each frame of point cloud is reduced by 48.89% compared with VoxelDNN, meeting the real-time requirements of multi-machine cooperative SLAM. In terms of compression performance, the method improves by 15.45% compared with the MPEG standard method. The method can compress each frame of point cloud in the KITTI dataset to an average of 53.66KB, which is reduced by 8.29KB compared with the MPEG standard method.
[0113] The embodiment is only one specific example and does not indicate that the present application is limited to such an implementation.
[0114] The above description is merely that of the preferred embodiments of the present application, and various changes or modifications can be made thereto without departing from the spirit and scope of the present application. In addition, the features and embodiments of the present application can be modified in various ways to adapt to specific conditions and materials without departing from the spirit and scope of the present application. Therefore, the present application is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present application are included in the scope of the present application.
Claims
1. A multi-machine collaborative SLAM point cloud data compression method based on attention networks, characterized in that, include: S100. Preprocess the raw point cloud data to obtain quantized point cloud data P. Q ; S200. Based on the octree coding principle, the quantized point cloud data P is processed... Q The conversion yields an octree point cloud; S300. The octree point cloud is processed using a hierarchical parity encoding and decoding architecture to predict the feature probability distribution of each node in the octree point cloud, and entropy encoding is performed based on the feature probability distribution of the nodes to obtain compressed point cloud data.
2. The multi-machine collaborative SLAM point cloud data compression method based on attention networks according to claim 1, characterized in that, Step S100 includes: In each original point cloud of the original point cloud data, obtain the minimum and maximum coordinates of the original point cloud in three-dimensional coordinates; Based on the minimum and maximum coordinates of each original point cloud, each original point cloud is standardized to obtain standardized point cloud data. The standardized point cloud data is quantized with a step size q. s The coordinate values of the standardized point cloud data are quantized, and then the coordinate values are rounded, deduplicated, and sorted to obtain the quantized point cloud data P. Q .
3. The multi-machine collaborative SLAM point cloud data compression method based on attention networks according to claim 1, characterized in that, Step S200 includes: S201, in the quantized point cloud data P Q Select point cloud data P, and divide the space where the point cloud data P is located into eight equal subspaces according to Cartesian coordinates; S202. Determine whether point cloud data exists in each of the subspaces; S203. If so, mark the node corresponding to the subspace containing the point cloud data as occupied, and encode the point cloud data in the subspace to obtain the node feature f of the node. S204. After obtaining the node feature f, determine whether the current subspace size is less than the threshold. S205. If the current subspace size is less than the threshold, then stop the algorithm to output the octree point cloud; S206. If the current subspace size is not less than the threshold, then the non-empty subspace is used as a new three-dimensional space to construct an octree subtree, and steps S201-S205 are repeated until the subspace is divided to a given depth.
4. The multi-machine collaborative SLAM point cloud data compression method based on attention networks according to claim 3, characterized in that, Step S300 includes: The octree point cloud is unfolded into a one-dimensional node sequence based on the width-first rule; The nodes to be encoded in the node sequence are extracted using a context window and their context information is integrated. By using a hierarchical and odd-even encoding / decoding architecture, the integrated contextual information is used for model training and testing to obtain the predicted node feature probability distribution. Each node is assigned an encoding based on the predicted node feature probability distribution, and the octree point cloud is compressed using an entropy encoding method to obtain compressed point cloud data.
5. The multi-machine collaborative SLAM point cloud data compression method based on attention networks according to claim 4, characterized in that, The step of extracting the nodes to be encoded from the node sequence through a context window and integrating the context information includes: The receptive field of each node to be encoded in the octree point cloud is expanded by using a context window of fixed size. Then, each node to be encoded is encoded and decoded in batches according to the hierarchical and odd-even method to obtain the node information of each node to be encoded. The input vector X is obtained by integrating the node information, wherein the node information includes the octree point cloud features and the absolute position information corresponding to the subspace.
6. The multi-machine collaborative SLAM point cloud data compression method based on attention networks according to claim 4, characterized in that, The method of training and testing the integrated context information using a hierarchical, odd-even encoding / decoding architecture to obtain the predicted node feature probability distribution includes: The node information in the input node sequence is converted into an embedding vector h in a high-dimensional space through an embedding layer; The attention score a is obtained based on the embedding vector h and the relationship between each node within the self-attention network calculation window; The embedding features of each node are input into the MLP layer, and the context fusion features C are determined based on the attention score a. i (t) Through the context fusion feature C i (t) Output the occupancy code probability distribution of even-numbered nodes and the occupancy code probability distribution of odd-numbered nodes to obtain the probability estimation model; The probability estimation model is used to predict the odd-numbered and even-numbered node information within the window to be predicted, thereby obtaining the predicted node feature probability distribution.
7. The multi-machine collaborative SLAM point cloud data compression method based on attention networks according to claim 6, characterized in that, The step of obtaining the attention score a based on the embedding vector h and the relationship between each node within the self-attention network calculation window includes: The occupancy code of the local node in the embedding vector h is removed to obtain the removed embedding vector h. The removed embedding vector h is input into a two-layer multi-head self-attention network that incorporates positional encoding and local enhancement modules. The attention score a is obtained by calculating the relationship between each node within the window through the self-attention network. Where CONV is a one-dimensional convolutional network, MLPi is a fully connected neural network, t is the number of heads in the self-attention mechanism, m is the node number currently being computed, and n is the nth sibling node in the window, and n≤m.
8. The multi-machine collaborative SLAM point cloud data compression method based on attention networks according to claim 6, characterized in that, The embedding features of each node are input into the MLP layer, and the context fusion features are determined based on the attention score 'a'. Through the context fusion feature C i (t) Output the occupancy code probability distributions for even-numbered nodes and odd-numbered nodes to obtain the probability estimation model, including: The embedding features of each node are input into the MLP layer, and the context fusion features are determined based on the attention score 'a'. The calculation formula is: The features of even-numbered nodes are input into a multilayer perceptron, which outputs the occupancy code probability distribution of the even-numbered nodes. Similarly, the features of odd-numbered nodes are input into the multilayer perceptron, which outputs the occupancy code probability distribution of the odd-numbered nodes, thereby obtaining the probability estimation model. The calculation formula for each node is as follows: Among them, P m i This represents the probability that a node occupies the code m+1.
9. The multi-machine collaborative SLAM point cloud data compression method based on attention networks according to claim 6, characterized in that, The step of predicting odd-numbered and even-numbered node information within the prediction window using the probability estimation model to obtain a detailed probability distribution for each node includes: The obtained even-numbered node occupancy codes are processed through an embedding layer to obtain a 128-dimensional vector, which is then concatenated with the segmented latent space vector to obtain h = {h1,...,h}. N / 2 }∈R N / 2×(K+1)×150 The prediction of information for odd-numbered nodes and even-numbered nodes is expressed by the following formula; Among them, A i For information about the parent node, x i2 For even-numbered node information, x i1 Information for odd-numbered nodes.
10. The multi-machine collaborative SLAM point cloud data compression method based on attention networks according to claim 4, characterized in that, The process of assigning an encoding to each node based on the predicted node feature probability distribution and compressing the octree point cloud using an entropy coding method to obtain compressed point cloud data includes: The octree point cloud is flattened in a width-first manner, and codes of different lengths are assigned according to the predicted node feature probability distribution. The point cloud data is then converted into a binary stream through the codes to complete the compression of the octree point cloud and obtain the compressed point cloud data.