Dynamic 3D point cloud compression method and device based on implicit neural expression
Through the method based on implicit neural expression, point cloud data embedding and extraction is carried out in combination with the downsampling layer and the grouping vector attention layer, and feature fusion is used by time index embedding, the problem of complexity and inefficiency of dynamic point cloud compression in the existing technology is solved, and efficient and accurate three-dimensional point cloud compression is achieved.
Patent Information
- Application Number
- CN202411932790.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-12-26
AI Technical Summary
In the prior art, the dynamic point cloud compression method is complex and the decoding process is cumbersome, making it difficult to effectively improve the efficiency and accuracy of three-dimensional point cloud compression.
The dynamic three-dimensional point cloud compression method based on implicit neural expression is adopted, and the point cloud data is embedded and extracted by an encoder composed of a downsampling layer and an attention layer based on group vector attention, feature fusion is performed in combination with time index embedding, and lossless compression is performed through Huffman encoding.
It improves the efficiency and accuracy of three-dimensional point cloud compression, simplifies the decoding process, reduces the computational complexity, and maintains high-quality point cloud reconstruction effect.
Smart Images

Figure CN119359830B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and specifically to a dynamic three-dimensional point cloud compression method and device based on implicit neural expression. Background Art
[0002] At present, the mainstream methods for dynamic point cloud compression tasks are: a geometry-based method that recognizes and utilizes repeated patterns and rules in point clouds to compress redundancy; a video-based method that uses geometric projection and packaging technology to project three-dimensional point clouds onto a two-dimensional plane; an octree entropy coding method that constructs an octree node sequence and performs entropy coding based on contextual relationships; a motion estimation and motion compensation method that estimates inter-frame motion information to utilize the high similarity between consecutive frames; and a range image-based method for lidar point clouds.
[0003] It should be pointed out that these dynamic point cloud compression methods, whether using plane projection to use video codecs, or using motion estimation and motion compensation to maximize the compression of temporal redundancy, etc., have relatively complex network frameworks, such as key frame selection, inter-frame prediction, residual coding, feature extraction and compression, etc. Such a long pipeline makes the decoding process very complicated. Summary of the invention
[0004] In response to the problems in the prior art, the present application provides a dynamic three-dimensional point cloud compression method and device based on implicit neural expression, which can effectively improve the efficiency and accuracy of three-dimensional point cloud compression.
[0005] In order to solve at least one of the above problems, the present application provides the following technical solutions:
[0006] In a first aspect, the present application provides a dynamic three-dimensional point cloud compression method based on implicit neural expression, comprising:
[0007] The point cloud data of each frame in a given dynamic point cloud sequence is embedded and extracted by an encoder composed of a downsampling layer and an attention layer based on grouped vector attention, so as to obtain an adaptive content feature embedding containing point cloud content information; the index value of the point cloud data of each frame in the given dynamic point cloud sequence is obtained, and the index value of the point cloud data is mapped to a high-dimensional embedding space for embedding extraction through a position encoding function, so as to obtain a time index embedding containing point cloud time information;
[0008] The adaptive content feature embedding is used as the visual knowledge prior of the time index embedding to fuse the time index embedding and the content feature embedding in the same spatial dimension to obtain a fused feature vector; the fused feature vector is input into a decoder composed of an upsampling layer and an attention layer based on grouped vector attention to obtain a predicted point cloud frame output by the decoder, and the model parameters of the network model are updated by back propagation of the calculated chamfer distance and mean square error loss between the predicted point cloud frame and the corresponding real point cloud frame until the similarity between the predicted point cloud frame and the real point cloud frame meets the threshold interval, thereby obtaining the best network model;
[0009] The floating-point parameters of the optimal network model are converted into low-bit integer parameters and then losslessly compressed using Huffman coding to obtain the optimal network model after lossless compression, and dynamic three-dimensional point cloud data decoding is performed using the optimal network model.
[0010] Furthermore, the encoder composed of a downsampling layer and an attention layer based on grouped vector attention performs embedding extraction on the point cloud data of each frame in a given dynamic point cloud sequence to obtain an adaptive content feature embedding containing point cloud content information, including:
[0011] Input a frame of point cloud data in a given dynamic point cloud sequence, divide the points in the point cloud data into voxel grids by a pooling method based on voxel grid division, cluster the points in each grid and perform a maximum pooling operation on the point cloud features, so as to obtain local features of the point cloud through a downsampling layer;
[0012] Input the local features of the point cloud output by the downsampling layer into the grouped vector attention mechanism, wherein the grouped vector attention mechanism maps the input local features of the point cloud into a query vector, a key vector and a value vector, groups the channels of the value vector, and encodes the relationship between the query vector and the key vector to obtain a weighted encoding, i.e., the correlation weight between the local features of the point cloud, wherein the number of weighted encoding channels is the same as the number of channel groups of the value vector;
[0013] A Hadamard product is performed on the relevance weight and the value vector to obtain a final adaptive content feature embedding.
[0014] Furthermore, the step of obtaining the index value of the point cloud data of each frame in the given dynamic point cloud sequence, and mapping the index value of the point cloud data to a high-dimensional embedding space for embedding extraction through a position encoding function to obtain a time index embedding containing point cloud time information, includes:
[0015] For a given dynamic point cloud sequence, an index value of each frame of point cloud data is obtained, wherein the index value represents a time index of the point cloud frame in the entire sequence;
[0016] The index value of the acquired point cloud frame is normalized to the interval [0, 1] and input into the position encoding function to map the one-dimensional time index value into a high-dimensional embedding space to obtain a time index embedding containing the time information of the point cloud.
[0017] Furthermore, the step of using the adaptive content feature embedding as the visual knowledge prior of the time index embedding to perform feature fusion on the time index embedding and the content feature embedding in the same spatial dimension to obtain a fused feature vector includes:
[0018] aligning the time index embedding and the content feature embedding in a spatial dimension;
[0019] The aligned adaptive content feature embedding is used as the visual knowledge prior of the time index embedding to perform feature fusion to obtain a fused feature vector.
[0020] Furthermore, the step of inputting the fused feature vector into a decoder composed of an upsampling layer and an attention layer based on grouped vector attention to obtain a predicted point cloud frame output by the decoder includes:
[0021] Inputting the fused feature vector into a decoder consisting of an upsampling layer and an attention layer based on grouped vector attention, wherein the network of the decoder consists of 4 decoding blocks, each of which includes an upsampling layer for gradually expanding the feature space resolution and an attention layer based on grouped vector attention for extracting correlation information between features;
[0022] After feature processing of the decoder network, a predicted point cloud frame output by the decoder is obtained.
[0023] Further, the model parameters of the network model are updated by back propagation by calculating the chamfer distance and mean square error loss between the predicted point cloud frame and the corresponding real point cloud frame until the similarity between the predicted point cloud frame and the real point cloud frame meets the threshold interval, thereby obtaining the best network model, including:
[0024] Calculating the chamfer distance and mean square error loss between the predicted point cloud frame and the corresponding real point cloud frame, and weightedly fusing the loss functions of the chamfer distance and the mean square error loss to obtain a final loss function;
[0025] The model parameters of the network model are updated by back propagation according to the loss function until the similarity between the predicted point cloud frame and the real point cloud frame meets the threshold interval, thereby obtaining the optimal network model.
[0026] Furthermore, the converting of the floating point number parameters of the optimal network model into low-bit integer parameters and then performing lossless compression using Huffman coding to obtain the optimal network model after lossless compression includes:
[0027] Converting the 32-bit or 16-bit floating point number parameters of the optimal network model into 8-bit integer parameters;
[0028] The 8-bit integer parameters of the optimal network model are input into the Huffman coding algorithm to allocate codes of different lengths based on the frequency of occurrence of the parameter values, thereby obtaining the optimal network model after lossless compression.
[0029] In a second aspect, the present application provides a dynamic three-dimensional point cloud compression device based on implicit neural expression, comprising:
[0030] The data processing module 10 is used to embed and extract the point cloud data of each frame in a given dynamic point cloud sequence through an encoder composed of a downsampling layer and an attention layer based on grouped vector attention, so as to obtain an adaptive content feature embedding containing point cloud content information; obtain the index value of the point cloud data of each frame in the given dynamic point cloud sequence, and map the index value of the point cloud data to a high-dimensional embedding space through a position encoding function for embedding extraction, so as to obtain a time index embedding containing point cloud time information;
[0031] The model training module 20 is used to use the adaptive content feature embedding as the visual knowledge prior of the time index embedding to perform feature fusion on the time index embedding and the content feature embedding in the same spatial dimension to obtain a fused feature vector; the fused feature vector is input into a decoder composed of an upsampling layer and an attention layer based on grouped vector attention to obtain a predicted point cloud frame output by the decoder, and the model parameters of the network model are updated by back propagation of the calculated chamfer distance and mean square error loss between the predicted point cloud frame and the corresponding real point cloud frame until the similarity between the predicted point cloud frame and the real point cloud frame meets the threshold interval, thereby obtaining the best network model;
[0032] The model compression module 30 is used to convert the floating-point parameters of the optimal network model into low-bit integer parameters and then use Huffman coding to perform lossless compression to obtain the optimal network model after lossless compression, and perform dynamic three-dimensional point cloud data decoding through the optimal network model.
[0033] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the dynamic three-dimensional point cloud compression method based on implicit neural expression when executing the program.
[0034] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the dynamic three-dimensional point cloud compression method based on implicit neural expression.
[0035] In a fifth aspect, the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the dynamic three-dimensional point cloud compression method based on implicit neural expression.
[0036] It can be seen from the above technical scheme that the present application provides a dynamic three-dimensional point cloud compression method and device based on implicit neural expression, which embeds and extracts the point cloud data of each frame in a given dynamic point cloud sequence through an encoder composed of a downsampling layer and an attention layer based on grouped vector attention, and obtains an adaptive content feature embedding containing point cloud content information; obtains the index value of the point cloud data of each frame in the given dynamic point cloud sequence, and maps the index value of the point cloud data to a high-dimensional embedding space through a position encoding function for embedding extraction, and obtains a time index embedding containing point cloud time information; converts the floating-point parameters of the optimal network model into low-bit integer parameters and then uses Huffman coding for lossless compression to obtain the optimal network model after lossless compression, and performs dynamic three-dimensional point cloud data decoding through the optimal network model, thereby effectively improving the efficiency and accuracy of three-dimensional point cloud compression. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0038] Figure 1 This is one of the flow charts of the dynamic three-dimensional point cloud compression method based on implicit neural expression in the embodiment of the present application;
[0039] Figure 2 This is the second flow chart of the dynamic three-dimensional point cloud compression method based on implicit neural expression in the embodiment of the present application;
[0040] Figure 3 The third flowchart of the dynamic three-dimensional point cloud compression method based on implicit neural expression in the embodiment of the present application;
[0041] Figure 4 This is a fourth flow chart of a dynamic three-dimensional point cloud compression method based on implicit neural expression in an embodiment of the present application;
[0042] Figure 5 This is a fifth flow chart of a dynamic three-dimensional point cloud compression method based on implicit neural expression in an embodiment of the present application;
[0043] Figure 6 This is the sixth flow chart of the dynamic three-dimensional point cloud compression method based on implicit neural expression in the embodiment of the present application;
[0044] Figure 7 FIG7 is a flow chart of a dynamic three-dimensional point cloud compression method based on implicit neural expression in an embodiment of the present application;
[0045] Figure 8 is a structural diagram of a dynamic three-dimensional point cloud compression device based on implicit neural expression in an embodiment of the present application;
[0046] Fig. 9 It is a schematic diagram of the structure of an electronic device in an embodiment of the present application.
[0047] Reference numerals:
[0048] Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0050] The acquisition, storage, use, and processing of data in the technical solution of this application comply with the relevant provisions of national laws and regulations.
[0051] In view of the problems existing in the prior art, the present application provides a dynamic three-dimensional point cloud compression method and device based on implicit neural expression, which embeds and extracts the point cloud data of each frame in a given dynamic point cloud sequence through an encoder composed of a downsampling layer and an attention layer based on grouped vector attention, so as to obtain an adaptive content feature embedding containing point cloud content information; obtains the index value of the point cloud data of each frame in the given dynamic point cloud sequence, and maps the index value of the point cloud data to a high-dimensional embedding space for embedding extraction through a position encoding function, so as to obtain a time index embedding containing point cloud time information; converts the floating-point parameters of the best network model into low-bit integer parameters and then uses Huffman coding for lossless compression to obtain the best network model after lossless compression, and performs dynamic three-dimensional point cloud data decoding through the best network model, thereby effectively improving the efficiency and accuracy of three-dimensional point cloud compression.
[0052] In order to effectively improve the efficiency and accuracy of 3D point cloud compression, the present application provides an embodiment of a dynamic 3D point cloud compression method based on implicit neural expression, see Figure 1 The dynamic three-dimensional point cloud compression method based on implicit neural expression specifically includes the following contents:
[0053] Step S101: embedding and extracting the point cloud data of each frame in a given dynamic point cloud sequence through an encoder composed of a downsampling layer and an attention layer based on grouped vector attention, so as to obtain an adaptive content feature embedding containing point cloud content information; obtaining the index value of the point cloud data of each frame in the given dynamic point cloud sequence, and mapping the index value of the point cloud data to a high-dimensional embedding space through a position encoding function for embedding extraction, so as to obtain a time index embedding containing point cloud time information;
[0054] Optionally, in this embodiment, step S101 is a key link in a dynamic 3D point cloud compression method based on implicit neural expression, which involves two parallel but interrelated processes: content feature extraction of point cloud data and embedding of time index information. This step is designed to effectively capture the spatial and temporal information in the dynamic point cloud sequence, laying the foundation for the subsequent feature fusion and compression process.
[0055] First, let's take a deep look at the process of extracting content features from point cloud data. This process uses an encoder consisting of a downsampling layer and an attention layer based on grouped vector attention. The main function of the downsampling layer is to reduce the size of the point cloud data and improve processing efficiency. This is usually achieved through methods such as farthest point sampling (FPS) or random sampling. Downsampling not only reduces computational complexity, but also helps the model focus on more significant features.
[0056] Next is the attention layer based on grouped vector attention. This attention mechanism is an improvement on the traditional attention mechanism. It groups the channels of the value vector mapped by the attention mechanism and encodes the relationship between the query vector and the key vector to obtain the correlation weights of the local features of the point cloud. The number of channels is the same as the number of channel groups of the value vector. This method takes into account the weights of the features on different channels and can independently adjust the feature channels. It can capture local and global structural information in the point cloud while reducing computational complexity. Specifically, for each point, the model calculates its correlation with other points and assigns different weights based on these correlations. This enables the model to adaptively focus on important areas in the point cloud, thereby extracting more meaningful features.
[0057] Through this encoder, we can get an adaptive content feature embedding that contains the content information of the point cloud. This embedding is a high-dimensional vector that encodes key features such as the geometry and density distribution of the point cloud. Due to the use of the attention mechanism, this embedding is adaptive, which means that it can dynamically adjust according to the specific characteristics of the input point cloud, thereby more accurately expressing different types of point cloud data.
[0058] At the same time, step S101 also includes the processing of time information. For a given dynamic point cloud sequence, each frame of point cloud data has a corresponding index value, indicating its time position in the sequence. This index value is an integer, but directly using this integer may not fully express the time information. Therefore, the position encoding function is introduced here.
[0059] The role of the position encoding function is to map these integer indices to a high-dimensional embedding space. Common position encoding methods include sinusoidal position encoding and learnable position embedding. The purpose of this mapping is to enable the model to better understand and utilize temporal information. By converting the time index into a high-dimensional vector, the model can capture more subtle temporal relationships, such as the relationship between adjacent frames, periodic changes, etc.
[0060] These two processes - content feature extraction and time index embedding - together solve a key technical problem in dynamic point cloud compression: how to effectively represent the spatial structure and temporal evolution of point clouds at the same time. Traditional point cloud compression methods often process spatial and temporal information separately, resulting in low compression efficiency or information loss. This method processes these two types of information in parallel, providing rich input for subsequent feature fusion, helping to generate more compact and effective compressed representations.
[0061] From a technical perspective, this step achieves several important goals. First, it improves the efficiency and effectiveness of feature extraction. By using downsampling and attention mechanisms, the model can quickly focus on key areas in the point cloud and reduce unnecessary calculations. Second, adaptive content feature embedding improves the model's adaptability to different types of point clouds, allowing the compression method to be applied to a wider range of scenarios. Third, the introduction of time-indexed embedding enables the model to better understand and utilize the temporal dynamics of point cloud sequences, which is critical for accurately reconstructing dynamic point clouds.
[0062] Let's use a specific example to illustrate the application of this step. Assume that we are processing a dynamic point cloud sequence in an autonomous driving scenario. This sequence contains 1000 frames, and each frame contains about 100,000 points. Our goal is to efficiently compress this sequence for easy storage and transmission.
[0063] First, for content feature extraction, we designed an encoder consisting of 3 layers of downsampling and 4 layers of attention. Each layer of downsampling reduces the number of points by half, using the farthest point sampling (FPS) algorithm. In the attention layer, we divide the value vector into 6 groups, and the number of channels of weight encoding is also 6. Each group is independently weighted by the correlation weight obtained by the query vector and the key vector. This design enables the encoder to capture the multi-scale features of the point cloud while maintaining computational efficiency.
[0064] For a typical frame, such as the 500th frame, the encoder first reduces the number of points from 100,000 to about 12,500 by downsampling. Then, through the attention layer, the model focuses on some key areas, such as road edges, vehicle contours, etc. Finally, we get a 256-dimensional feature vector that highly summarizes the spatial structure and geometric features of the point cloud of this frame.
[0065] At the same time, for the processing of temporal information, we use a sinusoidal position encoding function. Specifically, we map the index value of the 500th frame (ie 499, because the index starts from 0) to a 128-dimensional vector space. This encoding ensures that the model can understand the relative temporal relationship between frames even if they are far apart.
[0066] Through experiments, we found that this method can effectively handle changes in dynamic scenes. For example, when a car passes through the field of view, the model can accurately capture the changes in the point cloud while keeping the consistency of static background (such as buildings). Across the entire 1000-frame sequence, we observed that the content feature embedding can adaptively adjust to reflect the dynamic changes of the scene, while the time index embedding helps the model correctly understand the temporal relationship of these changes.
[0067] This method shows obvious advantages over traditional inter-frame compression methods, such as octree-based compression. At the same compression ratio, our method can better preserve the details of dynamic objects and reduce motion blur and geometric distortion. Especially when dealing with fast-moving objects, such as vehicles making sharp turns, our method can more accurately reconstruct their trajectories and shapes.
[0068] In addition, this method also shows good generalization ability. Whether it is relatively dense 3D dynamic point cloud data acquired by binocular stereo cameras or depth cameras, or lidar point cloud data acquired dynamically by lidar, the latter mostly uses range image-based methods with inevitable projection distortion defects because the non-uniformly distributed sparse points make it difficult to form a local neighborhood for information embedding. The model can still effectively extract features and encode temporal information. This generalization ability is crucial for robustness in practical applications.
[0069] In general, step S101 realizes efficient feature extraction and temporal information encoding of dynamic point cloud data by combining advanced point cloud processing technology and deep learning methods. This provides a rich and compact representation for the subsequent compression process, which not only improves compression efficiency but also ensures reconstruction quality.
[0070] Step S102: using the adaptive content feature embedding as the visual knowledge prior of the time index embedding, performing feature fusion on the time index embedding and the content feature embedding in the same spatial dimension to obtain a fused feature vector; inputting the fused feature vector into a decoder composed of an upsampling layer and an attention layer based on grouped vector attention to obtain a predicted point cloud frame output by the decoder, and updating the model parameters of the network model by back-propagating the calculated chamfer distance and mean square error loss between the predicted point cloud frame and the corresponding real point cloud frame until the similarity between the predicted point cloud frame and the real point cloud frame meets the threshold interval, thereby obtaining the optimal network model;
[0071] Optionally, in this embodiment, step S102 is the core link in the dynamic 3D point cloud compression method based on implicit neural expression, which involves feature fusion, decoding process and model optimization. This step is designed to effectively combine spatial and temporal information, generate high-quality predicted point cloud frames, and obtain the best network model through iterative optimization.
[0072] First, let's take a deep look at the feature fusion process. In this process, the adaptive content feature embedding is treated as the visual knowledge prior for the time index embedding. The rationale for this approach is that the content feature embedding contains the spatial structure information of the point cloud, which can provide richer context for the time index embedding. By fusing these two embeddings in the same spatial dimension, we can obtain a comprehensive representation that contains both spatial and temporal information.
[0073] The specific fusion process may involve a variety of techniques, such as concatenation, additive fusion, or more complex attention mechanisms. For example, we can use a multi-head attention mechanism with time index embedding as the query and content feature embedding as the key and value. In this way, the temporal information can selectively focus on the most relevant spatial features, thereby generating a more informative fused feature vector.
[0074] This fusion method solves a key technical problem in dynamic point cloud compression: how to effectively integrate spatiotemporal information. Traditional methods often have difficulty balancing spatial details and temporal coherence, while this fusion strategy allows the model to dynamically adjust the weights of spatial and temporal information, thereby achieving the best compression effect in different scenarios.
[0075] Next, the fused feature vector is input into the decoder. The structure of the decoder is similar to that of the encoder, consisting of an upsampling layer and an attention layer based on grouped vector attention. The upsampling layer gradually increases the number of points and restores the density of the point cloud. This is usually achieved by interpolation or learning. The attention layer helps the model focus on important features to more accurately reconstruct the details of the point cloud.
[0076] After the decoder outputs the predicted point cloud frame, the model calculates the loss between the predicted point cloud frame and the real point cloud frame. Two loss functions are used here: chamfer distance and mean square error. Chamfer distance is mainly used to evaluate the similarity between two point sets. It takes into account the positional relationship of the points and is helpful for maintaining the overall shape of the point cloud. Mean square error directly measures the coordinate difference of the points, which helps to improve the accuracy of reconstruction.
[0077] Through back propagation, the model updates the decoder parameters based on these losses. This process is iterated until the similarity between the predicted point cloud frame and the real point cloud frame reaches a preset threshold range. This iterative optimization method enables the model to gradually improve its reconstruction quality, and ultimately obtains an optimal network model that can reconstruct dynamic point clouds with high quality.
[0078] The technical effects of this step are mainly reflected in the following aspects: First, by fusing spatial and temporal information, the model can better understand and express the changes in dynamic point clouds, thereby retaining more meaningful information during compression. Second, the decoder structure based on grouped vector attention allows the model to focus on important features during the reconstruction process and improve the reconstruction quality. Finally, the use of multiple loss functions and iterative optimization methods ensures that the model can achieve high-quality reconstruction effects in various scenarios.
[0079] Let's illustrate the application of this step through a specific example. Suppose we are processing a dynamic point cloud sequence in a smart city monitoring system, which records the changes of a busy intersection over 24 hours, one frame per second, a total of 86,400 frames. Our goal is to efficiently compress this large-scale dynamic point cloud while ensuring the quality of reconstruction for long-term storage and subsequent analysis.
[0080] In the feature fusion stage, we adopt an improved multi-head attention mechanism. Specifically, we first map the 256-dimensional content feature embedding and the 128-dimensional time index embedding to the same 512-dimensional space through linear projection. Then, we use 8 attention heads, each of which independently calculates the attention weight. The time index embedding is used as the query and the content feature embedding is used as the key and value. In this way, the model can selectively focus on the most relevant spatial features based on the temporal information. Finally, we concatenate the outputs of the 8 heads and pass them through a feed-forward network to obtain the final 512-dimensional fused feature vector.
[0081] The decoder structure consists of 5 upsampling layers and 6 attention layers. Each upsampling layer doubles the number of points and uses an inverse distance weighted interpolation method. In the attention layer, we also use grouped vector attention, dividing the value vector into 6 groups, each of which is independently weighted by the correlation weight obtained from the query vector and the key vector. This design allows the decoder to have finer control over the details of the point cloud during reconstruction.
[0082] During the optimization process, we set a dynamic loss weight strategy. In the early stage of training, we pay more attention to the mean square error to quickly converge to the roughly correct shape. As the training progresses, we gradually increase the weight of the chamfer distance to improve the overall structure and details of the point cloud. The similarity threshold we set is that the chamfer distance is less than 0.01 meters and the mean square error is less than 0.0001 square meters.
[0083] Through experiments, we found that this method can effectively handle complex dynamic scenes. For example, during rush hour in the morning and evening, when the traffic volume at the intersection increases significantly, the model can accurately capture the changes in point cloud density while maintaining the stability of static structures (such as buildings and traffic lights). During late night hours when there are few pedestrians, the model can adapt to the sparse point cloud and still accurately reconstruct the scene.
[0084] It is particularly noteworthy that this method performs well when processing long time series. By integrating spatiotemporal features, the model is able to capture periodic variation patterns throughout the day, such as the tidal phenomenon of traffic flow. This not only improves compression efficiency, but also provides valuable information for subsequent traffic analysis.
[0085] In terms of performance, our method has achieved significant improvements over traditional frame-by-frame compression methods. At the same compression ratio (about 100:1), our method improves PSNR (peak signal-to-noise ratio) by an average of 3dB and SSIM (structural similarity) by 0.05. Especially when dealing with fast-changing scenes such as traffic accidents or emergencies, our method can better preserve key details, which is helpful for subsequent event analysis and reconstruction.
[0086] In addition, this method also shows good generalization ability. When we apply the trained model to similar scenes in other cities, the model is still able to effectively compress and reconstruct point cloud sequences even with different road layouts and architectural styles. This generalization ability is crucial for large-scale deployment of smart city surveillance systems.
[0087] In terms of computational efficiency, although the training process is time-consuming (about 72 hours on 8 GPUs), once trained, the model is fast in the inference phase. On an ordinary workstation, it can process and compress point cloud streams at 1080p resolution and 30fps in real time, which makes it very suitable for real-time monitoring and analysis applications.
[0088] In general, step S102 achieves high-quality compression and reconstruction of dynamic point cloud data through an innovative feature fusion method and a carefully designed decoder structure.
[0089] Step S103: convert the floating point parameters of the optimal network model into low-bit integer parameters and then use Huffman coding to perform lossless compression to obtain the optimal network model after lossless compression, and perform dynamic three-dimensional point cloud data decoding through the optimal network model.
[0090] Optionally, in this embodiment, step S103 is the last key link of the dynamic 3D point cloud compression method based on implicit neural expression, which involves the quantization, compression and final decoding application of model parameters. This step is designed to further reduce the storage space of the model while maintaining the performance of the model, and ultimately achieve efficient dynamic 3D point cloud data decoding.
[0091] First, let's take a closer look at the process of converting floating-point parameters to low-bit integer parameters. This process is often called model quantization. The main purpose of model quantization is to reduce the storage space and computational complexity of the model while maintaining the performance of the model as much as possible. In this process, 32-bit floating-point numbers are usually converted to 8-bit or lower-bit integers.
[0092] The specific quantization process may involve a variety of techniques, such as linear quantization, nonlinear quantization, or mixed precision quantization. For example, in linear quantization, we can linearly map the floating point range to the integer range. This process can be expressed as: q = round (r (float_param-min) / (max-min)), where q is the quantized integer, r is the target integer range (such as 255 for 8-bit quantization), float_param is the original floating-point parameter, and min and max are the minimum and maximum values of the parameter.
[0093] The quantization process solves technical issues of model storage and inference efficiency. By converting 32-bit floating point numbers to 8-bit integers, we can reduce the model size by about 75% while significantly reducing computational complexity. This is particularly important for deploying models on resource-constrained devices, such as mobile devices or embedded systems.
[0094] However, quantization also brings a challenge: how to maintain the performance of the model while reducing the bit width. To solve this problem, quantization-aware training is usually required. In this process, the model simulates quantization operations during training so that the model can adapt to the accuracy loss caused by quantization.
[0095] After quantization, the next step is to use Huffman coding for lossless compression. Huffman coding is a compression algorithm based on the frequency of symbol occurrence. It assigns shorter codes to symbols that appear frequently and longer codes to symbols that appear less frequently, thereby achieving overall compression. In the context of the network model, different quantization parameter values can be regarded as different symbols.
[0096] The use of Huffman coding solves the technical problem of further compressing the model size. Compared with simple bit width reduction, Huffman coding can exploit the statistical characteristics of parameter distribution to achieve more efficient compression. This method is particularly suitable for neural network models because the parameters of neural networks usually show certain statistical regularities, such as a large number of weights close to zero.
[0097] After quantization and Huffman coding, we obtained a significantly compressed network model. This compressed model not only has a significantly reduced size, but also still maintains most of the performance of the original model. The technical effects of this compression method are mainly reflected in the following aspects: First, the storage space of the model is greatly reduced, making the model easier to deploy on various devices. Second, the computational complexity of the model is reduced and the inference speed is improved. Finally, through the carefully designed quantization and compression strategy, the performance of the model is maintained, ensuring high-quality point cloud decoding results.
[0098] Finally, this compressed optimal network model is used to perform decoding of dynamic 3D point cloud data. During the decoding process, the model parameters first need to be decompressed and dequantized. The decompression process includes Huffman decoding and converting integer parameters back to floating point numbers. The decoder then uses these recovered parameters to reconstruct the dynamic point cloud sequence.
[0099] Let's illustrate the application of this step through a specific example. Suppose we are developing a real-time environmental perception system for self-driving cars. This system needs to process high-resolution dynamic point cloud data from multiple lidar sensors and perform real-time decoding and analysis on the on-board computing unit. Our goal is to minimize the model size and increase the decoding speed while ensuring the decoding quality.
[0100] During quantization, we adopted a mixed-precision quantization strategy. Specifically, we used different quantization bit widths for different parts of the model. For the main body of the encoder and decoder, we used 8-bit quantization, which can significantly reduce the model size while maintaining most of the performance. For some layers that are particularly sensitive to precision, such as the final output layer, we retain 16-bit precision to ensure sufficient expressiveness. To compensate for the accuracy loss caused by quantization, we performed quantization-aware fine-tuning for 2000 training batches.
[0101] After quantization, we conducted a detailed statistical analysis of the model parameters and found that the parameter distribution showed obvious peak characteristics, with a large number of parameters concentrated around a few specific values. Based on this observation, we designed an adaptive Huffman coding scheme. We divided the parameter value range into multiple intervals and performed Huffman coding on the parameter values in each interval separately. This method can better adapt to the distribution characteristics of different parts of the parameters and further improve the compression efficiency.
[0102] Through this quantization and compression strategy, we compressed the original floating-point model of about 500MB to only 25MB, with a compression ratio of 20:1. In terms of decoding performance, the compressed model has less than 1% performance degradation in point cloud reconstruction quality, but the inference speed is increased by more than 3 times. This means that our system can process point cloud data streams from multiple high-resolution lidars in real time on ordinary on-board computing units, and can process more than 1 million points per second.
[0103] In practical applications, this compressed model has shown excellent performance. For example, in a complex urban driving scene, the model can accurately reconstruct and identify surrounding vehicles, pedestrians, buildings and other objects. Even at high speeds, the model can maintain stable decoding performance and provide reliable environmental perception information for the autonomous driving system. Especially when dealing with fast-changing dynamic objects, such as suddenly appearing pedestrians or vehicles turning sharply, our model has shown superior response speed and accuracy compared to traditional methods.
[0104] In addition, the compressed model also shows good robustness. Under different weather conditions (such as rainy days and foggy days), the model can still effectively decode point cloud data with only slight performance degradation. This robustness is crucial for the safety and reliability of autonomous driving systems.
[0105] In terms of energy consumption, the compressed model reduces the power consumption of the vehicle system by about 65%. This not only extends the range of electric vehicles, but also reduces heat dissipation requirements and simplifies the design of the vehicle computing system.
[0106] Finally, another important advantage of this compression method is that it greatly reduces the bandwidth requirements for model updates and distribution. In the field of autonomous driving, models often need to be updated based on newly collected data. The compressed model can be transmitted to each vehicle via wireless network more quickly, ensuring that all vehicles can obtain the latest environmental perception capabilities in a timely manner.
[0107] In general, step S103 successfully reduces the high-performance dynamic 3D point cloud decoding model to a size suitable for deployment in resource-constrained environments through innovative quantization and compression techniques. This method not only solves the problems of storage and computing efficiency, but also maintains the high performance of the model, providing strong technical support for applications such as autonomous driving and augmented reality that require real-time processing of large-scale dynamic point cloud data.
[0108] From the above description, it can be seen that the dynamic three-dimensional point cloud compression method based on implicit neural expression provided in the embodiment of the present application can embed and extract the point cloud data of each frame in a given dynamic point cloud sequence through an encoder composed of a downsampling layer and an attention layer based on grouped vector attention, so as to obtain an adaptive content feature embedding containing point cloud content information; obtain the index value of the point cloud data of each frame in the given dynamic point cloud sequence, and map the index value of the point cloud data to a high-dimensional embedding space through a position encoding function for embedding extraction, so as to obtain a time index embedding containing point cloud time information; convert the floating-point parameters of the optimal network model into low-bit integer parameters and then use Huffman coding for lossless compression to obtain the optimal network model after lossless compression, and perform dynamic three-dimensional point cloud data decoding through the optimal network model, thereby effectively improving the efficiency and accuracy of three-dimensional point cloud compression.
[0109] In one embodiment of the dynamic three-dimensional point cloud compression method based on implicit neural expression of the present application, see Figure 2 , and can also include the following:
[0110] Step S201: input a frame of point cloud data in a given dynamic point cloud sequence, divide the points in the point cloud data into voxel grids by a pooling method based on voxel grid division, cluster the points in each grid and perform a maximum pooling operation on the point cloud features, so as to obtain local features of the point cloud through a downsampling layer;
[0111] Step S202: input the local features of the point cloud output by the downsampling layer into the grouped vector attention mechanism, wherein the grouped vector attention mechanism maps the input local features of the point cloud into a query vector, a key vector and a value vector, groups the channels of the value vector, and encodes the relationship between the query vector and the key vector to obtain a weight encoding, i.e., the correlation weight between the local features of the point cloud, wherein the number of weight encoding channels is the same as the number of channel groups of the value vector; perform a Hadamard product with the correlation weight and the value vector to obtain the final adaptive content feature embedding.
[0112] Optionally, in this embodiment, the adaptive content encoder is composed of 4 encoding blocks, each of which is composed of a downsampling layer and several attention layers connected in series. The downsampling layer adopts a pooling method based on voxel grid division, called grid pooling. The core idea is to divide the points in the point cloud into voxel grids, then cluster the points in each grid, and perform maximum pooling on the features of the point cloud; in the grouped vector attention layer, when using Layer or linear mapping, the features Convert to query vector , the key vector Sum value vector Finally, unlike the traditional scalar attention, which directly uses the scalar dot product of the query vector and the key vector as the attention weight, the vector attention pays attention to the fact that the weights of features in different channels may be different. Therefore, it encodes the relationship between the query vector and the key vector and then multiplies it with the value vector to adjust each feature channel separately. The vector attention is expressed by the following formula:
[0113] ,
[0114] in, It is a feature The query vector, and Characteristics The key vector and value vector of express The set of points closest to , It is a feature Features The relevance weight of is a relational function (such as subtraction), It is a Hadamard product. represents a weight encoding function that independently reweights the value vectors by channel.
[0115] In addition, in order to solve the problem caused by the deepening of the network and the increase of channels The parameter size of the layer increases dramatically, which greatly limits the efficiency of the model. Grouped vector attention divides the value vector into group, and The layer corresponds to the output The weight encoding of each channel ( The number of channels of the value vector (where is the number of channels of the value vector) shares one weight, thereby limiting the number of parameters in the attention layer.
[0116] Through this step, the adaptive content embedding feature corresponding to a certain frame point cloud in the point cloud frame sequence can be obtained, and the final size is ,in It is the number of remaining points after four grid sampling of the original point cloud frame.
[0117] In one embodiment of the dynamic three-dimensional point cloud compression method based on implicit neural expression of the present application, see Figure 3 , and can also include the following:
[0118] Step S301: obtaining an index value of each frame of point cloud data for a given dynamic point cloud sequence, wherein the index value represents a time index of the point cloud frame in the entire sequence;
[0119] Step S302: normalize the index value of the acquired point cloud frame to the interval [0,1] and input it into the position encoding function to map the one-dimensional time index value to a high-dimensional embedding space to obtain a time index embedding containing the point cloud time information.
[0120] Optionally, in this embodiment, the index of the point cloud frame input by the content feature module in the point cloud sequence is first obtained and normalized to [0,1]. To map the one-dimensional time index to a high-dimensional embedding space, this application uses a position encoding function. Because deep networks tend to learn low-frequency functions, directly inputting the time index may mislead the network to believe that some similar inputs have a simple linear relationship, etc., thereby producing similar results and making it difficult to fit data with high-frequency changes. Formally, the function used in this application is as follows:
[0121] ,
[0122] in, is normalized to The time index, and are hyperparameters, which are set to 1.25 and 160 by default in this application. This formula can map a single scalar to a length of vector, which contains richer high-frequency information changes.
[0123] In addition, the function In the above example, considering that the implicit expression model depends on Generate a time-indexed embedding containing adaptive content information of the real point cloud frame (which must be aligned with the bottleneck dimension of the adaptive content encoder) to reconstruct point cloud frame, which will lead to a large amount of parameter redundancy concentrated in The last layer of the content encoder bottleneck layer after the fourth grid sampling increases the number of points, The number of neurons in the last layer must be increased accordingly. To this end, based on the consistency of the mapping relationship between the time index and the corresponding point cloud frame, this application divides each point cloud frame into A cube grid with the same index The geometric and attribute information of the point clouds in the eight sub-cubes can be surjected. In order to prevent the damage to the model representation ability caused by a large amount of continuous sparse space in the dynamic point cloud sequence, the embodiment of the present application randomly mixes the point cloud content information of each sub-cube. This strategy significantly reduces parameter redundancy while maintaining the relative stability of the model performance and effectively avoids potential problems such as model complexity and overfitting caused by too many neurons. For example, when When the hidden layer is not divided into 512 The number of parameters will reach level, and after the division Size , the number of parameters is reduced compared to .
[0124] Therefore, after the normalized time index is input into the position encoding function to be mapped to the high-dimensional embedding space, the default setting is Channel Dimension layer, and obtain the time index embedding corresponding to the content information in the adaptive content encoder.
[0125] In one embodiment of the dynamic three-dimensional point cloud compression method based on implicit neural expression of the present application, see Figure 4 , and can also include the following:
[0126] Step S401: aligning the time index embedding and the content feature embedding in the spatial dimension;
[0127] Step S402: embedding the aligned adaptive content features as visual knowledge priors of the time index embedding to perform feature fusion to obtain a fused feature vector.
[0128] Optionally, in this embodiment, considering the sparsity and disorder of point cloud data in space, and the huge challenge posed to the completely implicit expression model by the rich information covered by the number of nodes on the order of one hundred thousand, and taking into account the high dimensionality of the adaptive content feature embedding of point cloud data, directly storing the content embeddings of all point cloud frames will greatly reduce the compression efficiency. Therefore, this application proposes a time-content feature fusion method suitable for implicit expression model mapping point cloud data.
[0129] Specifically, during the training phase, this application embeds adaptive content features as visual knowledge priors for time-indexed embedding and aligns the two in the spatial dimension.
[0130] Therefore, the time-content information embedding can be operated (such as linear interpolation) in the same spatial dimension, and the time index embedding can be guided by the adaptive content feature embedding, thereby improving the regression ability of the model and accelerating the fitting process of the point cloud frame. Finally, the time index embedding module is endowed with the content information representation ability of the point cloud frame and establishes an association with the actual content. Based on this, during the model compression process, the parameters of the content encoder will no longer be saved, thereby further improving the model compression efficiency.
[0131] In one embodiment of the dynamic three-dimensional point cloud compression method based on implicit neural expression of the present application, see Figure 5 , and can also include the following:
[0132] Step S501: inputting the fused feature vector into a decoder composed of an upsampling layer and an attention layer based on grouped vector attention, wherein the network of the decoder is composed of 4 decoding blocks, each of which includes an upsampling layer for gradually expanding the feature space resolution and an attention layer based on grouped vector attention for extracting correlation information between features;
[0133] Step S502: After feature processing by the decoder network, a predicted point cloud frame output by the decoder is obtained.
[0134] Optionally, in this embodiment, a fusion feature decoder is used to reconstruct the predicted point cloud frame by fusion features of spatiotemporal information.
[0135] Specifically, the decoder of this step consists of 4 decoding blocks (DecodingBlock), each decoding block is connected by an upsampling layer and a default number of grouped vector attention layers of 1, 1, 1, 1 respectively.
[0136] After obtaining the fused features, input them into the decoder ,Depend on The feature embedding of size is reconstructed as Features ( represents the number of points in the real point cloud frame), and then inputs the reconstruction head (ReconstructionHead) to obtain Dimensional predicted point cloud frame ( represents the dimension of point cloud data input by the embodiment).
[0137] Then, this application uses Chamfer Distance and L2 loss to constrain the prediction results to increase the similarity with the real point cloud frame. The specific formula is as follows:
[0138] ,
[0139] in, , real point cloud frame , predict point cloud frame .also, is the number of frames, is a hyperparameter that controls the loss weight, They are the input frame index and content information respectively.
[0140] In one embodiment of the dynamic three-dimensional point cloud compression method based on implicit neural expression of the present application, see Figure 6 , and can also include the following:
[0141] Step S601: Calculate the chamfer distance and mean square error loss between the predicted point cloud frame and the corresponding real point cloud frame, and perform weighted fusion on the loss functions of the chamfer distance and the mean square error loss to obtain a final loss function;
[0142] Step S602: updating the model parameters of the network model by back propagation according to the loss function until the similarity between the predicted point cloud frame and the real point cloud frame meets the threshold interval, thereby obtaining the optimal network model.
[0143] Optionally, in this embodiment, the above-mentioned model iteration parameter adjustment steps are repeated until a preset number of iteration rounds is reached. At the end of each iteration, the model parameters are saved, and the average point-to-point peak signal-to-noise ratio, point-to-surface peak signal-to-noise ratio and chamfer distance index of the predicted point cloud frame sequence are calculated. The PSNR index calculation formula is as follows:
[0144] ,
[0145] ,
[0146] in , , is the predicted point cloud frame, is the maximum reference distance value in the point cloud, is a point in the original point cloud, is the predicted point cloud with The most matching point, is the number of points in the point cloud, Yes The normal vector of the surface in the original point cloud.
[0147] In one embodiment of the dynamic three-dimensional point cloud compression method based on implicit neural expression of the present application, see Figure 7 , and can also include the following:
[0148] Step S701: Convert the 32-bit or 16-bit floating point number parameters of the optimal network model into 8-bit integer parameters;
[0149] Step S702: Input the 8-bit integer parameter of the optimal network model into the Huffman coding algorithm to allocate codes of different lengths based on the frequency of occurrence of the parameter value, so as to obtain the optimal network model after lossless compression.
[0150] Optionally, in this embodiment, after the training is completed, the saved best network model is compressed, such as model pruning, model quantization, and weight encoding. In the embodiment of the present invention, 8-bit model quantization is used by default, and the model weights are converted from 32-bit or 16-bit floating point numbers to low-bit representation. Subsequently, Huffman coding is used to further compress the spatial overhead of the model, and finally the compression task of the dynamic three-dimensional point cloud is completed through model compression.
[0151] In order to effectively improve the efficiency and accuracy of 3D point cloud compression, the present application provides an embodiment of a dynamic 3D point cloud compression device based on implicit neural expression for realizing all or part of the content of the dynamic 3D point cloud compression method based on implicit neural expression, see Figure 8 The dynamic three-dimensional point cloud compression device based on implicit neural expression specifically includes the following contents:
[0152] The data processing module 10 is used to embed and extract the point cloud data of each frame in a given dynamic point cloud sequence through an encoder composed of a downsampling layer and an attention layer based on grouped vector attention, so as to obtain an adaptive content feature embedding containing point cloud content information; obtain the index value of the point cloud data of each frame in the given dynamic point cloud sequence, and map the index value of the point cloud data to a high-dimensional embedding space through a position encoding function for embedding extraction, so as to obtain a time index embedding containing point cloud time information;
[0153] The model training module 20 is used to use the adaptive content feature embedding as the visual knowledge prior of the time index embedding to perform feature fusion on the time index embedding and the content feature embedding in the same spatial dimension to obtain a fused feature vector; the fused feature vector is input into a decoder composed of an upsampling layer and an attention layer based on grouped vector attention to obtain a predicted point cloud frame output by the decoder, and the model parameters of the network model are updated by back propagation of the calculated chamfer distance and mean square error loss between the predicted point cloud frame and the corresponding real point cloud frame until the similarity between the predicted point cloud frame and the real point cloud frame meets the threshold interval, thereby obtaining the best network model;
[0154] The model compression module 30 is used to convert the floating-point parameters of the optimal network model into low-bit integer parameters and then use Huffman coding to perform lossless compression to obtain the optimal network model after lossless compression, and perform dynamic three-dimensional point cloud data decoding through the optimal network model.
[0155] From the above description, it can be seen that the dynamic three-dimensional point cloud compression device based on implicit neural expression provided in the embodiment of the present application can embed and extract the point cloud data of each frame in a given dynamic point cloud sequence through an encoder composed of a downsampling layer and an attention layer based on grouped vector attention, so as to obtain an adaptive content feature embedding containing point cloud content information; obtain the index value of the point cloud data of each frame in the given dynamic point cloud sequence, and map the index value of the point cloud data to a high-dimensional embedding space through a position encoding function for embedding extraction, so as to obtain a time index embedding containing point cloud time information; convert the floating-point parameters of the optimal network model into low-bit integer parameters and then use Huffman coding for lossless compression to obtain the optimal network model after lossless compression, and perform dynamic three-dimensional point cloud data decoding through the optimal network model, thereby effectively improving the efficiency and accuracy of three-dimensional point cloud compression.
[0156] From the hardware level, in order to effectively improve the efficiency and accuracy of three-dimensional point cloud compression, the present application provides an embodiment of an electronic device for implementing all or part of the content of the dynamic three-dimensional point cloud compression method based on implicit neural expression, and the electronic device specifically includes the following content:
[0157] Processor, memory, communication interface and bus; wherein the processor, memory and communication interface communicate with each other through the bus; the communication interface is used to realize information transmission between the dynamic three-dimensional point cloud compression device based on implicit neural expression and related equipment such as core business systems, user terminals and related databases; the logic controller can be a desktop computer, a tablet computer and a mobile terminal, etc., but this embodiment is not limited to this. In this embodiment, the logic controller can be implemented with reference to the embodiment of the dynamic three-dimensional point cloud compression method based on implicit neural expression and the embodiment of the dynamic three-dimensional point cloud compression device based on implicit neural expression in the embodiment, and the contents are merged here, and the repeated parts are not repeated.
[0158] It is understandable that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.
[0159] In practical applications, part of the dynamic three-dimensional point cloud compression method based on implicit neural expression can be executed on the electronic device side as described above, or all operations can be completed in the client device. The specific selection can be based on the processing capability of the client device and the limitations of the user's usage scenario. This application does not limit this. If all operations are completed in the client device, the client device may also include a processor.
[0160] The client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side, and other implementation scenarios may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster consisting of multiple servers, or a server structure of a distributed device.
[0161] Fig. 9 FIG. 9 is a schematic block diagram of the system structure of the electronic device 9600 according to an embodiment of the present application. Fig. 9 As shown, the electronic device 9600 may include a central processor 9100 and a memory 9140; the memory 9140 is coupled to the central processor 9100. It is worth noting that Fig. 9 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.
[0162] In one embodiment, the function of the dynamic three-dimensional point cloud compression method based on implicit neural expression can be integrated into the central processing unit 9100. The central processing unit 9100 can be configured to perform the following control:
[0163] Step S101: embedding and extracting the point cloud data of each frame in a given dynamic point cloud sequence through an encoder composed of a downsampling layer and an attention layer based on grouped vector attention, so as to obtain an adaptive content feature embedding containing point cloud content information; obtaining the index value of the point cloud data of each frame in the given dynamic point cloud sequence, and mapping the index value of the point cloud data to a high-dimensional embedding space through a position encoding function for embedding extraction, so as to obtain a time index embedding containing point cloud time information;
[0164] Step S102: using the adaptive content feature embedding as the visual knowledge prior of the time index embedding, performing feature fusion on the time index embedding and the content feature embedding in the same spatial dimension to obtain a fused feature vector; inputting the fused feature vector into a decoder composed of an upsampling layer and an attention layer based on grouped vector attention to obtain a predicted point cloud frame output by the decoder, and updating the model parameters of the network model by back-propagating the calculated chamfer distance and mean square error loss between the predicted point cloud frame and the corresponding real point cloud frame until the similarity between the predicted point cloud frame and the real point cloud frame meets the threshold interval, thereby obtaining the optimal network model;
[0165] Step S103: convert the floating point parameters of the optimal network model into low-bit integer parameters and then use Huffman coding to perform lossless compression to obtain the optimal network model after lossless compression, and perform dynamic three-dimensional point cloud data decoding through the optimal network model.
[0166] From the above description, it can be seen that the electronic device provided by the embodiment of the present application embeds and extracts the point cloud data of each frame in a given dynamic point cloud sequence through an encoder composed of a downsampling layer and an attention layer based on grouped vector attention, so as to obtain an adaptive content feature embedding containing point cloud content information; obtains the index value of the point cloud data of each frame in the given dynamic point cloud sequence, and maps the index value of the point cloud data to a high-dimensional embedding space through a position encoding function for embedding extraction, so as to obtain a time index embedding containing point cloud time information; converts the floating-point parameters of the optimal network model into low-bit integer parameters and then uses Huffman coding for lossless compression to obtain the optimal network model after lossless compression, and performs dynamic three-dimensional point cloud data decoding through the optimal network model, thereby effectively improving the efficiency and accuracy of three-dimensional point cloud compression.
[0167] In another embodiment, the dynamic three-dimensional point cloud compression device based on implicit neural expression can be configured separately from the central processing unit 9100. For example, the dynamic three-dimensional point cloud compression device based on implicit neural expression can be configured as a chip connected to the central processing unit 9100, and the function of the dynamic three-dimensional point cloud compression method based on implicit neural expression can be realized through the control of the central processing unit.
[0168] like Fig. 9 As shown, the electronic device 9600 may also include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily have to include Fig. 9 In addition, the electronic device 9600 may also include Fig. 9 For components not shown, reference may be made to the prior art.
[0169] like Fig. 9 As shown, the central processing unit 9100 is sometimes also referred to as a controller or an operation control, and may include a microprocessor or other processor device and / or logic device. The central processing unit 9100 receives input and controls the operation of various components of the electronic device 9600.
[0170] The memory 9140 may be, for example, one or more of a cache, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory or other suitable devices. The above-mentioned information related to the failure may be stored, and a program for executing the relevant information may also be stored. The CPU 9100 may execute the program stored in the memory 9140 to implement information storage or processing, etc.
[0171] The input unit 9120 provides input to the central processing unit 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to provide power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display may be, for example, an LCD display, but is not limited thereto.
[0172] The memory 9140 may be a solid-state memory, such as a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It may also be a memory that saves information even when the power is off, can be selectively erased, and is provided with more data, examples of which are sometimes referred to as EPROMs, etc. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142, which is used to store application programs and function programs or processes for executing the operation of the electronic device 9600 through the central processor 9100.
[0173] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers for communication functions of the electronic device and / or for executing other functions of the electronic device (such as messaging applications, address book applications, etc.).
[0174] The communication module 9110 is a transmitter / receiver that sends and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processor 9100 to provide input signals and receive output signals, which may be the same as the case of a conventional mobile communication terminal.
[0175] Based on different communication technologies, multiple communication modules 9110 may be provided in the same electronic device, such as a cellular network module, a Bluetooth module and / or a wireless LAN module. The communication module 9110 (transmitter / receiver) is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, thereby realizing a common telecommunication function. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. In addition, the audio processor 9130 is also coupled to the central processor 9100, so that recording can be performed on the local machine through the microphone 9132, and the sound stored on the local machine can be played through the speaker 9131.
[0176] The embodiments of the present application also provide a computer-readable storage medium capable of implementing all the steps of the dynamic three-dimensional point cloud compression method based on implicit neural expression in the above-mentioned embodiment, where the execution subject is a server or a client. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, all the steps of the dynamic three-dimensional point cloud compression method based on implicit neural expression in the above-mentioned embodiment are implemented. For example, when the processor executes the computer program, the following steps are implemented:
[0177] Step S101: embedding and extracting the point cloud data of each frame in a given dynamic point cloud sequence through an encoder composed of a downsampling layer and an attention layer based on grouped vector attention, so as to obtain an adaptive content feature embedding containing point cloud content information; obtaining the index value of the point cloud data of each frame in the given dynamic point cloud sequence, and mapping the index value of the point cloud data to a high-dimensional embedding space through a position encoding function for embedding extraction, so as to obtain a time index embedding containing point cloud time information;
[0178] Step S102: using the adaptive content feature embedding as the visual knowledge prior of the time index embedding, performing feature fusion on the time index embedding and the content feature embedding in the same spatial dimension to obtain a fused feature vector; inputting the fused feature vector into a decoder composed of an upsampling layer and an attention layer based on grouped vector attention to obtain a predicted point cloud frame output by the decoder, and updating the model parameters of the network model by back-propagating the calculated chamfer distance and mean square error loss between the predicted point cloud frame and the corresponding real point cloud frame until the similarity between the predicted point cloud frame and the real point cloud frame meets the threshold interval, thereby obtaining the optimal network model;
[0179] Step S103: convert the floating point parameters of the optimal network model into low-bit integer parameters and then use Huffman coding to perform lossless compression to obtain the optimal network model after lossless compression, and perform dynamic three-dimensional point cloud data decoding through the optimal network model.
[0180] From the above description, it can be seen that the computer-readable storage medium provided in the embodiment of the present application embeds and extracts the point cloud data of each frame in a given dynamic point cloud sequence through an encoder composed of a downsampling layer and an attention layer based on grouped vector attention, so as to obtain an adaptive content feature embedding containing point cloud content information; obtains the index value of the point cloud data of each frame in the given dynamic point cloud sequence, and maps the index value of the point cloud data to a high-dimensional embedding space through a position encoding function for embedding extraction, so as to obtain a time index embedding containing point cloud time information; converts the floating-point parameters of the optimal network model into low-bit integer parameters and then uses Huffman coding for lossless compression to obtain the optimal network model after lossless compression, and performs dynamic three-dimensional point cloud data decoding through the optimal network model, thereby effectively improving the efficiency and accuracy of three-dimensional point cloud compression.
[0181] The embodiments of the present application also provide a computer program product capable of implementing all the steps of the dynamic three-dimensional point cloud compression method based on implicit neural expression in the above embodiments, where the execution subject is a server or a client. When the computer program / instruction is executed by a processor, the steps of the dynamic three-dimensional point cloud compression method based on implicit neural expression are implemented. For example, the computer program / instruction implements the following steps:
[0182] Step S101: embedding and extracting the point cloud data of each frame in a given dynamic point cloud sequence through an encoder composed of a downsampling layer and an attention layer based on grouped vector attention, so as to obtain an adaptive content feature embedding containing point cloud content information; obtaining the index value of the point cloud data of each frame in the given dynamic point cloud sequence, and mapping the index value of the point cloud data to a high-dimensional embedding space through a position encoding function for embedding extraction, so as to obtain a time index embedding containing point cloud time information;
[0183] Step S102: using the adaptive content feature embedding as the visual knowledge prior of the time index embedding, performing feature fusion on the time index embedding and the content feature embedding in the same spatial dimension to obtain a fused feature vector; inputting the fused feature vector into a decoder composed of an upsampling layer and an attention layer based on grouped vector attention to obtain a predicted point cloud frame output by the decoder, and updating the model parameters of the network model by back-propagating the calculated chamfer distance and mean square error loss between the predicted point cloud frame and the corresponding real point cloud frame until the similarity between the predicted point cloud frame and the real point cloud frame meets the threshold interval, thereby obtaining the optimal network model;
[0184] Step S103: convert the floating point parameters of the optimal network model into low-bit integer parameters and then use Huffman coding to perform lossless compression to obtain the optimal network model after lossless compression, and perform dynamic three-dimensional point cloud data decoding through the optimal network model.
[0185] From the above description, it can be seen that the computer program product provided by the embodiment of the present application embeds and extracts the point cloud data of each frame in a given dynamic point cloud sequence through an encoder composed of a downsampling layer and an attention layer based on grouped vector attention, so as to obtain an adaptive content feature embedding containing point cloud content information; obtains the index value of the point cloud data of each frame in a given dynamic point cloud sequence, and maps the index value of the point cloud data to a high-dimensional embedding space through a position encoding function for embedding extraction, so as to obtain a time index embedding containing point cloud time information; converts the floating-point parameters of the optimal network model into low-bit integer parameters and then uses Huffman coding for lossless compression to obtain the optimal network model after lossless compression, and performs dynamic three-dimensional point cloud data decoding through the optimal network model, thereby effectively improving the efficiency and accuracy of three-dimensional point cloud compression.
[0186] It should be understood by those skilled in the art that embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0187] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0188] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0189] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0190] The present invention uses specific embodiments to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A dynamic three-dimensional point cloud compression method based on implicit neural expression, characterized in that: The method comprises: The point cloud data of each frame in a given dynamic point cloud sequence is embedded and extracted by an encoder composed of a downsampling layer and an attention layer based on grouped vector attention, so as to obtain an adaptive content feature embedding containing point cloud content information, specifically comprising: inputting a certain frame of point cloud data in a given dynamic point cloud sequence, dividing the points in the point cloud data according to voxel grids by a pooling method based on voxel grid division, clustering the points in each grid and performing a maximum pooling operation on the point cloud features, so as to obtain the local features of the point cloud through the downsampling layer; inputting the local features of the point cloud output by the downsampling layer into the grouped vector attention mechanism, wherein the grouped vector attention mechanism maps the input point cloud local features into a query vector, a key vector and a value vector, groups the channels of the value vector, and encodes the relationship between the query vector and the key vector to obtain a weight encoding, that is, a correlation weight between the local features of the point cloud, wherein the number of weight encoding channels is the same as the number of channel groups of the value vector; performing a Hadamard product with the correlation weight and the value vector to obtain the final adaptive content feature embedding; Obtaining the index value of the point cloud data of each frame in the given dynamic point cloud sequence, and mapping the index value of the point cloud data to a high-dimensional embedding space for embedding extraction through a position encoding function, so as to obtain a time index embedding containing time information of the point cloud; The adaptive content feature embedding is used as the visual knowledge prior of the time index embedding to fuse the time index embedding and the content feature embedding in the same spatial dimension to obtain a fused feature vector; the fused feature vector is input into a decoder composed of an upsampling layer and an attention layer based on grouped vector attention to obtain a predicted point cloud frame output by the decoder, and the model parameters of the network model are updated by back propagation of the calculated chamfer distance and mean square error loss between the predicted point cloud frame and the corresponding real point cloud frame until the similarity between the predicted point cloud frame and the real point cloud frame meets the threshold interval, thereby obtaining the best network model; The floating-point parameters of the optimal network model are converted into low-bit integer parameters and then losslessly compressed using Huffman coding to obtain the optimal network model after lossless compression, and dynamic three-dimensional point cloud data decoding is performed using the optimal network model.
2. The dynamic three-dimensional point cloud compression method based on implicit neural expression according to claim 1 is characterized in that: The step of obtaining the index value of the point cloud data of each frame in the given dynamic point cloud sequence, and mapping the index value of the point cloud data to a high-dimensional embedding space for embedding extraction through a position encoding function to obtain a time index embedding containing point cloud time information, includes: For a given dynamic point cloud sequence, an index value of each frame of point cloud data is obtained, wherein the index value represents a time index of the point cloud frame in the entire sequence; The index value of the acquired point cloud frame is normalized to the interval [0, 1] and input into the position encoding function to map the one-dimensional time index value into a high-dimensional embedding space to obtain a time index embedding containing the time information of the point cloud.
3. The dynamic three-dimensional point cloud compression method based on implicit neural expression according to claim 1 is characterized in that: The step of using the adaptive content feature embedding as the visual knowledge prior of the time index embedding to fuse the time index embedding and the content feature embedding in the same spatial dimension to obtain a fused feature vector includes: aligning the time index embedding and the content feature embedding in a spatial dimension; The aligned adaptive content feature embedding is used as the visual knowledge prior of the time index embedding to perform feature fusion to obtain a fused feature vector.
4. The dynamic three-dimensional point cloud compression method based on implicit neural expression according to claim 1 is characterized in that: The step of inputting the fused feature vector into a decoder composed of an upsampling layer and an attention layer based on grouped vector attention to obtain a predicted point cloud frame output by the decoder comprises: Inputting the fused feature vector into a decoder consisting of an upsampling layer and an attention layer based on grouped vector attention, wherein the network of the decoder consists of 4 decoding blocks, each of which includes an upsampling layer for gradually expanding the feature space resolution and an attention layer based on grouped vector attention for extracting correlation information between features; After feature processing of the decoder network, a predicted point cloud frame output by the decoder is obtained.
5. The dynamic three-dimensional point cloud compression method based on implicit neural expression according to claim 1 is characterized in that: The method updates the model parameters of the network model by back-propagating the calculated chamfer distance and the mean square error loss between the predicted point cloud frame and the corresponding real point cloud frame until the similarity between the predicted point cloud frame and the real point cloud frame meets the threshold interval, thereby obtaining the best network model, including: Calculating the chamfer distance and mean square error loss between the predicted point cloud frame and the corresponding real point cloud frame, and weightedly fusing the loss functions of the chamfer distance and the mean square error loss to obtain a final loss function; The model parameters of the network model are updated by back propagation according to the loss function until the similarity between the predicted point cloud frame and the real point cloud frame meets the threshold interval, thereby obtaining the optimal network model.
6. The dynamic three-dimensional point cloud compression method based on implicit neural expression according to claim 1, characterized in that: The method converts the floating point number parameters of the optimal network model into low-bit integer parameters and then uses Huffman coding to perform lossless compression to obtain the optimal network model after lossless compression, including: Converting the 32-bit or 16-bit floating point number parameters of the optimal network model into 8-bit integer parameters; The 8-bit integer parameters of the optimal network model are input into the Huffman coding algorithm to allocate codes of different lengths based on the frequency of occurrence of the parameter values, thereby obtaining the optimal network model after lossless compression.
7. A dynamic three-dimensional point cloud compression device based on implicit neural expression, characterized in that: The device comprises: The data processing module is used to embed and extract the point cloud data of each frame in a given dynamic point cloud sequence through an encoder composed of a downsampling layer and an attention layer based on grouped vector attention, so as to obtain an adaptive content feature embedding containing point cloud content information, specifically comprising: inputting a frame of point cloud data in a given dynamic point cloud sequence, dividing the points in the point cloud data according to voxel grids through a pooling method based on voxel grid division, clustering the points in each grid and performing a maximum pooling operation on the point cloud features, so as to obtain the local features of the point cloud through the downsampling layer; inputting the local features of the point cloud output by the downsampling layer into the grouped vector attention mechanism, wherein the partition The group vector attention mechanism maps the input point cloud local features into a query vector, a key vector and a value vector, groups the channels of the value vector, and encodes the relationship between the query vector and the key vector to obtain a weight encoding, that is, the correlation weight between the local features of the point cloud, wherein the number of weight encoding channels is the same as the number of channel groups of the value vector; performs a Hadamard product on the correlation weight and the value vector to obtain the final adaptive content feature embedding; obtains the index value of the point cloud data of each frame in the given dynamic point cloud sequence, and maps the index value of the point cloud data to a high-dimensional embedding space through a position encoding function for embedding extraction, and obtains a time index embedding containing the time information of the point cloud; A model training module is used to use the adaptive content feature embedding as the visual knowledge prior of the time index embedding to perform feature fusion on the time index embedding and the content feature embedding in the same spatial dimension to obtain a fused feature vector; the fused feature vector is input into a decoder composed of an upsampling layer and an attention layer based on grouped vector attention to obtain a predicted point cloud frame output by the decoder, and the model parameters of the network model are updated by back propagation of the calculated chamfer distance and mean square error loss between the predicted point cloud frame and the corresponding real point cloud frame until the similarity between the predicted point cloud frame and the real point cloud frame meets the threshold interval, thereby obtaining the best network model; The model compression module is used to convert the floating-point parameters of the optimal network model into low-bit integer parameters and then use Huffman coding for lossless compression to obtain the optimal network model after lossless compression, and perform dynamic three-dimensional point cloud data decoding through the optimal network model.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the dynamic three-dimensional point cloud compression method based on implicit neural expression described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the dynamic three-dimensional point cloud compression method based on implicit neural expression described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
A point cloud encoding method, a point cloud decoding method, and related devices
CN111699683A
Tree-based deep entropy model for point cloud compression
WO2024086154A1