Data compression transmission method for vehicle and road sensing system based on point cloud cluster

Through the data compression transmission method based on point cloud clusters, combined with point aggregation, space, timing and feature compression technologies, the balance of communication efficiency and perceptual performance in vehicle-road collaborative autonomous driving systems is solved, and efficient data transmission and perception accuracy are achieved.

CN120567940APending Publication Date: 2025-08-29BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510609604.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

In the existing vehicle-road collaborative autonomous driving system, the communication bandwidth and computing resources of the early fusion method have high demands, while the perception accuracy and robustness of the later fusion method are insufficient, making it difficult to meet the perception needs in complex dynamic scenarios.

Method used

The data compression transmission method based on point cloud cluster is adopted, through point aggregation coding, spatial compression, timing compression and feature compression, point cloud cluster data of roadside and vehicle side data is obtained, and fused to generate target fusion data for vehicle road perception tasks.

Benefits of technology

It realizes the improvement of perceptual performance while reducing communication costs. It is suitable for autonomous driving applications in complex dynamic scenarios, solves the problems of low communication efficiency and limited perceptual performance, and improves perception accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120567940A_ABST
    Figure CN120567940A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle and road sensing system data compression transmission method based on a point cloud cluster, and the method comprises the steps: obtaining road side data and vehicle side data, and carrying out the point aggregation coding of the road side data and the vehicle side data, so as to obtain road side point cloud cluster data and vehicle side point cloud cluster data corresponding to a foreground object; performing space compression, time sequence compression and feature compression on the roadside point cloud cluster data to obtain roadside target point cloud cluster data; and fusing the roadside target point cloud cluster data and the vehicle side point cloud cluster data to obtain target fusion data for the vehicle and road sensing task. According to the method, the communication efficiency and the sensing performance can be balanced at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data compression and transmission, and in particular to a data compression and transmission method for a vehicle-road perception system based on point cloud clusters. Background Art

[0002] Vehicle-road cooperative perception technology can be mainly divided into early fusion technology and late fusion technology according to different stages of data sharing and collaboration.

[0003] Early-stage fusion technology, also known as data-level collaboration or early collaborative technology, focuses on the comprehensive fusion of raw data collected by vehicles and roadside infrastructure. This stage is characterized by an intelligent agent (such as an autonomous vehicle) receiving raw sensor data from other agents (such as other vehicles or roadside sensors) and combining this data with its own collected data to obtain comprehensive data information, thereby improving perception accuracy. Currently, point cloud data collected by intelligent agents is fused for perception by leveraging their irregularities and clustering. For example, collaborative 3D object detection system models and collaborative perception models based on 3D point clouds achieve fused prediction of point cloud data by reconstructing and spatially connecting LiDAR data, achieving promising results. However, sharing raw sensor data requires a large amount of communication bandwidth and is prone to congestion in communication networks due to excessive data load, which in most cases hinders its practical application. Long processing and transmission times also make early-stage fusion technology difficult to meet the requirements of real-time perception.

[0004] Late-stage fusion technology, also known as result-level collaborative perception technology, focuses on fusing the perception results independently generated by each intelligent agent. Vehicle-side and road-side sensors each perform local processing on the collected raw data to extract key features or generate perception results. These processed features or results are then transmitted to the central processing unit for fusion to generate comprehensive perception results for autonomous driving decisions. This approach significantly reduces communication bandwidth requirements because the amount of data transmitted is much smaller than the original data. However, because the processed information is transmitted, late-stage fusion methods may lose details in the original data, resulting in insufficient perception accuracy and robustness. Especially in complex and dynamic scenes, this information loss can seriously affect the accuracy of perception tasks, limiting its use in autonomous driving applications that require high-precision perception.

[0005] Although early fusion and late fusion methods have been widely used in vehicle-road cooperative autonomous driving systems, they each have some significant shortcomings, which limit the overall performance of the system and the feasibility of practical applications.

[0006] While early fusion methods can provide highly accurate perception results, they place extremely high demands on communication bandwidth and computing resources. Due to the large amount of raw data required, early fusion methods are prone to transmission delays and resource waste. Furthermore, the central processing unit must process a massive amount of raw data, which not only increases computational complexity but also places higher demands on hardware. In practical applications, these high bandwidth and computational complexity requirements make early fusion methods difficult to scale up.

[0007] While late-stage fusion methods significantly reduce communication bandwidth requirements by transmitting processed features or results, their perception performance is limited. Because processed information is transmitted, details from the original data may be lost, resulting in insufficient perception accuracy and robustness. This information loss can severely impact the accuracy of perception tasks, particularly in complex and dynamic scenes, limiting their use in autonomous driving applications that require high-precision perception. Furthermore, the models deployed on edge devices of different agents vary, making it difficult for vehicles to integrate features from other agents.

[0008] In summary, the shortcomings of existing technologies primarily focus on communication efficiency, computing resources, perception accuracy, and model compatibility. Early fusion methods face bottlenecks in communication and computing resources, while late fusion methods face challenges in perception accuracy and model compatibility. These shortcomings indicate the need for a new technical solution that balances communication efficiency and perception performance to meet the requirements of vehicle-infrastructure cooperative autonomous driving systems in complex and dynamic scenarios. Summary of the Invention

[0009] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.

[0010] To this end, the first objective of the present invention is to propose a data compression and transmission method for a vehicle-road perception system based on point cloud clusters, so as to balance communication efficiency and perception performance.

[0011] The second object of the present invention is to propose a data compression and transmission system for a vehicle-road perception system based on point cloud clusters.

[0012] A third object of the present invention is to provide an electronic device.

[0013] A fourth object of the present invention is to provide a computer-readable storage medium.

[0014] To achieve the above objectives, the first aspect of the present invention proposes a method for compressing and transmitting data in a vehicle-road perception system based on point cloud clusters, comprising:

[0015] Acquire roadside data and vehicle-side data, and perform point aggregation encoding on the roadside data and the vehicle-side data to obtain roadside point cloud cluster data and vehicle-side point cloud cluster data corresponding to the foreground object;

[0016] Performing spatial compression, temporal compression, and feature compression on the roadside point cloud cluster data to obtain roadside target point cloud cluster data;

[0017] The roadside target point cloud cluster data is fused with the vehicle side point cloud cluster data to obtain target fusion data for vehicle-road perception tasks.

[0018] In the method of the first aspect of the present invention, the spatial compression, temporal compression and feature compression of the roadside point cloud cluster data to obtain roadside target point cloud cluster data include: using a key point based sampling method to spatially compress the roadside point cloud cluster data to obtain a compressed point cloud cluster data.

[0019] In the method of the first aspect of the present invention, the spatial compression, temporal compression and feature compression of the roadside point cloud cluster data to obtain roadside target point cloud cluster data also includes: using a method based on point cloud flow estimation to temporally compress the once compressed point cloud cluster data to obtain secondary compressed point cloud cluster data.

[0020] In the method of the first aspect of the present invention, the spatial compression, temporal compression and feature compression of the roadside point cloud cluster data to obtain roadside target point cloud cluster data also includes: obtaining a trained encoder, and using the trained encoder to perform feature compression on the secondary compressed point cloud cluster data to obtain roadside target point cloud cluster data.

[0021] In the method of the first aspect of the present invention, the key point based sampling method refers to a point cloud downsampling method based on farthest point sampling for point cloud selection, and semantic importance and shape importance are combined in the selection process.

[0022] In the method of the first aspect of the present invention, the method based on point cloud flow estimation refers to retaining the dynamic moving objects in the foreground objects by calculating the motion changes of the point clouds of the previous and next frames.

[0023] In the method of the first aspect of the present invention, the method for obtaining a trained encoder includes: obtaining a compression reconstruction model, the compression reconstruction model including an encoder, a codebook and a decoder; training the compression reconstruction model based on reconstruction pre-training to obtain a trained compression reconstruction model, and the encoder in the trained compression reconstruction model is the trained encoder.

[0024] To achieve the above objectives, the second aspect of the present invention provides a data compression and transmission system for a vehicle-road perception system based on point cloud clusters, comprising:

[0025] an encoding module, configured to obtain roadside data and vehicle-side data, and perform point aggregation encoding on the roadside data and the vehicle-side data to obtain roadside point cloud cluster data and vehicle-side point cloud cluster data corresponding to foreground objects;

[0026] A compression module, configured to perform spatial compression, temporal compression, and feature compression on the roadside point cloud cluster data to obtain roadside target point cloud cluster data;

[0027] A fusion module is used to fuse the roadside target point cloud cluster data with the vehicle-side point cloud cluster data to obtain target fusion data for vehicle-road perception tasks.

[0028] To achieve the above-mentioned purpose, the third aspect of the present invention proposes an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method proposed in the first aspect of the present invention.

[0029] To achieve the above-mentioned purpose, the fourth aspect of the present invention proposes a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the method proposed in the first aspect of the present invention.

[0030] The present invention provides a point cloud cluster-based vehicle-road perception system data compression and transmission method, system, electronic device, and storage medium. The method acquires roadside and vehicle-side data, performs point aggregation encoding on the roadside and vehicle-side data to obtain roadside and vehicle-side point cloud cluster data corresponding to foreground objects. The roadside point cloud cluster data is then spatially, temporally, and feature-compressed to obtain roadside target point cloud cluster data. The roadside target point cloud cluster data is then fused with the vehicle-side point cloud cluster data to obtain target fused data for vehicle-road perception tasks. In this case, the roadside point cloud cluster data is compressed using a combination of spatial, temporal, and feature compression to obtain roadside target point cloud cluster data. This fused data is then fused with the vehicle-side point cloud cluster data to obtain target fused data, which serves as input data for subsequent vehicle-road perception tasks and corresponding perception predictions. The target fused data retains key information, effectively improving subsequent perception performance, while the roadside target point cloud cluster data reduces communication costs, achieving efficient data compression and transmission. Therefore, the present invention balances communication efficiency and perception performance.

[0031] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:

[0033] Figure 1 A schematic flow chart of a method for data compression and transmission of a vehicle-road perception system based on point cloud clusters provided by an embodiment of the present invention;

[0034] Figure 2 A specific flow chart of the data compression and transmission method for a vehicle-road perception system based on point cloud clusters provided by an embodiment of the present invention;

[0035] Figure 3 A block diagram of a vehicle-road perception system data compression and transmission system based on point cloud clusters provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0036] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0037] The following describes a method and system for data compression and transmission of a vehicle-road perception system based on point cloud clusters according to an embodiment of the present invention with reference to the accompanying drawings.

[0038] Currently, data transmission efficiency and perception performance are two key challenges in vehicle-road cooperative autonomous driving systems. With the development of autonomous driving technology, vehicles and roadside equipment need to share large amounts of perception data in real time to achieve collaborative perception. However, existing data transmission and processing methods have the following problems: if the complete raw data collected by the sensor is directly transmitted to the central processing unit for fusion processing, although all the original information can be retained, the amount of data transmitted is extremely large, the communication bandwidth requirements are extremely high, and it is easy to cause transmission delays and resource waste; if the processed data such as target detection results are transmitted instead of the original data, such an operation reduces the amount of data transmitted, but because the processed results are transmitted instead of the original data, the system cannot fully utilize the detailed information in the original data, resulting in insufficient perception accuracy and robustness. Especially in complex dynamic scenes, the perception performance degradation is more obvious.

[0039] Based on this, an embodiment of the present invention provides a data compression and transmission method for a vehicle-road perception system based on point cloud clusters, so as to balance communication efficiency and perception performance.

[0040] Figure 1 A schematic flow chart of a method for data compression and transmission of a vehicle-road perception system based on point cloud clusters provided by an embodiment of the present invention. Figure 2This is a specific flow chart of the data compression and transmission method for the vehicle-road perception system based on point cloud clusters provided by an embodiment of the present invention.

[0041] like Figure 1 As shown, the data compression and transmission method of the vehicle-road perception system based on point cloud clusters includes the following steps:

[0042] Step S101 : Acquire roadside data and vehicle-side data, and perform point aggregation coding on the roadside data and vehicle-side data to obtain roadside point cloud cluster data and vehicle-side point cloud cluster data corresponding to foreground objects.

[0043] In step S101, corresponding raw point cloud data can be collected by vehicle-side and road-side devices respectively. The raw point cloud data on the road side is roadside data, and the raw point cloud data on the vehicle side is vehicle-side data. In addition, the raw point cloud data is scene point cloud data, which includes information about foreground objects and background noise.

[0044] In step S101, the raw point cloud data is processed using point aggregation coding to achieve foreground segmentation, remove background noise, and retain point cloud cluster data representing foreground objects. Specifically, the input raw point cloud data is first downsampled, and the sampled points are used as the center points of local regions. Then, for each center point, the neighboring points within a certain radius are determined to form a local region. Within each local region, a multi-layer perceptron is used to extract the features of each point, and all processed features within the local region are aggregated using maximum pooling, thereby completing a set feature extraction. By stacking multiple such set feature extractions, gradually increasing the radius of the local region and reducing the number of sampling points, the downsampled point cloud features of the scene point cloud data are obtained. By predicting the probability that the point belongs to a foreground object, a foreground point cloud is obtained, and spatial distance is used to aggregate the foreground point cloud clusters representing the foreground objects. In this step, point aggregation coding is performed on the roadside and vehicle-side data to obtain roadside and vehicle-side point cloud clusters corresponding to the foreground objects.

[0045] In step S101 , the roadside point cloud cluster data carries the probability information of each point belonging to a foreground object (ie, the foreground probability of each point).

[0046] In step S101 , after obtaining the roadside point cloud cluster data, it is necessary to calculate the sparsity of the area where the point cloud is located.

[0047] Step S102 : performing spatial compression, temporal compression, and feature compression on the roadside point cloud cluster data to obtain roadside target point cloud cluster data.

[0048] In step S102, spatial compression, temporal compression, and feature compression are performed on the roadside point cloud cluster data to obtain the roadside target point cloud cluster data, including: using a key-point based sampling method to perform spatial compression on the roadside point cloud cluster data to obtain the first compressed point cloud cluster data (see Figure 2 ). Among them, the key-point based sampling method refers to a point cloud downsampling method based on farthest point sampling for point cloud selection, and semantic importance and shape importance are combined during the selection process.

[0049] Specifically, in order to further reduce the data volume while retaining the key structural information of the target, this step uses a key-point based sampling method to compress the number of point clouds. Specifically, a point cloud downsampling method based on farthest point sampling is used, and dual guidance of semantic importance and shape importance is introduced during the selection process. The semantic importance is based on the foreground probability predicted by point aggregation coding, and the shape importance is based on the sparsity of the region where the point cloud is located. Points with higher foreground probability and lower sparsity of the region are preferentially retained during the downsampling process. Sparse key points representing object features are selected in the above manner. These sparse key points constitute the first compressed point cloud cluster data. For example, the compressed roadside point cloud cluster data is a point cloud of size [N1, C], where N1 is the number of point clouds in the roadside point cloud cluster data, and C is the dimension of the point clouds in the roadside point cloud cluster data. The obtained first compressed point cloud cluster data is a point cloud of size [N2, C]. The dimension of the point clouds in the first compressed point cloud cluster data is equal to the dimension of the point clouds in the roadside point cloud cluster data, and N2 is the number of point clouds in the first compressed point cloud cluster data, and N2 << N1. This step significantly reduces the data size while retaining the key point cloud structural information, laying a foundation for subsequent processing.

[0050] In step S102, spatial compression, temporal compression, and feature compression are performed on the roadside point cloud cluster data to obtain the roadside target point cloud cluster data, and it also includes: using a point cloud flow estimation based method to perform temporal compression on the first compressed point cloud cluster data to obtain the second compressed point cloud cluster data (see Figure 2 ). Among them, the point cloud flow estimation based method refers to calculating the motion changes of the point clouds in the front and rear frames to retain the dynamic moving objects in the foreground objects.

[0051] Specifically, the method based on point cloud flow estimation can achieve information compression in streaming information transmission by identifying dynamic objects and filtering out static objects that have been transmitted in historical frames during transmission. First, a point cloud flow prediction model is used to obtain the predicted future velocity vector of each point based on historical point cloud frames. Then, the average flow vector of the points is obtained based on the point cloud clusters, which helps to filter out minor movements or noises within static clusters and maintain the structural integrity of moving objects by considering the movement of the entire cluster. By setting a threshold for the magnitude of the flow vector, dynamically significant point cloud clusters can be identified for incremental data transmission. Thus, by calculating the movement changes of the point clouds in consecutive frames, regions with dynamic changes exceeding a certain proportion are selected for transmission, and only the dynamic moving objects in the foreground objects are retained. During the temporal compression process, the once-compressed point cloud cluster data to be compressed is a point cloud of size [N2, C], and the twice-compressed point cloud cluster data obtained is a point cloud of size [N3, C]. The dimension of the point cloud in the twice-compressed point cloud cluster data is equal to that of the point cloud in the once-compressed point cloud cluster data. N3 is the number of point clouds in the twice-compressed point cloud cluster data, and N3 < N2. This step further reduces the amount of transmitted data, focuses on the dynamically changing regions, avoids repeated transmission of the static background, and improves the efficiency and pertinence of data transmission.

[0052] In step S102, spatial compression, temporal compression, and feature compression are performed on the roadside point cloud cluster data to obtain the roadside target point cloud cluster data, and it further includes: obtaining a trained encoder, and using the trained encoder to perform feature compression on the twice-compressed point cloud cluster data to obtain the roadside target point cloud cluster data (see Figure 2 ). Among them, the method for obtaining the trained encoder includes: obtaining a compression and reconstruction model, where the compression and reconstruction model includes an encoder, a codebook, and a decoder; training the compression and reconstruction model based on reconstruction pre-training to obtain a trained compression and reconstruction model, and the encoder in the trained compression and reconstruction model is the trained encoder.

[0053] Among them, when training the compression and reconstruction model based on reconstruction pre-training, the encoder maps the input point cloud cluster features to a continuous latent space to obtain continuous latent vectors. The codebook contains a set of learnable compressed feature vectors, which have fewer channels (i.e., dimensions) than the original features. The compression and reconstruction model uses nearest neighbor search to replace the continuous latent vectors with the closest compressed features in the codebook to achieve quantization representation. The decoder reconstructs from the quantization representation to obtain the reconstructed point cloud cluster features. When the loss function between the reconstructed point cloud cluster features and the input point cloud cluster features is less than the set threshold, the training is completed. Thus, the trained encoder can retain key information.

[0054] The trained encoder is used to replace the secondarily compressed point cloud cluster data with the compressed features in the codebook of the trained compression and reconstruction model to obtain the roadside target point cloud cluster data, thereby mapping the high-dimensional point cloud features to a low-dimensional space and achieving the feature compression of the point cloud cluster. During the feature compression process, the secondarily compressed point cloud cluster data to be compressed is a point cloud of size [N3, C], and the obtained roadside target point cloud cluster data is a point cloud of size [N3, D]. The number of points in the roadside target point cloud cluster data is equal to the number of points in the secondarily compressed point cloud cluster data. D is the dimension of the points in the roadside target point cloud cluster data, and D << C. This step not only reduces the feature dimension and the transmission burden but also improves the feature expression ability through pre-training.

[0055] Step S103: Fuse the roadside target point cloud cluster data with the vehicle-side point cloud cluster data to obtain the target fusion data for the vehicle-road perception task.

[0056] In step S103, perception enhancement is achieved through roadside data fusion. Specifically, the compressed data at the roadside (i.e., the roadside target point cloud cluster data) is fused with the data at the vehicle end (i.e., the vehicle-side point cloud cluster data) to generate fused features (i.e., the target fusion data), and then it is input into the perception head for perception tasks such as target detection and tracking. Through this fusion method, the system can make full use of the data advantages of the vehicle end and the roadside end, improve the perception accuracy and robustness, and at the same time maintain a low communication cost. In this step, the vehicle-road data fusion is carried out at the level of point cloud clusters. The compressed data received from the roadside can be used to enhance the features of the corresponding point cloud clusters detected by the vehicle-side sensors. The matching of the point cloud clusters on the vehicle side and the roadside can be based on the spatial proximity of their cluster centers, and the fusion method can be achieved through feature connection. By integrating complementary information from different sensors and perspectives (vehicle and roadside), the accuracy of subsequent target detection and classification can be improved.

[0057] To implement the above embodiments, the present invention also proposes a data compression and transmission system for a vehicle-road perception system based on point cloud clusters.

[0058] Figure 3 It is a block diagram of a data compression and transmission system for a vehicle-road perception system based on point cloud clusters provided by an embodiment of the present invention.

[0059] As Figure 3 shown, the data compression and transmission system for a vehicle-road perception system based on point cloud clusters includes an encoding module 11, a compression module 12, and a fusion module 13, where:

[0060] The encoding module 11 is configured to obtain roadside data and vehicle-side data, and perform point aggregation encoding on the roadside data and vehicle-side data to obtain the roadside point cloud cluster data and vehicle-side point cloud cluster data corresponding to foreground objects;

[0061] A compression module 12 is used to perform spatial compression, temporal compression and feature compression on the roadside point cloud cluster data to obtain roadside target point cloud cluster data;

[0062] The fusion module 13 is used to fuse the roadside target point cloud cluster data with the vehicle side point cloud cluster data to obtain target fusion data for the vehicle-road perception task.

[0063] Furthermore, in a possible implementation of an embodiment of the present invention, in the compression module 12, the roadside point cloud cluster data is spatially compressed, temporally compressed, and feature compressed to obtain roadside target point cloud cluster data, including: using a key point-based sampling method to spatially compress the roadside point cloud cluster data to obtain once-compressed point cloud cluster data.

[0064] Furthermore, in a possible implementation of an embodiment of the present invention, in the compression module 12, the key point-based sampling method refers to a point cloud downsampling method based on farthest point sampling for point cloud selection, and semantic importance and shape importance are combined in the selection process.

[0065] Furthermore, in a possible implementation of an embodiment of the present invention, in the compression module 12, spatial compression, temporal compression and feature compression are performed on the roadside point cloud cluster data to obtain roadside target point cloud cluster data, and it also includes: using a method based on point cloud flow estimation to temporally compress the once compressed point cloud cluster data to obtain secondary compressed point cloud cluster data.

[0066] Furthermore, in a possible implementation of the embodiment of the present invention, in the compression module 12, the method based on point cloud flow estimation refers to retaining dynamic moving objects in the foreground objects by calculating the motion changes of point clouds of previous and next frames.

[0067] Furthermore, in a possible implementation of an embodiment of the present invention, in the compression module 12, spatial compression, temporal compression and feature compression are performed on the roadside point cloud cluster data to obtain roadside target point cloud cluster data, and it also includes: obtaining a trained encoder, and using the trained encoder to perform feature compression on the secondary compressed point cloud cluster data to obtain roadside target point cloud cluster data.

[0068] Furthermore, in a possible implementation of an embodiment of the present invention, in the compression module 12, the method for obtaining a trained encoder includes: obtaining a compression reconstruction model, the compression reconstruction model including an encoder, a codebook and a decoder; training the compression reconstruction model based on reconstruction pre-training to obtain a trained compression reconstruction model, and the encoder in the trained compression reconstruction model is a trained encoder.

[0069] It should be noted that the aforementioned explanation of the embodiment of the method for compressing and transmitting data of a vehicle-road perception system based on point cloud clusters is also applicable to the data compression and transmission system of a vehicle-road perception system based on point cloud clusters in this embodiment, and will not be repeated here.

[0070] In an embodiment of the present invention, roadside data and vehicle-side data are acquired and point aggregation encoding is performed on the roadside data and vehicle-side data to obtain roadside point cloud cluster data and vehicle-side point cloud cluster data corresponding to the foreground object; spatial compression, temporal compression, and feature compression are performed on the roadside point cloud cluster data to obtain roadside target point cloud cluster data; and the roadside target point cloud cluster data are fused with the vehicle-side point cloud cluster data to obtain target fusion data for the vehicle-road perception task. In this case, the roadside point cloud cluster data is compressed by combining spatial compression, temporal compression, and feature compression to obtain roadside target point cloud cluster data, and then the roadside target point cloud cluster data is fused with the vehicle-side point cloud cluster data to obtain target fusion data. The target fusion data is used as input data for subsequent vehicle-road perception tasks to perform corresponding perception predictions. The target fusion data retains key information and can effectively improve subsequent perception performance, while the roadside target point cloud cluster data reduces communication costs and achieves efficient data compression and transmission. Therefore, the present invention can strike a balance between communication efficiency and perception performance.

[0071] The method and system of the present invention are a method and system for compressing and transmitting point cloud data in a vehicle-road perception system. Through key technical modules such as spatial compression, temporal compression, feature compression, and data fusion, efficient compression and transmission of data in the vehicle-road perception system are achieved, and the perception performance is significantly improved. This method reduces communication costs while retaining key information. It is suitable for autonomous driving applications in complex dynamic scenes and solves the problems of low communication efficiency, waste of computing resources, and limited perception performance in the existing technology. It can balance communication efficiency and perception performance (i.e., high-performance joint perception can be achieved under low bandwidth constraints) to overcome the limitations of existing technologies and meet the needs of vehicle-road collaborative autonomous driving systems in complex dynamic scenes. It provides an efficient and reliable solution to the problem of data transmission compression in vehicle-road collaborative autonomous driving systems.

[0072] It has the following beneficial effects: 1) Reduced communication costs: Through the three stages of spatial compression, temporal compression, and feature compression, the amount of data that needs to be transmitted is significantly reduced. This not only reduces the pressure on communication bandwidth, but also reduces the delay in data transmission and improves communication efficiency. 2) Improved perception performance: By retaining key point cloud structures and dynamic change information, as well as enhanced perception through road-side data fusion, this method can improve perception accuracy and robustness while reducing the amount of data. 3) Improved autonomous driving safety: Efficient data compression and transmission methods can ensure that autonomous driving systems can obtain and process key information in a timely and accurate manner, thereby improving the safety and reliability of autonomous driving.

[0073] In order to implement the above embodiments, the present invention also proposes an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided by the above embodiments.

[0074] In order to implement the above embodiments, the present invention further provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the methods provided in the above embodiments.

[0075] In order to implement the above embodiments, the present invention further provides a computer program product, including a computer program, which implements the methods provided in the above embodiments when executed by a processor.

[0076] In the descriptions of the foregoing embodiments, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and features of different embodiments or examples, unless they are mutually inconsistent.

[0077] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0078] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0079] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.

[0080] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement the present invention: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0081] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0082] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.

[0083] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and are not to be construed as limiting the present invention. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A data compression and transmission method for a vehicle-road perception system based on point cloud clusters, characterized in that: include: Acquire roadside data and vehicle-side data, and perform point aggregation encoding on the roadside data and the vehicle-side data to obtain roadside point cloud cluster data and vehicle-side point cloud cluster data corresponding to the foreground object; Performing spatial compression, temporal compression, and feature compression on the roadside point cloud cluster data to obtain roadside target point cloud cluster data; The roadside target point cloud cluster data is fused with the vehicle side point cloud cluster data to obtain target fusion data for vehicle-road perception tasks.

2. The method for data compression and transmission of a vehicle-road perception system based on point cloud clusters according to claim 1, characterized in that: The performing spatial compression, temporal compression, and feature compression on the roadside point cloud cluster data to obtain roadside target point cloud cluster data includes: The roadside point cloud cluster data is spatially compressed using a key point-based sampling method to obtain primary compressed point cloud cluster data.

3. The method for data compression and transmission of a vehicle-road perception system based on point cloud clusters according to claim 2, characterized in that: The spatial compression, temporal compression and feature compression of the roadside point cloud cluster data to obtain roadside target point cloud cluster data further includes: The first compressed point cloud cluster data is temporally compressed using a method based on point cloud flow estimation to obtain second compressed point cloud cluster data.

4. The method for data compression and transmission of a vehicle-road perception system based on point cloud clusters according to claim 3, characterized in that: The spatial compression, temporal compression and feature compression of the roadside point cloud cluster data to obtain roadside target point cloud cluster data further includes: A trained encoder is obtained, and feature compression is performed on the secondary compressed point cloud cluster data using the trained encoder to obtain roadside target point cloud cluster data.

5. The method for data compression and transmission of a vehicle-road perception system based on point cloud clusters according to claim 2, characterized in that: The key point based sampling method refers to the point cloud downsampling method based on the farthest point sampling to select point clouds, and combines semantic importance and shape importance in the selection process.

6. The method for data compression and transmission of a vehicle-road perception system based on point cloud clusters according to claim 3, characterized in that: The method based on point cloud flow estimation refers to retaining the dynamic moving objects in the foreground objects by calculating the motion changes of the point clouds of the previous and next frames.

7. The method for data compression and transmission of a vehicle-road perception system based on point cloud clusters according to claim 4, characterized in that: The methods for obtaining the trained encoder include: Acquire a compression reconstruction model, where the compression reconstruction model includes an encoder, a codebook, and a decoder; The compression reconstruction model is trained based on reconstruction pre-training to obtain a trained compression reconstruction model, and the encoder in the trained compression reconstruction model is the trained encoder.

8. A data compression and transmission system for a vehicle-road perception system based on point cloud clusters, characterized in that: include: an encoding module, configured to obtain roadside data and vehicle-side data, and perform point aggregation encoding on the roadside data and the vehicle-side data to obtain roadside point cloud cluster data and vehicle-side point cloud cluster data corresponding to foreground objects; A compression module, configured to perform spatial compression, temporal compression, and feature compression on the roadside point cloud cluster data to obtain roadside target point cloud cluster data; A fusion module is used to fuse the roadside target point cloud cluster data with the vehicle-side point cloud cluster data to obtain target fusion data for vehicle-road perception tasks.

9. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.