Self-adaptive multi-vehicle cooperative sensing method and device based on intelligent distributed decision

Through the adaptive multi-vehicle collaborative perception method of intelligent distributed decision-making, combined with voxel feature coding and multi-agent reinforcement learning, the perceived data transmission is optimized, and the problem of low efficiency of communication resource utilization in the existing technology is solved, and the perception accuracy and communication efficiency are improved.

CN120302255APending Publication Date: 2025-07-11XI AN JIAOTONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510454792.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

While ensuring perceptual performance, the existing collaborative perception technology has low efficiency in the utilization of communication resources, especially in high-density traffic environments, communication delay and bandwidth resources are severely wasted, making it difficult to make independent decisions in dynamic environments.

Method used

Adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making is adopted, and the perceived data transmission strategy is optimized to achieve a balance between perceived accuracy and communication efficiency through voxel feature coding, signal-to-noise ratio adaptive compression and multi-agent reinforcement learning.

Benefits of technology

Effectively reduce communication delay, save bandwidth resources, improve perception accuracy and communication efficiency in dynamic environments, especially in high vehicle density or resource-constrained scenarios to avoid network overload and improve system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302255A_ABST
    Figure CN120302255A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive multi-vehicle cooperative sensing method and device based on intelligent distributed decision, and the method comprises the steps: carrying out the voxelization of unified point cloud data for each vehicle, carrying out the high-dimensional voxel feature extraction of the voxelized point cloud data through a voxel feature coding algorithm, and carrying out the high-dimensional feature extraction of the voxelized point cloud data; based on the extracted high-dimensional voxel features, vehicle local high-dimensional perception features are generated; each vehicle adaptively determines a compression ratio according to the signal-to-noise ratio of the current channel, and dynamic compression coding is performed on the local high-dimensional perception features of the vehicle by using the compression ratio to obtain local compression features of the vehicle; each vehicle iteratively solves a pre-constructed multi-vehicle cooperative sensing optimization problem by adopting a multi-agent reinforcement learning method based on a global state to obtain a communication sub-channel for transmitting the local compression characteristics of the vehicle and a target vehicle for receiving the local compression characteristics of the vehicle; wherein the optimization objective of the multi-vehicle cooperative sensing optimization problem is to maximize the sensing precision and minimize the communication time delay; each vehicle sends the vehicle local compression characteristics to the target vehicle through the communication sub-channel. According to the method, the sensing accuracy can be ensured, the communication delay is reduced to the maximum extent, and bandwidth resources are saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle networking, and in particular to an adaptive multi-vehicle collaborative perception method and device based on intelligent distributed decision-making. Background Art

[0002] With the rapid development of intelligent transportation systems and vehicle networking technologies, networked autonomous vehicles have shown great potential in improving traffic efficiency and driving safety. However, due to the limitations of the field of view and resolution of on-board sensors, the perception ability of a single vehicle is easily affected by environmental occlusion and dynamic changes, resulting in blind spots and information loss. This not only limits the widespread application of autonomous driving technology, but may also increase the risk of traffic accidents. To this end, collaborative perception technology has emerged. Through information sharing between networked autonomous vehicles, it expands the perception range, makes up for the shortcomings of single-vehicle perception, and improves the accuracy and robustness of overall perception.

[0003] Existing collaborative perception technologies mainly include three strategies: raw data sharing, feature-level data sharing, and detection result sharing. Raw data sharing can retain the most comprehensive environmental information, but due to the large amount of data, it leads to high communication bandwidth requirements and large transmission delays, limiting its feasibility in actual vehicle networks. Although detection result sharing significantly reduces the amount of transmitted data, due to the high degree of information compression, the perception accuracy is low and the sensitivity to single-vehicle detection errors is strong, resulting in unstable performance in high-density traffic environments. In contrast, feature-level data sharing achieves a good balance between communication overhead and perception performance improvement, but its data volume is still at the megabit level. Especially when there are a large number of connected autonomous driving vehicles, the demand for communication resources increases significantly, which may lead to low communication efficiency or even infeasibility when transmission resources are limited. In addition, the value of perception information obtained by different vehicles is not the same. Blindly broadcasting large amounts of perception data to all vehicles can easily cause network overload and resource waste, and lead to a sharp increase in latency. Therefore, how to optimize the use of communication resources while ensuring perception performance has become a key issue that needs to be solved in collaborative perception technology.

[0004] Based on the above problems, if adaptive decisions can be made for each vehicle under different environments and bandwidth conditions, allowing the vehicle to flexibly choose whether to initiate the transmission of perception information, and decide which perception features to share with which vehicles, and optimize the data compression and transmission strategy on this basis, it can not only improve the effectiveness of multi-vehicle collaborative perception, but also significantly reduce the communication burden. It can be seen that there is an urgent need for a distributed multi-vehicle collaborative perception method that can combine communication and perception collaborative optimization and make autonomous decisions in a dynamic environment, so as to minimize communication delay and save bandwidth resources while ensuring perception accuracy. Summary of the invention

[0005] In view of the problems existing in the prior art, the present invention provides an adaptive multi-vehicle collaborative perception method and device based on intelligent distributed decision-making, which can optimize the combination of communication and perception collaboration and make autonomous decisions in a dynamic environment, so as to minimize communication delay and save bandwidth resources while ensuring perception accuracy.

[0006] To solve the above technical problems, the present invention is implemented through the following technical solutions:

[0007] According to a first aspect of the present invention, there is provided an adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making, which is applied to the multi-vehicle collaborative perception scenario of the vehicle-to-everything (V2X) network, and includes:

[0008] Each vehicle collects the original point cloud data of the surrounding environment, and preprocesses the original point cloud data and then unifies it to the global coordinate system;

[0009] Each vehicle voxelizes the unified point cloud data, extracts high-dimensional voxel features from the voxelized point cloud data through a voxel feature encoding algorithm, and generates a vehicle-local high-dimensional perception feature based on the extracted high-dimensional voxel features;

[0010] Each vehicle adaptively determines a compression ratio according to the signal-to-noise ratio of the current channel, and uses the compression ratio to perform dynamic compression encoding on the vehicle-local high-dimensional perception feature to obtain a vehicle-local compressed feature;

[0011] Each vehicle uses a multi-agent reinforcement learning method based on the global state to iteratively solve a pre-constructed multi-vehicle collaborative perception optimization problem, and obtains a communication sub-channel for transmitting the vehicle-local compressed feature and a target vehicle for receiving the vehicle-local compressed feature; wherein, the optimization objective of the multi-vehicle collaborative perception optimization problem is to maximize perception accuracy and minimize communication delay;

[0012] Each vehicle sends the vehicle-local compressed feature to the target vehicle through the communication sub-channel.

[0013] In a possible implementation manner of the first aspect, the preprocessing the original point cloud data and then unifying it to the global coordinate system includes:

[0014] Preprocess the original point cloud data by means of random shuffling, voxel downsampling, and distance filtering;

[0015] Unify the preprocessed point cloud data to the global coordinate system through an attitude transformation matrix, specifically:

[0016]

[0017] In the formula, is the point cloud data set of vehicle i in the global coordinate system, that is, the unified point cloud data; [p; 1] represents the homogeneous coordinate of point p; T i represents the pose transformation matrix of vehicle i.

[0018] In a possible implementation of the first aspect, generating the local high-dimensional perception feature of the vehicle based on the high-dimensional voxel feature includes:

[0019] Generating a sparse and learnable 2D pseudo-map according to the high-dimensional voxel feature, specifically:

[0020] F BEV,i = Scatter({f i,k})

[0021] where F BEV,i is the 2D pseudo-map of vehicle i; Scatter(·) is the scattering operation; {f i,k} is the vector set of voxel k in the point cloud data of vehicle i;

[0022] Using a convolutional backbone network to refine the 2D pseudo-map to generate the local high-dimensional perception feature of the vehicle.

[0023] In a possible implementation of the first aspect, each vehicle adaptively determines the compression ratio according to the signal-to-noise ratio of the current channel, specifically:

[0024] Considering the signal-to-noise ratio SNR i,j between vehicle i and vehicle j, the following multi-segment mapping is adopted:

[0025]

[0026] In the formula, η i,j is the compression ratio between vehicle i and vehicle j; γ1, γ2, γ3, γ4, γ5 are channel thresholds.

[0027] In a possible implementation of the first aspect, the multi-vehicle collaborative perception optimization problem is specifically:

[0028]

[0029] Among them,

[0030]

[0031] R i,j = B i,j log2(1 + SNR i,j )

[0032]

[0033]

[0034] Where, ΔAP imp,i represents the improved value of the perception accuracy of vehicle i after fusion; D i represents the communication delay of vehicle i on the communication links of each target vehicle; α is the coefficient for weighing the perception gain; β is the coefficient for weighing the communication delay overhead; V is the set of vehicles; t i,j ∈{0,1} indicates whether vehicle i sends local compressed features to vehicle j. When t i,j =0, it does not send. When t i,j =1, it sends; s i,m ∈{0,1} indicates whether vehicle i selects sub-channel m. When s i,m =0, it does not select. When s i,m =1, it selects;

[0035] P t is the transmission power of vehicle i; P max is the maximum transmission power; B i is the total bandwidth allocated by vehicle i on all sub-channels; B total is the total bandwidth; B i,m is the bandwidth allocated by vehicle i on sub-channel m; R i,j is the communication rate between vehicle i and vehicle j; R min is the minimum communication rate; D max is the maximum communication delay; K is the number of sub-carriers of each sub-channel; AP i (t) is the perception accuracy of vehicle i; AP min is the minimum perception accuracy; C i (t) is the processing cost of vehicle i; C max is the maximum processing cost; k i,m is the number of sub-carriers allocated by vehicle i on sub-channel m; is the set of integers; M is the number of sub-channels obtained by dividing the total bandwidth; B c is the bandwidth of each sub-channel; g i,j is the gain; N0 is the noise power spectral density; B i,j is the total bandwidth allocated to the communication link between vehicle i and vehicle j; PL i,j is the path loss between vehicle i and vehicle j; d i,j is the distance between vehicle i and vehicle j; f is the carrier frequency; D i,j is the transmission delay from vehicle i to vehicle j; is the set of vehicles used to transmit compressed features; is the size of the effective point cloud data of vehicle i after compression; S i is the size of the original point cloud data of vehicle i; n m is the number of vehicles selecting the same sub-channel m.

[0036] In a possible implementation of the first aspect, the multi-agent reinforcement learning method based on global state is specifically the proximal policy optimization method enhanced by global state information;

[0037] Each vehicle i with sensing and communication capabilities is regarded as an agent, and the state vector of vehicle i at time slot t is defined as:

[0038]

[0039] where is the position of vehicle i at time slot t; is the speed of vehicle i at time slot t; is the size of the sensing data of vehicle i at time slot t; is the set of the positions, speeds, and sensing data sizes of the other vehicles at time slot t;

[0040] The action of vehicle i at time slot t is defined as:

[0041]

[0042] where is the action of vehicle i at time slot t; is the target vehicle selected by vehicle i at time slot t; is the communication sub-channel selected by vehicle i at time slot t;

[0043] The global reward r of each vehicle i at time slot t is defined t i as:

[0044]

[0045] where is the global sensing accuracy; is the global communication delay.

[0046] In a possible implementation of the first aspect, after each vehicle sends the local compressed feature of the vehicle to the target vehicle through the communication sub-channel, it further includes:

[0047] After the target vehicle receives the compressed features from multiple source vehicles, it decodes the compressed features and performs a fusion operation on the decoded features and the features of its own vehicle to obtain fused features;

[0048] The target vehicle inputs the fused features into the detection head, outputs the target detection box and confidence, and calculates the intersection over union and average precision based on the target detection box and confidence, and evaluates the collaborative sensing performance using the intersection over union and average precision.

[0049] In a possible implementation of the first aspect, the operation of fusing the decoded features is specifically as follows:

[0050] When the compressed features of multiple source vehicles arrive synchronously, the decoded features are fused by using element-wise maximum operation, specifically as follows:

[0051]

[0052] where F fused is the fused feature; are multiple source vehicles that send perception information to the target vehicle; F decoded,i is the feature obtained by decoding the compressed feature of vehicle i.

[0053] According to the second aspect of the present invention, there is provided a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making as described above is implemented.

[0054] According to the third aspect of the present invention, there is provided a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making as described above is implemented.

[0055] Compared with the prior art, the present invention has at least the following beneficial effects:

[0056] An adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making provided by the present invention extracts high-dimensional features from the preprocessed point cloud data through a voxel feature encoding algorithm, effectively stripping redundant information, retaining key environmental features, and enhancing the representation ability of perception data. Compared with traditional methods, this feature extraction mechanism improves the perception accuracy, especially showing stronger robustness in high-dynamic and complex occlusion scenarios, effectively making up for the blind spots and information loss problems of single-vehicle perception. Based on a dynamic compression strategy adaptive to the signal-to-noise ratio, an intelligent matching between the perception data compression ratio and the channel quality is achieved. When the channel quality is good, a low compression ratio is adopted to retain data details and ensure perception accuracy; when the channel quality deteriorates, the compression ratio is automatically increased to reduce the amount of transmitted data and reduce communication latency. Experiments show that this strategy reduces the average communication latency by 18.3%, realizes the full utilization of bandwidth resources, and improves communication efficiency. Through a multi-agent reinforcement learning algorithm, each vehicle can autonomously make decisions on communication sub-channel allocation and target vehicle selection based on global state information, achieving optimal resource allocation. Compared with traditional fixed allocation strategies, this method improves the overall perception performance by 28%. Especially in scenarios with a high vehicle density or limited communication resources, it can effectively avoid network overload and resource waste and improve system stability. The distributed decision-making mechanism of the present invention endows the system with the ability of autonomous adjustment in a dynamic environment, can real-time sense traffic scene changes (such as vehicle quantity, communication quality fluctuations) and dynamically optimize the collaborative strategy. Experimental verification shows that it performs well when the vehicle quantity fluctuates between 5 and 20 vehicles, can maintain the stability of perception accuracy and communication efficiency, and has good adaptability and robustness.

[0057] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] To more clearly illustrate the technical solutions in the specific embodiments of the present invention, the following will briefly introduce the drawings required for use in the description of the specific embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0059] Figure 1 It is a flowchart of an adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making according to an embodiment of the present invention;

[0060] Figure 2 It is a flowchart of an adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making according to an embodiment of the present invention;

[0061] Figure 3 It is a schematic diagram of a multi-vehicle collaborative perception scenario involved in an embodiment of the present invention;

[0062] Figure 4 This is a schematic diagram of the intelligent distributed decision-making system framework according to the embodiments of the present invention. Specific implementation manners

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0064] Combined with Figure 1 and Figure 2 As shown, the embodiments of the present invention provide an adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making, which is applied to the multi-vehicle collaborative perception scenario of the vehicle network. For multiple connected autonomous vehicles equipped with lidar and communication units, by jointly optimizing at the perception and communication levels, intelligent decision-making is realized for the active transmission strategy of perception data and sub-channel allocation, so as to improve the perception accuracy and effectively control the communication delay in a dynamic and bandwidth-limited network environment.

[0065] The adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making specifically includes the following steps:

[0066] Step 1: Each vehicle collects the original point cloud data of the surrounding environment, and after preprocessing the original point cloud data, it is unified to the global coordinate system.

[0067] Specifically, each vehicle uses a lidar sensor to collect the original point cloud data of the surrounding environment in real time.

[0068] In a feasible implementation manner, after preprocessing the original point cloud data, it is unified to the global coordinate system as follows:

[0069] First, the original point cloud data is preprocessed by means of random shuffling, voxel downsampling, and distance filtering.

[0070] Specifically, the collected original point cloud data is randomly shuffled to break the sequential correlation between the original point cloud data. Exemplarily, random shuffling can be achieved by randomly arranging the points in the point cloud data.

[0071] The voxel downsampling algorithm is used to downsample the shuffled point cloud data. Voxel downsampling divides the three-dimensional space into a series of small cubic cells (voxels), and then selects a representative point within each voxel to represent all the points within that voxel. This can reduce the data volume and computational complexity while preserving the main features of the point cloud data.

[0072] Distance filtering is to perform distance filtering on the downsampled point cloud data according to a preset distance threshold. Exemplarily, distance filtering can remove the point cloud data that is too far or too close to the vehicle, and these data usually have less impact on the perception task. By distance filtering, the data volume can be further reduced and the efficiency of perception processing can be improved.

[0073] Finally, the preprocessed point cloud data is unified to the global coordinate system through the pose transformation matrix. Specifically:

[0074]

[0075] In the formula, is the set of point cloud data of vehicle i in the global coordinate system, that is, the unified point cloud data; [p; 1] represents the homogeneous coordinate of point p; T i represents the pose transformation matrix of vehicle i.

[0076] In one embodiment, the connected and autonomous vehicle collects the original LiDAR point cloud data where each point p i,n =[x n , y n , z n , ρ n contains the spatial coordinates (x n , y n , z n ) and the intensity ρ n . To efficiently process the point cloud data for data fusion, while improving the data quality and controlling the computational load, the following preprocessing steps are performed:

[0077] Shuffling and downsampling: To ensure uniform distribution of points and reduce computational complexity, the original point cloud data is randomly shuffled and voxel-based downsampling is used. The point cloud after voxel downsampling is represented as:

[0078]

[0079] The voxel downsampling process aggregates points within voxels of a predetermined size, reducing the point cloud to a manageable size N down .

[0080] Self-vehicle CAV masking: To avoid false detection caused by the reflection of the self-vehicle CAV, remove the points from the self-vehicle. Let Let the point set belong to the ego vehicle CAV, then:

[0081]

[0082] Wherein, It can be determined based on the geometric characteristics of the ego vehicle CAV.

[0083] Range filtering: To focus on the region of interest and eliminate distant points that may contribute limitedly to perception, a range filter is adopted. Only the points within the maximum perception radius r max are retained, that is:

[0084]

[0085] Range filtering confines the point cloud within a circular area centered on the connected autonomous vehicle with a radius of r max .

[0086] After preprocessing the original point cloud data, each connected autonomous vehicle uses the pose transformation matrix to transform the filtered point cloud data into a common reference frame (such as the global coordinate frame) to account for the position and orientation of vehicle i. The formula is as follows:

[0087]

[0088] Wherein, [p; 1] represents the homogeneous coordinates of point p.

[0089] Step 2: Each vehicle voxelizes the unified point cloud data, extracts high-dimensional voxel features from the voxelized point cloud data through a voxel feature encoding algorithm, and generates vehicle-local high-dimensional perception features based on the extracted high-dimensional voxel features.

[0090] It should be understood that voxelizing the unified point cloud data means dividing the three-dimensional space into a series of small cubic cells (voxels), and the point cloud data within each voxel is treated as a whole.

[0091] In one implementable manner, each vehicle voxelizes the unified point cloud data as follows: The unified point cloud data is divided into a number of voxel grids with a fixed resolution, and each voxel grid contains a number of original point clouds

[0092] That is to say, the unified point cloud data is divided into vertical columns on the ground plane, that is, voxels. By discretizing the 3D space into a voxel grid, each voxel covers a specific area in the x and y dimensions. For voxel k, the set of points falling within its spatial boundaries is collected Then, a neural network based on voxel feature encoding is used to extract high-dimensional voxel features from each voxel. The neural network based on voxel feature encoding processes the points within each pillar to generate a feature vector f of a fixed size i,k , which is expressed as:

[0093]

[0094] The neural network based on voxel feature encoding usually consists of layers that capture the geometric and intensity information of the points, thereby generating discriminative features for each voxel.

[0095] In one implementable manner, vehicle local high-dimensional perception features are generated based on the high-dimensional voxel features, specifically as follows:

[0096] First, a sparse and learnable 2D pseudo-map is generated according to the high-dimensional voxel features, which is represented in the bird's-eye view framework as:

[0097] F BEV,i = Scatter({f i,k})

[0098] The 2D pseudo-map is used to represent the spatial distribution of the high-dimensional voxel features on the ground plane. Among them, F BEV,i is the 2D pseudo-map of vehicle i; Scatter(·) is the scattering operation; {f i,k} is the set of vectors of voxel k in the point cloud data of vehicle i.

[0099] Finally, a convolutional backbone network is used to refine the 2D pseudo-map to generate vehicle local high-dimensional perception features. The convolutional backbone network consists of multiple convolutional layers with batch normalization and activation functions, which is described as:

[0100] F backbone,i = Backbone(F BEV,i )

[0101] Among them, F backbone,i is the vehicle local high-dimensional perception feature of vehicle i; Backbone(·) is the convolutional backbone network for feature extraction.

[0102] Step 3: Each vehicle adaptively determines the compression ratio according to the signal-to-noise ratio of the current channel, and uses the compression ratio to perform dynamic compression encoding on the vehicle local high-dimensional perception features to obtain vehicle local compressed features.

[0103] Specifically, the higher the signal-to-noise ratio, the lower the compression ratio is used. When the channel condition deteriorates, the compression ratio is increased accordingly to relieve the bandwidth pressure and reduce the impact of excessive transmission delay on collaborative perception. In other words, a channel with a high signal-to-noise ratio can adopt a lower compression ratio to retain more detailed information, while a channel with a low signal-to-noise ratio adopts a higher compression ratio to reduce the data volume, enabling both data fidelity and communication bandwidth limitations to be taken into account at different signal-to-noise ratio levels. The adaptive compression strategy ensures the efficient utilization of available bandwidth and adapts to changing channel conditions.

[0104] In an implementable manner, each vehicle adaptively determines the compression ratio according to the signal-to-noise ratio of the current channel, specifically:

[0105] Considering the signal-to-noise ratio SNR i,j between vehicle i and vehicle j, the compression ratio η i,j between vehicle i and vehicle j is dynamically adjusted according to SNR i,j and is defined as:

[0106] η i,j = f compression (SNR i,j )

[0107] Specifically, the function f compression (·) maps SNR i,j to the compression ratio according to predefined thresholds and is defined as:

[0108]

[0109] where η i,j is the compression ratio between vehicle i and vehicle j; γ1, γ2, γ3, γ4, γ5 are channel thresholds.

[0110] This piecewise function ensures that when the channel condition is poor (i.e., SNR i,j is low), a higher compression ratio is adopted to reduce the data volume, thereby achieving successful transmission under limited bandwidth. On the contrary, when the channel condition is excellent (i.e., SNR i,j is high), transmission can be carried out without compressing the data to maintain the data quality. Therefore, η i,j is set to 0.

[0111] To achieve efficient V2V communication, an encoder-decoder architecture is deployed at the sender for feature compression. The encoding process at the sender can be expressed as:

[0112] F encoded,i = Encoder(F backbone,i )

[0113] where F encoded,iis the compressed feature of vehicle i; Encoder(·) consists of 2D convolutional and max pooling layers, which compress the backbone features into a compact representation.

[0114] Step 4: Each vehicle iteratively solves the pre-constructed multi-vehicle collaborative perception optimization problem using a multi-agent reinforcement learning method based on the global state, to obtain a communication sub-channel for transmitting the local compressed feature of the vehicle and a target vehicle for receiving the local compressed feature of the vehicle; wherein, the optimization objective of the multi-vehicle collaborative perception optimization problem is to maximize the perception accuracy and minimize the communication delay.

[0115] It should be understood that a multi-agent reinforcement learning framework based on the global state is constructed, where each vehicle acts as an agent, and the agent learns to make optimal decisions in a dynamic environment through interaction and communication with other agents. The perception accuracy is measured by evaluating the consistency between the fused perception data and the real environment, and the communication delay is evaluated by measuring the transmission time of data from the sender to the receiver.

[0116] In one realizable manner, an overall optimization problem is constructed based on the perception and communication models. The goal is to jointly optimize the data sharing decision and communication resource allocation to maximize the perception accuracy and minimize the communication delay. The multi-vehicle collaborative perception optimization problem is formulated as follows:

[0117]

[0118] The constraints are as follows:

[0119] C1: Limit the transmission power of each connected autonomous vehicle not to exceed P max .

[0120] C2: Limit the total bandwidth usage of all connected autonomous vehicles not to exceed B total .

[0121] C3: Set a minimum communication rate R for any pair of connected autonomous vehicles min .

[0122] C4: Ensure the communication delay D of each connected autonomous vehicle i does not exceed D max .

[0123] C5: Limit the total number of subcarrier allocations on each sub-channel m not to exceed K.

[0124] C6: Enforce the perception accuracy AP of each connected autonomous vehicle i (t) is not lower than AP min .

[0125] C7: Ensure the processing cost C of each connected autonomous vehicle i(t) does not exceed its maximum processing cost C max 。

[0126] C8: Define the value range of relevant variables, and specify t i,j and s i,m as binary variables, and k i,m as a non-negative integer.

[0127] In the formula, ΔAP imp,i represents the improved value of the perception accuracy of vehicle i after fusion; D i represents the maximum communication delay of vehicle i on each target vehicle link; α is the coefficient for weighing the perception gain; β is the coefficient for weighing the communication delay overhead; V is the set of vehicles; t i,j is whether vehicle i sends local compressed features to vehicle j. When t i,j = 0, it does not send; when t i,j = 1, it sends; s i,m is whether vehicle i selects sub-channel m. When s i,m = 0, it does not select; when s i,m = 1, it selects.

[0128] A more detailed description of the multi-vehicle collaborative perception optimization problem is as follows:

[0129] The standard path loss model is adopted to characterize the signal attenuation between two connected autonomous vehicles. The path loss PL i,j between connected autonomous vehicle i and connected autonomous vehicle j is:

[0130]

[0131] where d i,j is the distance between connected autonomous vehicle i and j, and f is the carrier frequency (Hz).

[0132] The signal-to-noise ratio between connected autonomous vehicle i and j is calculated as:

[0133]

[0134] where P t is the transmit power of connected autonomous vehicle i, N0 is the noise power spectral density (W / Hz), and B i,j is the total bandwidth allocated to the communication link between connected autonomous vehicle i and j. The gain represents the ratio of the received power to the transmit power due to path loss.

[0135] To achieve the goal of balancing perception accuracy and communication delay, the Shannon-Hartley theorem is introduced to calculate the communication rate R i,j between vehicle i and vehicle j, and let:

[0136] R i,j = B i,j log2(1 + SNR i,j )

[0137] If the size of the original point cloud data of vehicle i is S i , under the compression ratio η i,j between vehicle i and vehicle j, the size of the effective point cloud data of compressed vehicle i is:

[0138]

[0139] The transmission delay (transmission delay of a single link) D from vehicle i to vehicle j i,j is:

[0140]

[0141] To reflect the worst-case delay experienced by connected and autonomous vehicle i when sharing data, the communication delay of connected and autonomous vehicle i is defined as the maximum transmission delay among all its communication links, that is:

[0142]

[0143] is the set of connected and autonomous vehicles used to transmit data.

[0144] The bandwidth division strategy is: divide the total bandwidth B total into M sub-channels, and each sub-channel m contains K sub-carriers; when n m vehicles select the same sub-channel m, first evenly divide the sub-carriers by K / n m , and the remaining part is randomly allocated. Let the binary variable s i,m indicate whether vehicle i selects sub-channel m. Assume that connected and autonomous vehicles select sub-channels from the set . When multiple connected and autonomous vehicles select the same sub-channel, the allocation of their sub-carriers is as follows:

[0145] Equal distribution: When sub-channel m is selected by n m connected and autonomous vehicles, each connected and autonomous vehicle is allocated k i,m sub-carriers

[0146] Random allocation of remaining sub-carriers: The remaining sub-carriers K - n m k m are randomly allocated to n m connected and autonomous vehicles.

[0147] Let ki,m Denote the number of sub - carriers allocated by the connected autonomous vehicle i on sub - channel m. Then, the bandwidth allocated by the connected autonomous vehicle i on sub - channel m is:

[0148]

[0149] where s i,m ∈{0, 1} indicates whether the connected autonomous vehicle i selects sub - channel m. When s i,m = 0, it does not select; when s i,m = 1, it selects. is the bandwidth of each sub - channel. The total bandwidth allocated by the connected autonomous vehicle i on all sub - channels is:

[0150]

[0151] In an implementable manner, the multi - agent reinforcement learning method based on global state is specifically the proximal policy optimization method enhanced by global state information. Each vehicle i with sensing and communication functions is regarded as an agent. Define the state vector of vehicle i at time slot t as:

[0152]

[0153] where, is the position of vehicle i at time slot t; is the speed of vehicle i at time slot t; is the amount of sensing data of vehicle i at time slot t; is the set of positions, speeds, and sizes of sensing data of the remaining vehicles at time slot t.

[0154] Define the action of vehicle i at time slot t as:

[0155]

[0156] where, is the action of vehicle i at time slot t; is the target vehicle selected by vehicle i at time slot t; is the communication sub - channel selected by vehicle i at time slot t.

[0157] Define the global reward r t i of each vehicle i at time slot t as:

[0158]

[0159] where, is the global sensing accuracy; is the global communication delay; α is the coefficient for weighing the perception gain; β is the coefficient for weighing the communication delay overhead.

[0160] Step 5: Each vehicle sends the locally compressed features of the vehicle to the target vehicle through the communication sub-channel.

[0161] By implementing the adaptive multi-vehicle collaborative perception method based on intelligent distributed decision provided by the present invention, networked autonomous vehicles can, while ensuring the perception accuracy, minimize the communication delay and save bandwidth resources to the greatest extent.

[0162] As a more preferable implementation manner, after each vehicle sends the locally compressed features of the vehicle to the target vehicle through the communication sub-channel, it further includes:

[0163] Step 6: After the target vehicle receives the compressed features from multiple source vehicles, it decodes the compressed features and performs a fusion operation on the decoded features and the features of its own vehicle to obtain the fused features.

[0164] Specifically, when the target vehicle decodes the compressed features, that is, the compressed feature reconstruction at the receiving end is as follows:

[0165] F decoded,i = Decoder(F encoded,i )

[0166] where, F decoded,i is the feature obtained after decoding the compressed features of vehicle i; Decoder(·) uses a transposed convolutional layer to restore the feature map. This compression scheme realizes efficient feature sharing while maintaining high perception accuracy.

[0167] The decoded features are then forwarded to the subsequent fusion module for a fusion operation on the decoded features. Specifically:

[0168] When the compressed features of multiple source vehicles arrive synchronously, that is, after receiving the decoded features of the source vehicles in the set , the decoded features are fused by using element-wise maximum operation (max pooling). Specifically:

[0169]

[0170] where, F fused is the fused feature; are multiple source vehicles that send perception information to the target vehicle.

[0171] Step 7: The target vehicle inputs the fusion feature into the detection head, outputs the target detection box and confidence, and calculates the intersection over union (IoU) and mean average precision (mAP) based on the target detection box and confidence, and evaluates the collaborative perception performance using the IoU and mAP.

[0172] Specifically, the fusion feature (feature map) F fused generates detection outputs through separate classification and regression heads.

[0173] Classification head: Calculates the target probability at each spatial location, representing the likelihood of the presence of a target:

[0174] p j = σ(Conv cls (F fused ))

[0175] where σ(·) represents the Sigmoid activation function, j represents the spatial location in the feature map, and Conv cls (·) represents the convolution layer dedicated to classification.

[0176] Regression head: Outputs the adjustment amounts for predefined anchor boxes to better fit the detected target. The regression output includes location offsets, dimensions, and orientation parameters, described as:

[0177] Δ j = Conv reg (F fused )

[0178] Using the regression output Δ j = (Δx j , Δy j , Δz j , Δl j , Δw j , Δh j , Δθ j ) and the predefined anchor box A = (x a , y a , z a , l a , w a , h a , θ a ), calculate the final bounding box parameters P j = (x j , y j , z j , l j , w j , h j , θ j ). The diagonal of the anchor box d a is calculated as follows:

[0179]

[0180] The update of the bounding box parameters is as follows:

[0181] x j = Δx j ·d a + x a , y j = Δy j ·d a + y a , z j = Δz j ·h a + z a ,

[0182] l j = exp(Δl j )·l a , w j = exp(Δw j )·w a , h j = exp(Δh j )·h a ,

[0183] θ j = Δθ j + θ a

[0184] This transformation adjusts the anchor boxes according to the predictions of the network to more accurately fit the detected objects. The detection network is trained using the following multi-task loss function, which combines classification and regression losses:

[0185]

[0186] where L cls and L reg represent the classification and regression losses respectively. γ and δ are weight factors. To evaluate the detection performance, the intersection over union (IoU) metric between the predicted box P j and the ground truth box G j is used, which is defined as:

[0187]

[0188] Based on the IoU threshold (e.g., IoU ≥ 0.7), true positives (TP), false positives (FP), and false negatives (FN) are determined. Then the precision and recall are calculated:

[0189]

[0190] where TP(k) and FP(k) are the cumulative true positive and false positive counts before rank k, respectively, and N gtis the total number of ground truth boxes. The average precision is calculated as the area under the precision-recall curve:

[0191]

[0192] This metric provides a comprehensive evaluation of the detection performance, and a higher AP value indicates better perception ability.

[0193] In one embodiment, as Figure 3 shown, this embodiment considers a typical road intersection scenario involving multiple connected and autonomous vehicles, denoted as Assume that each connected and autonomous vehicle is equipped with a LiDAR sensor and a wireless communication module to achieve real-time data sharing capabilities. In this case, each connected and autonomous vehicle has two functions simultaneously: as a data source, it broadcasts its perception data, and at the same time, it is also a target that may require the perception data of other connected and autonomous vehicles. Specifically, each source connected and autonomous vehicle dynamically identifies and selects the target vehicle with the highest demand for its perception data according to the current traffic conditions, channel quality, and potential perception requirements, and at the same time determines the best communication channel for data transmission.

[0194] The perception ability of each connected and autonomous vehicle is measured by AP@IoU70, denoted as AP70, as Figure 3 shown. The perception accuracy can be improved through data fusion, and AP70 will increase after receiving valuable data from other connected and autonomous vehicles. For example, CAV1 has a low initial perception accuracy due to the occlusion of surrounding vehicles. After receiving the data from CAV5, the perception accuracy of CAV1 is significantly improved, demonstrating the effectiveness of selecting CAV5 as a data source. In contrast, although CAV3 receives information from CAV2, the accuracy only increases slightly due to the similarity between its perception data and the received data, highlighting the importance of selecting the appropriate data source. The perception information shared among CAVs mainly includes the intermediate features preprocessed and extracted by the neural network.

[0195] In addition to selecting targets, each connected and autonomous vehicle must also determine which sub-channels to use for communication. As Figure 3 shown, consider a set of sub-channels Each sub-channel is further divided into K sub-carriers. When multiple CAVs select the same sub-channel, the K sub-carriers are first evenly distributed to these CAVs. If the distribution is uneven, the remaining sub-carriers are randomly distributed. Let s i ={s i,1 ,…,s i,M} be the sub-channel selection vector of , where the binary variable s i,m = 1 indicates that CAVi occupies sub-channel m, si,m = 0 means not occupied. The total available bandwidth B total is evenly allocated to M sub-channels, and each sub-channel occupies a bandwidth of B c = B total / M.

[0196] In one embodiment, to achieve the optimization goal, this embodiment proposes a collaborative sensing framework based on multi-agent reinforcement learning, as Figure 4 shown. When a connected autonomous vehicle captures the original point cloud data, the framework preprocesses it through a voxel feature network. This network converts the original point cloud data into structured features by stacking voxel modules, and then extracts meaningful features through deep learning to generate a representative 2D pseudo-image. The preprocessed data and the environmental state information are then input into the decision-making module, which generates a policy for selecting data transmission targets and allocating sub-channels. Once the data transmission policy is determined, the sensed data is sent through the selected sub-channel. Once other connected autonomous vehicles receive the sensed data, the features are first extracted through a sparse convolutional layer to generate preliminary spatial features. The fused feature map then generates detection outputs through independent classification and regression heads. The entire framework continuously optimizes the decision-making policy through a designed reward mechanism, and after training converges, it can make reasonable decisions based on the real-time environmental state to achieve a dynamic balance between communication efficiency and sensing performance.

[0197] To integrate the intelligent decision-making component into the collaborative sensing framework, the decision-making process is modeled as a Markov decision process, defined as a tuple where is the state space, is the action space, R is the reward function, and γ ∈ [0,1] is the discount factor.

[0198] At time step t, the state of agent i (i.e., connected autonomous vehicle i) is defined as:

[0199]

[0200] where, represents the three-dimensional position vector of agent i, represents the speed scalar of agent i, represents the size of the sensor data generated by agent i. The matrix contains the information of the other N - 1 agents in the system. Each row corresponds to the state of another agent j, defined as where and represent the three-dimensional position vector and speed scalar of agent j respectively, represents the size of the sensor data generated by agent j.

[0201] Agent i takes an action at time step t It consists of two parts, namely transmission target selection and sub-channel selection The action is defined as:

[0202]

[0203] Target vehicle selection: Consider two different target vehicle selection methods to meet the requirements of different scenarios.

[0204] One is single-target discrete selection. Define a scalar indicating the index of the selected target vehicle. For the single-target discrete selection method:

[0205]

[0206] This means that data can only be transmitted to one agent at a time. For example, if then i will transmit data to agent 2 at time step t.

[0207] The other method is multi-target binary vector, defined as:

[0208]

[0209] In this formula, is a binary vector, where indicates that agent i decides to transmit data to agent j, and 0 indicates not to transmit. This method allows data to be transmitted to multiple connected autonomous vehicles simultaneously. For example, means that agent i will transmit data to agent 0 and agent 2 at time step t, and not to agent 1 and agent 3.

[0210] Sub-channel selection: Similarly, consider two sub-channel selection methods: single sub-channel discrete selection and multi-sub-channel binary vector. For the first method, define a scalar indicating the index of the selected sub-channel, defined as:

[0211]

[0212] This means that only one sub-channel can be used for data transmission at a time. For example, if then agent i will transmit only using sub-channel 3 at time step t.

[0213] For the second method, the multi-sub-channel binary vector is defined as:

[0214]

[0215] where is a binary vector, indicating that agent i selects subchannel m, and 0 indicates non - selection. This method allows CAVs to transmit using multiple subchannels simultaneously. For example, if then agent i will transmit using subchannel 0 and subchannel 2 at time step t.

[0216] Based on the global performance metric to emphasize the cooperation among agents in the multi - agent system, the reward r of each agent at time step t t i is defined as:

[0217]

[0218] where represents the global communication delay, and N is the total number of agents. is the global AP improvement, that is, the sum of the AP improvements of all agents:

[0219]

[0220] where F i global includes all agents that transmit data to agent i at time step t. To ensure stable learning, the reward is normalized based on the minimum and maximum values of the global perception improvement and the global delay. The normalized reward function is expressed as:

[0221]

[0222] where AP max and AP min represent the maximum and minimum values of the global perception improvement respectively, and D max and D min represent the maximum and minimum values of the global communication delay. This normalization ensures that the reward remains within a consistent range, facilitating smooth learning and more stable convergence during the training process.

[0223] To achieve the optimal policy in the multi - agent environment, a proximal policy optimization algorithm based on global information is adopted. This algorithm is constructed based on the proximal policy optimization algorithm for independent learning. The global - information variant of proximal policy optimization is used to perform policy search for the transmission target selection and subchannel allocation of each vehicle at different time slots, including:

[0224] Sampling phase: Collect the state, action, and return sequence of each vehicle in the actual or simulation environment;

[0225] Advantage estimation: Use Generalized Advantage Estimation (GAE) to calculate the advantage function of the vehicle at each time slot

[0226] Policy update: Introduce a clipping parameter ∈ into the loss function to limit the large changes in the probability ratio f i (t), i.e., clip(f i (t), 1 - ∈, 1 + ∈). By minimizing the truncated policy update term, suppress the drastic gradient changes;

[0227] Repeated training and iteration: Continuously approximate the optimal solution of the joint optimization problem under the premise of satisfying the sensing and communication processes, so that the vehicle can obtain higher sensing accuracy and lower communication delay in a network environment with limited bandwidth and dynamic channels.

[0228] In the proposed algorithm, each agent uses proximal policy optimization to learn a decentralized policy, and at the same time uses the shared global state information to enhance coordination and training stability. As a deep reinforcement learning algorithm based on policy gradients, this algorithm introduces a policy clipping mechanism to prevent large policy updates, thus ensuring the stability of training. For agent i, its policy parameters are represented by θ i , i.e., π i (θ i ). The goal is to find the optimal policy parameters to maximize the expected cumulative reward. The corresponding optimization problem is defined as:

[0229]

[0230] where is the state value function of agent i, estimating the expected return starting from state s(t) under policy π i (θ i ). is defined as:

[0231]

[0232] Similarly, the action value function is defined as:

[0233]

[0234] where R i (t) is the discounted cumulative reward of agent i at time step t. Define γ ∈ [0, 1) as the discount factor, which is used to balance the immediate reward and the future reward r i (t) is the immediate reward of agent i at time step t. R i (t) is defined as:

[0235]

[0236] According to the policy gradient theorem, the policy gradient can be derived for agent i:

[0237]

[0238] where is the advantage function, which measures how good the action a i (t) is relative to the average action under the policy π i (θ i ) in state s(t). is defined as:

[0239]

[0240] To enhance the stability of training, Generalized Advantage Estimation (GAE) is incorporated into the algorithm. The GAE-based advantage estimate is defined as:

[0241]

[0242] where is the TD(0) error. λ ∈ [0, 1] is the GAE parameter, which is used to control the trade-off between bias and variance in the advantage estimate. When λ = 0, the GAE advantage estimate degenerates to a one-step TD error; when λ = 1, it is close to the advantage estimate of the full trajectory. To ensure the stability of training and improve sample efficiency, PPO introduces a surrogate objective with a clipping strategy, which is defined as:

[0243]

[0244] where C(s(t), a i (t)) represents the clipping function, which is defined as:

[0245]

[0246] where f i (t) is the probability ratio of the new policy to the old policy, and ∈ is the clipping parameter. Define as the policy parameter of the previous iteration, and f i (t) is calculated as follows:

[0247]

[0248] The clipping function clip(f i (t), 1 - ∈, 1 + ∈) helps the stability of the training process by restricting the policy change within a small range controlled by ∈, preventing the policy from deviating too much.

[0249] Based on GAE, the target value of the critic network is calculated as:

[0250]

[0251] The critic network loss function of agent i is defined as the mean squared error between the state value estimate and the target value:

[0252]

[0253] The gradient of the critic network with respect to its parameters is:

[0254]

[0255] In practice, mini-batch sampling is used to approximate the gradient, i.e., the gradient estimate of the critic network is:

[0256]

[0257] This gradient guides the update of the critic network parameters to minimize the prediction error of the state value function. To maximize the clipped agent objective and balance exploration and exploitation, the critic network is updated by:

[0258]

[0259] where, is the learning rate of the critic network. Similarly, the gradient estimate of the actor network of agent i is:

[0260]

[0261] Through mini-batch stochastic gradient ascent, the actor network is updated by:

[0262]

[0263] where, is the learning rate of the actor network and is updated to maximize the clipped surrogate objective and balance exploration and exploitation. In another embodiment of the present invention, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store a computer program. The computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used for the operation of an adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making.

[0264] In another embodiment of the present invention, the present invention also provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space that stores the operating system of the terminal. And in this storage space, one or more instructions suitable for being loaded and executed by the processor are also stored. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the above-mentioned embodiment of an adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making.

[0265] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0266] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0267] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0268] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0269] In the present invention, terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0270] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, which are used to illustrate the technical solutions of the present invention rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. An adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making, which is applied to the multi-vehicle collaborative perception scenario of the vehicle network, and is characterized in that, Including: Each vehicle collects the original point cloud data of the surrounding environment, and preprocesses the original point cloud data and then unifies it to the global coordinate system; Each vehicle voxelizes the unified point cloud data, extracts high-dimensional voxel features from the voxelized point cloud data through a voxel feature encoding algorithm, and generates local high-dimensional perception features of the vehicle based on the extracted high-dimensional voxel features; Each vehicle adaptively determines the compression ratio according to the signal-to-noise ratio of the current channel, and uses the compression ratio to dynamically compress and encode the local high-dimensional perception features of the vehicle to obtain local compressed features of the vehicle; Each vehicle uses a multi-agent reinforcement learning method based on global state to iteratively solve a pre-constructed multi-vehicle collaborative perception optimization problem to obtain a communication sub-channel for transmitting the local compressed features of the vehicle and a target vehicle for receiving the local compressed features of the vehicle; wherein, the optimization objective of the multi-vehicle collaborative perception optimization problem is to maximize the perception accuracy and minimize the communication delay; Each vehicle sends the local compressed features of the vehicle to the target vehicle through the communication sub-channel.

2. The adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making according to claim 1, wherein The preprocessing the original point cloud data and then unifying it to the global coordinate system includes: Preprocessing the original point cloud data by means of random shuffling, voxel downsampling and distance filtering; Unifying the preprocessed point cloud data to the global coordinate system through an attitude transformation matrix, specifically: In the formula, is the point cloud data set of vehicle i in the global coordinate system, that is, the unified point cloud data; [p; 1] represents the homogeneous coordinates of point p; T i represents the attitude transformation matrix of vehicle i.

3. An adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making according to claim 1, characterized in that, The generating local high-dimensional perception features of the vehicle based on the high-dimensional voxel features includes: Generating a sparse and learnable 2D pseudo-map according to the high-dimensional voxel features, specifically: F BEV,i = Scatter({f i,k [[ID=4}]}) Among them, F BEV,i is the 2D pseudo-image of vehicle i; Scatter(·) is the scattering operation; {f i,k} is the vector set of voxel k in the point cloud data of vehicle i; Refining the 2D pseudo-map by using a convolutional backbone network to generate local high-dimensional perception features of the vehicle.

4. An adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making according to claim 1, characterized in that, Each vehicle adaptively determines the compression ratio according to the signal-to-noise ratio of the current channel, specifically: Consider the signal-to-noise ratio SNR between vehicle i and vehicle j i,j , and adopt the following multi-segment mapping: where η i,j is the compression ratio between vehicle i and vehicle j; γ1, γ2, γ3, γ4, γ5 are channel thresholds.

5. An adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making according to claim 1, characterized in that, The multi-vehicle collaborative perception optimization problem is specifically: Wherein, where ΔAP imp,i represents the improvement value of the perception accuracy of vehicle i after fusion; D i represents the communication delay of vehicle i on each target vehicle link; α is the coefficient for weighing the perception gain; β is the coefficient for weighing the communication delay overhead; V is the vehicle set; t i,j ∈ {0, 1} indicates whether vehicle i sends local compressed features to vehicle j. When t i,j = 0, it does not send. When t i,j = 1, it sends; s i,m ∈ {0, 1} indicates whether vehicle i selects sub-channel m. When s i,m = 0, it does not select. When s i,m = 1, it selects; P t is the transmission power of vehicle i; P max is the maximum transmission power of the vehicle; B i is the total bandwidth allocated by vehicle i on all sub-channels; B total is the total bandwidth; B i,m is the bandwidth allocated by vehicle i on sub-channel m; R i,j is the communication rate between vehicle i and vehicle j; R min is the minimum communication rate; D max is the maximum communication delay; K is the number of sub-carriers per sub-channel; AP i (t) is the sensing accuracy of vehicle i; AP min is the minimum sensing accuracy; C i (t) is the processing cost of vehicle i; C max is the maximum processing cost; k i,m is the number of sub-carriers allocated by vehicle i on sub-channel m; is the set of integers; M is the number of sub-channels obtained by dividing the total bandwidth; B c is the bandwidth of each sub-channel; g i,j is the gain; N0 is the noise power spectral density; B i,j is the total bandwidth allocated to the communication link between vehicle i and vehicle j; PL i,j is the path loss between vehicle i and vehicle j; d i,j is the distance between vehicle i and vehicle j; f is the carrier frequency; D i,j is the transmission delay from vehicle i to vehicle j; is the set of vehicles used to transmit compressed features; is the size of the effective point cloud data of vehicle i after compression; S i is the size of the original point cloud data of vehicle i; n m is the number of vehicles selecting the same sub-channel m.

6. An adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making according to claim 5, characterized in that, The multi-agent reinforcement learning method based on global state is specifically the proximal policy optimization method enhanced by global state information; Each vehicle \(i\) with sensing and communication capabilities is regarded as an agent, and the state vector of vehicle \(i\) at time slot \(t\) is defined as follows: Among them, is the position of vehicle i at time slot t; is the speed of vehicle i at time slot t; is the size of the sensed data of vehicle i at time slot t; is the set of the positions, speeds, and sizes of the sensed data of the remaining vehicles at time slot t; Define the action of vehicle i at time slot t as follows: wherein, the action of vehicle i at time slot t; the target vehicle selected for vehicle i at time slot t; the communication sub-channel selected for vehicle i at time slot t; Define the global reward r of each vehicle i at time slot t t i as follows: Among them, is the global perception accuracy; is the global communication delay.

7. An adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making according to claim 1, characterized in that After each vehicle sends the local compressed features of the vehicle to the target vehicle through the communication sub-channel, it further includes: After receiving the compressed features from multiple source vehicles, the target vehicle decodes the compressed features and fuses the decoded features with its own vehicle features to obtain fused features; The target vehicle inputs the fused features into a detection head, outputs a target detection box and a confidence level, and calculates the intersection over union and the average precision according to the target detection box and the confidence level, and evaluates the collaborative perception performance by using the intersection over union and the average precision.

8. An adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making according to claim 7, characterized in that The fusing operation on the decoded features is specifically: When the compressed features of multiple source vehicles arrive synchronously, the decoded features are fused using an element-wise maximum operation, specifically: Among them, F fused is the fusion feature; are multiple source vehicles that send perception information to the target vehicle; F decoded,i is the feature obtained by decoding the compressed feature of vehicle i.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements an adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements an adaptive multi-vehicle collaborative perception method based on intelligent distributed decision-making as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Multi-agent cooperative sensing method and system for Internet of Vehicles

    CN121837878A

  • Underwater acoustic sensor network intelligent compression method based on reinforcement learning

    CN122458092A