Multi-vehicle collaborative fusion sensing method based on spatio-temporal context

Through the multi-vehicle collaborative fusion perception method based on spatiotemporal context, the transmission delay and bandwidth limitation problems in multi-vehicle collaborative perception are solved, efficient and accurate perception information transmission and fusion are achieved, and the perception ability of the autonomous driving system in complex traffic scenarios is improved.

CN120612562AActive Publication Date: 2025-09-09HANGZHOU DIANZI UNIV +1

Patent Information

Application Number
CN202511122181.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-09-09
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Existing multi-vehicle collaborative perception technology faces challenges in transmission delay, communication bandwidth limitation and data fusion, which leads to increased system burden, inference delay and memory overflow, limiting the system's scalability and practical application capabilities.

Method used

A multi-vehicle collaborative fusion perception method based on spatiotemporal context is adopted. Through spatiotemporal alignment and lightweight communication strategies, combined with robust feature fusion methods, it ensures that the perception data is accurately aligned in the spatiotemporal dimensions, and prioritizes the transmission of the most valuable information. The neighborhood spatiotemporal attention LSTM module is used for prediction compensation and spatial decoding and communication module for sparse transmission, realizing vehicle-centered collaborative feature fusion.

Benefits of technology

It significantly reduces the spatiotemporal asynchronous errors caused by communication delays, breaks through the bottleneck of information transmission due to limited bandwidth, improves the accuracy and real-time performance of perception, enhances the ability of autonomous driving systems to cope with complex traffic scenarios, and reduces the safety risks caused by perception errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612562A_ABST
    Figure CN120612562A_ABST
Patent Text Reader

Abstract

The invention discloses a spatio-temporal context-based multi-vehicle collaborative fusion sensing method. The method comprises the following steps of: firstly, acquiring a large-scale open collaborative sensing data set; secondly, inputting metadata in the data set into a metadata sharing and feature extraction module to obtain feature representation of each vehicle; and then the self-vehicle feature representation and the feature representation of each cooperative vehicle are respectively input into corresponding neighborhood space-time modules, two hidden states are output, the hidden states pass through a position embedding input space decoding and communication module, and self-vehicle features, cooperative vehicle features, features of the cooperative vehicles in the shared space and features of the self-vehicle in the shared space are output. And finally, aligning and fusing the obtained features through a spatial adaptive fusion mechanism, outputting fused features, inputting the fused features into a regression head and a classification head, and outputting bounding box prediction information and classification prediction information. According to the invention, the automatic driving vehicle can realize more reliable and timely environment perception in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of autonomous driving and intelligent transportation, and in particular to a multi-vehicle collaborative fusion perception method based on spatiotemporal context. Background Art

[0002] Multi-vehicle collaborative perception technology aims to improve the accuracy and real-time performance of overall environmental perception by sharing sensory information among multiple autonomous vehicles. It has garnered widespread attention in the autonomous driving field in recent years. Currently, research on multi-vehicle collaborative perception is primarily applied to scenarios such as safe driving in complex traffic environments, traffic management, and intelligent connected vehicles. By sharing sensor data from each vehicle, a global perception model is constructed in real time. However, effectively addressing issues such as transmission delays, communication bandwidth limitations, and the fusion of heterogeneous data remains a technical challenge.

[0003] In multi-vehicle collaborative perception research, most existing work focuses on improving perception accuracy and real-time performance, employing complex algorithms for data fusion and optimization. However, overly complex models and computationally intensive operations significantly increase the system burden. In environments with limited hardware resources (such as edge devices or on-board computing platforms), these can easily lead to problems such as inference delays and memory overflows, severely limiting the system's scalability and practical application capabilities. Summary of the Invention

[0004] In response to the above problems, the present invention proposes a multi-vehicle collaborative fusion perception method based on spatiotemporal context. This method solves the transmission delay, communication bandwidth limitation and data fusion problems in multi-vehicle collaborative perception through spatiotemporal context perception and efficient communication technology. In this method, spatiotemporal alignment and lightweight communication strategies are introduced to ensure that the perception data of different vehicles can be accurately aligned in the spatiotemporal dimensions, and the most valuable information is transmitted preferentially under limited bandwidth conditions. By designing a robust feature fusion method, this method can effectively integrate the perception data of different vehicles, improve the accuracy and real-time performance of the multi-vehicle collaborative perception system, thereby enhancing the response capability of the autonomous driving system in complex traffic scenarios and effectively reducing the safety risks caused by perception errors.

[0005] The present invention solves the technical problem by comprising the following steps:

[0006] S1. Obtain a large-scale open collaborative perception simulation dataset and a real dataset, mix them, and divide them into training, validation, and test data.

[0007] S2. The metadata (such as pose and external parameters) in the mixed dataset are input into the metadata sharing and feature extraction module. The ego vehicle builds a collaborative graph and broadcasts the metadata. After receiving the metadata, the collaborative vehicle projects its point cloud data into the ego vehicle's coordinate system and aligns it with the ego vehicle's current LiDAR point cloud data. The point cloud after alignment with each collaborative vehicle is encoded into a bird's-eye view feature through the point cloud encoder, and the output of each vehicle in Feature representation of time , the feature dimension is ,in 、 and Represent the height, width and number of channels respectively, Indicates the index of the vehicle.

[0008] S3. The car (index is ) Feature Representation and each cooperative vehicle (indexed by ) Feature Representation Input them into the corresponding neighborhood spatiotemporal LSTM modules respectively, and output the hidden state after spatiotemporal information refinement. The specific steps of the neighborhood spatiotemporal attention LSTM module are as follows:

[0009] S31. Input the hidden state of the previous moment and current moment features , first concatenated in the channel dimension and then normalized.

[0010] S32. Input the normalized result output by S31 into the neighborhood attention module, calculate the attention only within the local window of each position, extract the local spatiotemporal context features, and obtain the neighborhood attention .

[0011] S33. Neighborhood attention obtained in step S32 Apply the Sigmoid activation function element by element , get the gate weight , used to control information retention and update.

[0012] S34. Calculation Cell state at any moment :Get the cell state at the last moment ( Cell state at any moment ),current Moment-by-moment neighborhood attention features and gate weights After that, first set the gate weight The cell state at the previous moment Multiply element by element. Through the tanh activation function, and then with Multiply element by element to extract the new features that need to be supplemented. Finally, add the results of the element-by-element multiplication of the two parts to get the updated cell state .

[0013] S35. Input cell status , first perform tanh activation on it to obtain The feature map within the interval, and then the activation result and the gate weight Multiply element by element to obtain the hidden state after spatiotemporal information refinement , as a spatiotemporal feature.

[0014] S4. The hidden state output of step S3 After position embedding, the input is spatial decoding and communication module, and the vehicle-specific features are output , Unique regional characteristics of cooperative vehicles , Characteristics of cooperative vehicles in shared spaces and the characteristics of the car in the shared space The specific steps of the spatial decoding and communication module are as follows:

[0015] S41. Input the spatiotemporal characteristics of each vehicle , encode the perceptual space, and obtain features containing position information by adding a learnable position embedding to the spatiotemporal feature .

[0016] S42. Input features containing location information , apply average pooling operations on the width and height dimensions respectively, encode the width and height dimensional information into the channel dimension, and obtain spatial features and .

[0017] S43. Spatial features and Align the dimensions by broadcasting and concatenate them, and then perform nonlinear transformations Processing is performed to obtain the intermediate vector. Then the spatial attention score vectors in the width and height dimensions are calculated and the spatial confidence map is generated. .

[0018] Get the spatial confidence map Finally, the important positions are filtered according to the set threshold through the screening function. If the confidence of a position is greater than the threshold, it is marked as 1, otherwise it is marked as 0. Then, we get the corresponding important features .

[0019] S44. Collaborative vehicles will place important spatial positions and corresponding important features After receiving the data, the car calculates the unique perception area features of the car. , Unique perception area characteristics of cooperative vehicles and features of the shared perceptual area 、 .

[0020] S5. Input common receptive area features and , align and fuse through the spatial adaptive fusion mechanism to output unified common perception area features . Then the vehicle’s unique perception area features are , Unique perception area characteristics of cooperative vehicles and unified shared perceptual region features Input to the collaborative feature fusion module centered on the vehicle, perform robust feature fusion, and finally output the fused feature The specific steps of the vehicle-centric collaborative feature fusion module are as follows:

[0021] S51. First, the shared perception area features of the self-vehicle and the cooperative vehicle are and The two features are spliced ​​together to form a joint feature tensor containing the information of both. Then, the spliced ​​tensor is subjected to maximum pooling and average pooling respectively, and the two pooling results are spliced ​​along the channel direction to obtain a composite feature with both local strong response and global average response. Finally, the composite feature is fused and reorganized in the spatial dimension through the sliding convolution operation to output a unified common perception area feature. .

[0022] S52. Spatiotemporal characteristics of the vehicle Each position in , first extract the query vector , Is a learnable weight matrix; and defines its neighborhood matrix , then based on the neighborhood matrix , calculate the corresponding neighborhood matrix in the collaborative space through coordinate alignment operation , and 、 and Neighborhood matrix Multiply the corresponding learnable matrix respectively and , get the corresponding three sets of key vectors Sum vector ; Then calculate the attention weight to obtain the unified representation of the characteristics of the vehicle's unique perception area from the perspective of the vehicle , Unified representation of the characteristics of the cooperative vehicle's unique perception area from the perspective of the vehicle Unified representation of features of the unified shared perception area from the perspective of the vehicle .

[0023] S53. Combine the fused vehicle-specific perception area features , the unique perception area characteristics of the fused cooperative vehicles and the unified shared perception area features after fusion The features of each pixel position are convolved to generate the corresponding original weight distribution map. Then, the Sigmoid activation function is applied to the generated weight map to map all values ​​to Within the range, the spatial importance map of the vehicle’s unique perception area is obtained. , Spatial importance diagram of cooperative vehicle-specific perception area and a unified shared perception region spatial importance map .

[0024] S54. Combine three spatial importance maps 、 and , the standardized attention map between different features is obtained through the softmax function Similarly, the calculation formula of the vehicle-specific perception space attention map is , the calculation formula of cooperative vehicle’s unique perception space attention map is .

[0025] S55. Spatial Attention Map 、 and Respectively with the corresponding feature maps 、 and Multiply element-wise in the spatial dimension to obtain three weighted feature maps. Then, these three weighted feature maps are directly added in the channel and space to obtain the final fusion feature.

[0026] S6. During the inference process, the final fused features output from step S5 are input to the regression head, and the bounding box prediction information is output, including the position, size, and yaw angle of the bounding box at each position.

[0027] The final fused features are input to the classification head, which outputs classification prediction information, indicating the confidence value of each predefined box as a target or background.

[0028] S7. During training, this method uses smooth L1 loss for regression and focal loss for classification.

[0029] The beneficial effects of the present invention are: the proposed multi-vehicle collaborative fusion perception method based on spatiotemporal context, through the overall technical path of "prediction compensation-lightweight communication-robust fusion", realizes the high efficiency and high precision of multi-vehicle collaborative perception in complex dynamic traffic environments: Overall, this method can significantly reduce the spatiotemporal asynchronous error caused by communication delay, break through the information transmission bottleneck of bandwidth limitation, and improve the perception robustness through cross-vehicle fusion strategy, so that autonomous driving vehicles can achieve more reliable and timely environmental perception in complex scenarios. Specifically, the metadata sharing and feature extraction module aligns the perception information of each vehicle in the same semantic space through point cloud projection and shared encoder extraction in a unified coordinate system, improving the prerequisite for multi-source data fusion. The neighborhood spatiotemporal attention LSTM module predicts and compensates for historical frame features based on the local spatiotemporal attention mechanism, effectively alleviating the perception offset problem caused by transmission delay and improving feature prediction accuracy while ensuring real-time performance. The spatial decoding and communication module uses position embedding and spatial attention screening to select only potential key area features for sparse transmission, significantly reducing the amount of communication data and ensuring efficient information exchange under limited bandwidth conditions. The self-vehicle-centric collaborative feature fusion module uses cross-neighborhood attention and spatial importance generators to specifically map the features of the self-vehicle and collaborative vehicles to the self-vehicle perspective, and suppresses noise interference based on a weighted fusion strategy to achieve robust fusion of multi-source features. The detection head module accurately decodes the fused global features and outputs high-confidence three-dimensional detection results, thereby further improving the accuracy and reliability of the final target detection task. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 This is a diagram of the multi-vehicle collaborative fusion perception method based on spatiotemporal context;

[0031] Figure 2 is the neighborhood attention map;

[0032] Figure 3 is the neighborhood LSTM graph;

[0033] Figure 4 Diagram of the vehicle-centric collaborative feature fusion module. DETAILED DESCRIPTION

[0034] The following figures further illustrate the multi-vehicle collaborative fusion perception method based on spatiotemporal context. Figure 1 As shown, the following steps are included:

[0035] S1. Two V2V collaborative perception benchmark datasets, the OPV2V simulation dataset (73 scenarios, an average of 2.89 collaborative vehicles, and a total of 11,464 frames) and the V2V4Real real-world dataset (410 km of real-vehicle road testing, 20,000 frames), are used as training, validation, and test data.

[0036] S2. Metadata sharing and feature extraction module: The metadata (such as pose and external parameters) in the dataset obtained in step S1 are input into the metadata sharing and feature extraction module. A connected vehicle is identified as the ego vehicle to construct the collaborative graph, and other vehicles t within the communication range will provide information as collaborators. The ego vehicle broadcasts metadata, and after the collaborative vehicle receives the metadata, it projects its original point cloud into the coordinate system of the ego vehicle. Similarly, the historical frame of each vehicle is also aligned with the current frame of the ego vehicle. Then, the point cloud of the ego vehicle and the point cloud of each collaborative vehicle after alignment are encoded into bird's-eye view features through the point cloud encoder. Given the converted point cloud of the kth vehicle at time t , the extracted features are ,in Represents the PointPillar point cloud encoder shared between all vehicles, , t represents the timestamp , H, W, and C represent the height, width, and number of channels, respectively.

[0037] S3. Neighborhood spatiotemporal attention LSTM module: The ego vehicle (retrieved as ) Feature Representation and each collaborative vehicle (retrieved as ) Feature Representation Input them into the corresponding neighborhood spatiotemporal attention LSTM module respectively, and output the hidden state after spatiotemporal information refinement By integrating "neighborhood attention" into the hidden state update of the traditional LSTM, the dynamic features of the target object in the local spatiotemporal range are efficiently captured. This enables the module to more quickly and accurately predict and align temporal features in multi-vehicle collaborative perception, effectively alleviating the impact of temporal asynchrony. The specific steps of the neighborhood spatiotemporal attention LSTM module are as follows:

[0038] S31. Input the hidden state at the previous moment (i.e. time t-1) and the current moment (i.e., time t) features , first concatenate in the channel dimension to get , and then normalized by layer normalization LayerNorm, and the normalized result is input into the neighborhood attention module, such as Figure 2 As shown, the top one is conventional self-attention and the bottom one is neighborhood attention.

[0039] S32. Perform neighborhood attention calculation on the normalized output of S31. Compared with the widely used self-attention mechanism, the neighborhood attention mechanism is an attention mode similar to the convolution operation. In neighborhood attention, the focus is limited to a local window centered on a specific query position, thereby extracting relevant keys and values ​​within the window for calculation. Figure 2 As shown in , the self-attention mechanism involves the interaction of global features, which considers the influence of all pixels in the input feature map on the current pixel. Figure 2 As shown in the figure, neighborhood attention limits the scope of attention operations to the neighborhood of each pixel. This approach simulates the "sliding window" effect in the convolution operation and only focuses on the feature interactions in the local area.

[0040] This design not only significantly reduces computational overhead but also better reflects the changing state characteristics of target objects in real-world traffic scenarios. In real-world environments, the motion state of target objects (such as position and velocity) typically changes more significantly within a local area, while global changes are often sparse and difficult to capture. Therefore, by limiting the attention range, the model can more efficiently capture the local dynamic characteristics of target objects, thereby improving perception accuracy and real-time performance.

[0041] Attention is only calculated within the local window of each position, thereby extracting local spatiotemporal context features and obtaining neighborhood attention .

[0042] S33. Although neighborhood attention reduces the computational overhead of global attention, the introduction of the attention mechanism inevitably increases the computational burden. In order to further reduce the computational overhead and reduce network latency, this paper merges the forget gate, input gate, and output gate in the original LSTM. Figure 3 As shown, the neighborhood attention obtained in step S32 Apply the Sigmoid activation function element by element , get the gate weight This gate is used to control information retention and update.

[0043] S34. Calculation Cell state at any moment :Get the cell state at the last moment ( Cell state at any moment ),current Moment-by-moment neighborhood attention features and gate weights After that, first and Multiply element by element, so that only the historical information that is allowed to continue to pass by the gate is retained. First perform tanh activation and map it to interval, and then with Multiply element by element to extract the new features that need to be supplemented. Finally, add the two parts of the results to get the updated cell state , thus integrating historical information and current information into the same tensor. The formula is as follows:

[0044]

[0045] Among them, the first Indicates that the cell information of the previous moment is retained, the second Indicates using the current features to supplement new information. Get the updated cell state .

[0046] S35. Get the updated cell state After that, we first perform tanh activation on it to obtain The feature map within the interval, and then the activation result and the gate weight Multiply element by element so that only the information allowed by the gate can continue to be transmitted, and the hidden state after the spatiotemporal information is refined is obtained. , as the spatiotemporal feature. The formula is as follows:

[0047]

[0048] S4. Spatial decoding and communication module: hidden state output of step S3 After the position is embedded, it is sent to the spatial decoding and communication module to filter out the vehicle (index is ) Unique characteristics , cooperative vehicles (indexed by )'s unique regional characteristics , Characteristics of cooperative vehicles in shared spaces and the characteristics of the car in the shared space The spatial decoding and communication module decodes the perception space of the collaborative vehicle according to the ego vehicle-specific area, the collaborative vehicle-specific area, and the shared area. It calculates the importance confidence map of each area through position encoding and spatial attention, selects high-value sparse features and packages them for transmission, so as to exchange only key perception information under limited bandwidth. This not only significantly reduces communication overhead, but also ensures that the ego vehicle can obtain key spatial clues from different perspectives, laying the foundation for subsequent collaborative feature fusion. The specific steps of the spatial decoding and communication module are as follows:

[0049] S41. Input the spatiotemporal characteristics of each vehicle calculated in step S3 , then encode the perceptual space by embedding a learnable position Added to the spatiotemporal feature ( is a learnable parameter), enhances the vehicle’s sensitivity to spatial position, and thus obtains features containing position information .

[0050] S42. Input the feature containing the location information obtained in step S41 , respectively in width and high Apply average pooling operation on the dimension to encode the width and height dimensional information into the channel dimension to obtain spatial features and .

[0051] S43. The spatial features obtained in step S42 and Align the dimensions by broadcasting and concatenate them, and then perform nonlinear transformations Processing to obtain the intermediate vector .

[0052] Then the spatial attention scores on the width and height dimensions are calculated and a spatial confidence map is generated The formula is as follows.

[0053]

[0054] in, and Represents the segmentation operation in the width and height dimensions, Represents the Sigmoid activation function. Get the spatial confidence map Then, the filter function is used according to the set threshold To filter important locations. If the confidence of a location is greater than the threshold, it is marked as 1, otherwise it is marked as 0. Get important spatial locations Then, we get the corresponding important features .

[0055] S44. Collaborative vehicles will place important spatial locations and corresponding important features After receiving the data, the ego vehicle calculates the features of its own unique perception area, the cooperative vehicle’s unique perception area, and the shared perception area:

[0056] in Indicates that the area covered by the cooperative vehicle is eliminated and only the mask of the area exclusively occupied by the vehicle is retained, combined with the sparse position of the vehicle itself and weighted features Get the unique effective features of the vehicle The formula is as follows:

[0057]

[0058] in Represents element-wise multiplication. Indicates that the area covered by the own car is removed and only the mask of the area exclusively occupied by the cooperative car is retained, combined with the sparse mask of the cooperative car. and weighted features Get the unique and effective features of cooperative vehicles The formula is as follows:

[0059]

[0060] The area shared by the autonomous vehicle and the cooperative vehicle is divided into two parts. Indicates the position that both cars can perceive, and the shared features are updated by the car. Updated features of the collaborative car Weighting is performed to ensure that the perception information of the common area between the two is retained. The formula is as follows:

[0061]

[0062]

[0063] Finally, we get the unique characteristics of our vehicle , Unique regional characteristics of cooperative vehicles , Characteristics of cooperative vehicles in shared spaces and the characteristics of the car in the shared space . It provides complete and hierarchical spatial information for subsequent fusion and detection modules.

[0064] S5. Vehicle-centric collaborative feature fusion module: Input the features of the common area output in step S4 and Align and fuse through spatial adaptive fusion mechanism to generate unified common regional features . Then the vehicle’s unique perception area features are , Unique perception area characteristics of cooperative vehicles and unified shared perceptual region features Input to the collaborative feature fusion module centered on the vehicle for robust feature fusion, and finally obtain the fused feature The vehicle-centric collaborative feature fusion module maps the perception features from surrounding vehicles or infrastructure to the vehicle coordinate system and fuses them to improve the accuracy and robustness of environmental perception. The specific steps of the vehicle-centric collaborative feature fusion module are as follows:

[0065] S51. Input feature maps of the self-vehicle and cooperative vehicles in the shared perception area 、 , aggregate the shared perception area features of the self-vehicle and the cooperative vehicle. First, the feature maps of the self-vehicle and the cooperative vehicle in the shared perception area are and Then, the spliced ​​tensor is subjected to maximum pooling. and average pooling , where max pooling is used to extract salient information from the feature map (i.e., capturing the most prominent activations), while average pooling is used to preserve the overall distribution of features. These two pooling results are concatenated along the channel direction to obtain a composite feature with both strong local responses and global average responses. Finally, a 3×3 convolution kernel is applied to the concatenated tensor. , through sliding convolution operation, the information is fused and reorganized in the spatial dimension, and the unified common perception area fusion feature is output. The formula is as follows:

[0066]

[0067] S52. This method realizes cross-neighborhood attention and unifies the perception areas of all perspectives into the representation of the vehicle’s perspective. Figure 4 As shown, for the spatiotemporal characteristics of the vehicle Each position in , first extract the query vector , Is a learnable weight matrix; and defines its neighborhood matrix , then based on the neighborhood matrix , calculate the corresponding neighborhood matrix in the collaborative space through coordinate alignment operation , and 、 and Neighborhood matrix Multiply the corresponding learnable matrix respectively and , get the corresponding three sets of key vectors Sum vector ; Then calculate the attention weight; Then, normalize the attention weight through the softmax function and calculate the value vector Perform weighted summation to obtain a unified representation of the features of the vehicle's unique perception area from the vehicle's perspective , Unified representation of the characteristics of the cooperative vehicle's unique perception area from the perspective of the vehicle Unified representation of features of the unified shared perception area from the perspective of the vehicle The formula is as follows:

[0068]

[0069] in, Indicates different perceptual areas.

[0070] S53. Use the importance generator to calculate the spatial importance map after different perception space regions are enhanced. First, the fused vehicle-specific perception area features obtained in step S52 are , the unique perception area characteristics of the fused cooperative vehicles and the unified shared perception area features after fusion Input into the same importance generator The importance generator processes the features of each pixel position and generates the corresponding original weight distribution map. Then, the Sigmoid activation function is applied to the generated weight map to map all values ​​to Within the range, the spatial importance map of the vehicle’s unique perception area is obtained. , Spatial importance diagram of cooperative vehicle-specific perception area and a unified shared perception region spatial importance map These three importance maps achieve quantitative labeling of the contribution of different perceptual spatial regions by indicating that the closer the value is to 1, the greater the contribution of the spatial position to the final perception result, and the closer it is to 0, the weaker the contribution of the position.

[0071] S54. Concatenate the three spatial importance maps obtained in S53 、 and , through the softmax function Get the normalized attention map between different features The formula is as follows:

[0072]

[0073] in, Represents exponential operation. Similarly, the calculation formula of the vehicle-specific perception space attention map is , the calculation formula of cooperative vehicle’s unique perception space attention map is .

[0074] S55. The spatial attention map obtained in step S54 、 and Respectively with the corresponding feature maps 、 and Multiplying the elements in the spatial dimension will result in three weighted feature maps. This step will amplify or suppress the feature response at each position according to the importance of the perceptual space to which it belongs. Then, these three weighted feature maps are directly added in the channel and space to obtain the final fusion feature. The purpose of this is to use the attention map to dynamically weight the three parts of information: public perception, ego-vehicle-specific perception, and cooperative vehicle-specific perception, to more strongly retain more important features at the same location, thereby generating a comprehensive feature map that takes into account the key information of the three perception spaces. The formula is as follows:

[0075]

[0076] in, Represents element-wise multiplication of spatial dimensions.

[0077] S6. During the inference process, the final fusion feature output of step S5 is Input to the regression head and output bounding box prediction information , including the position, size, and yaw angle of the bounding box at each location.

[0078] The final fusion feature Input to the classification head and output classification prediction information , represents the confidence value of each predefined box as a target or background,

[0079] S7. During training, this method uses smooth L1 loss for regression and focal loss for classification. The total training loss formula is:

[0080]

[0081] The specific training details are as follows:

[0082] S71. The smooth L1 loss formula is as follows:

[0083]

[0084] in, is the true value of the 𝑖th sample, is the predicted value of the 𝑖th sample. The formula is as follows:

[0085]

[0086] in, It is a parameter that defines the smoothing area and is used to control the smooth transition range.

[0087] The focal classification loss formula is as follows:

[0088]

[0089] in is the predicted probability of the 𝑖th sample to its true label; Used to balance positive and negative samples, Used to focus on difficult samples.

[0090] S72. Training begins with a "single vehicle warm-up": freeze the collaborative feature fusion module and train for several epochs on the independent vehicle data using the above loss to stabilize the regression-classification head. Then, unfreeze the attention layer and enter "cooperative fine-tuning." In each batch, randomly apply communication packet loss, occlusion, and temporal jitter to improve robustness to dynamic topology and bandwidth fluctuations. The optimizer uses AdamW (initial learning rate , weight decay , cosine annealing), batch normalization and gradient clipping jointly stabilize the training; the super parameter is .

[0091] Example:

[0092] The steps of this embodiment are the same as those of the specific implementation method, and will not be repeated here. The implementation process and results are shown below.

[0093] To comprehensively evaluate model performance, this paper sets up two benchmark comparison groups: No Fusion (a single-vehicle perception baseline without cooperative communication) and Late Fusion (a traditional late fusion solution). Also included in the comparison are five advanced cooperative perception models: F-Cooper (3D point cloud-based cooperative perception for connected autonomous vehicles), V2VNet (vehicle-to-vehicle communication for joint perception and prediction), V2X-ViT (Visual Transformer-based cooperative perception for connected vehicles), Where2comm (Efficient Communication via Spatial Confidence Maps), and ERMVP (Efficient Communication and Robust Cooperation for Multi-Vehicle Perception in Complex Environments). ERMVP is the baseline solution in the V2V4Real dataset.

[0094] Table 1 shows the performance comparison of 3D collaborative detection methods at different distances on the V2V4Real dataset. Experimental results show that on the V2V4Real dataset, our multi-vehicle collaborative fusion perception method based on spatiotemporal context achieves an overall performance of 66.1% and 45.2% in AP@0.5 and AP@0.7, respectively. Overall is the average of the AP@0.5 and AP@0.7 values ​​for all detected objects. Compared with No Fusion, our method improves AP@0.5 and AP@0.7 by 26.9% and 23.2%, respectively; compared with Late Fusion, our method improves by 11.1% and 18.5%, respectively. Notably, compared with ERMVP, a leading intermediate fusion solution, our method achieves improvements of 1.2% and 1.9% in AP@0.5 and AP@0.7, respectively.

[0095] Table 1 Performance comparison with baseline methods on the V2V4Real data test set

[0096]

Claims

1. A multi-vehicle collaborative fusion perception method based on spatiotemporal context, characterized by: The following steps are involved: S1. Obtain a large-scale open collaborative perception simulation dataset and a real dataset and mix them; S2. Input the metadata in the mixed dataset into the metadata sharing and feature extraction module to obtain the feature representation of each vehicle; S3. Input the feature representation of the ego vehicle and each cooperative vehicle into the corresponding neighborhood spatiotemporal LSTM module, and output two hidden states after spatiotemporal information refinement. S4. After position embedding, the hidden state is input into the spatial decoding and communication module, which outputs the characteristics of the ego vehicle, the characteristics of the cooperative vehicle, the characteristics of the cooperative vehicle in the shared space, and the characteristics of the ego vehicle in the shared space. S5. Align and fuse the features obtained in S4 through a spatial adaptive fusion mechanism, and output the fused features; S6. During the inference process, the fused features are input into the regression head and the classification head, and the bounding box prediction information and the classification prediction information are output. During the training process, the smooth L1 loss is used for regression and the focal loss is used for classification.

2. The multi-vehicle collaborative fusion perception method based on spatiotemporal context according to claim 1 is characterized in that: Step S2 is specifically implemented as follows: the metadata in the mixed dataset is input into the metadata sharing and feature extraction module; the ego vehicle constructs a collaborative graph and broadcasts the metadata; after receiving the metadata, the collaborative vehicle projects its point cloud data into the ego vehicle's coordinate system and aligns it with the ego vehicle's current LiDAR point cloud data; The point cloud of the ego vehicle and each cooperative vehicle is aligned and encoded into a bird’s-eye view feature through the point cloud encoder, outputting the point cloud of each vehicle at Feature representation of time , the feature dimension is ,in 、 and Represent the height, width and number of channels respectively, Indicates the index of the vehicle.

3. The multi-vehicle collaborative fusion perception method based on spatiotemporal context according to claim 2 is characterized in that: The specific implementation process of step S3 is as follows: S31. Input the hidden state of the previous moment and current moment features , concatenated in the channel dimension and then normalized; S32. The normalized result is input into the neighborhood attention module, which only calculates attention within the local window of each position, extracts local spatiotemporal context features, and obtains neighborhood attention; S33. Apply the Sigmoid activation function to the neighborhood attention element by element to obtain the gate weight ; S34. Multiply the gate weight by the cell state at the previous moment element by element; pass the neighborhood attention through the tanh activation function, and then Multiply element by element; finally, add the results of the element-by-element multiplication of the two parts to get the updated cell state; S35. Input the cell state, first perform tanh activation on it, and then combine the activation result with the gate weight Multiply element by element to obtain the hidden state after spatiotemporal information refinement , as a spatiotemporal feature.

4. The multi-vehicle collaborative fusion perception method based on spatiotemporal context according to claim 3 is characterized in that: The specific implementation process of step S4 is as follows: S41. Input the spatiotemporal characteristics of each vehicle , encode the perceptual space, and obtain features containing position information by adding a learnable position embedding to the spatiotemporal feature ; S42. Input features containing location information , apply average pooling in width and height dimensions respectively to obtain spatial features and ; S43. Spatial features and After aligning the dimensions through broadcasting operations, the concatenation is performed and processed through nonlinear transformations to obtain intermediate vectors. The spatial attention score vectors in the width and height dimensions are then calculated and a spatial confidence map is generated. After obtaining the spatial confidence map, the important positions are screened according to the set threshold through the screening function. If the confidence of a position is greater than the threshold, it is marked as 1, otherwise it is marked as 0; important spatial positions Point product contains the features of position information to obtain the corresponding important features ; S44. Collaborative vehicles will place important spatial positions and corresponding important features Send to the ego vehicle; after the ego vehicle receives the data, it calculates the characteristics of the ego vehicle’s perception area , cooperative vehicle perception area characteristics and features of the shared perceptual area 、 , specifically: and Respectively and corresponding and Perform element-wise multiplication to obtain and ; and Respectively Perform element-wise multiplication to obtain and , Represents element-wise multiplication.

5. The multi-vehicle collaborative fusion perception method based on spatiotemporal context according to claim 4 is characterized in that: The specific process of generating the spatial confidence map is as follows: the intermediate vector is subjected to wide-dimensional and high-dimensional segmentation operations respectively, and then subjected to nonlinear transformation and Sigmoid activation function successively, and the two results are multiplied to obtain the spatial confidence map.

6. The multi-vehicle collaborative fusion perception method based on spatiotemporal context according to claim 5 is characterized in that: The specific implementation process of step S5 is as follows: S51. The characteristics of the shared perception area of ​​the self-vehicle and the cooperative vehicle and Perform splicing, perform maximum pooling and average pooling on the spliced ​​tensor respectively, and splice the two pooling results along the channel direction to obtain a composite feature with both local strong response and global average response; The composite features are fused and reorganized in the spatial dimension through sliding convolution operations to output unified common perceptual area features. ; S52. For each position in the spatiotemporal features of the vehicle , extract the query vector by multiplying the spatiotemporal features by the weight matrix; delineate Neighborhood Matrix , then based on the neighborhood matrix , calculate the corresponding neighborhood matrix in the collaborative space through coordinate alignment operation , and 、 and Neighborhood matrix Multiply the corresponding learnable matrices respectively and , and obtain the corresponding three sets of key vectors and value vectors; then calculate the attention weights to obtain the unified feature representation of the ego vehicle perception area from the perspective of the ego vehicle, the unified feature representation of the cooperative vehicle perception area from the perspective of the ego vehicle, and the unified feature representation of the unified shared perception area from the perspective of the ego vehicle; S53. The three fused features are uniformly represented and fed into the same importance generator. The features at each pixel are convolved to generate a corresponding original weight distribution map. Next, a Sigmoid activation function is applied to the generated weight map to obtain a spatial importance map of the ego vehicle's perception area, a spatial importance map of the cooperative vehicle's perception area, and a unified spatial importance map of the shared perception area. S54. Concatenate the three spatial importance maps and use the softmax function to obtain a normalized spatial attention map between different features. Similarly, obtain the ego vehicle perception spatial attention map and the collaborative vehicle perception spatial attention map. S55. Combine the three spatial attention maps with the corresponding feature maps 、 and Multiply element-by-element in the spatial dimension to obtain three weighted feature maps; These three weighted feature maps are then directly added together in channels and spaces to obtain the final fusion features.

Citation Information

Patent Citations

  • BEVFormer and neighborhood attention Transform fused end-to-end automatic driving trajectory planning system, method and training method

    CN117516581A

  • New energy automobile automatic driving fusion perception method based on improved multi-head attention mechanism

    CN118627018A

  • Video instance segmentation method and apparatus based on spatio-temporal memory information

    WO2023116632A1

Cited By

  • Multi-agent cooperative sensing method and system for Internet of Vehicles

    CN121837878A

  • An internet of vehicles multi-agent cooperative perception method and system

    CN121837878B