Multi-vehicle cooperative sensing method based on attention mechanism
By adopting an attention-based multi-vehicle collaborative perception method, the problems of information redundancy and asynchrony in multi-vehicle collaborative perception are solved, achieving efficient feature transmission and fusion, and improving perception accuracy and robustness.
Patent Information
- Application Number
- CN202511076016.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies lack selective attention to information from different vehicles in multi-vehicle cooperative perception, resulting in high communication overhead, huge consumption of computing resources, and difficulty in handling timestamp inconsistencies caused by asynchronous perception, which affects perception accuracy and robustness.
A multi-vehicle collaborative perception method based on attention mechanism is adopted, which optimizes the feature transmission and fusion process by filtering redundant information between and within classes, selecting cooperative vehicles, calibrating sparse consensus foreground features, attention fusion, and aligning distillation learning time.
It significantly reduces unnecessary data transmission and computational overhead, improves the communication efficiency and scalability of multi-vehicle cooperative perception systems, and enhances the fusion effect and the perception accuracy and robustness of dynamic targets.
Smart Images

Figure CN120953944A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mobile communication technology and relates to a multi-vehicle collaborative perception method based on an attention mechanism. Background Technology
[0002] In recent years, connected and automated vehicles (CAVs) have been rapidly integrating into urban road networks and achieving large-scale deployment. Collaborative perception is gradually becoming a core driving force for improving environmental understanding and driving safety. Multi-vehicle collaboration effectively alleviates problems such as single-vehicle field-of-view occlusion and limited perception range by sharing perception data in real time. Traditional collaborative perception methods typically employ global feature stitching or weighted averaging fusion strategies to stitch or weight and integrate the perception features of each C-CAV along the channel dimension, effectively supporting improved object detection and perception performance in multi-vehicle collaborative scenarios. However, this method lacks selective attention to information from different vehicles and regions, suffers from high communication overhead, insufficient flexibility, and struggles to address the issue of inconsistent BEV feature timestamps caused by asynchronous C-CAV perception.
[0003] Traditional methods typically collect and transmit the perception features of all neighboring vehicles in a unified manner, lacking effective filtering of C-CAVs and redundant information. This results in huge consumption of communication bandwidth and computing resources, especially in densely populated areas. The system does not distinguish between high-confidence targets and invalid background information, and transmits and fuses a large number of irrelevant or redundant features, which limits the scalability of the cooperative perception system.
[0004] Furthermore, traditional feature fusion mainly relies on global feature splicing or weighted averaging fusion strategies, which fail to distinguish the differences in contributions from different vehicles and make it difficult to adjust the fusion weights based on the complementarity of perspectives between vehicles or the confidence of observations. For example, in occluded areas, features of distant vehicles with low confidence may be treated the same as clear observation features from nearby vehicles, weakening the fusion effect.
[0005] Meanwhile, traditional methods generally assume strict time synchronization of C-CAV, which lacks robustness when dealing with the time delay of asynchronous perception and information transmission among multiple vehicles. They ignore the differences in timestamps of data collection and uploading between different vehicles. Especially in high-speed or dynamic scenarios, this can easily cause problems such as inconsistent dynamic target positions and distorted motion trajectories, which seriously affect the spatiotemporal consistency and detection accuracy of collaborative perception.
[0006] To address the aforementioned issues, this invention designs a multi-vehicle collaborative perception method based on an attention mechanism. First, redundant information in the multi-vehicle perception features is filtered from both inter-class and intra-class perspectives to reduce unnecessary communication and computational overhead. Simultaneously, a collaborative vehicle selection model is constructed based on the intersection-union ratio (IUU) of bird's-eye view features, selecting key C-CAVs to participate in information sharing. Second, spatial calibration of the multi-vehicle perception features is completed by combining sparse consensus foreground features and geometric consistency verification to ensure accurate spatial alignment. Then, efficient feature fusion is achieved through an attention mechanism. Furthermore, based on a selective cross-attention mechanism, the spatial consistency and semantic expressiveness of the fusion result are improved through the positional accuracy of R-CAV features. Finally, a temporal alignment strategy based on distillation learning is employed to align dynamic targets, enhancing the system's adaptability to asynchronous perception and dynamic changes, ultimately improving the perception accuracy and robustness of multi-vehicle collaboration. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to provide a multi-vehicle collaborative perception method based on an attention mechanism to solve the problem of inefficient feature utilization caused by the lack of selective attention to different value perception information, as well as the problem of decreased dynamic target perception accuracy caused by communication delay and asynchronous feature alignment.
[0008] To achieve the above objectives, the present invention provides the following technical solution:
[0009] In a first aspect, based on the real-time and perception accuracy requirements of multi-vehicle cooperative perception scenarios, embodiments of the present invention employ a lightweight perception architecture that allows direct communication between the autonomous vehicle and cooperating vehicles to achieve high-precision multi-vehicle cooperative perception. This method includes the following steps:
[0010] S1: Feature filtering model based on inter-class and intra-class redundancy;
[0011] S2: Collaborative vehicle selection strategy based on feature intersection-union ratio;
[0012] S3: Spatial feature calibration method based on sparse consensus foreground;
[0013] S4: Feature fusion model based on attention mechanism;
[0014] S5: Dynamic target temporal alignment strategy based on distillation learning;
[0015] Secondly, in embodiment S1 of this invention, a feature selection and aggregation method based on modeling the relationship between redundant information between classes and within classes is designed. This method is jointly optimized from two aspects, with the design of FS and MS modules respectively. In FS, a confidence generation network is constructed to generate a confidence mapping in the feature space based on the class confidence output by the target detection head, used to characterize the importance of features at different locations. Then, foreground and background are distinguished according to the confidence threshold, retaining only sparse foreground features with high confidence as subsequent shared information, reducing inter-class redundancy to avoid transmitting large areas of irrelevant background information. In MS, addressing local redundancy in the foreground region, a similarity-based feature grouping method is adopted. By calculating spatial distance and semantic similarity, spatially adjacent and class-similar foreground features are automatically aggregated into a group. Then, attention-weighted features within the group are fused, retaining only the most representative and efficient fusion results, thereby reducing intra-class redundancy. The two modules work together to effectively improve the communication efficiency and feature expression quality of multi-vehicle collaborative perception while ensuring the complete expression of key target information.
[0016] Thirdly, in embodiment S2 of this invention, a collaborative vehicle screening method based on the intersection-union ratio of BEV features between vehicles is designed. The R-CAV receives perception features from the C-CAV and, within its own constructed BEV feature space, calculates the spatial intersection-union ratio between these external perception results and its own perception results. The system quantitatively evaluates the overlap and importance of the information provided by each C-CAV. To further consider the total coverage area of the perception field of view, the coverage area of each C-CAV's perception region is introduced as a supplementary indicator to reflect the perception field of view completion capability brought by the C-CAV. The above two indicators are jointly embedded into a utility function, forming a multi-objective screening standard oriented towards accuracy and coverage area. Based on this, C-CAVs with advantages in both information value and spatial coverage are selected through iterative loops, ensuring that the accuracy and sparsity of the fusion result achieve an optimal balance.
[0017] Fourthly, in embodiment S3 of this invention, a feature space calibration model based on sparse consensus foreground features and geometric consistency verification is constructed to address the cross-vehicle feature misalignment problem caused by positioning errors in multi-vehicle collaboration. This model first selects foreground regions with high confidence and sparse spatiality from the perception data of each vehicle as potential alignment anchors. By constructing a weighted bipartite graph and defining a matching cost function combining spatial distance and semantic similarity, the model solves for the matching result with the minimum total cost to establish a preliminary feature correspondence. Subsequently, a random sampling consensus algorithm is used to verify the geometric consistency of the matching results, eliminating pseudo-matching pairs with abnormal positions. Finally, an error adjustment mechanism is introduced. By comparing the consistency of local regions before and after alignment, the final calibration result is adaptively adjusted to improve the spatial consistency of features under conditions of positioning deviation and incomplete observation, providing a robust spatial foundation for subsequent collaborative perception.
[0018] Fifthly, in embodiment S4 of this invention, a feature fusion method based on a local-global attention mechanism combined with a precision-enhancing cross-attention mechanism is designed. During the fusion stage, local and global attention mechanisms are introduced for alternating modeling. The local attention module focuses on detailed features at spatially adjacent locations to capture local detail information, while the global attention module models the global context of the entire feature to obtain global semantic relationships, jointly enhancing the multi-scale expressive power of the fused features. In the precision-enhancing feature interaction stage, the BEV features of R-CAV are used as the query, and the fused multi-vehicle features are used as the key. Selective feature interaction is performed through the cross-attention mechanism, using the spatial precision of the R-CAV features as a spatial reference benchmark to guide the adaptive adjustment of the fused features in spatial position. Ultimately, the overall fusion result possesses both fine-grained local details and rich global semantic information, while maintaining accurate spatial alignment.
[0019] Sixthly, in embodiment S5 of this invention, a temporal alignment model based on distillation learning and combined with scene flow technology is constructed. The system estimates the motion offset of C-CAV features at the current moment by constructing a short-time scene flow, and uses this to guide its historical features back to the target frame. The teacher model obtains high-quality BEV features with spatiotemporal consistency through real trajectory optimization, and uses this as a supervision signal to guide the student model to align and refine its features, thereby improving temporal consistency and dynamic robustness on the basis of spatial registration, effectively enhancing fusion accuracy and system stability.
[0020] The beneficial effects of this invention are as follows:
[0021] This invention significantly reduces unnecessary data transmission and computational overhead by filtering redundant information between and within classes of the perception features of collaborative vehicles and selecting key collaborative vehicles to participate in information sharing. This effectively reduces the consumption of communication bandwidth and computing resources, and improves the communication efficiency and scalability of the multi-vehicle collaborative perception system.
[0022] This invention employs a feature fusion method based on an attention mechanism, which can distinguish the differences in the value and contribution of perception information from different vehicles, avoiding the problem of treating low-confidence features and clear observation features equally. Through selective attention and high-precision feature guidance, it improves the spatial consistency and semantic expressive ability of the fused features, thereby enhancing the fusion effect.
[0023] This invention constructs a dynamic target time alignment strategy based on distillation learning, which solves the problem that traditional methods are difficult to handle asynchronous perception and information delay due to the assumption of strict time synchronization. It effectively addresses the position offset and trajectory distortion of dynamic targets caused by inconsistent timestamps, enhances the system's adaptability to dynamic changes, and improves the perception accuracy of dynamic targets and the system's robustness.
[0024] This invention introduces a spatial feature calibration model and uses sparse consensus foreground features and geometric consistency verification to solve the feature misalignment problem caused by positioning errors in multi-vehicle collaboration. This provides a robust and accurately aligned spatial foundation for subsequent collaborative perception tasks, further ensuring perception accuracy.
[0025] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0026] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:
[0027] Figure 1 A lightweight perception architecture diagram for direct communication between autonomous vehicles and collaborating vehicles;
[0028] Figure 2 This is a schematic diagram of a feature fusion model based on an attention mechanism.
[0029] Figure 3 This is a flowchart illustrating the execution of a multi-vehicle collaborative perception scheme based on an attention mechanism. Detailed Implementation
[0030] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0031] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0032] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0033] Figure 1 A possible structural schematic diagram of a communication system according to an embodiment of the present invention is shown. For example... Figure 1 As shown, this architecture divides vehicles into R-CAVs and C-CAVs. C-CAVs, acting as auxiliary nodes, are composed of cooperating CAVs surrounding the R-CAVs. They provide supplementary information to the R-CAVs, filling in their perception blind spots and improving overall perception performance. They primarily perceive the environment through their own perception system and generate point clouds. These point clouds are then projected onto coordinates based on metadata broadcast by the R-CAVs and encoded using a shared feature encoding model to extract BEV features and metadata. Furthermore, the features are filtered and aggregated from both inter-class and intra-class perspectives to obtain representative feature information. Finally, the information is uploaded through direct communication with the R-CAVs.
[0034] As the core node of the system, R-CAV is the initiator of the entire task. Besides environmental perception, after receiving the perception information uploaded by C-CAV, it needs to filter C-CAVs based on the intersection-over-union ratio (IoU) and perception coverage area. After determining the C-CAVs, their features are spatially calibrated, and then fused and refined using an attention mechanism. Then, temporal alignment is achieved through a method based on distillation learning combined with scene flow prediction, improving the perception accuracy of dynamic targets. Finally, the fused and calibrated features are output as the overall perception result via the detection head.
[0035] 1. Feature filtering model based on inter-class and intra-class redundancy
[0036] In multi-vehicle cooperative perception, the perception data collected by C-CAV contains a large amount of redundant information, including two aspects: inter-class redundancy and intra-class redundancy. Inter-class redundancy refers to the redundancy between foreground and background features. In perception tasks, background areas such as roads and skies not only contain a large amount of information but also have very low value for the perception task and do not need to be uploaded to R-CAV. Intra-class redundancy refers to the same target being described repeatedly from different perspectives, resulting in information redundancy. This redundancy not only increases communication overhead but also increases the computational burden of R-CAV in the feature fusion stage. Especially in dynamic environments, vehicle localization errors and low information aggregation efficiency further exacerbate these problems.
[0037] Therefore, in order to reduce communication redundancy, reduce communication overhead, and improve the overall performance of the model, the FS module and MS module are used in the C-CAV terminal to handle inter-class redundancy and intra-class redundancy, respectively.
[0038] In FS, a confidence map is first generated for each feature using a confidence mapping, reflecting the perceptual importance of different spatial regions. Higher confidence indicates potential object regions, while lower confidence indicates redundant background regions. For the features F of the cooperative vehicle i... i Its two-dimensional confidence plot C i Defined as:
[0039] C i =φ conf (F i )∈[0,1] H×W (1)
[0040] Where φ conf This represents a confidence generation network with a detector-decoder structure; H×W is the confidence map C. i The dimensions are represented by their height and width, respectively.
[0041] Next, set the confidence threshold θ conf For values ∈(0,1), confidence thresholding is performed. The specific operation is shown in the following formula:
[0042]
[0043] Where B(h,w)∈{0,1} is a binary mask with dimension H×W, representing whether each position is preserved. B(h,w) is further optimized using Non-Maximum Suppression (NMS) to ensure the sparsity and independence of features.
[0044] Finally, the original feature F is analyzed based on the binary mask B. i Perform element-wise multiplication to retain high-confidence foreground features and filter out low-confidence background features. This can be expressed as Fi s =B⊙F i
[0045] To cope with dynamically changing environmental conditions and fully ensure the robustness of the system, the filtering threshold can be dynamically changed according to sensor data and network status.
[0046] In MS, the redundancy between the filtered forward features is optimized and aggregated. This involves the following three steps:
[0047] 1) Feature grouping based on similarity: First, the nearest neighbor clustering algorithm is applied to initially group the feature vector set. For a given set of feature vectors F i s =[f1,f2,...,f N ] T and cluster center F c The similarity between each feature vector and all cluster centers is calculated, and the most similar cluster centers are selected. Similarity is measured by a combination of feature distance and positional distance. The mathematical expression for the calculation is:
[0048]
[0049] Where δ i Represents the eigenvector f i The corresponding clustering; p(·) represents the function to obtain the position of the feature vector; λ is a hyperparameter. This divides all feature vectors into K clusters, defined by G = {G1, G2, ..., G...}. K}express.
[0050] The total number of clusters is adjusted by the dynamic clustering ratio β, i.e., K = β × N.
[0051] 2) Attention-Inspired Feature Merging: Within each cluster, an attention mechanism is used to weight and merge feature vectors to reduce the impact of outliers. The weight of each feature vector is determined by its confidence score. The specific formula for merging feature vectors is as follows:
[0052]
[0053] in For the k-th cluster G k The p-th eigenvector in This represents the confidence score corresponding to the feature vector.
[0054] This yields the final feature vector.
[0055] 3) Index-based Feature Reconstruction: In the feature grouping and merging stage, each feature vector is assigned to a specific cluster, and each cluster is represented by a merged feature vector. During this process, an index mapping record is constructed to accurately reflect the correspondence between the original feature vectors and the merged feature vectors. With this index mapping, R-CAV can accurately remap the merged feature vectors to their corresponding positions in the original features, thus efficiently completing the feature reconstruction.
[0056] This index-based feature reconstruction mechanism not only ensures the precise consistency of features in the spatial dimension, but also provides high-quality and consistent feature inputs for subsequent perception tasks, thereby effectively guaranteeing the performance and reliability of the entire multi-vehicle cooperative perception system.
[0057] 2. Collaborative vehicle selection strategy based on feature intersection-union ratio
[0058] With a large number of candidate C-CAVs and varying values of perception information among them, using all of them without screening would lead to excessive communication overhead and reduced fusion performance. Therefore, a C-CAV screening strategy is adopted to improve the complementarity of perception information while ensuring coverage. The C-CAV screening strategy is executed locally on the R-CAV, comprehensively considering the cross-border ratio of candidate vehicle BEV views and the coverage gain of the perception area, maximizing the overall value through a multi-objective utility function trade-off.
[0059] A higher intersection-over-union (IoU) ratio indicates a greater spatial overlap between the perception results of the two vehicles, meaning their perceptions of the surrounding environment are more consistent and provide more reliable fusion data. However, relying solely on the IoU ratio may lead to excessive overlap; therefore, it is also necessary to consider the coverage area of the perception range to comprehensively optimize data fusion efficiency. The specific utility function is as follows:
[0060]
[0061] Where, r i Indicates the communication transmission rate of the cooperating vehicle i; s i ∈{0,1} represents the decision variable for whether to establish a cooperative connection; ω1 and ω2 are weight coefficients; J i The cross-union ratio (C-CAVi) fusion value score between C-CAVi and R-CAV is expressed as follows:
[0062]
[0063] Wherein, BEV0 represents the BEV characteristic of R-CAV; BEV i Represents the BEV characteristics of the cooperative vehicle i; [·] π This indicates the extraction of the shared target region; This represents the total area of the sensing region ultimately obtained through collaborative sensing; V = {v i |s i =1} is the set of cooperative vehicles that establish communication connections with R-CAV; Y i A is the sensing region of the cooperative vehicle i itself; A(·) is a function for calculating the area of the joint sensing region.
[0064] In the iterative process of finding the maximum utility function, the connection matrix between vehicles is first fixed, and the optimal transmission rate that satisfies the constraints is found using a standard linear programming solver. Then, the transmission rate is fixed again, and the optimal link establishment problem is solved by progressively selecting the links that contribute the most to the utility function. The priority weights are then updated according to the new allocation. Simultaneously, to ensure the stability of the algorithm and avoid large jumps that could cause the loop to get trapped in local optima, when the change in priority weights exceeds a set threshold, initialization and recalculation are performed.
[0065] This allows for the selection of the optimal C-CAV group, improving the efficiency and spatial completeness of multi-vehicle feature fusion from the source.
[0066] 3. Spatial Feature Calibration Method Based on Sparse Consensus Foreground
[0067] In multi-vehicle collaboration, sensor calibration errors and local positioning deviations exist between different vehicles. Direct fusion can lead to feature ambiguity and boundary misalignment. To address this, a spatial feature calibration model is introduced. By extracting sparse consensus foreground points and performing geometric consistency verification, an accurate spatial transformation matrix is estimated to achieve cross-vehicle feature alignment.
[0068] First, the BEV features F, after filtering and aggregating vehicle i and vehicle j, are analyzed. i agg and High-confidence foreground points are extracted, and a sparse consensus point pair set is constructed based on feature similarity. The specific definition is as follows:
[0069]
[0070] in, and These represent the nth foreground feature points of vehicle i, respectively. The corresponding feature vector, and the m-th foreground feature point of vehicle j. The corresponding feature vector; sim(·,·) represents the similarity function; θ sim The similarity threshold controls the semantic consistency of consensus point pairs.
[0071] Then, in the point-to-point set Q ij Construct a weighted bipartite graph M = (U) based on this. i Uj E), where the set of points is... The edge weights of edge set E are feature similarities:
[0072]
[0073] Then, the maximum weight matching is solved using the Hungarian algorithm, resulting in a set of highly reliable matching point pairs. Used for subsequent geometric consistency verification.
[0074] To further enhance the robustness of spatial calibration and suppress interference from noise and outlier pairs, a random sample consensus algorithm is used to perform geometric verification of the matching points, thereby estimating the optimal spatial transformation matrix M of vehicle j relative to vehicle i. ij .
[0075] Firstly, from Randomly select a subset Singular value analysis is used to estimate a transformation matrix that minimizes residuals. This transformation matrix is then used to calculate the error for all matching point pairs, and points with errors less than the error tolerance threshold ε are considered "interiors." Through repeated calculations, the transformation matrix with the most "interiors" is finally selected. The mathematical expression is:
[0076]
[0077] Where M represents the rigid transformation matrix; Γ[·] is an indicator function used to determine whether the condition of "interior point" is satisfied.
[0078] The advantage of the random sample consensus algorithm lies in reducing sensitivity to a small number of outliers, thereby mitigating the impact of noise on model performance. It can stably fit the global geometric transformation relationship and obtain the optimal spatial transformation matrix M. ij .
[0079] The spatial transformation matrix obtained by the above method is applied to the BEV features of C-CAV to align them to the coordinate system of R-CAV, thus obtaining aligned features and ensuring semantic matching and spatial consistency.
[0080] 4. Feature fusion model based on attention mechanism
[0081] Traditional attention mechanisms in multi-vehicle collaborative perception often employ a unified attention modeling approach. However, this approach has certain limitations in real-world scenarios: while global attention mechanisms possess the ability to model long-distance dependencies, they are susceptible to error accumulation and noise interference, leading to a decrease in the accuracy of local spatial position perception; and while local attention mechanisms can accurately model spatial proximity relationships, they struggle to integrate semantic contextual information across vehicles or large-scale scenes.
[0082] like Figure 2 As shown, in order to balance perception accuracy and semantic integrity, this system deploys a feature fusion module with a local-global alternating attention mechanism on the R-CAV side. It fuses the features of multiple vehicles from two dimensions: local detail capture and global semantic aggregation, thereby improving the spatial consistency of feature representation and scene understanding capabilities.
[0083] First, the features of BEVs from N C-CAVs are stacked into a four-dimensional tensor: H, W, and Z represent the height, width, and number of channels of the feature, respectively. Next, the tensor is divided into multiple non-overlapping local windows, and local self-attention is applied within each window. This mechanism explicitly enhances local target information with spatial adjacency and semantic similarity by calculating the point-to-point correlation between features within the window, effectively preserving fine-grained features such as target boundaries, local contours, and occlusion contour changes, thereby improving the accuracy of local perception.
[0084] On a larger scale, to avoid the problems of isolated perceptual information and scene fragmentation between local windows and to enhance global semantic consistency, a shared global query vector g is introduced, which can be obtained by applying the following formula to the feature tensor:
[0085] g(x)=P(C 1×1 (SE(G(D(x)))+x)) (10)
[0086] Among them, SE(GE(D(x))) from the inside out are: depthwise separable convolution (extracting long-range spatial features), GELU activation function, and channel attention mechanism; C 1×1 (·) represents a dimension-reducing convolution, and P(·) is a function that performs max pooling.
[0087] Subsequently, cross-attention is calculated between the global query vector g and the key and value val in the local window. The specific calculation formula is as follows:
[0088]
[0089] Where d key The dimension of the key; key T This is the transpose of the key. This step is a globally semantically guided multi-vehicle feature aggregation process. It guides information exchange between key regions in the guided window through a global query, thereby establishing cross-view semantic consensus and reducing semantic drift caused by observation offsets between vehicles. This is then combined with the classic Transformer architecture, including normalization layers, multilayer perceptrons, and residual connections, to achieve effective integration of multi-vehicle features.
[0090] To further improve the spatial accuracy and robustness of the perceived output, a precision-enhanced feature interaction was introduced. Through a selective cross-attention mechanism, the spatial credibility of R-CAV features was fully utilized to guide the fused features for local reconstruction and correction.
[0091] Specifically, the BEV features of R-CAV are used as the query source, and the fused features are used as the context source. The cross-attention mechanism is activated only in high-confidence regions. This strategy has two key advantages: First, R-CAV features have higher local positioning accuracy and can be used as a spatial reference to guide the adaptive adjustment of the fused feature position, improving the structural consistency of multi-vehicle information for the same target; second, by selectively applying attention operations, interaction is only performed on perceived salient target areas, avoiding the introduction of noise from redundant backgrounds and reducing the consumption of ineffective computing resources.
[0092] 5. Dynamic target time alignment strategy based on distillation learning
[0093] Due to practical issues such as asynchronous perception by collaborative vehicles and communication latency between vehicles, the perception features transmitted from C-CAV to R-CAV often have timestamps earlier than the current frame of the R-CAV. If these historical frame features are fused into the current frame without processing, it can easily lead to feature misalignment and dynamic target offset, significantly affecting the fusion effect and target localization accuracy of the perception system in complex dynamic environments. Therefore, a dynamic target time alignment strategy based on distillation learning is introduced to achieve accurate alignment of historical features in the time dimension.
[0094] For the perception features F uploaded by collaborative vehicle i at timestamp t i t The core task of this method is to learn an alignment function and map it to the R-CAV time frame t0, thereby overcoming the effects of time synchronization errors and latency. The expression is as follows:
[0095]
[0096] Where Δt = t0 - t represents the historical delay length. During the training phase, the model uses real features as a reference target to perform semantic correction on the above output, but during the inference phase, it only relies on the historical input and Δt to achieve unlabeled mapping.
[0097] During training, a teacher-student distillation learning framework was employed to train the alignment function. The teacher model uses real trajectories to generate accurately aligned BEV features. The features generated by the teacher model are accurate and will serve as supervisory signals to guide the training of the student model. The specific steps are as follows:
[0098] 1) The fused BEV features are processed using bilinear interpolation. The (x,y) coordinates of each point are interpolated within the BEV features to extract the features of that point. This process maps the spatial coordinates of each point to the corresponding position in the BEV features, thereby ensuring that the data of each point can be unified to the same image format.
[0099] 2) Generate 3D bounding boxes using an object detection head. By processing interpolated point features in the BEV features, the object detection head network predicts the 3D position, size, and orientation of each object, including the object's center position, size, and orientation.
[0100] 3) Scene flow is a vector field describing the motion of points over time. In LiDAR point cloud processing, scene flow refers to the displacement of each point from one time step to another, typically used to compensate for point cloud misalignment caused by dynamic objects or R-CAV motion. To estimate scene flow, it is necessary to understand the motion of each point between different time steps. The point feature processing module analyzes these point features and predicts the scene flow for each point based on a deep neural network. Then, the scene flow is used to correct the detection boxes, eliminating positional offsets during the perception of dynamic objects.
[0101] 4) The corrected data is remapped into BEV features using a point cloud mapping method. At this point, the influence of dynamic objects has been eliminated, but the corrected point cloud is a relatively sparse BEV feature that can represent the correct structure of the environment.
[0102] 5) Next, the corrected sparse BEV features are fused with the original BEV features. While the original BEV features have higher density and more detailed information, they are affected by the movement of dynamic targets. The corrected BEV features, although more accurate, are relatively sparse. The fused BEV features combine the advantages of both, retaining high-density features while reducing interference from movement.
[0103] The BEV features obtained after fusing the student models will be used as the learning target, mimicking the teacher's parameter adjustments. Specifically, the difference between the student and teacher models will be quantified using distillation loss, as shown in the following expression:
[0104]
[0105] The loss is calculated as the square of the Euclidean distance between the student model output and the teacher model output, which is the sum of squared point-to-point differences between them.
[0106] 6. System Flowchart
[0107] Figure 3 The diagram shows the multi-vehicle collaborative perception execution process based on the attention mechanism. The specific steps are as follows:
[0108] S601-S602: The R-CAV initiates a collaboration request by broadcasting a small amount of metadata to nearby CAVs willing to participate in collaborative perception. Upon receiving the information, each C-CAV independently perceives its environment based on its own perception system, generating its own point cloud set. S603-S604: Based on the broadcast metadata, each C-CAV projects its point cloud set at its own coordinates onto the coordinates of the R-CAV. Simultaneously, features are extracted from the point cloud sets using a shared BEV feature encoding model.
[0109] S605-S607: C-CAV filters and aggregates extracted features from both inter-class and intra-class redundancy perspectives to obtain representative, high-quality BEV features, which are then compressed and uploaded to R-CAV. R-CAV then receives the information, decodes it based on the index information, and further filters C-CAV using a priority evaluation mechanism based on intersection-union ratio (IU) to obtain high-quality C-CAVs for subsequent collaborative perception tasks.
[0110] S608: Spatial calibration of uploaded BEV features is performed using a spatial feature calibration method based on sparse consensus foreground and geometric consistency verification, thereby aligning static targets.
[0111] S609-610: Feature fusion is achieved using attention mechanisms from both local and global perspectives. The local attention mechanism focuses on capturing spatial details, while the global attention mechanism effectively extracts global contextual information. Next, a selective cross-attention mechanism is employed to interact the R-CAV features with the fused multi-vehicle features, thereby enhancing the spatial consistency of semantic expression based on the high positional accuracy of the R-CAV features.
[0112] S611: Short-term motion offsets of C-CAV are estimated through scene flow and predicted to the current time step. Based on a teacher-student distillation mechanism, high-quality teacher BEV features are used as the supervision target to guide the student model to optimize output features. Ultimately, the fused BEV features are ensured to improve temporal consistency and dynamic scene perception accuracy while maintaining spatial alignment.
[0113] S612: If the R-CAV receives broadcast metadata (i.e., cooperation request) from other vehicles during the process of completing its own perception task, it will convert the obtained perception results into a point cloud and return to S603, so that the C-CAV can continue to perform subsequent steps; otherwise, it will end and obtain the final perception result.
[0114] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0115] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network-side device, etc.) to execute the cell handover method described in the various embodiments of the present invention.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A multi-vehicle cooperative perception method based on an attention mechanism, applied to a cooperative perception system consisting of a requesting-connected and automated vehicle (R-CAV) and at least one cooperating-connected and automated vehicle (C-CAV), characterized in that: The method includes: The at least one C-CAV filters and aggregates its own perceptual features to generate features to be shared; The R-CAV receives the features to be shared from the at least one C-CAV and selects the target C-CAV according to a preset strategy; The R-CAV performs spatial feature calibration on the features to be shared from the target C-CAV to achieve spatial alignment of multi-vehicle features; The R-CAV is based on an attention mechanism to fuse features after spatial feature calibration to generate fused features; The R-CAV performs dynamic target temporal alignment on the fused features to obtain the final perception result.
2. The multi-vehicle cooperative perception method based on attention mechanism according to claim 1, characterized in that: The steps for C-CAV to filter and aggregate its respective perceptual features include: Feature selection is based on inter-class redundancy, and foreground and background are distinguished by confidence mapping to retain foreground features with high confidence; and, The foreground features are aggregated based on intra-class redundancy, and the foreground features are grouped by calculating spatial distance and semantic similarity, and the features within the group are weighted and fused.
3. The multi-vehicle cooperative perception method based on attention mechanism according to claim 1, characterized in that: The steps for R-CAV to select target C-CAV according to a preset strategy include: Construct a utility function that incorporates the intersection-union ratio of features from the Bird's Eye View (BEV) and the coverage area of the perceived region; and, The utility function is solved iteratively to select the target C-CAV that is optimal in terms of information value and spatial coverage.
4. The multi-vehicle cooperative perception method based on attention mechanism according to claim 1, characterized in that: The step of performing spatial feature calibration on the features to be shared from the target C-CAV by the R-CAV includes: High-confidence foresight points are extracted from the features of the R-CAV and the target C-CAV to construct consensus point pairs; and, The consensus point pairs are geometrically validated using a random sampling consensus algorithm to estimate the spatial transformation matrix and align the features to be shared.
5. The multi-vehicle cooperative perception method based on attention mechanism according to claim 1, characterized in that: The steps of fusing spatially calibrated features using the R-CAV based on an attention mechanism include: The features are modeled alternately by local attention and global attention mechanisms to capture local details and obtain global semantic relationships, thereby enhancing the multi-scale expressive power of the fused features.
6. The multi-vehicle cooperative perception method based on attention mechanism according to claim 5, characterized in that: The step of fusing spatially calibrated features based on the attention mechanism in the R-CAV further includes: Using the bird's-eye view BEV features of the R-CAV as the query and the fused multi-vehicle features as the key, the fused features are adaptively adjusted by using the spatial accuracy of the R-CAV features through a selective cross-attention mechanism.
7. The multi-vehicle cooperative perception method based on attention mechanism according to claim 1, characterized in that: The steps of the R-CAV for dynamically aligning the fused features to a target time include: A time alignment model based on distillation learning is constructed, in which the teacher model uses spatiotemporally consistent high-quality bird's-eye view BEV features generated from real trajectories as a supervision signal to guide the student model to align and refine the fused features.
8. The multi-vehicle cooperative perception method based on attention mechanism according to claim 7, characterized in that: The time alignment model also incorporates scene flow technology to estimate the motion offset of the target C-CAV features at the current moment, thereby assisting the student model in backtracking its historical features to the target frame.
9. A multi-vehicle cooperative perception system based on an attention mechanism, characterized in that: It includes at least one cooperating-connected and automated vehicle (C-CAV) and one requesting-connected and automated vehicle (R-CAV); The C-CAV includes a feature filtering and aggregation module, used to filter and aggregate the perceptual features of each of the C-CAVs to generate features to be shared; The R-CAV includes: A receiving module is configured to receive the feature to be shared from the at least one C-CAV; A collaborative vehicle screening module is used to screen out target C-CAVs from the at least one C-CAV according to a preset strategy; A spatial feature calibration module is used to perform spatial feature calibration on the shared features from the target C-CAV to achieve spatial alignment of multi-vehicle features; The feature fusion module is used to fuse spatially calibrated features based on an attention mechanism to generate fused features; and, The time alignment module is used to dynamically align the fused features to the target time to obtain the final perception result.
10. The multi-vehicle cooperative perception system based on an attention mechanism according to claim 9, characterized in that: The feature fusion module is specifically used for: By alternating between local and global attention modeling, the multi-scale expressive power of fused features is enhanced; and, Using the bird's-eye view BEV features of the R-CAV itself as the query, the features after multi-vehicle fusion are adaptively adjusted through a selective cross-attention mechanism to improve the spatial consistency and semantic expressiveness of the fusion results.
Citation Information
Cited By
Delay-constrained Internet of Vehicles multi-mode cooperative sensing method and system
CN121392543A
A time delay constrained vehicle networking multi-modal cooperative perception method and system
CN121392543B