Intelligent network-connected vehicle cooperative sensing system and method based on Mamba2 model
Through the intelligent connected vehicle collaborative perception system based on the Mamba2 model, the multi-head state space model and linear attention mechanism are used to solve the problems of high computing complexity and poor robustness in the existing technology, and low-latency and efficient multi-vehicle collaborative perception is achieved, and the target detection accuracy and environmental adaptability are improved.
Patent Information
- Application Number
- CN202510333780.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
The existing collaborative perception model is difficult to balance between computational complexity and real-time perception performance, especially the method based on the Transformer architecture, whose computational complexity is O(N2). As the number of vehicles participating in collaborative perception increases, the computing burden and delay are significantly improved, and the robustness is poor, making it difficult to achieve efficient and low-latency feature fusion in a multi-vehicle collaborative environment.
The intelligent connected vehicle collaborative perception system based on the Mamba2 model is adopted. Through the multi-head state space model combined with the bidirectional feature scanning strategy, a linear attention and channel-space fusion mechanism is introduced to achieve the enhancement and fusion of coordinated perception characteristics within the vehicle, including preprocessing modules, bidirectional state space modules, Mamba2 linear attention modules and channel-space fusion modules, reducing the computational complexity and improving perception performance.
It effectively reduces the computational complexity, improves perception performance, enhances the perception ability of multiple vehicles in complex environments, realizes low-latency feature fusion and efficient information processing, and significantly improves the target detection accuracy and perception adaptability.
Smart Images

Figure CN120258043A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cooperative perception of intelligent connected vehicles, and in particular to a cooperative perception method for intelligent connected vehicles for autonomous driving. Background Art
[0002] In recent years, cooperative perception has gradually become a disruptive approach in the field of autonomous driving, with the goal of enhancing the perception capabilities of individual vehicles by integrating the perception information of multiple agents. Cooperative perception leverages collective intelligence to overcome the inherent limitations of single-vehicle perception, such as limited field of view, occlusion problems, and challenges in long-range perception. Many recent studies have demonstrated the strong performance of cooperative perception. Current research classifies the feature map sharing methods between collaborators into three data fusion methods: Early Fusion: The raw sensor data is processed locally and then transmitted to other vehicles for further joint processing. Although this method has advantages in data integrity, it has extremely high requirements for communication bandwidth. Intermediate Fusion: The perception data is partially processed before data transmission, reducing the amount of transmitted data while retaining key perception information, and it is the fastest-growing fusion method in recent years. Late Fusion: The final results are shared after each vehicle has completed its perception processing, effectively reducing the communication load, but due to the lack of raw data sharing, its perception accuracy is relatively low.
[0003] Mamba2 is a model based on the Structured State Space Duality (SSD) framework, aiming to improve training efficiency and performance. With the rapid development of state space models, Mamba has attracted much attention due to its excellent performance in long sequence modeling and linear complexity. Based on the improvement of S4, Mamba introduces a selective mechanism to enhance the attention to input data, and improves selectivity through hardware-aware state expansion and parallel algorithms, replacing convolution with a scan operation to achieve parallel processing. Although numerous studies have explored its applications in the field of vision, the potential of the SSM model in cooperative perception has not been fully exploited. Different from simple two-dimensional image processing, cooperative perception requires fusing real-time motion features within an agent while handling the dynamic interaction features between multiple agents. How to achieve efficient and low-latency feature fusion while maintaining the low computational cost of Mamba remains an important challenge.
[0004] The present invention urgently needs to solve the following technical problems: The current cooperative perception models, especially the methods based on the Transformer architecture, have a computational complexity of O(N 2) As the number of vehicles participating in cooperative perception increases, the computational burden and latency increase significantly. This complexity severely limits the real-time performance of the model in a multi-vehicle cooperative environment. Existing methods are less robust in the face of noise in the actual environment (such as position noise and orientation noise). As the number of cooperative vehicles increases, the parameter scale and computational complexity of many models grow exponentially, restricting their scalability and practical deployment. Summary of the Invention
[0005] Aiming at the problem that it is difficult to balance the high computational complexity and real-time perception performance of existing cooperative perception systems, the present invention proposes an intelligent connected vehicle cooperative perception system and method based on the Mamba2 model. Through feature space enhancement of the internal cooperative perception features of vehicles by combining a multi-head state space model with a bidirectional feature scanning strategy, based on the Mamba2 architecture and the integration of multi-head attention of the internal cooperative perception features of vehicles, the internal cooperative perception features of the channel dimension of the internal cooperative perception features integrated by multi-head attention are spliced and fused in the spatial dimension to obtain the internal cooperative perception features of channel-space fusion, realizing the cooperative perception of intelligent connected vehicles.
[0006] The present invention is realized by the following technical solutions:
[0007] In a first aspect, the present invention proposes an intelligent connected vehicle cooperative perception system based on the Mamba2 model. The system includes a preprocessing module, a bidirectional state space module, a Mamba2 linear attention module, and a channel-space fusion module: wherein:
[0008] The preprocessing module fuses the ego-vehicle BEV feature and the other-vehicle BEV feature to obtain a BEV fusion feature, compresses the BEV fusion feature, and outputs the compressed BEV fusion feature as the internal cooperative perception feature of the vehicle;
[0009] The bidirectional state space module performs a linear transformation on the received compressed BEV fusion feature as the internal cooperative perception feature of the vehicle through a multi-head state space model combined with a bidirectional feature scanning strategy to obtain the core part A, respectively representing the processed feature tensor and the forget gate, splitting the core feature tensor into three vectors: a state vector matrix, an input matrix, and an output matrix, modeling to obtain a forward feature sequence and a backward feature sequence, and realizing the splicing operation and fusion of the forward feature sequence and the backward feature sequence through a state space dual module; sharing the forget gate and the Mamba2 linear attention module, fusing the forward sequence feature of the bidirectional BEV scanning feature and the backward sequence feature of the bidirectional BEV scanning feature, realizing the enhancement and fusion of the internal cooperative perception feature of the vehicle, and outputting the spatially enhanced internal cooperative perception feature;
[0010] The Mamba2 linear attention module is designed based on the Mamba2 architecture, introducing a linear attention mechanism to calculate the linear attention of the received spatially enhanced in-vehicle collaborative perception features, performing multi-head linear attention integration, dynamically screening key information using the forget gate mechanism shared with the bidirectional state space module, selectively enhancing and globally aggregating the spatially enhanced in-vehicle collaborative perception features according to the linear attention, and outputting the in-vehicle collaborative perception features with multi-head attention integration;
[0011] The channel-spatial fusion module introduces a channel attention mechanism, obtains the in-vehicle collaborative perception features in the channel dimension for the received in-vehicle collaborative perception features with multi-head attention integration through the channel attention mechanism, and performs spatial dimension splicing and fusion on the downsampled in-vehicle collaborative perception features in the channel dimension, outputting the in-vehicle collaborative perception features with channel-spatial fusion.
[0012] In some embodiments, the bidirectional state space module further includes:
[0013] Fusing the forward sequence features and the backward sequence features of the bidirectional BEV scan features to enhance and fuse the in-vehicle collaborative perception features;
[0014] Decomposing and integrating the in-vehicle perception features through a result splicing operation using the multi-head attention mechanism.
[0015] In some embodiments, the Mamba2 linear attention further includes:
[0016] (i) For the input features To uniformly model the cross-vehicle features, first reorganize the features in the channel dimension, merge the feature data of all vehicles, and obtain a new shape of where C' is the number of merged channels; simply transform A, B, D in the bidirectional state space module into A', B', D' to adapt to the unified input of cross-vehicle features; then, initialize the query vector Q, key vector K, and value vector V according to A', B', D', and these core components are generated from the input features through linear transformation:
[0017]
[0018] S = V·D',
[0019] K = B',
[0020] Q = C',
[0021] Among them, A is the shared forget gate in the bidirectional state space module, which combines with the dynamic adjustment value vector to adjust the importance of information in
[0022] (ii) The multi-head attention mechanism is adopted to decompose the input features into multiple subspaces, and each subspace independently generates corresponding query, key, and value vectors:
[0023] Q i = Q · W Qi
[0024] K i = K · W Ki
[0025] V i = V · W Vi
[0026] S i = S · W Si
[0027] where W Q , W K , W V , are trainable weight matrices;
[0028] Subsequently, the interaction between the key and the value is completed through the linear attention mechanism to obtain the output of each head. The calculation formula is:
[0029] O i = Q i · (K i · V i )
[0030] where the key-value interaction is completed through the calculation of K i · V i . This operation significantly reduces the computational complexity, which is much lower than the quadratic complexity of the traditional attention mechanism;
[0031] (iii) For the result O i of each head, add the skip connection result of the original feature to stabilize the attention feature to obtain To further integrate the feature outputs of each head, a concatenation operation is used to combine the results of all heads into an overall representation, and finally the output is represented as follows:
[0032]
[0033] The described channel-space fusion module finally generates a global feature representation from the perspective of the ego vehicle through the dual operations of channel attention and spatial feature fusion.
[0034] In some embodiments, the channel-space fusion module further includes:
[0035] Extract channel information from the multi-head attention integrated vehicle internal collaborative perception features output by the Mamba2 linear attention module, generate channel-level feature representations through global average pooling, and then calculate channel attention weights using two layers of 1×1 convolutions; in terms of spatial dimension fusion, extract the most significant spatial features in the scene through global max pooling, and at the same time capture the overall scene information using global average pooling, combine the two to generate a unified spatial feature representation, combine the channel and spatial features to generate integrated multi-vehicle perception features, and output the multi-head attention integrated vehicle internal collaborative perception features through the classification head and regression head.
[0036] In a second aspect, the present invention proposes an intelligent connected vehicle collaborative perception method based on the Mamba2 model, including:
[0037] S1, perform feature fusion on the BEV features of the ego vehicle and other vehicles to obtain BEV fusion features, compress the BEV fusion features, and output the compressed BEV fusion features as vehicle internal collaborative perception features;
[0038] S2, perform a linear transformation on the received compressed BEV fusion features as vehicle internal collaborative perception features through a multi-head state space model combined with a bidirectional feature scanning strategy to obtain the core part A, respectively represent the processed feature tensor and the forgetting gate, split the core feature tensor into three vectors: a state vector matrix, an input matrix, and an output matrix, model to obtain a forward feature sequence and a backward feature sequence, and perform a splicing operation on the forward feature sequence and the backward feature sequence through a state space dual module for fusion; share the forgetting gate with the Mamba2 linear attention module, fuse the forward sequence features and the backward sequence features of the bidirectional BEV scanning features, enhance and fuse the vehicle internal collaborative perception features, and output spatially enhanced vehicle internal collaborative perception features;
[0039] S3, based on the Mamba2 architecture design, introduce a linear attention mechanism, calculate the linear attention of the received spatially enhanced vehicle internal collaborative perception features, perform multi-head linear attention integration, use the forgetting gate mechanism shared with the bidirectional state space module to dynamically filter key information, and selectively enhance and globally aggregate the spatially enhanced vehicle internal collaborative perception features according to the linear attention, and output multi-head attention integrated vehicle internal collaborative perception features;
[0040] S4. Introduce the channel attention mechanism. The vehicle interior collaborative perception features integrated by the multi-head attention are processed through the channel attention mechanism to obtain the vehicle interior collaborative perception features in the channel dimension. The downsampled vehicle interior collaborative perception features in the channel dimension are then extracted and fused through spatial dimension splicing to output the vehicle interior collaborative perception features with channel-space fusion.
[0041] In some embodiments, S2 further includes:
[0042] Fuse the forward sequence features and the backward sequence features of the bidirectional BEV scan features to enhance and fuse the vehicle interior collaborative perception features;
[0043] Decompose and integrate the vehicle interior perception features using the result splicing operation through the multi-head attention mechanism.
[0044] In some embodiments, S3 further includes:
[0045] (i) For the input features To unify the modeling of cross-vehicle features, first reorganize the features in the channel dimension, merge the feature data of all vehicles, and obtain a new shape of where C' is the number of merged channels; perform simple transformations on A, B, and D in the bidirectional state space module to A', B', and D' to adapt to the unified input of cross-vehicle features; then, initialize the query vector Q, key vector K, and value vector V according to A', B', and D'. These core components are generated from the input features through linear transformations:
[0046]
[0047] S = V·D',
[0048] K = B',
[0049] Q = C',
[0050] where A is the shared forget gate in the bidirectional state space module, which dynamically adjusts the importance of the information in by combining with ;
[0051] (ii) Adopt the multi-head attention mechanism to decompose the input features into multiple subspaces, and each subspace independently generates the corresponding query, key, and value vectors:
[0052] Q i = Q·W Qi
[0053] K i = K·WKi
[0054] V i V = V·W Vi
[0055] S i S = S·W Si
[0056] Wherein, W Q , W K , W V , is a trainable weight matrix;
[0057] Subsequently, the interaction between the key and the value is completed through the linear attention mechanism to obtain the output of each head, and the calculation formula is:
[0058] Q i Q = Q i ·(K i ·V i )
[0059] Wherein, the key-value interaction is completed through the calculation of K i ·V i , and this operation significantly reduces the computational complexity, which is much lower than the quadratic complexity of the traditional attention mechanism;
[0060] (iii) For the result O of each head i add the skip connection result of the original feature to stabilize the attention feature to obtain To further integrate the feature outputs of each head, a concatenation operation is used to merge the results of all heads into a single overall representation, and finally an output is represented as follows:
[0061]
[0062] The described channel-spatial fusion module finally generates a global feature representation from the perspective of the vehicle itself through the dual operations of channel attention and spatial feature fusion.
[0063] In some embodiments, the S4 further includes:
[0064] Extract channel information from the multi - head attention integrated vehicle - interior collaborative perception features output by the Mamba2 linear attention module, generate channel - level feature representations through global average pooling, and then calculate channel attention weights using two - layer 1×1 convolutions. In terms of spatial - dimension fusion, extract the most prominent spatial features in the scene through global max pooling, and at the same time capture the overall scene information using global average pooling. Combine the two to generate a unified spatial feature representation. Combine the channel and spatial features to generate integrated multi - vehicle perception features, and output the vehicle - interior collaborative perception features integrated by multi - head attention through the classification head and regression head.
[0065] Compared with the prior art, the positive technical effects achieved by the present invention are as follows:
[0066] 1) Utilize the bidirectional state - space feature enhancement module to expand the receptive field of the multi - head state - space model through the bidirectional feature scanning strategy, solving the problems of limited long - distance dependence modeling and incomplete perception feature extraction in traditional collaborative perception methods; and, in order to further reduce the computational cost, this module uses result splicing instead of point - to - point multiplication operations, thereby retaining the feature integrity of the front - and - back scans while optimizing the computational efficiency. Decompose and integrate the vehicle - interior perception features through the multi - head attention mechanism, effectively expanding the perception range and maintaining low computational complexity;
[0067] 2) Utilize the Mamba2 linear attention module to solve the problems of high computational complexity and low cross - vehicle information interaction efficiency in traditional Transformer methods;
[0068] 3) Utilize the multi - head state - space model (SSM) combined with the bidirectional scanning strategy to enhance the feature learning ability without increasing parameters, solving the problem of insufficient feature reception range.
[0069] 4) The core of Mamba2 is to improve the SSM (state - space model) in Mamba, enabling it to be 2 - 8 times faster while performing equivalently to Transformer. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 It is a module diagram of the intelligent connected vehicle collaborative perception system based on the Mamba2 model of the present invention;
[0071] Figure 2 It is a specific implementation process diagram of the collaborative perception data of the pre - processing module of the present invention;
[0072] Figure 3 It is a specific implementation process diagram of the bidirectional state - space module and the Mamba2 linear attention module of the present invention; Figure (a) bidirectional state - space module, (b) Mamba2 linear attention module;
[0073] Figure 4 It is a specific implementation process diagram of the channel-space fusion module of the present invention;
[0074] Figure 5 It is a flow chart of the collaborative perception method for intelligent connected vehicles based on the Mamba2 model of the present invention. Specific implementation manners
[0075] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0076] As Figure 1 shown, a collaborative perception system for intelligent connected vehicles based on the Mamba2 model of the present invention includes a preprocessing module 100, a bidirectional state space module 200, a Mamba2 linear attention module 300, and a channel-space fusion module 400.
[0077] As Figure 2 shown, the preprocessing module 100 extracts the BEV features of the ego vehicle through a backbone network, and then obtains the BEV features of the vehicles ahead of the ego vehicle and the BEV features of other vehicles to perform feature fusion to obtain BEV fusion features, compresses the BEV fusion features, and uses the compressed BEV fusion features as the collaborative perception features inside the vehicle, and outputs the vehicle collaborative perception features. Among them, the features inside the vehicle refer to the features obtained by the vehicle itself, and the cross-vehicle features refer to the features mixed by multiple vehicles.
[0078] As Figure 3 shown, the bidirectional state space module 200 combines a multi-head state space model (SSM) with a bidirectional feature scanning strategy to enhance and fuse the collaborative perception features inside the vehicle, and outputs the spatially enhanced collaborative perception features inside the vehicle. Specifically, this module generates bidirectional BEV scanning features of the ego vehicle and adjacent vehicles through a bidirectional feature scanning strategy, and uses a linear transformation method to achieve feature transformation to obtain the core feature part A, which respectively represent the processed feature tensor and the forget gate. The multi-head state space model fuses the forward sequence features of the bidirectional BEV scanning features and the backward sequence features of the bidirectional BEV scanning features, effectively expanding the perception range and maintaining low computational complexity. To further reduce the computational cost, this module uses result concatenation instead of point-to-point multiplication operations, thereby retaining the feature integrity of the front and back scans while optimizing the computational efficiency.
[0079] Specifically, the module divides the received in-vehicle perception features, including the BEV features of the host vehicle and the BEV features received by the host vehicle from neighboring vehicles, into 9 blocks, and performs forward 1→2→3→4→5→6→7→8→9 and reverse 9→8→7→6→5→4→3→2→1 serialization processing respectively. It decomposes and integrates the in-vehicle perception features through the multi-head attention mechanism, and uses result concatenation instead of point-to-point multiplication operations. And it uses a bidirectional state space model to compress the visual representation, so as to perform standard normalization on the features after ImageNet classification and COCO object detection conversion, making the feature value distribution more uniform.
[0080] The Mamba2 linear attention module 300 is based on the Mamba2 architecture design, introduces the linear attention mechanism, generates query, key, and value, calculates the linear attention of the spatially enhanced in-vehicle collaborative perception features, performs multi-head linear attention integration, uses the forget gate mechanism to dynamically screen key information, and selectively enhances and globally aggregates the spatially enhanced in-vehicle collaborative perception features according to the linear attention, and outputs the in-vehicle collaborative perception features integrated by multi-head attention. This module reduces computational redundancy and storage overhead by sharing parameters; at the same time, it integrates the skip connection strategy to further stabilize the feature processing process.
[0081] As Figure 4 shown, the channel-space fusion module 400 enhances and selects the in-vehicle collaborative perception features integrated by multi-head attention through the channel attention mechanism to obtain the in-vehicle collaborative perception features in the channel dimension, uses the combination of global max pooling and average pooling to extract the pooled features of the in-vehicle collaborative perception features in the channel dimension, performs dimensionality reduction on the in-vehicle collaborative perception features in the channel dimension, and performs spatial dimension splicing and fusion on the extracted in-vehicle collaborative perception features in the channel dimension after dimensionality reduction, and outputs the in-vehicle collaborative perception features of channel-space fusion, realizing the lightweight fusion of the final features between vehicles.
[0082] As Figure 5 shown, the specific process of an intelligent connected vehicle collaborative perception method based on the Mamba2 model of the present invention includes:
[0083] Step 1: Perform preprocessing. Obtain the BEV fusion feature by fusing the BEV feature of the host vehicle with that of other vehicles, compress the BEV fusion feature, and output the compressed BEV fusion feature as the in-vehicle collaborative perception feature; including extracting the BEV feature of the host vehicle through a backbone network, then obtaining the BEV feature of the host vehicle and that of other vehicles for feature fusion to obtain the BEV fusion feature, and compressing the BEV fusion feature; among them, establish a network connection between vehicles, convert sensor data into a standardized feature representation, and share it through the V2X communication network; each vehicle encodes the point cloud data of its own lidar into a BEV feature through PointPillars, and the host vehicle receives the BEV features of other vehicles through the vehicle-to-vehicle communication network, shares the BEV features with surrounding vehicles, and after feature fusion and feature compression, the in-vehicle collaborative perception feature can be obtained;
[0084] Further, the formal definition of the in-vehicle collaborative perception feature specifically includes: the BEV feature represented by the sequence received by the host vehicle from neighboring vehicles is where N represents the total number of all vehicles, Channel is the number of feature channels, and H and W correspond to the height and width of the BEV feature in the vertical and horizontal directions; convert the feature into a compact BEV feature represented by a sequence through two convolutional operations where L is the length of the compressed sequence; for the converted compact feature perform standard normalization to make the feature value distribution more uniform;
[0085] Step 2: Use the multi-head state space model in combination with the bidirectional feature scanning strategy for the compressed BEV fusion feature received as the in-vehicle collaborative perception feature to achieve feature transformation to obtain the core feature part, which respectively includes the processed feature tensor and the forget gate, fuse the forward sequence feature and the backward sequence feature of the bidirectional BEV scanning feature, enhance and fuse the in-vehicle collaborative perception feature, and output the spatially enhanced in-vehicle collaborative perception feature; including decomposing and integrating the in-vehicle perception feature through the multi-head state space model, using result splicing instead of point-to-point multiplication operations. And use the bidirectional state space model to compress the visual representation, so as to perform standard normalization on the features after ImageNet classification and COCO object detection conversion to make the feature value distribution more uniform; this step also includes inputting the compressed BEV feature into the bidirectional state space module for processing; the specific description is as follows:
[0086] Step 2-1: Based on the bidirectional state space model as shown in Figure 3 (a), perform a linear transformation on the feature after standard normalization obtained in Step 1-1 to obtain the core feature part A, respectively representing the processed feature tensor and the forget gate; a separable convolution operation is used to extract the core feature tensor It is split into three vectors, namely the state vector matrix X, the input matrix B, and the output matrix C. The expressions are as follows:
[0087]
[0088] Among them, N represents the number of vehicles, L represents the length of the serialized features, h represents the number of attention heads, d s represents the dimensions of features C and B, and d represents the dimension of feature X. To alleviate the problem of incomplete vehicle information fusion caused by the unidirectional SSM, X, B, and C are further modeled to obtain the forward feature sequence and the backward feature sequence dependence, denoted as X fwd / X back , B fwd / B back , C fwd / C back ; furthermore, a Structured State Space Duality (SSD) module is used to perform feature modeling on the forward feature sequence and the backward feature sequence respectively. SSD is a module that improves the computational efficiency of the SSM operation by Mamba2. The forward feature sequence feature modeling F fwd , and the backward feature sequence feature modeling F back are expressed as follows:
[0089] F fwd = SSD(X fwd , B fwd , C fwd )
[0090] F back = SSD(X back , B back , C back )
[0091] The multi-head mechanism is adopted to simultaneously focus on different feature dimensions, further enhancing the ability to capture complex cross-vehicle features. For the input state vector matrix X, input matrix B, and output matrix C, the output Y is:
[0092] h t = Ah t-1 + Bx t
[0093] y t = Ch t
[0094] Y = (y0, y1…y end )
[0095] Among them, x t is the single-step input of the state vector matrix X, and h t is the hidden state of the state update, so as to obtain the final output Y;
[0096] Step 2-2, finally, the forward feature F fwd and the backward feature F back are fused through a concatenation operation, rather than element-wise multiplication, to reduce the computational complexity while retaining more feature details:
[0097]
[0098] Among them, represents the concatenation operation. is the output result of the bidirectional state space module.
[0099] This module is different from the unidirectional scanning strategy of the traditional Mamba2. Instead, it rearranges the order of the segmented features and scans them in two directions: forward (from start to end) and backward (from end to start). First, the input BEV features are divided into multiple sub-blocks, and each sub-block represents a subspace, preparing for subsequent parallel processing; then, the feature blocks are gradually scanned in forward and backward orders respectively to extract local and global feature information; then, the multi-head mechanism is used to perform block parallel analysis on the scanning results to further refine the subspace features and enhance the feature capture ability; finally, the features obtained from the forward and backward scans are integrated through a concatenation operation to completely retain the bidirectional feature information. Specifically, this module divides the received vehicle internal perception features, including the BEV features of the ego vehicle and the BEV features received by the ego vehicle from neighboring vehicles, into 9 blocks each, and performs forward 1→2→3→4→5→6→7→8→9 and backward 9→8→7→6→5→4→3→2→1 serialization processing respectively;
[0100] Step 3, based on the Mamba2 architecture design, introduce a linear attention mechanism, calculate the linear attention of the received spatially enhanced vehicle internal collaborative perception features, perform multi-head linear attention integration, use the forget gate mechanism shared with the bidirectional state space module to dynamically filter key information, and selectively enhance and globally aggregate the spatially enhanced vehicle internal collaborative perception features according to the linear attention, and output the vehicle internal collaborative perception features integrated by multi-head attention; the specific description is as follows:
[0101] Step 3-1, as Figure 3 (b) shows, for the input feature In order to uniformly model the cross-vehicle features, first reorganize the features in the channel dimension, merge the feature data of all vehicles, and obtain a new shape of Where C' is the number of merged channels, ensuring that the dimension of the input features is adapted to subsequent attention calculations. On this basis, sharing parameters with the bidirectional state space module strengthens the extraction and fusion of target features while reducing the computational load. A, B, and D in the bidirectional state space module are simply transformed into A', B', and D' to adapt to the unified input of cross-vehicle features. Then, the query vector Q, key vector K, and value vector V are initialized according to A', B', and D'. These core components are generated from the input features through linear transformation:
[0102]
[0103] S = V·D',
[0104] K = B',
[0105] Q = C',
[0106] where A is the shared forget gate in the bidirectional state space module, which dynamically adjusts the importance of information in by combining with . S is the dynamically enhanced value vector, used to dynamically control the importance of information.
[0107] Step 3-2: To further enhance the model's ability to capture complex features, the Mamba2 linear attention module adopts a multi-head attention mechanism. Specifically, the input features are decomposed into multiple subspaces, and corresponding query, key, and value vectors are independently generated for each subspace:
[0108] Q i = Q·W Qi
[0109] K i = K·W Ki
[0110] V i = V·W Vi
[0111] S i = S·W Si
[0112] where W Q , W K , W V , are trainable weight matrices, i is the subspace index, and S i serves as the independently adjusted value vector for each space;
[0113] Subsequently, the interaction between the key and the value is completed through the linear attention mechanism to obtain the output of each head. The calculation formula is:
[0114] O i= Q i · (K i · V i )
[0115] Wherein, the key-value interaction is completed through the calculation of K i · V i . This operation significantly reduces the computational complexity, which is much lower than the quadratic complexity of the traditional attention mechanism.
[0116] Step 3-3. Finally, for the result O i of each head, add the skip connection result of the original feature to stabilize the attention feature to obtain To further integrate the feature outputs of each head, a concatenation operation is used to merge the results of all heads into an overall representation, and finally an output is generated as follows:
[0117]
[0118] Wherein, N is the number of vehicles, L is the length of the serialized feature, and C is the dimension of the feature channel to realize the output integration of the attention mechanism of each head; is the attention output of the i-th subspace after residual connection, and there is
[0119] The described channel-space fusion module finally generates a global feature representation from the perspective of the ego vehicle through dual operations of channel attention and spatial feature fusion.
[0120] Step 4. Introduce a channel attention mechanism. The vehicle internal collaborative perception feature integrated by the multi-head attention received is passed through the channel attention mechanism to obtain the vehicle internal collaborative perception feature in the channel dimension. The vehicle internal collaborative perception feature in the dimension after dimensionality reduction is extracted and fused by spatial dimension concatenation to output the vehicle internal collaborative perception feature of channel-space fusion; specifically, extract the channel information from the vehicle internal collaborative perception feature integrated by the multi-head attention output by the Mamba2 linear attention module, generate a channel-level feature representation through global average pooling, and then calculate the channel attention weight using two layers of 1×1 convolution. In terms of spatial dimension fusion, the module extracts the most significant spatial features in the scene through global max pooling, and at the same time uses global average pooling to capture the overall scene information, combines the two to generate a unified spatial feature representation; finally combines the channel and spatial features to generate an integrated multi-vehicle perception feature; finally outputs the vehicle internal collaborative perception feature integrated by the multi-head attention through the classification head and the regression head;
[0121] Furthermore, the enhanced selection of the vehicle interior collaborative perception features integrated by the multi-head attention is achieved through the channel attention mechanism to obtain the vehicle interior collaborative perception features in the channel dimension. The pooling feature extraction of the vehicle interior collaborative perception features in the channel dimension is realized by combining global max pooling and average pooling, and the dimensionality reduction of the vehicle interior collaborative perception features in the channel dimension is performed. The dimensionality-reduced vehicle interior collaborative perception features in the channel dimension are extracted for spatial dimension splicing and fusion, and the vehicle interior collaborative perception features of channel-space fusion are output; specifically, it includes merging the vehicle interior collaborative perception features with enhanced multi-vehicle space output by the bidirectional state space module into a global feature representation, dynamically screening the key information in the input features by using the forgetting gate mechanism to suppress redundant features, and constructing the feature representations of key, query, and value through the linear attention mechanism, performing a linear transformation to calculate the global relevance of the cross-vehicle features, introducing the multi-head mechanism through the bidirectional state space module to divide the global features into multiple sub-spaces for parallel processing, refining the feature interaction, and fusing the features into the final result by combining the skip connection strategy to enhance the model stability and reduce information loss. The results after multi-head processing are integrated into a unified feature representation and output for use by subsequent modules to achieve selective enhancement and global aggregation of cross-vehicle features, while reducing the computational complexity and storage overhead; the specific description is as follows:
[0122] Step 4-1, as Figure 4 shown, first reshape the input feature into its original form for global channel feature extraction. By performing global average pooling on the height H and width W, a channel-level feature representation is generated Then input the channel feature F gap into two consecutive 1×1 convolutional layers, and use the activation function σ for non-linear mapping respectively to obtain the channel attention weight
[0123]
[0124] M ch = Conv2(σ(Conv1(F gap )))
[0125] Then perform the channel enhancement operation, recalibrate the input features using the channel attention weight, and achieve feature enhancement through channel-wise weighting. The formula is as follows:
[0126] F ch = M ch ⊙F (ego,neb)
[0127] Among them, ⊙ represents the per-channel multiplication operation. This step effectively highlights the key channel features and suppresses redundant information.
[0128] Step 4-2: Perform max-pooling and average-pooling operations on the enhanced channel feature F ch along the vehicle dimension N respectively. Max-pooling retains the most prominent features in the scene, while average-pooling captures the global context information.
[0129]
[0130] Step 4-3: Concatenate the max-pooling result F max and the average-pooling result F avg to obtain the final global spatial feature representation The expression is as follows:
[0131]
[0132] Among them, F out represents the concatenation operation, which fuses the significant features and the global background information, providing the ego vehicle with comprehensive spatial perception ability.
[0133] Compared with the prior art, the fusion method proposed by the present invention achieves a better balance between computational efficiency and perception performance. It significantly improves the object detection accuracy compared with the prior art, and demonstrates stronger perception ability and adaptability in complex traffic environments. At the same time, it effectively reduces the computational cost and communication bandwidth requirements. Different from the traditional SSM module, a multi-head mechanism is introduced in the SSM, which not only reduces the computational cost but also can more effectively capture the internal features and relationships of the vehicle. Moreover, by exploring the potential of the linear growth model in cooperative perception, the present invention focuses on solving the contradiction between real-time performance and computational complexity in the cooperative perception system, and realizes the comprehensive integration and efficient processing of multi-vehicle perception information in complex environments.
[0134] For those skilled in the art, improvements and changes can be made according to the above invention content. Any modifications, improvements, and changes made based on the present invention shall fall within the protection scope of the present invention.
[0135] In addition, based on a similar inventive concept, an embodiment of the present invention further provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above method is implemented.
[0136] In addition, based on a similar inventive concept, an embodiment of the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented.
[0137] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the function specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the flowchart.
[0138] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the flowchart.
Claims
1. An intelligent networked vehicle collaborative perception system based on the Mamba2 model, characterized in that It includes a preprocessing module, a bidirectional state space module, a Mamba2 linear attention module, and a channel-space fusion module: Among them: The preprocessing module fuses the ego-vehicle BEV features and the other-vehicle BEV features to obtain BEV fusion features, compresses the BEV fusion features, and outputs the compressed BEV fusion features as the in-vehicle collaborative perception features. The two-way state space module linearly transforms the received in-vehicle collaborative perception features through a multi-head state space model combined with a two-way feature scanning strategy to obtain the core part respectively represent the processed feature tensor and the forget gate. Split the core feature tensor into three vectors: a state vector matrix, an input matrix, and an output matrix. Model the forward feature sequence and the backward feature sequence, and realize the splicing operation of the forward feature sequence and the backward feature sequence through the state space dual module for fusion; Share the forget gate with the Mamba2 linear attention module, fuse the forward sequence features of the two-way BEV scanning features and the backward sequence features of the two-way BEV scanning features, enhance and fuse the in-vehicle collaborative perception features, and output the spatially enhanced in-vehicle collaborative perception features; The Mamba2 linear attention module is designed based on the Mamba2 architecture, introduces a linear attention mechanism, calculates the linear attention of the received spatially enhanced in-vehicle collaborative perception features, performs multi-head linear attention integration, dynamically filters key information using the forget gate shared with the bidirectional state space module, and selectively enhances and globally aggregates the spatially enhanced in-vehicle collaborative perception features according to the linear attention, and outputs the in-vehicle collaborative perception features integrated by multi-head attention. The channel-space fusion module introduces a channel attention mechanism, obtains the in-vehicle collaborative perception features in the channel dimension for the received in-vehicle collaborative perception features integrated by multi-head attention through the channel attention mechanism, extracts and performs spatial dimension splicing fusion on the in-vehicle collaborative perception features in the channel dimension after dimensionality reduction, and outputs the in-vehicle collaborative perception features fused by channel-space.
2. The intelligent networked vehicle collaborative perception system based on the Mamba2 model according to claim 1, wherein, The bidirectional state space module further includes: Fusing the forward sequence features and the backward sequence features of the bidirectional BEV scan features to enhance and fuse the in-vehicle collaborative perception features; Decomposing and integrating the in-vehicle perception features by using a result splicing operation through a multi-head attention mechanism.
3. The intelligent networked vehicle collaborative perception system based on the Mamba2 model according to claim 1, wherein, The Mamba2 linear attention further includes: (i) For the input features First, the features are reorganized in the channel dimension to merge the feature data of all vehicles, resulting in a new shape of where C' is the number of channels after merging; perform simple transformations on A, B, D in the bidirectional state space module to A', B', D' to adapt to the unified input of cross-vehicle features; then, initialize the core components including the query vector Q, key vector K, and value vector V according to A', B', D'. The query vector Q, key vector K, and value vector V are generated from the input features through linear transformation, and the expressions are as follows: S = V·D′, K = B′, Q = C′, Among them, A is the shared forget gate in the bidirectional state space module, which dynamically adjusts the importance of information in by combining with ; S is the dynamic enhancement value vector. (ii) Using a multi-head attention mechanism, the input features are decomposed into multiple subspaces, and each subspace independently generates the corresponding query vector Q of each head attention mechanism i , the key vector K of each head attention mechanism i and the value vector V of each head attention mechanism i , and the expression is as follows: Q i = Q · W Qi K i = K·W Ki V i = V · W Vi S i = S·W Si Among them, W Q , W K , W V , is a trainable weight matrix, i is a subspace index, and S i is a value vector for independent adjustment of each space; Subsequently, the interaction between the key and the value is completed through the linear attention mechanism to obtain the output O of each head attention mechanism i , and the calculation formula is: O i = Q i · (K i · V i ); (iii) For the output O of each head attention mechanism i Add the skip connection result of the query vector Q of each head attention mechanism i To obtain stable attention features Use a concatenation operation to merge the results of all heads into an overall representation, and finally generate the output of the overall attention mechanism Where N represents the number of vehicles, L represents the length of the serialized features, and C is the dimension of the feature channels; The output expression of the integrated multi-head attention mechanism is as follows: Among them, represents the attention output of the $i$-th subspace after residual connection; The channel-space fusion module finally generates a global feature representation from the ego-vehicle perspective through dual operations of channel attention and spatial feature fusion.
4. The intelligent networked vehicle collaborative perception system based on the Mamba2 model according to claim 1, wherein, The channel-space fusion module further includes: Extract the channel information from the in-vehicle collaborative perception features integrated by multi-head attention output by the Mamba2 linear attention module, generate a channel-level feature representation through global average pooling, and then calculate the channel attention weights using two layers of 1×1 convolution; in terms of spatial dimension fusion, extract the most significant spatial features in the scene through global max pooling, and at the same time capture the overall scene information using global average pooling, combine the two to generate a unified spatial feature representation, combine the channel and spatial features to generate the integrated multi-vehicle perception features, and output the in-vehicle collaborative perception features integrated by multi-head attention through a classification head and a regression head.
5. An intelligent networked vehicle collaborative perception method based on the Mamba2 model, characterized in that, The system includes a preprocessing module, a bidirectional state space module, a Mamba2 linear attention module, and a channel-space fusion module: Among them: S1. Feature fusion is performed on the BEV features of the host vehicle and other vehicles to obtain BEV fusion features, and compression is performed on the BEV fusion features, and the compressed BEV fusion features are output as the vehicle internal collaborative perception features; S2, the compressed BEV fusion feature received is used as the in-vehicle collaborative perception feature, and is linearly transformed through a multi-head state space model combined with a bidirectional feature scanning strategy to obtain the core part They respectively represent the processed feature tensor and the forget gate. The core feature tensor is split into three vectors: a state vector matrix, an input matrix, and an output matrix. A forward feature sequence and a backward feature sequence are modeled, and the splicing operation of the forward feature sequence and the backward feature sequence is realized through a state space dual module for fusion; the forget gate is shared with the Mamba2 linear attention module, and the forward sequence feature of the bidirectional BEV scanning feature and the backward sequence feature of the bidirectional BEV scanning feature are fused to enhance and fuse the in-vehicle collaborative perception feature, and the spatially enhanced in-vehicle collaborative perception feature is output; S3. Based on the Mamba2 architecture design, a linear attention mechanism is introduced to calculate the linear attention of the received spatially enhanced vehicle internal collaborative perception features, perform multi-head linear attention integration, dynamically filter key information using the forget gate shared with the bidirectional state space module, and selectively enhance and globally aggregate the spatially enhanced vehicle internal collaborative perception features according to the linear attention, and output the vehicle internal collaborative perception features integrated by multi-head attention; S4. A channel attention mechanism is introduced, and the received vehicle internal collaborative perception features integrated by multi-head attention are processed through the channel attention mechanism to obtain vehicle internal collaborative perception features in the channel dimension, and the downsampled vehicle internal collaborative perception features in the channel dimension are extracted and fused by spatial dimension splicing, and vehicle internal collaborative perception features fused by channel-space are output.
6. The intelligent networked vehicle collaborative perception method based on the Mamba2 model according to claim 5, wherein, S2 further includes: Fuse the forward sequence features of the bidirectional BEV scan features and the backward sequence features of the bidirectional BEV scan features to achieve enhancement and fusion of the vehicle internal collaborative perception features; Decompose and integrate the vehicle internal perception features by using the result splicing operation through the multi-head attention mechanism.
7. The intelligent networked vehicle collaborative perception system based on the Mamba2 model according to claim 5, wherein, The Mamba2 linear attention further includes: (i) For the input features First, the features are reorganized in the channel dimension to merge the feature data of all vehicles, resulting in a new shape of where C' is the number of channels after merging; A, B, and D in the bidirectional state space module are simply transformed into A', B', and D' to adapt to the unified input of cross-vehicle features; then, the core components including the query vector Q, key vector K, and value vector V are initialized according to A', B', and D'. The query vector Q, key vector K, and value vector V are generated from the input features through linear transformation, and the expressions are as follows: S = V·D′, K = B′, Q = C′, Among them, A is the shared forget gate in the bidirectional state space module, which dynamically adjusts the importance of the information in by combining with ; S is the dynamic enhancement value vector. (ii) Using the multi-head attention mechanism, the input features are decomposed into multiple subspaces, and each subspace independently generates the corresponding query vector Q of each head attention mechanism i , the key vector K of each head attention mechanism i and the value vector V of each head attention mechanism i , and the expression is as follows: Q i = Q · W Qi K i = K·W Ki V i = V · W Vi S i = S · W Si Among them, W Q , W K , W V , is a trainable weight matrix, i is a subspace index, and S i is a value vector for independent adjustment of each space; Subsequently, the interaction between the keys and values is completed through the linear attention mechanism to obtain the output O of each head attention mechanism i , and the calculation formula is as follows: O i = Q i ·(K i ·V i ); (iii) For the output O of each head attention mechanism i Add the skip connection result of the query vector Q of each head attention mechanism i To obtain stable attention features Use a concatenation operation to merge the results of all heads into an overall representation, and finally generate the output of the overall attention mechanism Where N represents the number of vehicles, L represents the length of the serialized features, and C is the dimension of the feature channels; The output expression of the integrated multi-head attention mechanism is as follows: Among them, represents the attention output of the i-th subspace after residual connection; The channel-space fusion module finally generates a global feature representation from the perspective of the host vehicle through dual operations of channel attention and spatial feature fusion.
8. The intelligent networked vehicle collaborative perception method based on the Mamba2 model according to claim 5, wherein, S4 further includes: Extract channel information from the vehicle internal collaborative perception features integrated by multi-head attention output from the Mamba2 linear attention module, generate a channel-level feature representation through global average pooling, and then calculate the channel attention weights using two layers of 1×1 convolutions; in terms of spatial dimension fusion, extract the most significant spatial features in the scene through global max pooling, and at the same time capture the overall scene information using global average pooling, combine the two to generate a unified spatial feature representation, combine the channel and spatial features to generate integrated multi-vehicle perception features, and output the vehicle internal collaborative perception features integrated by multi-head attention through the classification head and regression head.
9. An electronic device, characterized in that, It includes: A memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the intelligent networked vehicle collaborative perception method based on the Mamba2 model according to any one of claims 5 to 8.
10. A non-transitory computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by the processor, it implements the intelligent networked vehicle collaborative perception method based on the Mamba2 model according to any one of claims 5 to 8.