Vehicle-road cloud cooperative intelligent vehicle environment sensing system and method

Through 5G cellular-satellite multi-mode information fusion and pyramid Mamba network, the problem of limited positioning accuracy and transmission bandwidth in advanced autonomous driving is solved, efficient vehicle-road cloud collaborative perception is achieved, and perception capabilities in the entire region are enhanced.

CN120356185APending Publication Date: 2025-07-22JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510491705.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art has problems in the advanced autonomous driving with insufficient positioning accuracy, limited transmission bandwidth and low vehicle-road information fusion efficiency, especially in areas where low-orbit satellite systems and 5G cellular networks are insufficient, resulting in unreliable perception results.

Method used

The unified position coding module of 5G cellular-satellite multi-mode information fusion is adopted, combined with the wide-area scene feature extraction network based on pyramid Mamba and the vehicle-road cloud multi-mode information differential attention fusion module to realize unified encoding, wide-area feature extraction and differential information fusion information of bicycles and other vehicle sensor information.

Benefits of technology

It improves the positioning accuracy and information transmission efficiency of the whole-domain perception, enhances the stability and information fusion capabilities of vehicle-road cloud collaborative perception, and adapts to the wide-area environment needs of advanced autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356185A_ABST
    Figure CN120356185A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle-road cloud cooperative intelligent vehicle environment sensing system and method, and the method comprises the steps: fusing the sensor information of a vehicle with the sensor information of other vehicles through a position transformation matrix, and forming a unified position code of the sensor information of multiple vehicles; a data preprocessing module is used for converting vehicle 3D point cloud data into a 2D pseudo graph, the 2D pseudo graph is input into a pyramid-based Mama feature coding and decoding module, and feature extraction of multi-vehicle wide area information is realized through a multi-layer block coding module, a Mama module and a residual error up-sampling module; inputting the multi-vehicle wide-area feature information into a vehicle-road cloud multimode information difference attention fusion module to realize effective fusion of difference information in environment information sensed by different vehicles so as to form wide-area fusion features; and outputting the prediction frame containing the surrounding vehicle information. According to the invention, high-order global automatic driving environment perception based on vehicle and road cloud cooperation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of transportation, and relates to an intelligent vehicle environment perception system and method for vehicle-road-cloud collaboration. Background Art

[0003] The positioning accuracy of satellite navigation is generally divided into centimeter level, decimeter level and meter level. The navigation accuracy required for existing L2-level autonomous driving is meter level. As the autonomous driving level is evolving from L2 level to L3 / L4 level, the positioning accuracy requirement is also continuously increasing. For high-level autonomous driving, the traditional meter-level positioning accuracy is difficult to meet the accuracy requirements of L3 and above autonomous driving, and the positioning system accuracy needs to be improved from meter level to decimeter level and centimeter level.

[0004] At present, the mainstream of the autonomous driving positioning system is the multi-sensor fusion positioning scheme, which forms the basic positioning ability of autonomous driving vehicles by combining high-precision positioning with high-precision maps. Among them, high-precision positioning forms the self-vehicle position perception ability through the combination of satellite navigation, vehicle-mounted inertial navigation and wheel speed sensors, etc.; high-precision map matching with vehicle-mounted lidar sensors obtains the relative environmental position. The two are independent and redundant for safety, which can effectively enhance the system robustness and ensure the effective perception of the autonomous driving vehicle's own position.

[0005] Meeting the above positioning requirements necessarily requires a powerful communication infrastructure as support. The autonomous driving perception system in areas lacking high-speed broadband network access will be in an unavailable state.

[0006] In response to the above problem of limited 5G mobile communication network, the low-Earth orbit satellite constellation system has become an effective supplementary communication method. The constellation system integrates functions such as communication, remote sensing and navigation, and neither requires a large number of ground communication base stations to be evenly deployed nor is restricted by space. In remote areas and even extreme autonomous driving scenarios such as natural disasters, the satellite system can still provide high-precision positioning and data transmission services.

[0007] After accessing the above-mentioned low-earth orbit satellite system and cooperating with the vehicle-road-cloud collaborative perception system, the existing small-scale perception scenario centered on the self-vehicle perception information will be upgraded to a wide-area perception scenario. In this regard, the existing backbone network for feature extraction based on the convolutional neural network (CNN) will face major challenges, especially manifested as problems such as blurred features of small targets in wide-area scenarios and insufficient semantic detection, which further lead to a decrease in generalization in real scenarios. At the same time, in the process of fusing vehicle-side data and cloud-side data, the current fusion method centered on Vision Transformer has shortcomings. It only focuses on the common information between the two, ignoring the decisive role of the differential information between the two in the perception result, ultimately resulting in the failure of the detection task. At the same time, with the continuous enhancement of the cloud-side information transmission and storage capabilities, the issue of the output weights of differential data between the vehicle side and the cloud side will also need to be explored. Summary of the Invention

[0008] Aiming at the deficiencies in the prior art, the present invention provides an intelligent vehicle environment perception system and method for vehicle-road-cloud collaboration.

[0009] The present invention achieves the above technical objectives through the following technical means.

[0010] The intelligent vehicle environment perception system for vehicle-road-cloud collaboration includes:

[0011] A unified position coding module for 5G cellular-satellite multimodal information fusion: In terms of information transmission, a hybrid information transmission method of 5G cellular network-low-earth orbit satellite is adopted in cities, and a single-mode information transmission method of low-earth orbit satellite is adopted in remote areas; in terms of position information coding, a position transformation matrix is used to fuse the sensor information of the self-vehicle and the sensor information of other vehicles to form a unified position coding of multi-vehicle sensor information;

[0012] A wide-area scene feature extraction network based on Pyramid Mamba, including a data preprocessing module and a Pyramid Mamba feature encoding and decoding module, where the Pyramid Mamba feature encoding and decoding module includes two processes: Mamba downsampling encoding and residual upsampling encoding; the function of the data preprocessing module is to convert the 3D point cloud data of the vehicle into a 2D pseudo-image; in the Mamba downsampling encoding, first use the block encoding module to convert the 2D pseudo-image into a planar two-dimensional pseudo-image image block, and then project the image block onto a vector with a hidden state dimension of D, and add the position encoding of each image block. The position encoding is input into the Mamba module to convert the block encoding feature into a scene feature; the above process is repeated four times to generate four scene features of different scales; the four-scale scene features are respectively input into four rounds of residual upsampling encoding, and finally generate the single-view scene features of the vehicle end and other vehicles;

[0013] The vehicle-road-cloud multi-modal information difference attention fusion module includes an information difference attention module DIAM based on the cross-attention mechanism and an optional information attention module AIAM. The DIAM calculates the difference features between the ego vehicle and other vehicles through the cross-attention mechanism, and the AIAM fuses the common features of the ego vehicle and other vehicles through the cross-attention mechanism. The difference features and common features are passed through the AIAM process again to generate fusion features. The fusion features are respectively input into two convolutional layers to generate classification and regression prediction results and form prediction boxes.

[0014] In the above technical solution, the position transformation matrix satisfies the following formula:

[0015]

[0016] Where x and y represent the position coordinates of the point cloud in the sensor coordinate system of other vehicles, and x' and y' represent the position coordinates of the point cloud in the ego vehicle coordinate system. t x 、t y represent the displacement vectors of the point cloud along the X-axis and Y-axis translation transformations, and θ y represents the rotation angle of the point cloud along the Y-axis rotation transformation.

[0017] In the above technical solution, the point cloud data of the vehicle is obtained by applying the position transformation matrix to the point cloud data collected by the sensors of the ego vehicle and other vehicles.

[0018] In the above technical solution, the Mamba downsampling includes N loops, that is, the block-encoded feature T n-1 is sent to the encoder to obtain the output sequence T n , specifically: the input block-encoded feature T n-1 , in the first loop, the input feature T0 is first normalized, and T0 is projected to a higher dimension using a linear layer to generate x and z features. Subsequently, x is input into the bidirectional SSM process. In the bidirectional SSM process, first, the options of forward propagation and backward propagation are determined, and then x is preprocessed using a 1×1 convolution and a ReLU activation function to generate x' o , and then x' o is linearly projected into the intermediate quantities B o , C o and Δ o to obtain y o , y o includes y forward and y backward . Based on y forward and y backward , finally, the SSM converts z into y' forward and y' backward ; y' forward and y' backwardThe input after addition is linearly transformed to Linear T (), and is added to the sequence T at the previous moment n-1 to perform a residual connection and output the sequence T n .

[0019] In the above technical solution, the SSM transforms z into y' forward and y' backward , specifically:

[0020] y' forward :(B,M,E)=y forward ☉ReLU(z)

[0021] y' backward :(B,M,E)=y backward ☉ReLU(z)

[0022] Among them, B represents the batch size, M represents the length of the sequence T n , and E represents the expanded state dimension;

[0023] After adding y' forward and y' backward , it undergoes a linear transformation Linear T (), and is added to the previous moment T n-1 to output the next sequence T n , specifically:

[0024] T n :(B,M,D)=Linear T (y' forward +y' backward )+T n-1 .

[0025] In the above technical solution, corresponding to the Mamba four-time downsampling encoding, the upsampling encoding is performed four times, and the k-th upsampling encoding is marked as k.

[0026] In the above technical solution, the first encoding in the processing process of the residual upsampling module is expressed as:

[0027] F RUB 1 =RUB(F Mamba 1 )+F Mamba 1

[0028] The k-th upsampling encoding after that is expressed as:

[0029] F RUB k =RUB(F RUB k-1)+F Mamba k-1

[0030] Among them, RUB() represents the upsampling encoding process, and F RUB 1 represents the feature after being processed by the residual upsampling module for the first time, and F Mamba 1 represents the feature block generated in the first cycle of the downsampling encoding process, and F Mamba n-1 represents the feature after Mamba downsampling encoding at the corresponding scale, and F RUB k-1 represents the feature after being processed by the residual upsampling module for the (k - 1)th time.

[0031] In the above technical solution, the DIAM calculates the difference feature between the host vehicle and other vehicles through the attention mechanism. Specifically:

[0032] Obtain the self-attention common feature C of the host vehicle and other intelligent agent vehicles QV :

[0033]

[0034] Remove the common feature information of the self-attention common feature and the host vehicle feature, and obtain the difference feature DI between the common feature and the host vehicle feature QV :

[0035] DI QV = Linear(V - C QV )

[0036] Inject the difference feature DI QV into the other vehicle feature Q to obtain the difference feature F between the host vehicle and other vehicles DIAM :

[0037] F DIAM = MLP(LN(DI QV + Q))+(DI QV + Q)

[0038] Among them, d k is the channel number ratio factor, Q represents the other vehicle feature, K and V both represent the host vehicle feature, Linear() represents the linear layer, MLP() represents the multi-layer perceptron, and LN() represents the linear layer.

[0039] In the above technical solution, the AIAM realizes the fusion of the common features of the host vehicle and other vehicles through the cross-attention information fusion mechanism. Specifically:

[0040] Obtain the self-attention common feature C of the host vehicle and other intelligent agent vehicles QV:

[0041]

[0042] Common feature C QV Inject it into other vehicle feature Q to obtain the common feature F between the host vehicle and other vehicles AIAM :

[0043] F AIAM = MLP(LN(C QV + Q))+(C QV + Q)

[0044] Finally, through another AIAM process, the differential feature F DIAM and the common feature F AIAM generate the fusion feature M ego :

[0045] M ego = AIAM(F AIAM , F DIAM ).

[0046] Intelligent vehicle environment perception method for vehicle-road-cloud collaboration:

[0047] Use the position transformation matrix to fuse the sensor information of the host vehicle and the sensor information of other vehicles to form a unified position encoding of multi-vehicle sensor information;

[0048] Use the data preprocessing module to convert the vehicle 3D point cloud data into a 2D pseudo-map;

[0049] Input the 2D pseudo-map into the pyramid Mamba feature encoding and decoding module, and through multiple layers of block encoding modules, Mamba modules and residual upsampling modules, realize the feature extraction of multi-vehicle wide-area information;

[0050] Input the multi-vehicle wide-area feature information into the vehicle-road-cloud multi-modal information difference attention fusion module to effectively fuse the difference information in the environment information perceived by different vehicles and form a wide-area fusion feature;

[0051] Output the prediction box containing the information of surrounding vehicles.

[0052] The beneficial effects of the present invention are:

[0053] (1) The present invention proposes a unified position coding module for 5G cellular-satellite multimodal information fusion. Based on the information transmission in traditional 5G cellular networks and the self-vehicle positioning of GNSS, a low-Earth orbit satellite constellation system is introduced to provide a wide-area large-bandwidth communication environment and high-precision positioning. The vehicle-road-cloud collaborative perception domain is extended from the urban environment to the global environment, greatly enhancing the drivable area for high-level autonomous driving. The large-bandwidth communication can directly transmit the original perception information, avoiding problems such as information attenuation and feature loss. The vehicle-road-cloud unified position coding module based on high-precision positioning improves the accuracy of perception information stitching and fusion, providing a data basis for subsequent wide-area feature extraction.

[0054] (2) The present invention proposes a wide-area scene feature extraction network based on Pyramid Mamba; due to the large scope of vehicle-road-cloud collaborative perception, traditional CNN feature extractors are prone to phenomena such as fuzzy and insufficient feature semantics; the Mamba network can efficiently capture global spatial information and improve the information extraction ability for real-scene point cloud features by virtue of its advantage in extracting semantic features of long sequence data; the introduction of the residual upsampling module further enhances the stability of the network by concatenating feature information of different spatial scales.

[0055] (3) Since the traditional Transformer-based fusion method has low fusion efficiency for feature information from different sources and cannot extract and fuse differential information. In response, the present invention proposes a vehicle-road-cloud multimodal information difference attention fusion module, which separates differential feature information and common feature information by designing an information difference attention module and an optional information attention module, solves the problem of extraction and utilization of differential information, and improves the fusion efficiency of vehicle-road feature information. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 Schematic diagram of the unified position coding module for 5G cellular-satellite multimodal information fusion according to the present invention;

[0057] Figure 2 Schematic diagram of the wide-area scene feature extraction network based on Pyramid Mamba according to the present invention;

[0058] Figure 3 Schematic diagram of the vehicle-road-cloud multimodal information difference attention fusion module according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but the protection scope of the present invention is not limited thereto.

[0060] The present invention mainly realizes the environmental perception task of high-level vehicle-road-cloud full-domain autonomous driving. Traditional L2 autonomous driving perception relies on the GNSS system to provide vehicle positioning, the 5G cellular network to provide communication support, and the scene understanding provided by the vehicle sensors. However, the above scheme also faces many problems. For example, the positioning accuracy of the GNSS system is at the meter level, and there is a problem of low positioning accuracy; there are problems of data disconnection and information transmission blockage in areas with low 5G cellular network coverage; there are perception blind spots in the scene understanding of the vehicle sensors, resulting in unreliable perception results.

[0061] In order to achieve global perception of high-level autonomous driving, three main problems need to be solved: 1. Centimeter-level high-precision positioning and global large-bandwidth communication; 2. Wide-area scene feature extraction in vehicle-road-cloud scenarios; 3. Vehicle-road-cloud multi-source heterogeneous information fusion output. In view of the three technical difficulties raised in the implementation of the above-mentioned high-level autonomous driving global perception system, the present invention proposes targeted solutions. The following will introduce three aspects of the present invention: 1. Unified position coding module for 5G cellular-satellite multi-mode information fusion; 2. Wide-area scene feature extraction network based on pyramid Mamba; 3. Vehicle-road-cloud multi-mode information difference attention fusion module.

[0062] (1) Unified location coding module for 5G cellular-satellite multi-mode information fusion

[0063] High-level autonomous driving has full coverage in all areas, providing information transmission and positioning navigation services for areas not covered by 5G signals. Existing high-level autonomous driving relies on 5G cellular network base stations to provide communication support, obtain high-precision maps and combine them with the information from the on-board lidar sensor to achieve high-precision positioning. The above conditions can be met in urban environments with high coverage of 5G base stations, but in areas such as larger rural areas, sparsely populated mountainous areas and even vast deserts, the communication needs are difficult to meet. The main reason is that the coverage range of 5G cellular network base stations is small, and the construction cost in the above-mentioned areas is high. A large number of base stations need to be built to meet the needs of full-area communications. However, the above-mentioned areas are often sparsely populated and lack user demand, but full-area autonomous driving must also cover the above-mentioned areas. In this regard, the present invention proposes a full-area communication technology that integrates 5G cellular networks and low-orbit satellite constellations, namely a 5G cellular-satellite system, such as Figure 1 As shown, it includes two communication modes: cellular network-satellite hybrid signal transmission under urban conditions and satellite single-mode signal transmission under remote conditions.

[0064] The urban cellular network-satellite hybrid signal transmission provides two highly redundant transmission modes: 5G cellular network and satellite. In urban scenarios, since high-precision maps need to depict accurate static target information, the amount of road information data is large, so a larger information transmission bandwidth is required to transmit high-precision map information. Therefore, 5G cellular network information is added on the basis of the basic satellite communication transmission mode. In remote areas with weak 5G network coverage such as forests, the information depicted in the high-precision maps is also less, and high information transmission bandwidth is not required. Relying solely on satellite communication bandwidth can meet the needs of high-level autonomous driving.

[0065] After meeting the above requirements for high-precision positioning and large-bandwidth communication, it is necessary to upload the environmental perception information of the vehicle's own sensors to the cloud, and after position conversion with roadside facilities and other vehicle information, it is sent back to the vehicle to achieve wide-area environmental perception. The traditional non-satellite transmission information fusion method has problems such as limited transmission bandwidth and misalignment of perception information fusion, and it needs to be solved by the method of information compression and post-fusion of feature transmission. The collaborative perception system empowered by the low-earth orbit satellite system has the advantages of more accurate centimeter-level positioning information and higher transmission bandwidth. Through the position modeling module, the original point cloud sensed by other vehicle sensors can be transformed into the space of the vehicle's own coordinate system through coordinate transformation, so as to unify the perception data under the vehicle's own coordinate system. The core of the vehicle-road-cloud collaborative perception architecture is that the vehicle's own vehicle expands the global perception information through the cloud, so as to enhance the vehicle's own understanding ability of the surrounding environment. Therefore, it is necessary to use the vehicle's own position as the coordinate origin of the coordinate system to obtain the wide-area environmental information empowered by the cloud.

[0066] The position modeling module needs to adapt to the information transmission requirements in multi-modal data scenarios. For this reason, a general perception perspective based on the bird's-eye view (BEV) is selected. The BEV perspective lacks height information, so the position transformation relationship coefficient is simplified from six degrees of freedom (t x , t y , t z , θ p , θ r , θ y ) to three degrees of freedom (t x , t y , θ y ). The introduction of the position transformation matrix transforms the information of other sensors within the perception range into the vehicle's own BEV space. The formula is as follows:

[0067]

[0068] Among them, x and y represent the position coordinates of the point cloud in the coordinate system of other vehicle sensors, x' and y' represent the position coordinates of the point cloud in the vehicle's own coordinate system, t x , t y represent the displacement vectors of the point cloud along the X-axis and Y-axis translation transformations, and θy Represents the rotation angle of the point cloud for rotation transformation along the Y-axis.

[0069] Generate a unified position encoding through the above steps to achieve spatial alignment of the original point cloud with position information before information transmission.

[0070] (2) Wide-area scene feature extraction network based on Pyramid Mamba

[0071] After obtaining the point cloud data L in the BEV view i (Specifically: the point cloud data sensed by the sensors of the ego vehicle and other vehicles, obtained through the position transformation matrix), input the point cloud data into the wide-area feature extraction network based on Pyramid Mamba. The process of feature extraction is as Figure 2 shown. First, the environmental data needs to be converted into Pillar tensors through the Pillar Feature Net network, and then projected into a 2D pseudo-image. The size of the pseudo-image is H×W×C. The specific process is expressed as:

[0072]

[0073] Among them, H and W represent the height and width of the pseudo-image canvas respectively, and C represents the number of channels of the pseudo-image.

[0074] Subsequently, input the pseudo-image F i 2d into the Pyramid Mamba feature encoding and decoding module. The Pyramid Mamba feature encoding and decoding module mainly includes two parts: the Mamba downsampling encoding and the residual upsampling module RUB (Residual Upsample Block). The Mamba downsampling encoding contains the block encoding module and the Mamba module, where the Mamba module adopts the vision Mamba structure. First, it is necessary to use the block encoding module patch embedding to transform the 2D pseudo-image F i 2d into planar two-dimensional pseudo-image patches. The process is expressed as:

[0075]

[0076] Among them, PE(·) represents the block encoding module, I represents the number of pseudo-image patches, S represents the size of the pseudo-image patches, and C represents the number of channels after encoding the patches.

[0077] Subsequently, project the patches onto a vector with a hidden state dimension of D, and add the patch position encoding E pos , and the process is expressed as:

[0078]

[0079] Among them, Denote the i-th image patch in the 2D pseudo-image, Denote the learnable matrix.

[0080] The core of the original Mamba structure is to output the cls-token category token for classification tasks. However, since the present invention requires global features, after encoding is completed, all feature dimensions need to be output to enhance the representation ability of the Mamba structure for global spatial features, that is, Figure 2 of the Mamba module.

[0081] Mamba downsampling contains N loops, and the sequence T of the N-th loop n-1 is sent to the encoder, and the output sequence T n can be obtained. The process is expressed as:

[0082] T n = Mamba(T n-1 ) + T n-1

[0083] The specific process of the encoder of the Mamba module is as follows: The input block encoded feature T n-1 , in the first loop, when inputting T0, first perform a normalization operation on the feature T0, and project T0 into a higher dimension using a linear layer to generate x and z features. The process is expressed as:

[0084] T' n-1 (B, M, D) = Norm(T n-1 )

[0085] x: (B, M, E) = Linear x (T' n-1 )

[0086] z: (B, M, E) = Linear z (T' n-1 )

[0087] Among them, B represents the batch size, M represents the length of the sequence T n , D represents the hidden state dimension, E represents the state dimension after linear transformation expansion, and T' n-1 represents the sequence of the n - 1-th loop.

[0088] Subsequently, x is input into the bidirectional SSM process. The specific process is shown in the following algorithm:

[0089] Algorithm1 Bidirection SSM process

[0090] for o in {forward, backward} do

[0091] x'o :(B, M, E) ← ReLU(Conv1d o (x))

[0092]

[0093]

[0094] end

[0095] return y o

[0096] Among them, N represents the dimension of the SSM process. For the bidirectional SSM process, it is first necessary to determine the options for forward propagation and backward propagation, and then perform preprocessing operations on the input x using 1×1 convolution and the ReLU activation function to generate x'. o , and then x' o is linearly projected into the intermediate quantities B o , C o and Δ o , and finally y is obtained o , y o includes y forward and y backward ; finally, the SSM converts z into y' forward and y' backward , and the process is expressed as:

[0097] y' forward : (B, M, E) = y forward ☉ ReLU(z)

[0098] y' backward : (B, M, E) = y backward ☉ ReLU(z)

[0099] Adding y' forward and y' backward and inputting the result into the linear transformation Linear T (), and adding it to the previous moment sequence T n-1 to achieve the residual connection, and outputting the labeled sequence T n , and the process is expressed as:

[0100] T n : (B, M, D) = Linear T (y' forward + y' backward ) + T n-1

[0101] The above process completes one downsampling encoding process. The entire downsampling process needs to be looped four times to generate different scale feature blocks of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original scene, which are respectively denoted as F Mamba 1 , F Mamba 2 , F Mamba 3 and F Mamba 4 , which are used to obtain the scene receptive fields of different scales.

[0102] Since the 3D object detection task needs to obtain small-scale object features, it is necessary to fuse multi-scale scenes to obtain local spatial features. After downsampling the features using Mamba, an upsampling operation is performed on the multi-scale scene features. The traditional Mamba algorithm is mostly used for classification tasks and cannot meet the above object detection requirements. Since Mamba downsampling generates four different scale scene features, correspondingly, four upsampling encodings are required, and the k-th upsampling feature is marked as k.

[0103] In the above technical solution, the first encoding of the processing process of the residual upsampling module is expressed as:

[0104] F RUB 1 = RUB(F Mamba 1 ) + F Mamba 1

[0105] The subsequent k-th upsampling encoding is expressed as:

[0106] F RUB k = RUB(F RUB k-1 ) + F Mamba k-1

[0107] Among them, F Mamba k-1 represents the feature after Mamba downsampling encoding of the corresponding scale, F RUB k represents the feature after being processed by the RUB module, and the superscript k represents the k-th upsampling encoding process.

[0108] Finally, F single is generated after four rounds of upsampling structures, and represents the point cloud data feature of the vehicle end and the roadside end from a single perspective. Subsequently, the F single extracted from multiple vehicle perspectivesTransmitted to the ego agent, where differential feature fusion is achieved within the ego agent. The use of the residual upsampling module RUB can significantly enhance the network's stability in target detection.

[0109] (3) Vehicle-road-cloud multi-modal information differential attention fusion module

[0110] Existing attention mechanisms, such as the cross-attention mechanism, only focus on common information while ignoring the extraction and utilization of differential information. At the same time, the Transformer fusion structure cannot fully extract the common information between the vehicle and the road. The above reasons lead to low efficiency in fusing vehicle-road point cloud information.

[0111] The ego vehicle feature F is obtained from the wide-area scene feature extraction network based on Pyramid Mamba single ego and other vehicle features F single other , and through linear transformation, features Q, K, and V are obtained respectively.

[0112] The vehicle-road-cloud multi-modal information differential attention fusion module contains two attention mechanisms: the information differential attention module (DIAM) based on the cross-attention mechanism and the alternate information attention module (AIAM). As Figure 3 shown, the DIAM module calculates the differential features between the ego vehicle and other vehicles through the cross-attention mechanism, and the AIAM module effectively fuses the common features through the cross-attention information fusion mechanism.

[0113] In the specific implementation process, the DIAM module obtains the self-attention common feature C of the ego vehicle and other intelligent agent vehicles QV , that is, using the other vehicle feature Q and the ego vehicle feature K in a dot product manner to obtain the common feature matrix between the ego vehicle and other vehicles, and multiplying by the ego vehicle feature V to obtain the self-attention common feature C between the two, QV , and the process is expressed as:

[0114]

[0115] where d k is the channel number scaling factor.

[0116] Subsequently, by removing the common feature information of the self-attention common feature and the ego vehicle feature, the differential feature DI of the two can be obtained QV , and the process is expressed as:

[0117] DI QV = Linear(V - CQV )

[0118] Among them, Linear() represents a linear layer, which can enhance the feature representation ability.

[0119] Subsequently, the differential feature DI QV is injected into other vehicle features Q, and the process is expressed as:

[0120] F DIAM = MLP(LN(DI QV + Q)) + (DI QV + Q)

[0121] Among them, MLP() represents a multi-layer perceptron, and LN() represents a linear layer;

[0122] Finally, the differential feature F between the host vehicle and other intelligent agent vehicles can be obtained DIAM .

[0123] Subsequently, the cross-attention mechanism in the AIAM module is used to merge the common features of the host vehicle and other vehicles. The specific process is similar to DIAM, but the differential feature DI QV acquisition part is deleted. Specifically:

[0124] F add = C QV + Q

[0125] F AIAM = MLP(LN(F add )) + F add .

[0126] Among them, F add is an intermediate quantity.

[0127] Finally, after going through the AIAM process again, the differential feature F DIAM and the common feature F AIAM are input to generate the fusion feature M ego :

[0128] M ego = AIAM(F AIAM , F DIAM )

[0129] Traditional information fusion modules based on cross-attention mechanisms can only simply merge the common features of two images and often do not have the ability to integrate the differential features between the vehicle and the road. The final output features only contain the information of a single vehicle-road image and lack specific details from other vehicle-road modality features, and are not suitable for image fusion in the case of vehicle-road feature differences. The application of the DIAM and AIAM modules in the differential attention fusion module proposed by the present invention can effectively make up for the defects of the previous attention mechanisms.

[0130] After obtaining the fused feature M ego two 1×1 convolutional layers are respectively used to generate classification and regression prediction results and form prediction boxes. The specific process is as follows:

[0131] Y class = ξ class (M ego )

[0132] Y reg = ξ reg (M ego )

[0133] where ξ class (·) represents the classification layer, and the classification Y class outputs a score to indicate whether the object in the preselected box is an object or a background; ξ reg (·) represents the regression layer, and the regression Y reg outputs in 7 dimensions of (x, y, z, w, l, h, θ). Among them, x, y, and z respectively represent the position of the prediction box in the space coordinate system, l, w, and h respectively represent the length, width, and height of the prediction box, and θ represents the heading angle of the prediction box.

[0134] Based on the above vehicle-road-cloud collaborative intelligent vehicle environment perception system, the present invention also provides a vehicle-road-cloud collaborative intelligent vehicle environment perception method, which specifically includes the following steps:

[0135] Step 1: The ego vehicle obtains high-precision positioning information, transmits it via the 5G cellular-satellite system, and broadcasts it to other vehicles within a certain perception range;

[0136] Step 2: After other vehicles obtain the ego vehicle positioning information, they combine it with their own positioning information to jointly calculate the position transformation matrix;

[0137] Step 3: Other vehicles combine the sensor environment perception information with the position transformation matrix and transmit it to the ego vehicle via the 5G cellular-satellite system; fuse their own sensor information with the sensor information of other vehicles to form a unified position encoding among multi-vehicle information;

[0138] Step 4: Use the Pillar Feature Net data preprocessing module to convert the vehicle 3D point cloud data into a 2D pseudo-map;

[0139] Step 5: Input the 2D pseudo-map into the pyramid Mamba feature encoding and decoding module, and realize the feature extraction of multi-vehicle wide-area information through multiple block encoding modules, Mamba downsampling modules, and residual upsampling modules;

[0140] Step 6: Input the multi-vehicle wide-area feature information into the vehicle-road-cloud multi-modal information difference attention fusion module to effectively fuse the difference information in different vehicle information and form wide-area fusion information;

[0141] Step 7: Output the prediction box containing the information of surrounding vehicles.

[0142] The described embodiments are the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Without departing from the essence of the present invention, any obvious improvements, substitutions or modifications that those skilled in the art can make all fall within the protection scope of the present invention.

Claims

1. An intelligent vehicle environment perception system with vehicle-road-cloud collaboration, characterized in that, Including: Unified position encoding module for 5G cellular-satellite multimodal information fusion: In terms of information transmission, a hybrid information transmission method of 5G cellular network-low earth orbit satellite is adopted in urban areas, and a single-mode information transmission method of low earth orbit satellite is adopted in remote areas; In terms of position information encoding, a position transformation matrix is used to fuse the sensor information of the host vehicle with the sensor information of other vehicles to form a unified position encoding of multi-vehicle sensor information. Wide-area scene feature extraction network based on Pyramid Mamba, including a data preprocessing module and a Pyramid Mamba feature encoding and decoding module, where the Pyramid Mamba feature encoding and decoding module includes two processes: Mamba downsampling encoding and residual upsampling encoding; The function of the data preprocessing module is to convert the 3D point cloud data of the vehicle into a 2D pseudo-image; In Mamba downsampling encoding, first use the block encoding module to convert the 2D pseudo-image into a planar two-dimensional pseudo-image block, then project the block onto a vector with a hidden state dimension of D, and add the position encoding of each block. The position encoding is input into the Mamba module to convert the block encoding feature into a scene feature; The above process is repeated four times to generate four scene features with different scales. The four-scale scene features are respectively input into four rounds of residual upsampling encoding, and finally generate the single-view scene features of the host vehicle and other vehicles. Vehicle-road-cloud multimodal information difference attention fusion module, including an information difference attention module DIAM based on a cross-attention mechanism and an optional information attention module AIAM. The DIAM calculates the difference features between the host vehicle and other vehicles through the cross-attention mechanism, and the AIAM fuses the common features between the host vehicle and other vehicles through the cross-attention mechanism. The difference features and common features are passed through the AIAM process again to generate fusion features. The fusion features are respectively input into two convolutional layers to generate classification and regression prediction results and form prediction boxes.

2. The intelligent vehicle environment perception system for vehicle-road-cloud collaboration according to claim 1, wherein The position transformation matrix satisfies the following formula: Among them, x and y represent the position coordinates of the point cloud in the coordinate system of other vehicle sensors, and x' and y' represent the position coordinates of the point cloud in the coordinate system of the host vehicle, t x , t y represent the displacement vectors of the point cloud for translational transformation along the X-axis and Y-axis, and θ y represents the rotation angle of the point cloud for rotational transformation along the Y-axis.

3. The vehicle-road-cloud collaborative intelligent vehicle environment perception system according to claim 2, characterized in that The point cloud data of the vehicle is obtained by applying the position transformation matrix to the point cloud data collected by the sensors of the host vehicle and other vehicles.

4. The intelligent vehicle environment perception system for vehicle-road-cloud collaboration according to claim 1, characterized in that, The Mamba downsampling involves N loops, that is, the block-encoded feature T n-1 is sent to the encoder to obtain the output sequence T n , specifically: the input block-encoded feature T n-1 , in the first loop, the input feature T0. First, T0 is normalized, and then projected to a higher dimension using a linear layer to generate x and z features. Subsequently, x is input into the bidirectional SSM process; In the two-way SSM process, first, the options of forward propagation and backward propagation are judged. Subsequently, after preprocessing the x using a 1×1 convolution and a ReLU activation function, x' is generated. o , and then x' o is linearly projected into intermediate quantities B o , C o and Δ o respectively to obtain y o , where y o includes y forward and y backward . Based on y forward and y backward , finally, the SSM converts z into y' forward and y' backward ; the sum of y' forward and y' backward is input into the linear transformation Linear T (), and added to the previous moment sequence T n-1 to achieve a residual connection, and the output sequence T n is obtained.

5. The intelligent vehicle environment perception system for vehicle-road-cloud collaboration according to claim 4, characterized in that, The SSM converts z into y' forward and y' backward , specifically as follows: y′ forward :(B,M,E)=y foward ☉ReLU(z) y′ backward :(B,M,E)=y backward ☉ReLU(z) Among them, B represents the batch size, M represents the length of sequence T n , and E represents the expanded state dimension; Add y' forward and y' backward After adding them, perform a linear transformation Linear T (), and add them at the previous moment T n-1 to output the next sequence T n , specifically: T n :(B,M,D) = Linear T (y' forward + y' backward ) + T n-1 。 6. The intelligent vehicle environment perception system for vehicle-road-cloud collaboration according to claim 1, characterized in that Corresponding to the four times of Mamba downsampling encoding, the upsampling encoding is performed four times, and the k-th upsampling encoding is marked as k.

7. The vehicle-road-cloud collaborative intelligent vehicle environment perception system according to claim 6, characterized in that, The first encoding of the processing process of the residual upsampling module is expressed as: F RUB 1 = RUB(F Mamba 1 ) + F Mamba 1 The k-th encoding thereafter is expressed as: F RUB k = RUB(F RUB k-1 ) + F Mamba k-1 Among them, RUB() represents the upsampling encoding process, and F RUB 1 represents the feature after being processed by the residual upsampling module for the first time, and F Mamba 1 represents the feature block generated in the first cycle of the downsampling encoding process, and F Mamba k-1 represents the feature after Mamba downsampling encoding at the corresponding scale, and F RUB k-1 represents the feature after being processed by the residual upsampling module for the (k - 1)-th time.

8. The intelligent vehicle environment perception system for vehicle-road-cloud collaboration according to claim 1, wherein The DIAM calculates the difference features between the host vehicle and other vehicles through the attention mechanism. Specifically: Obtain the self-attention common feature C of the vehicle itself and other intelligent agent vehicles QV : Remove the common feature information of the self-attention common feature and the ego vehicle feature, and obtain the differential feature DI of the common feature and the ego vehicle feature QV : DI QV = Linear(V - CQ V ) Inject the differential feature DI QV into other vehicle features Q to obtain the differential feature F between the host vehicle and other vehicles DIAM : F DIAM = MLP(LN(DI QV + Q))+(DI QV + Q) where d k is the channel number scaling factor, Q represents other vehicle features, K and V both represent ego-vehicle features, Linear() represents a linear layer, MLP() represents a multi-layer perceptron, and LN() represents a linear layer.

9. The intelligent vehicle environment perception system for vehicle-road-cloud collaboration according to claim 8, characterized in that, The AIAM realizes the fusion of the common features between the host vehicle and other vehicles through the cross-attention information fusion mechanism. Specifically: Obtain the self-attention co-feature C of the vehicle itself and other intelligent agent vehicles QV : Common feature C QV Inject it into other vehicle features Q to obtain the common feature F between the host vehicle and other vehicles AIAM : F AIAM = MLP(LN(CQ V + Q))+(CQ V + Q) Finally, after going through the AIAM process one more time, the differential feature F DIAM and the common feature F AIAM generate the fused feature M ego : M ego = AIAM(F AIAM , F DIAM ).

10. A method for an intelligent vehicle environment perception system with vehicle-road-cloud collaboration according to any one of claims 1-9, characterized in that: Use the position transformation matrix to fuse the sensor information of the host vehicle with the sensor information of other vehicles to form a unified position encoding of multi-vehicle sensor information. Use the data preprocessing module to convert the 3D point cloud data of the vehicle into a 2D pseudo-image. Input the 2D pseudo-image into the Pyramid Mamba feature encoding and decoding module, and through multiple layers of block encoding modules, Mamba modules and residual upsampling modules, realize the feature extraction of multi-vehicle wide-area information. Input the multi-vehicle wide-area feature information into the vehicle-road-cloud multi-modal information difference attention fusion module to effectively fuse the difference information in the environmental information perceived by different vehicles and form wide-area fusion features. Output prediction boxes containing information about surrounding vehicles.