Air-ground cooperative intelligent connected vehicle complex intersection environment perception method and system
Patent Information
- Application Number
- CN202510531107.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2045-04-25
AI Technical Summary
在传统的单车感知中,由于远距离区域点云数据稀疏以及相机图像深度估计固有的误差,单车往往只能部分检测到附近目标的遮挡情况
[0028] The beneficial effects of this invention are as follows: The air-ground cooperative intelligent connected vehicle complex intersection environment perception method of this invention can significantly improve the perception capability and safety of intelligent connected vehicles in complex intersection environments. Firstly, environmental features are collected through ground-based multi-vehicle perception devices, and combined with low-orbit satellite constellation network wireless communication technology to achieve efficient data transmission and on-orbit information fusion with the cloud server platform, enabling real-time information sharing with other intelligent agents. This effectively overcomes the limitations of single-vehicle intelligent agents in complex intersection environments when facing occlusion and long-distance target detection tasks. Furthermore, by combining the top-view features provided by remote sensing satellites, the panoramic outline of complex intersections is comprehensively supplemented, further enhancing the completeness of perception. The cloud server platform utilizes a lightweight Transformer model based on an attention mechanism to dynamically fuse heterogeneous modal features from the vehicle end, and accurately aligns the BEV features fused by vehicle-end sensors with the satellite top-view features through a flow field displacement mapping model to generate air-road fusion features, optimizing the fusion effect of information from different sources. The introduction of a temporal cross-attention mechanism further enhances the perception capability of dynamic targets in complex intersections. Ultimately, by transmitting the fusion results in real time via low-orbit satellite communication, a comprehensive and accurate perception field of vision is provided for ground vehicles, significantly improving the perception accuracy, efficiency, and robustness in complex intersection environments, ensuring the safety and reliability of intelligent connected vehicles, and comprehensively enhancing the driving experience and traffic safety.
Smart Images

Figure CN120431545B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of vehicle engineering and transportation engineering, and specifically relates to a method for intelligent connected vehicles to perceive complex intersection environments using air-ground collaborative technology. Background Technology
[0002] Intelligent connected vehicles can improve urban traffic efficiency and reduce energy consumption. Achieving reliable and safe autonomous driving hinges on accurately understanding the surrounding environment—that is, comprehensively perceiving dynamic traffic scenarios. With the development of deep learning and computer vision, environmental perception technology for intelligent connected vehicles has made significant progress. However, widespread deployment of driverless cars has not yet been achieved. Currently, one of the main reasons hindering high-level intelligent driving is the insufficient accuracy of single-vehicle perception in estimating environmental states in complex traffic scenarios. Specifically, single-vehicle perception is often affected by occlusion when facing complex environments (such as complex urban intersections). Secondly, onboard sensors have physical limitations when perceiving distant targets. Furthermore, noise from single-vehicle sensors also degrades the performance of the perception system. These problems make it difficult for existing methods to simultaneously meet the requirements of high accuracy, high reliability, and high safety in real-world driving scenarios.
[0003] Air-ground cooperative perception is a multi-agent collaborative system where agents share sensory information to compensate for the limitations of single-vehicle perception and effectively overcome problems such as limited visual range. Low-Earth orbit (LEO) satellite constellation networks play a crucial role in this scenario. By deploying dense satellite arrays in low Earth orbit, LEO satellites provide seamless global communication coverage for connected vehicles, ensuring stable communication connections even in remote areas, highways, or extreme environments. With the continuous development of LEO satellite constellation technology, efficient interconnection between "low-altitude, low-orbit" and "ground" has become possible. In traditional single-vehicle perception, due to the sparse point cloud data in distant areas and inherent errors in camera image depth estimation, a single vehicle can often only partially detect occlusion of nearby targets. However, in air-ground cooperative perception scenarios, through the seamless integration of "ground + low-altitude + LEO," the main vehicle can receive sensory information from other agents, thereby significantly expanding its field of view. This collaborative approach enables the main vehicle to not only accurately detect distant targets and occluded objects, but also significantly improve detection accuracy in complex and dynamic traffic scenarios (such as complex urban intersections), providing strong support for the safety and efficiency of intelligent connected vehicles.
[0004] How to effectively utilize the integrated communication network of vehicles, roads, satellites, and the cloud to build a complete and efficient perception and information interaction system has become a key issue in improving the safety and comfort of high-level autonomous driving. Autonomous driving, commercial spaceflight, and the low-altitude economy are strategic emerging industries in various countries and important representatives of new productive forces. Building a future-oriented integrated space-ground mobility solution through seamless integration of "ground + low-altitude + low-orbit" will not only optimize autonomous driving technology but also comprehensively promote the upgrading and transformation of the overall efficiency of intelligent transportation systems. Summary of the Invention
[0005] In view of this, the present invention provides an air-ground collaborative intelligent connected vehicle complex intersection environment perception method. Through the seamless connection of "ground + low air + low orbit", it realizes information interconnection and interoperability between vehicles, roads, cloud and low orbit satellites, multi-domain information sharing expands the perception field of the main vehicle, overcomes the visual limitations of single-vehicle perception, and improves the detection performance of long-distance and occluded targets.
[0006] The present invention achieves the above-mentioned technical objectives through the following technical means.
[0007] Air-ground collaborative intelligent connected vehicle environment perception method at complex intersections:
[0008] The vehicle extracts the initial latent features from the camera-perceived images and the LiDAR-perceived point clouds.
[0009] Initial potential features are transmitted to a cloud server platform using a low-Earth orbit satellite constellation network; simultaneously, satellite overhead view features are obtained from satellite imagery.
[0010] The cloud server platform first dynamically fuses heterogeneous modal features from the vehicle end through a lightweight attention mechanism to obtain vehicle-side sensor-fused BEV features. Then, it fuses satellite top-view features and vehicle-side sensor-fused BEV features to generate air-road fusion features. Finally, it uses a temporal feature fusion encoder to consider temporal information and generate a comprehensive representation of dynamic targets or attributes from the air-road fusion features.
[0011] The results of air-ground collaborative perception from the cloud service platform are transmitted to vehicles traveling at complex intersections via low-orbit satellites.
[0012] A further technical solution involves the initial latent feature extraction of camera-perceived images, including: extracting 2D image features from images captured by surround-view cameras set around the vehicle through a shared backbone network, predicting depth features through a deep network, multiplying the image features and depth features element-wise to obtain initial view frustum features, and then flattening them to generate BEV image features that integrate rich semantic information and accurate depth information.
[0013] A further technical solution involves dynamically fusing heterogeneous modal features from the vehicle side using a lightweight attention mechanism. Specifically, this involves stitching together BEV image features and BEV point cloud features from different vehicles, and then fusing them using an element-level maximization operation to obtain multi-vehicle camera fused features. Multi-vehicle radar fusion features Then, and Aligning along the dimensions, we obtain the following: and Obtained from the Sigmoid function and The weights are used to adaptively fuse the data. and The fused features are used as queries in the attention mechanism. As values and keys in the attention mechanism, cross-modal attention is cross-fused to obtain BEV features fused from vehicle-side sensors. .
[0014] A further technical solution involves fusing satellite top-down view features and vehicle-mounted sensor BEV features. Specifically, this is achieved by using several convolutions to fuse the vehicle-mounted sensor BEV features. Generate a flow field and fuse BEV features from vehicle-side sensors. To achieve "flow," bilinear interpolation is used to spatially align BEV features fused from vehicle-side sensors with satellite top-down view features, resulting in aligned features. Finally and satellite top-down view features The features are then spliced together to generate the air-path fusion feature F.
[0015] A further technical solution employs a temporal feature fusion encoder, considering temporal information to generate a comprehensive representation of dynamic targets or attributes from the spatial fusion features. Specifically, the temporal feature fusion encoder consists of three layers, each containing a temporal cross-attention network and a feedforward network. In the first layer, a query is obtained by initializing the spatial fusion features of the current frame. Then, a temporal cross-attention mechanism is used to recursively establish a correspondence with the spatial fusion features of the previous frame. The generated query is updated through the feedforward network and used as the input to the next layer. After three layers of fusion encoding, the temporal fusion BEV feature is obtained. The The cloud service platform uses a central heatmap head to predict the center position of all objects and estimates the size, rotation, and velocity of the objects through regression.
[0016] A further technical solution uses a temporal cross-attention mechanism to recursively establish a correspondence with the air path fusion features of the previous frame. Specifically, the query features are repeatedly passed at each layer, the information of the current frame and the previous frame are compared frame by frame, and the correlation between each frame is dynamically adjusted by calculating attention weights.
[0017] A further technical solution is the BEV image features. Where f represents voxel pooling and d represents depth features. Representing image features, This indicates the outer product operation.
[0018] A further technical solution, after alignment, has the following features:
[0019]
[0020] Where W' represents the width of the flow field, H' represents the height of the flow field, and (h,w) represents the learned 2D transformation offset at the feature map location. Used to index satellite top-view features at the height and width positions. The eigenvalue at the index, △ hw1 and △ hw2 This represents the 2D transformation offset learned at the feature map location.
[0021] A further advanced technical solution: time-fusion BEV features ,in, This represents the air-path fusion characteristics at time t. DefAttn represents the air path fusion feature at time t-1, and DefAttn represents temporal cross attention.
[0022] An air-ground collaborative intelligent connected vehicle environment perception system for complex intersections includes:
[0023] The vehicle-side sensor feature extraction module extracts the initial potential features from the camera and lidar sensing information;
[0024] The low-orbit satellite communication and feature extraction module is used to transmit initial potential features to the cloud server platform and obtain satellite top-down view features;
[0025] A cloud-based heterogeneous feature fusion module performs heterogeneous modal feature fusion for vehicle terminals;
[0026] The cloud-based air-road feature fusion module fuses satellite overhead view features and vehicle-side sensor BEV features.
[0027] The cloud-based temporal feature fusion module uses a temporal feature fusion encoder to generate a comprehensive representation of dynamic targets or attributes from the air-path fusion features.
[0028] The beneficial effects of this invention are as follows: The air-ground cooperative intelligent connected vehicle complex intersection environment perception method of this invention can significantly improve the perception capability and safety of intelligent connected vehicles in complex intersection environments. Firstly, environmental features are collected through ground-based multi-vehicle perception devices, and combined with low-orbit satellite constellation network wireless communication technology to achieve efficient data transmission and on-orbit information fusion with the cloud server platform, enabling real-time information sharing with other intelligent agents. This effectively overcomes the limitations of single-vehicle intelligent agents in complex intersection environments when facing occlusion and long-distance target detection tasks. Furthermore, by combining the top-view features provided by remote sensing satellites, the panoramic outline of complex intersections is comprehensively supplemented, further enhancing the completeness of perception. The cloud server platform utilizes a lightweight Transformer model based on an attention mechanism to dynamically fuse heterogeneous modal features from the vehicle end, and accurately aligns the BEV features fused by vehicle-end sensors with the satellite top-view features through a flow field displacement mapping model to generate air-road fusion features, optimizing the fusion effect of information from different sources. The introduction of a temporal cross-attention mechanism further enhances the perception capability of dynamic targets in complex intersections. Ultimately, by transmitting the fusion results in real time via low-orbit satellite communication, a comprehensive and accurate perception field of vision is provided for ground vehicles, significantly improving the perception accuracy, efficiency, and robustness in complex intersection environments, ensuring the safety and reliability of intelligent connected vehicles, and comprehensively enhancing the driving experience and traffic safety. Attached Figure Description
[0029] Figure 1 This is a diagram of the intelligent connected vehicle complex intersection environment perception architecture based on air-ground cooperation proposed in this invention.
[0030] Figure 2 This is a diagram of the BEV feature alignment and fusion model for the cloud server platform of this invention;
[0031] Figure 3 This is a time-series feature fusion model diagram of the cloud server platform of the present invention. Detailed Implementation
[0032] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to embodiments and accompanying drawings. The content mentioned in the embodiments is not intended to limit the present invention.
[0033] This invention discloses an air-ground cooperative intelligent connected vehicle complex intersection environment perception system, the overall network architecture of which is as follows: Figure 1 As shown, it includes a vehicle-side sensor feature extraction module, a low-orbit satellite communication and feature extraction module, a cloud-based heterogeneous feature fusion module, a cloud-based air-to-ground feature fusion module, and a cloud-based time-series feature fusion module.
[0034] The vehicle-side sensor feature extraction module extracts the initial potential features from the camera and lidar sensing information.
[0035] The low-orbit satellite communication and feature extraction module is designed with a feature extractor for processing satellite images to obtain satellite top-down view features.
[0036] The cloud-based heterogeneous feature fusion module designs a learnable vehicle terminal heterogeneous modal feature fusion model to perform heterogeneous modal feature fusion of vehicle terminals.
[0037] The cloud-based air-road feature fusion module designs a BEV feature alignment and fusion model to fuse satellite top-view features and vehicle-side sensor-fused BEV features.
[0038] The cloud-based temporal feature fusion module specifically employs a temporal feature fusion encoder, which takes into account temporal information and generates a comprehensive representation of dynamic targets or attributes from the air-path fusion features.
[0039] This invention discloses a method for intelligent connected vehicles to perceive complex intersection environments using a coordinated air-ground approach. First, initial latent features are extracted from camera and LiDAR data on the vehicle. Then, multimodal features are uploaded using a low-Earth orbit (LEO) satellite network, supplementing the overhead view traffic environment perception with remote sensing satellites. A cloud server platform dynamically fuses heterogeneous multimodal features from the vehicle using a lightweight attention mechanism, designs a flow field displacement mapping principle to address the spatial misalignment between vehicle and remote sensing satellite features, and further designs a temporal feature fusion encoder to enhance the perception of dynamic targets. Finally, the collaborative perception results are transmitted to the vehicle in real time via LEO satellite communication, providing an accurate and comprehensive environmental perception field of view, significantly improving perception accuracy and driving safety in complex intersection environments. The specific implementation of this invention includes the following steps:
[0040] Step 1: Extract initial potential features from heterogeneous sensor information at the vehicle end.
[0041] As a crucial component of air-ground collaboration, vehicle-mounted terminals are equipped with various environmental perception sensors, including cameras and LiDAR, for efficient perception and comprehensive description of complex intersection environments. Different sensors reflect environmental information through their unique data formats; for example, cameras record the surrounding scene using RGB images, while LiDAR captures spatial features of the environment through high-resolution 3D point cloud data. Furthermore, different vehicle models may be equipped with different types of environmental perception sensors. In vehicle-mounted terminals, to accommodate this diverse sensor configuration and the data formats they generate, feature extraction of multimodal information needs to be designed as a relatively independent processing flow. This independence not only ensures that the feature extraction processes of different sensors do not interfere with each other but also lays the foundation for the vehicle-mounted terminal to flexibly adjust its interaction mode with the ground-based or satellite-based systems.
[0042] Camera image feature extraction: vehicle-mounted terminal reception The images are taken by surround-view cameras placed around the car, forming a collection of images. , Indicates by the first Images captured by a surround-view camera. The surround-view images are then processed by a shared backbone network to extract 2D image features. A deep network predicts deep features In this dimension representation, H, W, and C represent the height, width, and number of channels of the 2D image feature, respectively, while D represents the number of depth features in the discrete interval. To make image features... with depth features To align in dimensions, first, expand both elements; then, combine the image features... with depth features Element-wise multiplication yields the initial frustum features. Finally, through flattening, BEV image features that integrate rich semantic information and accurate depth information are generated. :
[0043] (1)
[0044] Where f represents voxel pooling and d represents depth features. This indicates the outer product operation.
[0045] LiDAR point cloud feature extraction: Following the voxelization method of sparse 3D convolutional networks, the LiDAR point cloud is processed to extract 3D features and fuse multi-scale features from different stages. The point cloud acquired by the LiDAR is... Composed of points: Each point is represented as a four-dimensional vector. Where x, y, and z are the coordinates of the point on the X, Y, and Z axes in the vehicle coordinate system, respectively. This refers to reflection intensity. The vehicle-mounted terminal first dynamically voxelizes the original point cloud, then sequentially applies multiple sparse 3D convolutional layers to the voxelized features to generate multi-scale 3D features. To enhance the LiDAR's ability to capture multi-scale objects, this invention introduces a multi-scale feature fusion strategy: 3D convolution is used to compress the multi-scale 3D features along the Z-axis, increasing the receptive field in the height direction, and the compressed multi-scale 3D features are converted into BEV features. Thus, multi-scale 3D features from different stages are transformed into multiple 2D BEV features. Finally, two-dimensional convolution is used to extract features from the BEV features, resulting in a dense BEV point cloud feature. The above process can be represented as:
[0046] (2)
[0047] in, MapToBEV represents a 2D convolution, MapToBEV represents the transformation of multi-scale 3D features into 2D BEV features, and Spconv represents a sparse 3D convolutional layer. This indicates a cascading operation.
[0048] Step 2: Use low-Earth orbit satellite constellation network communication to transmit the multimodal latent features extracted in Step 1; and obtain satellite top-down view features.
[0049] This invention provides reliable communication and accurate time synchronization between the vehicle, the cloud, and satellites through a low-Earth orbit satellite constellation network. The constellation network also provides additional data support for remote sensing satellite perception.
[0050] After the onboard terminal completes the initial extraction of multimodal potential features in step one, it feeds back the data to the low-Earth orbit satellite constellation via uplink. Information fusion is then performed on the cloud server platform, achieving collaborative perception through heterogeneous information fusion. This on-orbit information fusion mode not only effectively reduces the computing power consumption of the onboard terminal and lowers its dependence on local hardware resources, but also leverages the powerful computing capabilities and global perspective of the cloud to improve the accuracy and efficiency of perception information fusion.
[0051] Before uploading the multimodal latent features extracted by the vehicle-mounted terminal to the low-Earth orbit satellite constellation, the data must be effectively compressed and encoded to optimize transmission efficiency and ensure data integrity and reliability. This invention combines lossy compression (such as autoencoders, variational autoencoders, and convolutional neural networks) and lossless compression (such as sparse matrix compression), achieving efficient data compression while preserving key feature information. Simultaneously, a suitable channel coding scheme is selected based on the characteristics of the satellite communication link (bandwidth, noise level), such as LDPC, which is suitable for efficient error correction.
[0052] After data compression and encoding, the next step is to fragment and package the data to ensure its orderliness and integrity during satellite transmission. The specific process is as follows: First, add header information, including packet number, checksum, timestamp, and protocol identifier. The packet number identifies the order of each packet, facilitating data reassembly at the receiving end; the checksum (such as cyclic redundancy check CRC) detects whether errors occurred during transmission; the timestamp records the time of data generation or transmission, facilitating synchronization and timing management; and the protocol identifier indicates the transmission protocol or service type to which the packet belongs, ensuring correct processing and routing. Second, select the transmission protocol, choosing an appropriate protocol layer based on the characteristics of satellite communication. The physical layer defines the modulation method and frequency of the signal; the data link layer uses protocols such as HDLC (High-Level DataLink Control) or CCSDS (Consultative Committee for Space Data Systems) to ensure reliable packet transmission; and the network layer uses the IP protocol for data routing and address management. Then, perform data encapsulation, packaging the fragmented data and header information according to the selected protocol to generate a complete data packet. In this process, the bandwidth occupied by the header information can be reduced by compressing the header information, and multiple small data packets can be encapsulated together by batch packaging to further reduce transmission overhead.
[0053] Through the aforementioned compression, encoding, fragmentation, and packaging processes, the vehicle-mounted terminal can efficiently and systematically transmit the BEV feature data, processed by the vehicle's algorithm, back to the low-Earth orbit satellite constellation. This not only optimizes the bandwidth utilization of data transmission but also ensures the integrity and reliability of the data in complex satellite communication environments, providing a solid data foundation for collaborative perception under heterogeneous information fusion.
[0054] In the context of integrated space-ground communication for intelligent connected vehicles, satellite networks not only need to support communication but also need to handle perception functions, which are typically handled by remote sensing satellites. Satellite imagery provides a comprehensive overview of traffic routes from a top-down perspective, effectively compensating for the limitations of vehicle-mounted perception in sensing the overall shape of complex intersections and capturing obscured areas within intersections. This invention deploys a feature extractor to process satellite images. The feature extractor uses a ResNet18 backbone network, with satellite images as input and satellite top-down view features as output. Since satellite imagery inherently provides a top-down view, it does not require the extraction of complex initial latent features as on-board platforms, and can be directly applied to cross-view feature fusion on cloud service platforms.
[0055] Step 3: Perform heterogeneous modal feature fusion of vehicle terminals on the cloud server platform.
[0056] In the air-ground cooperative intelligent connected vehicle complex intersection environment perception method of the present invention, the cloud server platform integrates multi-source heterogeneous data from the vehicle end to achieve efficient fusion and deep perception of multimodal information, thereby improving the accuracy and robustness of complex intersection environment perception. The cloud server platform receives multimodal latent features from the vehicle terminal and satellite top-view features from the low-Earth orbit satellite constellation, and realizes real-time data upload and synchronization through the high-speed communication link of the low-Earth orbit satellite.
[0057] Because isomorphic (camera or radar) features originate from the same data source and possess shared and similar BEV features, this invention, on a cloud server platform, first performs simple stitching of isomorphic features from different vehicles, then employs element-level maximization operations to fuse the isomorphic BEV features from different vehicles, resulting in multi-vehicle camera fusion features. Multi-vehicle radar fusion features It can effectively express observation information from the same sensors of multiple vehicles at complex traffic intersections, which also lays the technical foundation for subsequent heterogeneous feature fusion.
[0058] To further integrate multimodal data from the vehicle, this invention designs a learnable heterogeneous modal feature fusion model for in-vehicle terminals. This model overcomes the limitations of simple feature fusion methods by dynamically adjusting the importance of different modal features using an attention mechanism. The Transformer used in this invention is a deep learning model based on an attention mechanism. Unlike traditional Transformer architectures, this invention does not use a multi-head attention module in the Transformer structure to achieve a lightweight design for cloud service platforms.
[0059] To simultaneously process the features fused from multimodal BEVs, this invention uses convolution to initially align the features fused from different BEV modalities in terms of dimensions, obtaining the following results: and In order to dynamically adjust the confidence level of different modes in complex intersections with varying environments, this invention obtains the confidence level from the Sigmoid function. and The weights are then determined; subsequently, the mixture is adaptively fused based on these weights. and The fused features are used as queries in subsequent attention mechanisms. These will serve as values and keys in subsequent attention mechanisms, enabling cross-modal attention fusion. The above process can be expressed by the following formula:
[0060] (3)
[0061] (4)
[0062] (5)
[0063] (6)
[0064] (7)
[0065] (8)
[0066] Here, query, key, and value represent the query, key, and value in the attention mechanism, respectively; , and W q These are learnable parameters during the training of the attention mechanism; Indicates a cross-merging operation; This represents the aligned image features. The aligned LiDAR point cloud features are represented by σ, where σ represents the Sigmoid function, conv represents the convolution operation, and concat represents the concatenation operation. This represents the fused feature after cross-modal attention optimization.
[0067] Step 4: Fuse satellite overhead view features and vehicle-mounted sensor BEV features on the cloud server platform.
[0068] The satellite top-down view features obtained in steps two and three respectively Fusion of BEV features with vehicle-side sensors Since these are images from the BEV (Battery Electric Vehicle) perspective, current methods directly stack and fuse multimodal feature maps. However, these heterogeneous features often suffer from misalignment, making them difficult for the network to understand. This is because there may still be differences between satellite maps and the actually generated BEV features, mainly due to positioning errors. Therefore, this invention designs a BEV feature alignment and fusion model in a cloud server platform, such as... Figure 2 As shown.
[0069] The model takes BEV features fused from vehicle-side sensors and satellite overhead view features as input, and performs several convolutions to process the BEV features fused from vehicle-side sensors. Generate a flow field (W' represents the width of the flow field, and H' represents the height of the flow field). Based on the principle of flow field displacement mapping, BEV features are fused from vehicle-side sensors. To achieve "flow," bilinear interpolation is used to spatially align BEV features fused from vehicle-side sensors with satellite overhead view features. The formula is as follows:
[0070] (9)
[0071] in, This represents the aligned eigenvalues on the feature map. Used to index satellite top-down features at height and width positions, where (h,w) represents the learned 2D transformation offset at the feature map position. The eigenvalue at the index, △ hw1 and △ hw2 This represents the 2D transformation offset learned at the feature map location.
[0072] Then and The features are then spliced together to generate the final air-path fusion feature F.
[0073] Step 5: Perform time-series feature fusion on the cloud server platform.
[0074] Performing time-series feature fusion on a cloud server platform aims to enhance the perception of dynamic targets or attributes at complex traffic intersections by integrating historical information. For example... Figure 3 As shown, the temporal feature fusion encoder consists of three layers, each containing a temporal cross-attention network and a feedforward network. In the first layer, the query is initialized from the spatial fusion features of the current frame, and then a temporal cross-attention mechanism is used to recursively establish a correspondence with the spatial fusion features of the previous frame. Specifically, by repeatedly passing the query features in each layer, the information of the current frame and the previous frame are compared frame by frame, and then the correlation between each frame is dynamically adjusted by calculating attention weights. This invention uses an attention mechanism to adaptively learn the receptive field, thereby effectively capturing the salient features of moving targets in the spatial fusion features. The generated query is updated through the feedforward network and used as the input to the next layer. After three layers of fusion encoding, the temporal fusion BEV features are finally obtained:
[0075] (10)
[0076] in, This represents the air-path fusion characteristics at time t. This represents the air-path fusion characteristics at time t-1. DefAttn represents the air path fusion feature after three-layer fusion encoding at time t, and DefAttn represents temporal cross attention.
[0077] This process combines air-to-ground fusion features and considers temporal information to generate a comprehensive representation of dynamic targets or attributes. The fusion process integrates relevant information from historical features with the current perceptual input, which can help detect distant, near, or occluded targets at complex intersections.
[0078] The air path fusion feature of this invention after three-layer fusion encoding The cloud service platform uses a central heatmap head to predict the center position of all objects and estimates the size, rotation, and velocity of the objects through regression.
[0079] Finally, this invention uses low-orbit satellite communication to transmit the air-ground collaborative perception results (center position, size, rotation, and speed of objects) from the cloud service platform to each vehicle driving at a complex intersection, thereby obtaining a more comprehensive perception field of view and significantly improving the user's driving experience and traffic safety.
[0080] The embodiments described above are preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Any obvious improvements, substitutions or modifications that can be made by those skilled in the art without departing from the essence of the present invention shall fall within the protection scope of the present invention.
Claims
1. A method for intelligent connected vehicles to perceive complex intersection environments using air-ground collaborative technology, characterized by: The vehicle extracts the initial latent features from the camera-perceived images and the LiDAR-perceived point clouds. Initial potential features are transmitted to a cloud server platform using a low-Earth orbit satellite constellation network; simultaneously, satellite overhead view features are obtained from satellite imagery. The cloud server platform first dynamically fuses heterogeneous modal features from the vehicle end through a lightweight attention mechanism to obtain vehicle-side sensor-fused BEV features. Then, it fuses satellite top-view features and vehicle-side sensor-fused BEV features to generate air-road fusion features. Finally, it uses a temporal feature fusion encoder to consider temporal information and generate a comprehensive representation of dynamic targets or attributes from the air-road fusion features. The results of air-ground collaborative perception from the cloud service platform are transmitted to vehicles traveling at complex intersections via low-orbit satellites. A lightweight attention mechanism is used to dynamically fuse heterogeneous modal features from the vehicle side. Specifically, BEV image features and BEV point cloud features from different vehicles are stitched together separately, and then fused using an element-level maximization operation to obtain multi-vehicle camera fusion features. Multi-vehicle radar fusion features Then, and Aligning along the dimensions, we obtain the following: and Obtained from the Sigmoid function and The weights are used to adaptively fuse the data. and The fused features are used as queries in the attention mechanism. As values and keys in the attention mechanism, cross-modal attention is cross-fused to obtain BEV features fused from vehicle-side sensors. ; The satellite top-view features and vehicle-mounted sensor BEV features are fused together. Specifically, the vehicle-mounted sensor BEV features are fused through several convolutions. A flow field is generated to "flow" the BEV features fused from the vehicle-side sensors. Bilinear interpolation is used to spatially align the BEV features fused from the vehicle-side sensors with the satellite top-view features, resulting in the aligned features. Finally and satellite top-down view features The data is then spliced together to generate the air-path fusion feature F; A temporal feature fusion encoder is employed, taking temporal information into account, to generate a comprehensive representation of dynamic targets or attributes from the spatial fusion features. Specifically, the temporal feature fusion encoder consists of three layers, each containing a temporal cross-attention network and a feedforward network. In the first layer, a query is obtained by initializing the spatial fusion features of the current frame. Then, a temporal cross-attention mechanism is used to recursively establish a correspondence with the spatial fusion features of the previous frame. The generated query is updated through the feedforward network and used as the input to the next layer. After three layers of fusion encoding, the temporal fusion BEV feature is obtained. The The cloud service platform uses a central heatmap head to predict the center position of all objects and estimates the size, rotation, and velocity of the objects through regression.
2. The air-ground cooperative intelligent connected vehicle complex intersection environment perception method according to claim 1, characterized in that, The initial latent feature extraction of the camera-perceived image includes: 2D image features are extracted from images taken by surround-view cameras set around the vehicle through a shared backbone network, depth features are predicted through a deep network, the image features and depth features are multiplied element-wise to obtain the initial view frustum features, and then flattened to generate BEV image features that integrate rich semantic information and accurate depth information.
3. The air-ground cooperative intelligent connected vehicle complex intersection environment perception method according to claim 1, characterized in that, The temporal cross-attention mechanism is used to recursively establish a correspondence with the air path fusion features of the previous frame. Specifically, the query features are repeatedly passed in each layer, the information of the current frame and the previous frame are compared frame by frame, and the correlation between each frame is dynamically adjusted by calculating attention weights.
4. The air-ground cooperative intelligent connected vehicle complex intersection environment perception method according to claim 2, characterized in that, The BEV image features Where f represents voxel pooling and d represents depth features. Representing image features, This indicates the outer product operation.
5. The air-ground cooperative intelligent connected vehicle complex intersection environment perception method according to claim 1, characterized in that, The aligned features are: Where W' represents the width of the flow field, H' represents the height of the flow field, and (h,w) represents the learned 2D transformation offset at the feature map location. Used to index satellite top-view features at the height and width positions. The eigenvalue at the index, △ hw1 and △ hw2 This represents the 2D transformation offset learned at the feature map location.
6. The air-ground cooperative intelligent connected vehicle complex intersection environment perception method according to claim 1, characterized in that, Time-fusion BEV features ,in, This represents the air-path fusion characteristics at time t. DefAttn represents the air path fusion feature at time t-1, and DefAttn represents temporal cross attention.
7. A system for implementing the air-ground cooperative intelligent connected vehicle complex intersection environment perception method according to any one of claims 1-6, characterized in that, include: The vehicle-side sensor feature extraction module extracts the initial potential features from the camera and lidar sensing information; The low-orbit satellite communication and feature extraction module is used to transmit initial potential features to the cloud server platform and obtain satellite top-down view features; A cloud-based heterogeneous feature fusion module performs heterogeneous modal feature fusion for vehicle terminals; The cloud-based air-road feature fusion module fuses satellite overhead view features and vehicle-side sensor BEV features. The cloud-based temporal feature fusion module uses a temporal feature fusion encoder to generate a comprehensive representation of dynamic targets or attributes from the air-path fusion features.