Self-driving automobile wide-area driving area detection method based on space-ground integrated communication

Through integrated communication in the world and multi-vehicle collaborative perception technology, the problem of insufficient detection of driving areas in ramp junction and occlusion scenarios of autonomous vehicles is solved, high-precision and low-latency traffic environment perception is achieved, and vehicle traffic safety and system robustness are improved.

CN120279513AActive Publication Date: 2025-07-08JIANGSU UNIV

Patent Information

Application Number
CN202510334518.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-08
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

In the high-speed ramp junction and occlusion scenarios, the driving area and lane line detection capabilities of autonomous vehicles are insufficient. The existing technology has problems such as limited field of view, poor lighting, high computational complexity, and poor real-time performance, which affects the safety of vehicle traffic.

Method used

The wide-area driving area detection method of autonomous driving vehicles based on integrated communication is adopted. Through collaborative vehicles to acquire single-view 2D images, low-orbit satellites transmit compressed feature tensors, and multi-vehicle features are fused with high-precision maps and spatial modulation cross attention mechanisms to generate high-precision semantic segmentation maps to achieve accurate identification of roads, vehicles, pedestrians and obstacles.

Benefits of technology

It significantly improves the system's real-time response capability and perception accuracy, reduces communication delay and computing complexity, enhances the perception robustness and stability in complex traffic environments, and supports efficient multi-vehicle collaborative perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279513A_ABST
    Figure CN120279513A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic driving vehicle wide-area driving area detection method based on space-ground integrated communication, which comprises the following steps: processing a single-view 2D image collected by a cooperative vehicle to obtain a feature tensor with a classification result, compressing the feature tensor in a channel dimension, transmitting the compressed feature tensor through a low-orbit satellite, sending the compressed feature tensor to the vehicle, and combining lane information to detect the wide-area driving area of the automatic driving vehicle. A self vehicle realizes feature fusion among multiple vehicles by using a spatial modulation cross attention mechanism to obtain global features, reconstructs the global features into a spatial feature map through a linear mapping layer, obtains a multi-scale fused feature map through an up-sampling operation, finally generates a category probability map through a convolution operation, obtains a single-channel semantic segmentation map through a pixel-by-pixel argmax operation, and finally obtains a multi-scale fused feature map through an up-sampling operation. And carrying out refined classification on each pixel in the single-view 2D image through a task head, and identifying roads, vehicles, pedestrians and obstacles. Under the condition of not depending on a large amount of annotated data, the overall perception performance is improved, and a new scheme is provided for further development of the intelligent network connection automobile perception technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent autonomous vehicle environment perception, and particularly designs a wide-area driving area detection method for autonomous vehicles based on space-ground integrated communication. Background Technique

[0002] With the rapid development of intelligent driving technology, autonomous vehicles are gradually becoming an important part of modern transportation. However, in the scenarios of highway ramp merging and occlusion, the safe driving of vehicles still faces many challenges. The technical problems in this scenario mainly focus on the insufficient detection capabilities of the driving area and lane lines, which are usually manifested in aspects such as short sight distance and limited field of view, directly affecting the traffic safety of vehicles. In response to these problems, there are currently various technical solutions, including vision detection based on cameras, lidar, high-precision map matching, millimeter-wave radar and vision fusion, and vehicle-road collaboration (V2X) technology, but each of these solutions has its own deficiencies.

[0003] Vision detection technology often adopts deep learning methods based on convolutional neural networks (CNNs), such as Faster R-CNN, YOLO, and SSD, to detect lane lines and obstacles through image feature extraction. However, in the scenarios of ramp merging and occlusion, the limited field of view of the camera or poor lighting will affect the detection accuracy, especially under night or strong backlight conditions.

[0004] Lidar provides high-precision environmental modeling through three-dimensional point cloud data, and deep learning methods such as PointNet and VoxelNet can accurately identify ramps and obstacles. However, the high cost and data processing complexity of lidar limit its popularization, and its detection accuracy and range will be affected under high-speed or adverse weather conditions.

[0005] The fusion of millimeter-wave radar and vision technology, such as DeepFusion and HVNet, improves the detection robustness by combining distance, speed, and semantic information. However, this method is computationally complex and has poor real-time performance, especially in high-speed dynamic scenarios, which may lead to insufficient response speed.

[0006] Vehicle-to-Everything (V2X) technology enables communication between vehicles and infrastructure, providing real-time road information for autonomous driving systems. This technology integrates various sensors on the vehicle and roadside ends, greatly enhancing the vehicle's perception ability in complex road conditions and occlusion scenarios. The V2X system mainly relies on in-vehicle communication modules, edge computing, and cloud computing platforms. Common technical standards include C-V2X (Cellular Vehicle-to-Everything) and DSRC (Dedicated Short-Range Communication). These technologies can share traffic information in real time, such as road conditions, traffic signal states, and dynamic information of surrounding vehicles, thereby improving driving safety and traffic efficiency.

[0007] In terms of the V2X network structure, in recent years, many studies have been dedicated to optimizing the network architecture to meet the requirements of intelligent driving for low latency and high reliability. For example, adopting a V2X architecture based on network slicing can divide the communication network into multiple virtual networks, optimize them according to different application requirements, and ensure the communication efficiency of autonomous driving vehicles in high-density environments. In addition, some deep learning methods, such as V2X communication optimization strategies based on graph neural networks (GNNs), have also been proposed to optimize the efficiency of multi-vehicle interaction and information transmission. These methods can improve the system's adaptability to dynamic traffic environments without increasing network load. However, these methods still face challenges such as imperfect infrastructure construction, limited network coverage, and data transmission delays.

[0008] Therefore, although traditional communication and perception technologies can provide effective support in some scenarios, existing solutions still have significant limitations in the face of increasingly complex application scenarios for intelligent driving. Summary of the Invention

[0009] Aiming at the deficiencies in the existing technology, the present invention provides a method for detecting the wide-area driving area of autonomous vehicles based on space-ground integrated communication.

[0010] The present invention achieves the above technical objectives through the following technical means.

[0011] Method for detecting the wide-area driving area of autonomous vehicles based on space-ground integrated communication:

[0012] Cooperative vehicles collect single-view 2D images of their surrounding environment. After processing, feature tensors with classification results are obtained;

[0013] After the feature tensors with classification results are gradually compressed in the channel dimension, they are transmitted via low-orbit satellites;

[0014] The vehicle receives the information transmitted by the low-earth orbit satellite, combines it with the lane information in the high-precision map, and uses the spatial modulation cross-attention mechanism to achieve feature fusion between multiple vehicles to obtain global features;

[0015] The global features are reconstructed into a spatial feature map through a linear mapping layer. The spatial feature map is upsampled to obtain a multi-scale fused feature map. The multi-scale fused feature map undergoes a convolution operation to finally generate a class probability map. The class probability map obtains a single-channel semantic segmentation map through a pixel-by-pixel argmax operation. The single-channel semantic segmentation map uses a task head to perform refined classification on each pixel in the single-view 2D image to accurately identify roads, vehicles, pedestrians, and obstacles.

[0016] Further, the global features are generated by optimizing the fused feature tensor X' through residual connection and a feed-forward network. The fused feature tensor X' = M ij ·X, where M ij is the spatial attention modulation weight, and X is the initial feature tensor transmitted by the low-earth orbit satellite.

[0017] Furthermore, the spatial attention modulation weight where PE and PE j are the unified position encodings of vehicle i and vehicle j respectively, ‖PE i -PE j ‖ is the spatial distance between vehicles, σ is the Gaussian kernel standard deviation, λ is to adjust the influence degree of the lane guidance factor on the spatial attention modulation weight, and R lane (i, j) is the lane guidance factor.

[0018] Furthermore, the unified position encoding satisfies: PE = f(P ego , L ego , Δx lane , Δy lane ), where P ego is the global position of the vehicle, L ego is the lane label where the vehicle is located, f is a non-linear mapping function, Δx lane represents the lateral offset of the vehicle in the lane, and Δy lane represents the longitudinal offset of the vehicle in the lane.

[0019] Further, the spatial feature map is upsampled to obtain a multi-scale fused feature map. Specifically: Multilevel low-resolution feature maps F1, F2, … F i … F nThe spatial resolution of each layer of feature maps gradually recovers. Then, a Feature Pyramid Network (FPN) is used to perform multi-scale fusion on feature maps of different layers. The FPN, through a top-down path, fuses the high-level feature maps F i …F n into the low-level feature maps F1 and F2 respectively. In the bottom-up path, the FPN gradually recovers the spatial resolution of the feature maps starting from F1 layer by layer.

[0020] Furthermore, the gradual recovery of the spatial resolution of each layer of feature maps is achieved by performing a transposed convolution operation on each layer of feature map F I to restore the resolution of the feature map from F i to the resolution of its upper-layer feature map. Then, after the transposed convolution operation, each layer of feature map is subjected to a skip connection with F i .

[0021] Further, the single-view 2D image is processed by a convolutional neural network to form a low-level feature image. The low-level feature image is fed into a multi-layer Transformer encoder, which is first cut into multiple small blocks of a fixed size. Each small block is flattened and projected into a d-dimensional space through a linear mapping to obtain an embedding vector z i , and a positional encoding E i is added to obtain an embedded feature vector z' i = z i +E i . The embedded feature vectors of each small block are respectively mapped into query vectors, key vectors, and value vectors through a linear transformation. The similarity between the query vector and the key vector of each pair of small blocks is calculated through a dot product operation to determine the attention weights between each small block and other small blocks. After normalization, a feature vector Z' is formed. The feature vector Z' is compressed and a discriminative global visual representation Z″ is extracted, and finally a feature tensor with classification results is formed where W is the weight matrix after training the perception model.

[0022] Furthermore, the training of the perception model includes the following stages: (1) Initialize the perception model of each cooperative vehicle and set the loss function for each training task; (2) Each cooperative vehicle trains its own perception through a self-supervised loss function to minimize the difference from the perception results of other vehicles; (3) Adjust the weights of each cooperative vehicle in cooperation by calculating the difference between the perception results of each cooperative vehicle and those of other vehicles; (4) According to the calculated weights, fuse the perception information of other vehicles to obtain the final perception result; (5) Optimize multiple training tasks simultaneously to improve the overall performance by sharing parameters; (6) Optimize the parameters of the perception model through the backpropagation algorithm according to the loss functions of each training task.

[0023] Furthermore, the calculation formula of the weight is as follows: Where, represents the difference in perception results between cooperative vehicle i and cooperative vehicle j, and w i is the weight of cooperative vehicle i, represents the set of cooperative vehicles communicating with cooperative vehicle i.

[0024] Furthermore, by fusing the perception information of other vehicles, the final perception result is obtained: Where, represents the perception output of cooperative vehicle j on the environment, and w h is the weight of cooperative vehicle j.

[0025] The beneficial effects of the present invention are as follows:

[0026] (1) By only transmitting the compressed feature tensor information with classification results, the present invention greatly reduces the data transmission volume, effectively reduces the communication delay, and thus significantly improves the real-time response ability of the system.

[0027] (2) In the feature extraction stage of the present invention, an architecture combining a deep convolutional neural network (CNN) and Vision Transformer (ViT) is adopted. The CNN is responsible for extracting low-level features of images, such as edges, textures, and local shapes, while the ViT captures the semantic relationships between long-distance targets through the multi-head self-attention mechanism, enhancing the global semantic expression ability. Through this feature extraction method, it is possible to efficiently identify and process targets in complex traffic scenarios, especially outstanding in the perception of long-distance targets and global interactions.

[0028] (3) In order to further optimize the cooperative perception among multiple vehicles, the present invention introduces a spatial modulation cross-attention mechanism and an adaptive weighting strategy, enabling dynamic adjustment of weights when multiple vehicles share features, preferentially fusing vehicle features related to themselves, and at the same time reducing the impact of features irrelevant to the perception task on the system. This guidance mechanism based on position and lane information effectively improves the spatial consistency of feature fusion, improves the perception accuracy and the robustness of the system. Especially in a complex urban traffic environment, it can effectively avoid interference and enhance the system stability.

[0029] (4) The present invention also adopts a self-supervised learning framework, enabling each vehicle to optimize through the comparison of self-perception and the perception of other vehicles without a large amount of labeled data, enhancing the generalization ability and robustness of the perception model. At the same time, through the multi-task learning framework, the vehicle can simultaneously perform tasks such as target detection, semantic segmentation, and obstacle recognition, further enhancing the overall perception ability.

[0030] (5) By organically combining low-earth orbit satellite communication, feature compression, Transformer architecture, spatial modulation cross-attention mechanism, and multi-task learning, the present invention provides an efficient, accurate, and low-latency multi-vehicle collaborative perception scheme, improving the perception accuracy and system robustness, and reducing the computational complexity and communication burden. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a schematic diagram of multi-vehicle collaboration based on the integration of space and ground of the present invention;

[0032] Figure 2 It is a flowchart of the implementation of the present invention;

[0033] Figure 3 It is an overall model architecture diagram of the present invention;

[0034] Figure 4 It is a decoder structure diagram of the present invention;

[0035] Figure 5 It is a fusion network structure diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but the protection scope of the present invention is not limited thereto.

[0037] The present invention proposes a space-ground integrated architecture based on low-earth orbit satellite communication, as Figure 1 shown, that is, the low-earth orbit satellite communicates with the autonomous vehicle. This architecture provides high-speed and low-latency communication globally through the collaborative work of low-earth orbit satellites and ground base stations, solving the coverage shortage problem of traditional ground communication systems under long distances, high speeds, and signal blind spots. Compared with the ground communication-based scheme, the space-ground integrated architecture can ensure stable communication quality in any environment and geographical location, especially suitable for complex conditions such as ramp merging and occlusion scenarios.

[0038] The advantage of the space-ground integrated architecture is that it can break through the limitations of traditional communication, does not rely on ground infrastructure, provides global seamless connection, and can still maintain efficient and low-latency data transmission in extreme weather and geographical environments. Through the direct communication between the low-earth orbit satellite and the vehicle, the space-ground integrated architecture provides more real-time and reliable road information support for the autonomous driving system, greatly improving driving safety. The introduction of this architecture makes the performance of intelligent driving more stable in dynamic complex environments and is an important guarantee for the future development of autonomous driving technology.

[0039] As Figure 2 , 3 shown, the method for detecting the wide-area driving area of an autonomous vehicle based on space-ground integrated communication of the present invention specifically includes the following steps:

[0040] Step 1, Cooperative vehicle feature extraction

[0041] Each cooperative vehicle collects single-view 2D images of its surrounding environment through the installed forward camera or surround-view camera. These images contain visual information of roads, vehicles, pedestrians, and other key objects. To efficiently extract features, the present invention combines the feature extraction capabilities of convolutional neural network (CNN) and Transformer, achieving an effective combination of local details and global semantics.

[0042] First, use the Residual Network (ResNet) in CNN to extract low-level features (such as edges, textures, and local shapes, etc.) of the single-view 2D images, and form a low-level feature image of H×W×C. Subsequently, introduce the Vision Transformer (ViT) architecture, including multiple layers of Transformer encoders, and capture the semantic relationships between distant objects through the multi-head self-attention mechanism of the Transformer encoder. Specifically: the low-level feature image is cut into multiple small patches p i , each small patch has a size of P×P, and the image size of each small patch is P 2 ×C. Each small patch is flattened and projected into a d-dimensional space through a linear mapping to obtain an embedding vector To provide position information, add position encoding so that the spatial position of each small patch can be encoded, obtaining an embedded feature vector with position encoding z i ′ = z i +E i ; The embedded feature vectors of each small patch are respectively mapped into three vectors of query (Q), key (K), and value (V) through a linear transformation. These vectors are used to represent the intrinsic semantic features between the image small patches. Subsequently, calculate the similarity between the query and the key of each pair of small patches through a dot product operation to determine the attention weights between each small patch and other small patches. After these weights are normalized by the softmax function, they reflect the semantic association strength between different small patches.

[0043] After multiple layers of Transformer encoding, a feature vector Z′ (the elements of which are normalized attention weights) is obtained, which fuses local features (such as edges, textures, etc.) and global semantic features (such as semantic relationships between distant objects); this feature vector not only contains the visual information of each image small patch itself but also reflects the semantic relevance between different spatial positions. Subsequently, through a linear mapping layer, the high-dimensional feature vector Z′ is further compressed and a discriminative global visual representation Z″ is extracted. This comprehensive feature vector Z″ reflects the visual features of each object in the entire 2D image and the spatial relationships between each object.

[0044]

[0045] Among them, W is the weight matrix after the perception model is trained, is the feature tensor with classification results. This operation converts the result of feature extraction into feature information with classification results, and the classification task further clarifies the extracted feature information.

[0046] Step2, Feature Compression and LEO Transmission

[0047] After feature extraction is completed, each cooperative vehicle shares the feature tensor with classification results collected and extracted by itself to other cooperative vehicles through a low-earth orbit satellite, and at the same time receives the feature tensors with classification results extracted by other vehicles. The high bandwidth and low latency characteristics of the low-earth orbit satellite effectively ensure the real-time nature of data transmission. At the same time, through the timestamp synchronization mechanism, the time alignment of multi-vehicle data is achieved, eliminating the errors that may be caused by time delay in dynamic scenarios.

[0048] Although low-earth orbit (LEO) satellite communication provides high transmission efficiency, in vehicle-to-vehicle (V2V) communication applications, the data transmission volume is still a key performance bottleneck. Considering that large bandwidth requirements may cause significant communication delays, therefore, reducing the transmission content becomes a necessary means to improve system efficiency. The present invention effectively reduces the volume of transmitted data and significantly reduces the transmission delay by only transmitting the compressed feature tensor information. Before data transmission, a series of 1×1 convolution operations are used to gradually compress the feature tensor with classification results in the channel dimension, thereby reducing the amount of feature information that needs to be transmitted.

[0049] Step3, Multi-view Feature Information Fusion of the Self-Vehicle

[0050] During the data transmission process, after the compressed feature tensor information with classification results is efficiently transmitted through low-earth orbit (LEO) satellite communication, the ego vehicle (EgoAV) at the receiving end will perform multi-view feature information fusion to further improve the accuracy and robustness of cooperative perception. Combining the lane information in the high-precision map, EgoAV uses an improved Transformer architecture and the spatial modulation cross-attention mechanism to achieve feature fusion between multiple vehicles. Compared with traditional methods, the present invention avoids the high computational complexity brought by the bird's-eye view (BEV) representation, completely relies on high-precision map data, and enhances the performance of multi-vehicle cooperative perception through an optimized attention mechanism, especially showing higher accuracy and robustness in complex traffic scenarios.

[0051] To optimize the spatial consistency of multi-vehicle features, such as Figure 5As shown in the figure, the present invention introduces lane information of a high-precision map in a unified position encoding, combines the global position of a vehicle with road topology information, and provides accurate spatial guidance for feature fusion. The high-precision map provides rich lane semantic information, including lane centerlines, widths, dividing lines, lane numbers, etc. By combining the positioning data of the vehicle (such as GPS, inertial navigation IMU, and visual odometer), the accurate position of the vehicle in the lane is calculated, such as lateral offset and longitudinal offset. The specific calculation formulas are as follows:

[0052] Δx lane =x ego -x center , Δy lane =y ego -y center (2)

[0053] Wherein, Δx lane and Δy lane respectively represent the lateral offset and longitudinal offset of the vehicle in the lane, x ego and y ego are the global coordinates of the vehicle, and x center and y center are the coordinates of the vehicle on the lane centerline in the high-precision map.

[0054] By combining the global position of the vehicle with lane information, the present invention generates a unified position encoding:

[0055] PE = f(P ego , L ego , Δx lane , Δy lane ) (3)

[0056] Wherein, P ego is the global position of the vehicle, L ego is the lane label where the vehicle is located, and f is a non-linear mapping function for integrating position and lane information into a unified feature representation.

[0057] This encoding not only reflects the global coordinates of the vehicle but also introduces rich lane semantic information, providing spatial guidance for subsequent feature fusion. Through the unified position encoding shared by low-orbit satellites, each cooperative vehicle can compare its own lane information with that of other vehicles to determine the relative spatial relationship, such as whether they are in the same lane or adjacent lanes. Guided by lane information, vehicles relevant to itself can be given priority in the feature fusion process to enhance spatial consistency, while ignoring features irrelevant to the perception task.

[0058] In the feature fusion stage, the present invention designs a feature fusion method based on a spatial modulation cross-attention mechanism, and further optimizes the accuracy and robustness of multi-vehicle collaborative perception through lane information guidance. The traditional multi-head self-attention mechanism is difficult to dynamically model the spatial relationship between vehicles when processing multi-vehicle shared features. Therefore, the present invention introduces spatial modulation attention weights, combines the position information of vehicles with lane relationships, and realizes dynamic weighting of feature fusion. The specific calculation formula is as follows:

[0059]

[0060] where M ij represents the spatial modulation attention weight between vehicle i and vehicle j, PE i and PE j are the unified position encodings of vehicle i and vehicle j respectively, ‖PE i -PE j ‖ is the spatial distance between vehicles, σ is the Gaussian kernel standard deviation, used to control the attenuation of the spatial distance, λ is the adjustment lane guidance factor to affect the attention weight, and R lane (i,j) is the lane guidance factor, indicating the semantic correlation between vehicle i and vehicle j on the lane, and the value range is [-1,1].

[0061] The spatial attention modulation weight M ij obtained through the spatial modulation attention mechanism is applied to the initial feature tensor to achieve adaptive weighted fusion of features. Specifically, the initial feature tensor transmitted by the low-earth orbit satellite is expressed as where N represents the number of collaborative vehicles, D is the dimension of the features (position and attitude) of each collaborative vehicle, and the fusion is achieved through the following formula:

[0062] X′ = M ij ·X (5)

[0063] M ij reflects the attention correlation based on spatial guidance between vehicles. After the fusion operation, the initial feature tensor is weighted and fused into a new feature tensor X′.

[0064] Through the above mechanism, the feature fusion process can adaptively adjust the weights: the features of vehicles in the same lane are given higher weights, the features of vehicles in adjacent lanes are adjusted according to lane information and distance, and the features of irrelevant lanes are weakened or even ignored, thereby reducing interference.

[0065] The fused feature tensor X′ is optimized through residual connection (Add and Norm) and feed-forward network (Feed Forward Network, FFN) to generate a high-level global feature X f, further improve the semantic integrity of features and reduce computational redundancy to meet the requirements of real-time processing.

[0066] Step4, Multi-view feature decoding of the ego vehicle

[0067] To effectively utilize the features extracted from multi-view information, the present invention designs an image semantic segmentation decoder, which can accurately decode and reconstruct image details and support the efficient execution of subsequent perception tasks. The global feature X f First, it is reconstructed into the form of a spatial feature map through a linear mapping layer, that is:

[0068]

[0069] where, Linear is the linear mapping layer, and the Reshape operation is used to restore the feature tensor into a feature map F of the corresponding spatial size i .

[0070] Subsequently, the spatial feature map is further decoded through an upsampling operation to match the spatial size of the feature map of the decoder: as Figure 4 shown, the spatial resolutions of the multi-level low-resolution feature maps F1, F2, … F i … F n in each layer are gradually restored (except for the last layer), and the feature maps contain semantic information of different scales; then, a Feature Pyramid Network (FPN) is used to perform multi-scale fusion on the feature maps of different levels. FPN, through the top-down path, fuses the high-level feature maps F i … F n into the low-level feature maps F1 and F2 respectively, thereby enhancing the expression ability of semantic features; in the bottom-up path, FPN restores the spatial resolution of the feature maps layer by layer starting from F1, gradually enhancing the detail information of the image to achieve a more accurate segmentation effect.

[0071] Among them, the gradual restoration of the spatial resolution of each layer of feature map is to perform an upsampling operation on each layer of feature map F i This operation is carried out through transposed convolution, and the transposed convolution operation restores the resolution of the feature map from F i to the resolution of its upper-layer feature map; each layer of feature map after upsampling i is skip-connected with F and directly performs feature fusion, thereby effectively enhancing detail restoration. The purpose of skip connection is to combine the low-level detail information in the encoder with the high-level semantic information in the decoder to avoid the loss of detail information during the upsampling process.

[0072] In the final stage of decoding, the feature maps of multi-scale fusion will undergo further convolutional operations to enhance the segmentation accuracy. After these processes, a class probability map with the same size as the original input image (i.e., a single-view 2D image) is finally generated. where H and W are the height and width of the image, N is the number of classes in the class probability map, each pixel in the class probability map corresponds to the probability distribution of N classes, and finally, P′ is obtained through a pixel-wise argmax operation to obtain a single-channel semantic segmentation map L∈R H×W , where each pixel corresponds to the specific class label of the target object or scene element in the single-view 2D image.

[0073] Through the above gradual restoration of spatial resolution, multi-scale fusion, and convolution, the present invention can restore the detailed information of the image while maintaining the semantic accuracy of the image, thereby generating a high-precision image semantic segmentation result.

[0074] Finally, the single-channel semantic segmentation map finely classifies each pixel in the original 2D image through the task head, accurately identifying targets such as roads, vehicles, pedestrians, and obstacles, and providing high-precision semantic information support for road planning and decision-making.

[0075] Step5, Perception Model Training and Loss Function

[0076] In the problem of cooperative perception in vehicle-to-vehicle (V2V) communication, the goal is to improve the overall perception ability by fusing sensor data from different vehicles. Considering the differences in sensor performance between vehicles, environmental interferences (such as noise, occlusion), and the influence of other external factors, traditional perception methods often face problems of accuracy and robustness. Therefore, the present invention proposes a training process that combines self-supervised learning and an adaptive weighting mechanism, aiming to improve the accuracy and reliability of the cooperative perception system. The training process is divided into the following stages: (1) Initialize the perception model of each cooperative vehicle (as an existing benchmark model) and set the loss function for each training task; (2) Each cooperative vehicle trains its own perception through the self-supervised loss function to minimize the difference from the perception results of other vehicles; (3) Adjust the weight of each cooperative vehicle in cooperation by calculating the difference between the perception results of each cooperative vehicle and other vehicles; (4) According to the calculated weights, fuse the perception information of other vehicles to obtain the final perception result; (5) Optimize multiple training tasks simultaneously to improve the overall performance by sharing parameters; (6) Optimize the parameters of the perception model through the backpropagation algorithm according to the loss functions of each training task. The following is a detailed elaboration:

[0077] Each collaborative vehicle optimizes itself by receiving data from other vehicles without relying on external labeled data, and this process is called self-supervised learning. Each collaborative vehicle compares its perception results with those of other vehicles and improves its perception ability by minimizing the prediction error. Suppose the perception result of vehicle i is while the perception result of vehicle j is The difference in perception results between vehicle i and vehicle j is optimized through a self-supervised loss function. The self-supervised loss function can be expressed as:

[0078]

[0079] where represents the set of vehicles communicating with vehicle i, and represent the perception outputs of vehicle i and vehicle j on the environment respectively.

[0080] By minimizing the above self-supervised loss function, collaborative vehicle i can adaptively adjust its perception model, thereby improving the accuracy of the perception result.

[0081] Due to the quality differences in sensor data among vehicles, how to effectively fuse the perception information from multiple vehicles is particularly important. To handle the heterogeneity of perception results from different vehicles, the present invention proposes an adaptive weighting mechanism, which makes the fusion of perception data more accurate by dynamically adjusting the weight of each vehicle in collaborative perception.

[0082] The weight w of each collaborative vehicle i i is adjusted according to the difference metric between its perception result and those of other collaborative vehicles. Specifically, the weight calculation formula is:

[0083]

[0084] Here, represents the difference in perception results between collaborative vehicle i and collaborative vehicle j;

[0085] The weight w i is calculated based on the similarity of perception data: when the difference in perception outputs between collaborative vehicle i and collaborative vehicle j is small, the weight of collaborative vehicle i is large, indicating that its perception result is relatively reliable, and vice versa.

[0086] After obtaining the weights of each collaborative vehicle, collaborative vehicle i will perform weighted fusion on the data of other collaborative vehicles according to these weights to obtain the final fused perception result. The process of fusing perception data can be expressed as:

[0087]

[0088] Through weighted fusion, cooperative vehicle i can reasonably combine the perception information from other cooperative vehicles according to the weights. The weight w j can be dynamically adjusted during the model training process to ensure that the perception system can adjust the fusion accuracy according to the contributions of vehicles in different scenarios.

[0089] To improve the overall performance of the collaborative perception system, each cooperative vehicle not only needs to complete perception tasks (such as object detection, obstacle recognition, etc.), but can also synchronously perform other related tasks (such as vehicle trajectory prediction, traffic flow prediction, etc.) in the same network. By sharing network parameters, the synergy between multiple tasks can improve the performance of the perception system. Each task has its own loss function, and the overall loss function is optimized in the form of weighted summation, specifically expressed as:

[0090]

[0091] where λ k is the weighting coefficient of task k, is the loss function of task k, and K represents the total number of tasks.

[0092] For example, for the perception task, the loss function can be the cross-entropy loss, while for the trajectory prediction task, the loss function can be the mean squared error loss.

[0093] Under the entire model training framework, the loss function of each task can reflect the specific requirements of the task. For each task k, its loss function is defined as:

[0094]

[0095] where and represent the predicted value and the true value of the nth sample respectively, and N′ is the number of samples. For the perception task, the loss function is usually calculated based on the difference between the predicted class label and the actual label.

[0096] Through the above training process, the goal is to maximize the accuracy and robustness of the collaborative perception system. The overall training goal can be expressed as:

[0097]

[0098] The described embodiments are the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Without departing from the essence of the present invention, any obvious improvements, substitutions, or variations that those skilled in the art can make all fall within the protection scope of the present invention.

Claims

1. A method for detecting a wide - area driving region of an autonomous vehicle based on space - ground integrated communication, characterized in that: Cooperative vehicles collect single - view 2D images of their surrounding environment. After processing, feature tensors with classification results are obtained; The feature tensors with classification results are gradually compressed in the channel dimension and then transmitted via low - earth - orbit satellites; The ego - vehicle receives the information transmitted by the low - earth - orbit satellites, combines the lane information in the high - definition map, and uses the spatial modulation cross - attention mechanism to achieve feature fusion between multiple vehicles, obtaining global features; The global features are reconstructed into a spatial feature map through a linear mapping layer. The spatial feature map is upsampled to obtain a multi - scale fused feature map. The multi - scale fused feature map undergoes a convolution operation to finally generate a class probability map. The class probability map undergoes a per - pixel argmax operation to obtain a single - channel semantic segmentation map. The single - channel semantic segmentation map finely classifies each pixel in the single - view 2D image through a task head to accurately identify roads, vehicles, pedestrians, and obstacles.

2. The method for detecting a wide driving area of an autonomous vehicle according to claim 1, wherein The global feature is generated by optimizing the fused feature tensor X' through residual connection and a feed-forward network. The fused feature tensor X' = M ij ·X, where M ij is the spatial attention modulation weight, and X is the initial feature tensor transmitted by the low-earth orbit satellite.

3. The method for detecting a wide driving area of an autonomous vehicle according to claim 2, wherein The spatial attention modulation weight where, PE i and PE j are the unified position encodings of vehicle i and vehicle j respectively, ‖PE i - PE j ‖ is the spatial distance between vehicles, σ is the Gaussian kernel standard deviation, λ is the degree of influence of the lane guidance factor on the spatial attention modulation weight, and R lane (i, j) is the lane guidance factor.

4. The method for detecting a wide driving area of an autonomous vehicle according to claim 3, wherein The unified position encoding satisfies: PE = f(P ego , L ego , Δx lane , Δy lane ), where P ego is the global position of the vehicle, L ego is the lane label where the vehicle is located, f is a non-linear mapping function, Δx lane represents the lateral offset of the vehicle in the lane, and Δy lane represents the longitudinal offset of the vehicle in the lane.

5. The method for detecting a wide driving area of an autonomous vehicle according to claim 1, characterized in that, The spatial feature map obtains a multi-scale fused feature map through an upsampling operation, specifically: the spatial resolutions of each layer of the multi-level low-resolution feature maps F1, F2,... F i ... F n are gradually restored. Then, a Feature Pyramid Network (FPN) is used to perform multi-scale fusion on the feature maps of different levels. The FPN, through a top-down path, fuses the high-level feature maps F i ... F n into the low-level feature maps F1 and F2 respectively. In the bottom-up path, the FPN restores the spatial resolution of the feature maps layer by layer starting from F1.

6. The method for detecting a wide driving area of an autonomous vehicle according to claim 5, characterized in that, The spatial resolution of each layer of feature maps is gradually restored by taking each layer of feature maps F i through a transposed convolution operation to restore the resolution of the feature maps from F i to the resolution of the previous layer of feature maps; Then, for each layer of feature maps after the deconvolution operation and F i perform skip connections.

7. The method for detecting a wide driving area of an autonomous vehicle according to claim 1, characterized in that The single-view 2D image is processed by a convolutional neural network to form a low-level feature image; the low-level feature image is fed into a multi-layer Transformer encoder, first cut into multiple small blocks of a fixed size, and each small block is flattened and projected into a d-dimensional space through a linear mapping to obtain an embedding vector z i , and the positional encoding E is added i , to obtain an embedding feature vector z' with positional encoding i = z i + E i ; the embedding feature vectors of each small block are respectively mapped into query vectors, key vectors and value vectors through linear transformation, the similarity between the query vector and the key vector of each pair of small blocks is calculated through dot product operation, the attention weights between each small block and other small blocks are determined, and after normalization, a feature vector Z' is formed; Compress the feature vector Z′ and extract the discriminative global visual representation Z″, and finally form a feature tensor with classification results The W is the weight matrix after training of the perception model.

8. The method for detecting a wide driving area of an autonomous vehicle according to claim 7, wherein The training of the perception model includes the following stages: (1) Initialize the perception model of each cooperative vehicle and set the loss function for each training task; (2) Each cooperative vehicle trains its own perception through a self - supervised loss function to minimize the difference from the perception results of other vehicles; (3) Adjust the weights of each cooperative vehicle in cooperation by calculating the difference between the perception results of each cooperative vehicle and those of other vehicles; (4) According to the calculated weights, fuse the perception information of other vehicles to obtain the final perception result; (5) Optimize multiple training tasks simultaneously to improve the overall performance by sharing parameters; (6) Optimize the parameters of the perception model through the backpropagation algorithm according to the loss functions of each training task.

9. The method for detecting a wide driving area of an autonomous vehicle according to claim 8, wherein The calculation formula for the weight is as follows: Among them, represents the difference in the perception results between cooperative vehicle i and cooperative vehicle j, and w i is the weight of cooperative vehicle i, represents the set of cooperative vehicles communicating with cooperative vehicle i.

10. The method for detecting a wide driving area of an autonomous vehicle according to claim 9, characterized in that, Fuse the perception information of other vehicles to obtain the final perception result: Among them, represents the perception output of cooperative vehicle j on the environment, and w j is the weight of cooperative vehicle j.

Citation Information

Patent Citations

  • Multi-sensor fusion vehicle-road collaborative sensing method for automatic driving

    CN114821507A

  • Multi-source heterogeneous sensing data fusion method for autonomous vehicle

    CN114912532A

  • Panoramic driving perception method based on deep learning

    CN117058641A

  • Heterogeneous traffic subject collaborative beyond visual range perception and swarm intelligence decision-making method

    CN117496713A

  • Multi-modal data fused low-cost high-performance vehicle-road cooperation method

    CN117915284A

Cited By

  • Internet of vehicles sensing data transmission method and device

    CN120935533A