An Autonomous Driving Cooperative Perception Method and System Based on Enhanced Perception with Few Vehicles
Through cross-scene differential modeling and lightweight neural network generation compensation features, the problem of collaborative perception performance degradation in few-vehicle perception scenarios is solved, and the perception ability and adaptability of autonomous driving are improved.
Patent Information
- Application Number
- CN202410760951.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-13
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-06-13
AI Technical Summary
In the few-vehicle perception scenario, the perception performance degradation of the existing collaborative perception methods is serious, affecting the safety and utility of autonomous driving.
The difference between the few-vehicle-aware scene and the multi-vehicle-aware scene is simulated through cross-scene difference modeling method, and compensated features are generated using lightweight neural networks to enhance collaborative perception performance. The specific steps include information exchange between vehicles, coordinate system alignment, feature extraction and fusion, feature modulation and 3D object detection.
On intelligent vehicles with limited computing power, the coordinated perception performance of fewer vehicles has been improved, and the perception ability, adaptability and universality of terminal equipment of autonomous driving have been enhanced.
Smart Images

Figure CN118779821B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more particularly to the field of autonomous driving technology. Specifically, the present invention relates to a method and system for collaborative perception of autonomous driving based on enhanced perception with few vehicles. Background Art
[0002] In recent years, with the development of wireless communication technology and intelligent vehicle sensors, autonomous driving, as a key technology of intelligent transportation systems, has encountered unprecedented development opportunities. Autonomous driving generally includes three important modules: perception, planning, and control. Among them, the perception module uses intelligent vehicle sensors (such as vehicle-mounted cameras, lidar) to scan the surrounding driving environment and establish overall environmental perception information. High-quality perception is one of the main technical bottlenecks of autonomous driving and is crucial for subsequent planning and control modules. Due to physical limitations of sensors and real-world problems such as external occlusion, single-vehicle perception cannot meet the requirements of high-level autonomous driving. By combining complementary perception information of other intelligent vehicles within the effective communication range, collaborative perception technology can assist intelligent vehicles in establishing a more comprehensive understanding of the surrounding environment, thereby improving the safety of autonomous driving to a certain extent.
[0003] The goal of collaborative perception technology is to jointly perceive the surrounding environment by fusing the perception information of the central intelligent vehicle and other connected vehicles within the communication range. According to the different collaborative stages, collaborative perception technology can be divided into three types: early collaboration, late collaboration, and mid-term collaboration. Among them, early collaboration directly performs coordinate system transformation and fusion on the original data to form comprehensive perception based on the overall original data. The main disadvantage is the high bandwidth requirement. Late collaboration generates the final perception result by fusing the perception results of different intelligent vehicles, and has the advantage of low bandwidth requirement, but the perception performance is not high. Mid-term collaboration fuses the intermediate features of different intelligent vehicles to form the final perception result. Considering the good balance achieved by mid-term collaboration in terms of bandwidth requirement and perception performance, the present invention mainly focuses on mid-term collaboration.
[0004] Effective collaborative perception depends on a sufficient number of intelligent vehicles participating in the collaborative perception process. However, in real-world autonomous driving scenarios, due to the mobility of intelligent vehicles, the volatility of communication networks, and the dynamics of the driving environment, it is impossible to ensure that the above assumptions hold. Existing technologies, such as the invention patent "A Collaborative Perception Method and Collaborative Perception Device Based on Spatiotemporal Feature Fusion" with the application publication number CN116992400 A and the invention patent "A Feature-Level Collaborative Perception Fusion Method and System for Vehicle-Road Collaboration" with the authorization announcement number CN 115578709 B, completely ignore the problem of perception performance degradation caused by insufficient numbers of intelligent vehicles participating in perception. Intuitively, in a multi-vehicle perception scenario with a sufficient number of intelligent vehicles participating, the perception performance of various collaborative perception methods is significantly better than that in a few-vehicle perception scenario. Therefore, in a few-vehicle perception scenario, the problem of perception performance degradation of existing collaborative perception methods has become the norm.
[0005] In summary, the problem of perception performance degradation in a few-vehicle perception scenario has seriously affected the effectiveness of autonomous driving collaborative perception technology. How to design enhanced collaborative perception methods for a few-vehicle perception scenario has become an urgent technical problem and is of great significance for high-level autonomous driving. Summary of the Invention
[0006] In view of the deficiencies of the existing technology, the present invention provides an autonomous driving collaborative perception method and system based on enhanced few-vehicle perception. The specific technical solutions are as follows:
[0007] An autonomous driving collaborative perception method based on enhanced few-vehicle perception, the method comprising the following steps:
[0008] Step 1, the intelligent vehicle collects perception data of the surrounding environment through in-vehicle sensors;
[0009] Step 2, the intelligent vehicle broadcasts basic information to surrounding intelligent vehicles and receives basic information from other intelligent vehicles;
[0010] Step 3, use the basic information to align the coordinate systems of each intelligent vehicle;
[0011] Step 4, the intelligent vehicle extracts intermediate features from the perception data using a feature extraction network;
[0012] Step 5, the intelligent vehicle broadcasts and sends the intermediate features to other intelligent vehicles;
[0013] Step 6, use a feature fusion network to aggregate the intermediate features of the central intelligent vehicle and the intermediate features of other intelligent vehicles to generate collaborative features containing different perspective information;
[0014] Step 7, the intelligent vehicle executes a unified cross-perception scenario difference modeling method to effectively model the performance difference from a multi-vehicle perception scenario to a few-vehicle perception scenario;
[0015] Furthermore, for a predefined hyperparameter k, the few-vehicle perception scenario is the set of scenarios where the number of vehicles participating in collaborative perception is less than or equal to k, and other scenarios are multi-vehicle perception scenarios. Usually, k is taken as 2;
[0016] Furthermore, step 7 includes the following two sub-steps:
[0017] Step 7.1, in the model training stage, execute a learning method without network bandwidth requirements to assist the learning process of the model. Specifically, for any intelligent vehicle i in the multi-vehicle perception scenario, the collaborative feature H generated by step 6 i is denoted as On the other hand, delete the intermediate features of some vehicles received from this intelligent vehicle from j ∈ N(i) to construct a simulated few-vehicle feature set j ∈ N(i)', where and the number of intelligent vehicles satisfies 1 + |N(i)'| ≤ k. and the collaborative feature H generated by j ∈ N(i)' through step 6 i is denoted as After that, use the learnable compensation feature T to model the difference between the few-vehicle perception scenario and the multi-vehicle perception scenario from a global perspective, that is where ⊕ is a relational operator implemented by step 8;
[0018] Step 7.2, in the model inference stage, perform feature modulation operations on intelligent vehicles in the few-vehicle perception scenario, while no feature modulation operations are required for intelligent vehicles in the multi-vehicle perception scenario;
[0019] Step 8, use the injection network to integrate the compensation feature T learned in the model training stage into the collaborative features related to the few-vehicle perception scenario, so as to achieve the purpose of enhancing collaborative perception in the few-vehicle perception scenario;
[0020] To additionally incorporate the context information of the driving environment where the intelligent vehicle is located, select the collaborative feature H generated by step 6 i as the context information for generating the fine-grained compensation feature T i . Considering the limited computing power of intelligent vehicles in autonomous driving, this process is implemented by a lightweight neural network:
[0021] T i = DConv s×s (Conv(CONCAT(T, H i )))
[0022] Among them, CONCAT is a feature concatenation operation acting on the dimension channels, Conv is a 1×1 convolutional layer used to halve the number of channels of the input features, and DConv s×s is a depthwise separable convolutional layer with a convolutional kernel size of s = 7. Then, the fine-grained compensation feature T i is used to perform feature modulation on the collaborative features related to the low-traffic scenarios, and the corresponding calculation is shown in the following formula:
[0023]
[0024] Among them, Conv is a 1×1 convolutional layer with the number of input channels equal to the number of output channels, ⊙ is the bitwise multiplication, is the collaborative feature after feature modulation.
[0025] Step 9, perform perception tasks such as 3D object detection based on the collaborative features.
[0026] An autonomous driving collaborative perception system based on low-traffic perception enhancement, including multiple vehicle-end computing units, is as follows:
[0027] The vehicle-end computing unit is used to collect perception information, send and receive basic information from other vehicle-end computing units, align the coordinate system, extract intermediate features, send and receive intermediate features from other vehicle-end computing units, extract collaborative features, execute the cross-perception scenario difference modeling method, execute feature modulation, and execute 3D object detection perception tasks. The vehicle-end computing unit generates compensation features by executing the cross-perception scenario difference modeling method, and then executes feature modulation to enhance the collaborative features of the low-traffic perception scenario, so as to achieve the purpose of low-traffic perception enhancement.
[0028] Compared with the prior art, the advantages of the present invention are as follows:
[0029] 1. The present invention simulates the differences between low-traffic perception scenarios and multi-vehicle perception scenarios through a general modeling method, and can support various collaborative perception methods in autonomous driving.
[0030] 2. The present invention enhances the collaborative perception in low-traffic perception scenarios through a lightweight method, and will not introduce excessive computational burden and network burden to intelligent vehicles with limited computing power.
[0031] 3. The method of the present invention has good general deployability and adaptability for terminal devices, the implementation method is simple, and it can support the requirements of autonomous driving for high-performance collaborative perception. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is the flowchart of the method of the present invention.
[0033] Figure 2Schematic diagram of the final test results of the embodiment.
[0034] Figure 3 Schematic diagram of the system structure of the present invention. Detailed implementation manners
[0035] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following further elaborates on the present invention in detail with reference to the accompanying drawings and by way of examples.
[0036] Aiming at the problem of degraded collaborative perception performance in scenarios with few vehicles, the present invention provides an autonomous driving collaborative perception method based on few-vehicle perception enhancement. The main idea is to simulate the missing information in the few-vehicle scenario through a general cross-scenario difference modeling method, and then fuse the compensation features into the collaborative features related to the few-vehicle perception scenario to achieve the purpose of enhancing collaborative perception in the few-vehicle perception scenario.
[0037] As Figure 1 shown, the specific steps of this method are as follows:
[0038] Step 1, the intelligent vehicle collects perception data of the surrounding environment through in-vehicle sensors. In autonomous driving, common in-vehicle sensors include lidar and multi-view cameras.
[0039] Taking the lidar sensor as an example, the lidar measures the distance to an object by periodically scanning the surrounding environment and receiving the reflected laser beam. For a certain intelligent vehicle i, the perception data of the lidar is modeled as point cloud data D i . Due to the limitations of the physical performance of in-vehicle sensors, an intelligent vehicle cannot individually perceive distant or occluded objects. By exchanging complementary perception information with other intelligent vehicles or intelligent infrastructure, collaborative perception can alleviate the limitations of individual perception of intelligent vehicles.
[0040] Step 2, the intelligent vehicle broadcasts basic information to the surrounding intelligent vehicles, mainly involving the pose information P=(x, y, z, θ, φ, ψ) of the vehicle and the timestamp information t, specifically including the position information (x, y, z) and three rotation angles: yaw angle θ, pitch angle φ, and roll angle ψ. At the same time, each intelligent vehicle also needs to receive the basic information from other intelligent vehicles.
[0041] Step 3, use the basic information to align the coordinate systems of each intelligent vehicle.
[0042] Since the collaborative perception technology in autonomous driving involves the interaction of multiple intelligent vehicles, it is necessary to align the coordinate systems of the perception information from different intelligent vehicles. As one of the implementation manners, Step 3 can be implemented by any of the following corresponding sub-steps.
[0043] Step 3.1, calculate the coordinate system transformation matrix based on the received pose information and the vehicle's own pose information, and directly transform the perception data D i to a unified coordinate system.
[0044] Step 3.2, calculate the coordinate system transformation matrix based on the received pose information and the vehicle's own pose information, and transform the intermediate features extracted in Step 4 to a unified coordinate system. Aligning the coordinate systems does not belong to the core content of the present invention, and those skilled in the art can choose a suitable matrix transformation method to align the coordinate systems. Those skilled in the art can refer to the coordinate system alignment method described in the paper "V2X-ViT: Vehicle-to-Everything Cooperative Perception with Vision Transformer" (Xu R, Xiang H, Tu Z, et al. V2x-vit: Vehicle-to-everything cooperative perception with vision transformer[C] / / European conference on computer vision. Cham: Springer Nature Switzerland, 2022: 107-124.).
[0045] Step 4, the intelligent vehicle uses a feature extraction network to extract intermediate features from the perception data. Each intelligent vehicle performs the feature extraction operation separately, and the vehicle-side computing unit generates the intermediate feature F i , where F i is a tensor with a shape of (C, H, W), and C, H, and W are the number of channels, feature height, and feature width respectively. To reduce the transmission bandwidth required for the intelligent vehicle to broadcast information to other intelligent vehicles, feature dimensionality reduction operations can be performed in the spatial dimension and channel dimension. Denote the reduced-dimensional intermediate feature as having a size of (C ↓ , H ↓ , W ↓ ). where C ↓ ≤C, H ↓ ≤H and W ↓ ≤W.
[0046] In some embodiments, the PointPillar network, which achieves a good balance between model performance and inference speed, is used as the feature extraction network, and the point cloud data D i is mapped to a high-dimensional intermediate feature F rich in semantic information i. PointPillar is divided into three parts in the network architecture: PillarEncoder network, Scatter network, and Backbone network. Among them, the PillarEncoder network is used to map the point cloud data D i into voxel features, the Scatter network is used to map the voxel features into a pseudo-image, and the Backbone network is usually a backbone network for converting the pseudo-image into intermediate features F i . For a more detailed description of the PointPillar network, those skilled in the art can refer to the method described in the paper "Pointpillars: Fast encoders for object detection from point clouds" (Lang A H, Vora S, Caesar H, et al. Pointpillars: Fast encoders for object detection from point clouds[C] / / IEEE / CVF conference on computer vision and pattern recognition. 2019: 12697-12705.).
[0047] In some embodiments, C 1×1 convolutional kernels are used to reduce the channel dimension of the intermediate feature F i , and a common convolutional layer with a downsampling rate of 2 is used to reduce the spatial dimension of the intermediate feature F i . After these two dimensionality reduction operations, F of size (C, H, W) i is converted into an intermediate feature of size
[0048] Step 5, the intelligent vehicle i broadcasts and sends to other intelligent vehicles j∈N(i) After the vehicle-side computing unit receives the complete information of other intelligent vehicles j∈N(i), it performs a feature decompression operation to generate an intermediate feature j∈N(i), where N(i) represents the set of intelligent vehicles centered on the intelligent vehicle i and interacting with each other.
[0049] In some embodiments, C 1×1 convolutional kernels are used to decompress the channel dimension of the intermediate feature , and a transposed convolutional layer with an upsampling rate of 2 is used to decompress the spatial dimension of the intermediate feature . After these two decompression operations, the intermediate feature is converted into intermediate features of size (C, H, W) For other intermediate features For j ∈ N(i), the same decompression operation is also performed, and the final result is denoted as j ∈ N(i).
[0050] Step 6: Use the feature fusion network to aggregate the of the central intelligent vehicle and the of other intelligent vehicles, where j ∈ N(i), to generate a collaborative feature H containing information from different perspectives i .
[0051] In some embodiments, the feature fusion network selects a fusion network f based on the self-attention mechanism fusion , and this process can be represented by the following formula:
[0052]
[0053] Step 7: The intelligent vehicle executes a unified cross-sensing scenario difference modeling method to effectively model the performance difference from the multi-vehicle sensing scenario to the few-vehicle sensing scenario.
[0054] In different collaborative sensing scenarios, the number of intelligent vehicles participating in the collaborative sensing process may be different. For a predefined hyperparameter k, the collaborative sensing scenarios can be divided into non-overlapping few-vehicle sensing scenarios and multi-vehicle sensing scenarios. Specifically, the few-vehicle sensing scenario is the set of scenarios where the number of vehicles participating in the collaborative sensing is less than or equal to k, and other scenarios are multi-vehicle sensing scenarios. Compared with the multi-vehicle sensing scenario, there are obvious performance degradation problems in various collaborative sensing methods in the few-vehicle sensing scenario. For example, in the few-vehicle sensing scenario, the performance of the intermediate features aggregated by the collaborative sensing method in the 3D object detection task is much lower than that in the multi-vehicle sensing scenario. The present invention does not impose any requirements on the value of k, and it is usually taken as 2.
[0055] Step 7.1: In the model training stage, a learning method without network bandwidth requirements is executed to assist the learning process of the model. Specifically, for any intelligent vehicle i in the multi-vehicle sensing scenario, the collaborative feature H generated in Step 6 i is denoted as On the other hand, the intermediate features of some vehicles are deleted from the received from this intelligent vehicle, where j ∈ N(i), to construct a simulated few-vehicle feature set j ∈ N(i)', where and the number of intelligent vehicles satisfies 1 + |N(i)'| ≤ k. and the collaborative feature H generated by j ∈ N(i)' through Step 6 i is denoted as
[0056] The collaborative features constructed by the above learning method without network bandwidth requirements and have a comparable relationship. The learnable compensation feature T is used to model the difference information between the few-vehicle perception scenario and the multi-vehicle perception scenario from a global perspective, that is where ⊕ is a relational operator, which is specifically implemented in step 8. Since the compensation feature T acts in the collaborative feature space, it has good generality and can be applied to various collaborative perception methods. After the training of the model, the compensation feature T can effectively model the differences between the two perception scenarios.
[0057] Step 7.2, in the model inference stage, perform feature modulation operations on the intelligent vehicles in the few-vehicle perception scenario, while no feature modulation operations are required for the intelligent vehicles in the multi-vehicle perception scenario. Specifically, if the intelligent vehicle i is currently in the few-vehicle perception scenario, it needs to go through feature modulation to obtain collaborative features Otherwise
[0058] It should be noted that step 7.1 and step 7.2 are only executed in the model training stage and the model inference stage respectively, and cannot be executed simultaneously.
[0059] Step 8, use the injection network to integrate the compensation feature T learned in the model training stage into the collaborative features related to the few-vehicle perception scenario, so as to achieve the purpose of enhancing collaborative perception in the few-vehicle perception scenario.
[0060] To additionally incorporate the context information of the driving environment where the intelligent vehicle is located, select the collaborative feature H generated in step 6 i as the context information for generating the fine-grained compensation feature T i . Considering the limited computing power of intelligent vehicles in autonomous driving, this process is implemented by a lightweight neural network:
[0061] T i = DConv s×s (Conv(CONCAT(T, H i )))
[0062] where CONCAT is a feature concatenation operation acting on the dimension channels, Conv is a 1×1 convolutional layer used to halve the number of channels of the input features, and DConv s×s is a depthwise separable convolutional layer with a convolutional kernel size of s = 7. Then, use the fine-grained compensation feature T i to perform feature modulation on the collaborative features related to the few-vehicle scenario, and the corresponding calculation is shown in the following formula:
[0063]
[0064] Among them, Conv is a 1×1 convolutional layer with the number of input channels equal to the number of output channels, and ⊙ is a bitwise multiplication. is the collaborative feature modulated by feature modulation.
[0065] Step 9: Based on the collaborative feature Perform perception tasks such as 3D object detection.
[0066] In the few-vehicle perception scenario, the collaborative feature Is modulated by the compensation feature T containing scene difference information, achieving the purpose of enhancing few-vehicle perception.
[0067] Figure 2 Shows the final detection result of 3D object detection, and the perception task can be used to support subsequent planning and control tasks in autonomous driving.
[0068] Such as Figure 3 As shown, taking the few-vehicle perception scenario participated by two intelligent vehicles as an example, the present invention provides an embodiment of an autonomous driving collaborative perception system 300 based on few-vehicle perception enhancement, including a first vehicle-end computing unit 310 and a second vehicle-end computing unit 320, specifically as follows:
[0069] The first vehicle-end computing unit 310 is used to collect perception information, send and receive basic information from the second vehicle-end computing unit 320, align the coordinate system, extract intermediate features, send and receive intermediate features from the second vehicle-end computing unit 320, extract collaborative features, execute a cross-perception scene difference modeling method, execute feature modulation, and execute a 3D object detection perception task.
[0070] The corresponding functions of this system correspond to steps 1, 2, 3, 4, 5, 6, 7, 8, and 9 of the method of the present invention, where executing the cross-perception scene difference modeling method and executing feature modulation are the core modules of the present invention. Similar to the first vehicle-end computing unit 310, the second vehicle-end computing unit 320 needs to complete similar operations. Specifically, the second vehicle-end computing unit 320 is used to collect perception information, send and receive basic information from the first vehicle-end computing unit 310, align the coordinate system, extract intermediate features, send and receive intermediate features from the first vehicle-end computing unit 310, extract collaborative features, execute a cross-perception scene difference modeling method, execute feature modulation, and execute a 3D object detection perception task. The first vehicle-end computing unit 310 and the second vehicle-end computing unit 320 generate a compensation feature by executing the cross-perception scene difference modeling method, and then execute feature modulation to enhance the collaborative feature of the few-vehicle perception scenario, so as to achieve the purpose of enhancing few-vehicle perception.
[0071] The method according to the present invention described above can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and to be downloaded via a network and stored in a local recording medium, so that the method described herein can be stored on such a software process on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, a collaborative perception method and system for autonomous driving based on less vehicle perception enhancement described herein are implemented. In addition, when a general-purpose computer accesses the code for implementing the processes shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the processes shown herein.
[0072] Those of ordinary skill in the art will realize that the embodiments described herein are to assist the reader in understanding the implementation methods of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A cooperative perception method for autonomous driving based on enhanced perception with fewer vehicles, characterized in that It includes the following steps: Step 1, the intelligent vehicle collects perception data of the surrounding environment through in-vehicle sensors; Step 2, the intelligent vehicle broadcasts basic information to surrounding intelligent vehicles and receives basic information from other intelligent vehicles; Step 3, use the basic information to align the coordinate systems of each intelligent vehicle; Step 4, the intelligent vehicle extracts intermediate features from the perception data using a feature extraction network; Step 5, the intelligent vehicle broadcasts and sends the intermediate features to other intelligent vehicles; Step 6, utilize a feature fusion network to aggregate the intermediate features of the central intelligent vehicle and the intermediate features of other intelligent vehicles to generate collaborative features containing information from different perspectives; Step 7, the intelligent vehicle executes a unified cross-perception scenario difference modeling method to effectively model the performance difference from a multi-vehicle perception scenario to a few-vehicle perception scenario; The said unified cross-perception scenario difference modeling method includes two sub-steps, specifically as follows: Sub-step 1, in the model training phase, execute a learning method without network bandwidth requirements to assist the learning process of the model: For any intelligent vehicle i in the multi-vehicle perception scenario, its corresponding collaborative feature H i is denoted as On the other hand, delete the intermediate features of some vehicles received from this intelligent vehicle to construct a simulated sparse-vehicle feature set F j ↑ , j ∈ N(i) to construct a simulated sparse-vehicle feature set F j ↑ , j ∈ N(i)′, where and satisfy 1 + |N(i)′| ≤ k in terms of the number of intelligent vehicles; F i ↑ and the collaborative feature H generated by F j ↑ , j ∈ N(i)′ is denoted as i is denoted as After that, use the learnable compensation feature T to model the difference between the sparse-vehicle perception scenario and the multi-vehicle perception scenario from a global perspective, that is where ⊕ is a relational operator; Sub-step 2, in the model inference stage, perform feature modulation operations on intelligent vehicles in a few-vehicle perception scenario, while no feature modulation operations are required for intelligent vehicles in a multi-vehicle perception scenario; Step 8, use an injection network to incorporate the compensation feature T learned in the model training stage into the collaborative features related to the few-vehicle perception scenario to achieve the purpose of enhancing collaborative perception in the few-vehicle perception scenario; The constructed injection network is implemented based on two lightweight neural networks, specifically as follows: To additionally incorporate the context information of the driving environment in which the intelligent vehicle is located, the collaborative feature H is selected i as the context information for generating the fine-grained compensation feature T i ; Considering the limited computing power of intelligent vehicles in autonomous driving, this process is implemented by a lightweight neural network: T i = DConv s×s (Conv(CONCAT(T, H i ))) where CONCAT is a feature concatenation operation acting on the dimension channel, Conv is a 1×1 convolutional layer used to halve the number of channels of the input features, and DConv s×s is a depthwise separable convolutional layer with a convolutional kernel size of s = 7; then, the fine-grained compensation feature T i is used to perform feature modulation on the collaborative features related to the less-vehicle scenario, and the corresponding calculation is shown in the following formula: where Conv is a 1×1 convolutional layer with the number of input channels equal to the number of output channels, and ⊙ is the element-wise multiplication, is the collaborative feature after feature modulation; Step 9, perform 3D object detection perception tasks based on the collaborative features.
2. An autonomous driving collaborative perception system based on enhanced perception with fewer vehicles, characterized in that, For implementing the method described in claim 1; the said system includes: A vehicle-end computing unit, which is used to collect perception information, send and receive basic information from other vehicle-end computing units, align coordinate systems, extract intermediate features, send and receive intermediate features from other vehicle-end computing units, extract collaborative features, execute a cross-perception scenario difference modeling method, execute feature modulation, and execute 3D object detection perception tasks; the vehicle-end computing unit generates a compensation feature by executing the cross-perception scenario difference modeling method, and then executes feature modulation to enhance the collaborative features in the few-vehicle perception scenario to achieve the purpose of enhancing few-vehicle perception.
Citation Information
Patent Citations
A feature-level cooperative perception fusion method and system for vehicle-road cooperation
CN115578709B
Collaborative perception method and collaborative perception device based on spatial-temporal feature fusion
CN116992400A