A target detection method for low-bandwidth vehicle-to-road feature fusion known to lighthouses
By employing target center point feature selection and Gaussian function fusion in the vehicle-road cooperative system, the problem of high communication bandwidth was solved, the perception capability and detection accuracy of the vehicle were improved, and the deployment of vehicle-road cooperative perception was realized.
Patent Information
- Application Number
- CN202310769437.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-27
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2043-06-27
AI Technical Summary
In existing vehicle-road cooperative perception systems, the communication bandwidth of multiple sensing entities is difficult to reduce effectively, resulting in the communication volume of medium-term cooperation methods being far higher than the requirements for deployment, thus affecting the practicality of vehicle-road cooperative perception.
A feature selection mechanism based on target center point features is adopted. By performing low-dimensional sparse feature selection at the time and space levels, the communication bandwidth is reduced. The sparse target center point feature information is generated and transmitted by the roadside lighthouse sensing subject, and feature fusion is performed by combining Gaussian function to improve the vehicle-side sensing capability.
This approach achieves improved vehicle-side perception performance, particularly in blind spot and long-distance detection capabilities, while reducing communication bandwidth, thus meeting the needs of deployment.
Smart Images

Figure CN116977957B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent driving technology, and in particular to a target detection method based on low-bandwidth vehicle-road feature fusion known to be a lighthouse. Background Technology
[0002] One of the primary challenges of vehicle-road cooperative perception is the communication bandwidth issue among multiple sensing entities, including roadside beacons and vehicles. Roadside beacons need to simultaneously establish networked cooperative relationships with multiple vehicles participating in cooperative perception, requiring the minimization of communication bandwidth. From the perspective of information communication stages, vehicle-road cooperative systems are mainly divided into early-stage cooperation, mid-stage cooperation, and late-stage cooperation. Early-stage cooperation involves communication of raw data, with raw data fusion performed at the input of the perception model. This stage offers the best cooperative perception performance, but the massive amount of raw data communication makes deployment difficult. In late-stage cooperation, each sensing entity performs 3D target detection independently, then shares the communication prediction results for result-level fusion cooperation. Although its perception performance and robustness are the worst, its low communication requirements make it the preferred choice for vehicle-side deployment at this stage. Mid-stage cooperation is a choice that balances perception performance and communication bandwidth. Typically, different sensing entities extract features using perception models and transmit feature-level information for feature fusion.
[0003] How to select features is the key to reducing communication bandwidth and improving perception performance through feature fusion in mid-term collaboration. Although existing mid-term collaboration methods have greatly reduced communication bandwidth compared to early-term collaboration, they are still hundreds to thousands of times more communication than late-term fusion. This is obviously not feasible for deployment under current communication conditions. Summary of the Invention
[0004] In view of this, this application provides a target detection method based on the well-known low-bandwidth vehicle-road feature fusion for lighthouses, in order to solve the above-mentioned technical problems.
[0005] In a first aspect, embodiments of this application provide a target detection method based on low-bandwidth vehicle-road feature fusion, known in lighthouse architecture, applied to a vehicle-side perception subject; including:
[0006] Obtain the perception data at the current moment;
[0007] The perception data at the current moment is processed to obtain the original BEV feature map at the current moment.
[0008] The roadside features of multiple detection boxes sent by the roadside lighthouse sensing subject at time t0 are obtained. The roadside features include: detection box information, channel features, and motion features; t0≤t, where t is the current time.
[0009] Using the motion features in the roadside features of each detection box at time t0, the spatial position of the geometric center point of each detection box at the current time is calculated, thereby obtaining the roadside features of each detection box at the current time.
[0010] The roadside features of each detection box at the current time are processed based on the Gaussian function to obtain the heatmap corresponding to each detection box; the original BEV feature map at the vehicle end and the heatmaps of all detection boxes are fused to obtain the BEV fused feature map at the current time.
[0011] The heatmaps of detection boxes belonging to the same category are processed to obtain several category feature maps;
[0012] The BEV fusion feature map at the current time step is superimposed with multiple category feature maps to obtain the final fusion feature at the current time step.
[0013] The detection head is used to process the final fusion features at the current moment to obtain the target detection result.
[0014] Furthermore, the perceived data is RGB image and / or 3D point cloud data.
[0015] Furthermore, the method also includes:
[0016] The original BEV feature map at the current time t is processed using the detection head to obtain multiple initial target detection boxes;
[0017] Multiple initial target detection boxes are sent to the roadside lighthouse sensing subject.
[0018] Furthermore, the roadside features of the multiple detection boxes are generated by the roadside lighthouse sensing entity, and the generation steps include:
[0019] The roadside sensing data at time t0 is processed using a depth sensing feature extractor to obtain the top-view BEV feature map at time t0.
[0020] The top-view BEV feature map at time t0 is processed using the detection head to obtain the target detection box set object0; the initial target detection box sets sent by M vehicle-side perception subjects are obtained: object1, object2, ... and object0. M ;
[0021] Calculate the set of target detection boxes G sent to the nth vehicle-side sensing subject. n :
[0022]
[0023] Among them, for the object detection box set object mFor any bounding box 'a' in the object, calculate IoU(a, b), where m ≠ n and b ∈ object. n If IoU(a, b) ≤ threshold, then a ∈ object m -object m ∩object n ; Object detection box set G n It includes P detection boxes, and the detection box information of the p-th detection box is represented by a seven-dimensional tensor Z. p =(x p y p , z p dx p dy p dz p , class p ) indicates that (x p y p , z p ) represents the spatial location of the geometric center point of the p-th detection box, (dx) p dy p dz p ) represents the length of the p-th detection box along the three directions, class p Let p be the target category of the detection box.
[0024] For the set of target detection boxes G n For the p-th detection box in the image, calculate the position of its location in the top-view BEV feature map based on the spatial position of its geometric center point, and use the feature at this position as the channel feature Q of the p-th detection box. p ;
[0025] The top-view BEV feature map at time t0 is processed using a motion prediction head to obtain a BEV motion prediction feature map, in which the feature value of each pixel includes the predicted velocity in the x and y directions.
[0026] For the set of target detection boxes G n For the p-th detection box in the BEV motion prediction feature map, the feature at that position is used as the motion feature M of the p-th detection box. p M p =(v p,x v p,y );v p,x v is the velocity in the x-direction; p,y The velocity is in the y-direction;
[0027] The detection box information Z of the p-th detection box p Channel characteristics Q p and motion characteristics M p By concatenating the data, the road-end features f of the p-th detection box are obtained.p Where 1≤p≤P.
[0028] Furthermore, using the motion features in the roadside features of each detection box at time t0, the spatial position of the geometric center point of each detection box at the current time is calculated, thereby obtaining the roadside features of each detection box at the current time; including:
[0029] Calculate the set of object detection boxes G n The spatial position (x) of the geometric center point of the p-th bounding box at the current time. p,t y p,t , z p,t ):
[0030]
[0031] Then the detection box information of the p-th detection box at the current time is Z. p,t =(x p,t y p,t , z p,t dx p dy p dz p , class p );
[0032] The detection box information Z at the current time of the p-th detection box. p,t Channel characteristics Q p and motion characteristics M p By splicing the data, we obtain the roadside features of the p-th detection box at the current time.
[0033] Furthermore, the roadside features of each detection box at the current time are processed based on a Gaussian function to obtain a heatmap corresponding to each detection box; the original BEV feature map at the vehicle end and the heatmaps of all detection boxes are fused to obtain the BEV fused feature map at the current time; including:
[0034] Based on the target detection box set G n The spatial position (x) of the geometric center point of the p-th bounding box at the current time. p,t y p,t , z p,t This allows us to obtain the BEV spatial position corresponding to the geometric center point of the p-th detection box at the current time.
[0035] Calculate the feature value of pixel (i, j) in the heatmap of the p-th detection box using the Gaussian function.
[0036]
[0037] Where δ is the area of the p-th detection box in the BEV space;
[0038] The feature value f1 of pixel (i, j) in the BEV fusion feature map f1 at the current time. ij for:
[0039]
[0040] Among them, f ij Let f be the feature value of pixel (i, j) in the original BEV feature map f at the vehicle end.
[0041] Furthermore, the heatmaps of detection boxes belonging to the same category are processed to obtain several category feature maps; including:
[0042] Based on the categories in the detection box information, several detection boxes belonging to the c-th category are obtained; 1≤c≤NUM, where NUM is the total number of categories;
[0043] The feature value of each pixel in the category feature map of the c-th category is the maximum feature value of the corresponding pixels in the heatmaps of several detection boxes belonging to the c-th category.
[0044] Secondly, embodiments of this application provide a target detection device based on low-bandwidth vehicle-road feature fusion, known to be used in lighthouses, applied to a vehicle-side perception subject; comprising:
[0045] The first acquisition unit is used to acquire the sensing data at the current moment;
[0046] The first processing unit is used to process the perception data at the current time t to obtain the original BEV feature map of the vehicle at the current time.
[0047] The second acquisition unit is used to acquire the roadside features of multiple detection boxes sent by the roadside lighthouse sensing subject at time t0. The roadside features include: detection box information, channel features, and motion features; t0≤t, where t is the current time.
[0048] The motion prediction unit is used to calculate the spatial position of the geometric center point of each detection box at the current time by utilizing the motion features in the roadside features of each detection box at time t0, thereby obtaining the roadside features of each detection box at the current time.
[0049] The first fusion unit is used to process the roadside features of each detection box at the current time based on the Gaussian function to obtain the heatmap corresponding to each detection box; and to fuse the original BEV feature map at the vehicle end and the heatmaps of all detection boxes to obtain the BEV fusion feature map at the current time.
[0050] The second processing unit is used to process the heatmaps of detection boxes belonging to the same category to obtain several category feature maps;
[0051] The second fusion unit is used to superimpose the BEV fusion feature map at the current time and multiple category feature maps to obtain the final fusion feature at the current time.
[0052] The detection unit is used to process the final fusion features at the current moment using the detection head to obtain the target detection result.
[0053] Thirdly, embodiments of this application provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method of embodiments of this application.
[0054] Fourthly, an embodiment of this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method of the embodiment of this application.
[0055] This application reduces the data transmission frame rate and transmission volume from the roadside lighthouse to the vehicle, and improves the detection capability of the vehicle-side sensing subject in blind spots and uncertain areas. Attached Figure Description
[0056] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0057] Figure 1 A schematic diagram of the technical route for a target detection method based on low-bandwidth vehicle-road feature fusion known in lighthouses, provided in an embodiment of this application;
[0058] Figure 2 A flowchart of a target detection method for low-bandwidth vehicle-road feature fusion based on lighthouses, provided in an embodiment of this application;
[0059] Figure 3 A functional structure diagram of a target detection device for low-bandwidth vehicle-road feature fusion based on lighthouses, provided in an embodiment of this application;
[0060] Figure 4 A functional structure diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0061] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0062] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0063] First, a brief introduction to the design concept of the embodiments of this application will be given.
[0064] Vehicle-road cooperative perception refers to the concept of real-time information exchange and cooperation between vehicles and roadside perception devices, aiming to improve road safety, traffic efficiency, and driving experience. Roadside perception devices, also known as roadside beacons, are infrastructure architectures with a "five-knowledge" functional system. These "five-knowledge" functions are: The known functional layer: Roadside beacons consider static maps of intersections, intersection wholesale models, and sensor information. This information provides the basic data for vehicle-road cooperative perception. The cognitive functional layer: Roadside beacons include modules for dynamic intersection maps and localization, traffic participant detection, and traffic participant intent recognition. They can sense and identify traffic participants such as vehicles, pedestrians, and bicycles within the intersection and infer their behavioral intentions. These cognitive functions provide real-time traffic scene information for cooperative perception. The predictive functional layer: Traffic scene prediction, traffic flow anomaly prediction, and traffic violation prediction modules can predict future traffic conditions, flow anomalies, and potential traffic violations based on historical data and real-time sensor information. These predictive capabilities help collaborative perception provide future, temporal traffic scene information; in the knowledge layer, the end-to-end perception model and sensor calibration module are responsible for integrating information from different sensors and data sources to build a comprehensive model for perceiving and understanding the traffic environment. These models are one of the sources of the roadside beacon system's perception capabilities; in the common knowledge layer, the adaptive fusion model module fuses and processes information from different functional layers to generate comprehensive perception information. This is one of the keys to collaborative perception.
[0065] The vehicle-road cooperative perception first involves the vehicle sending a request to the roadside beacon. The roadside beacon's public knowledge layer integrates its own perception with the vehicle's perception results, upgrading the independent cognition of all vehicles participating in the cooperative perception to public knowledge, thereby improving the perception capability of the entire vehicle-road cooperative perception system. Then, the roadside beacon sends effective perception information such as occlusion / beyond line of sight to the vehicle. The vehicle integrates its own perception with the comprehensive perception results, thus enabling the roadside beacon to empower the vehicle.
[0066] Vehicle-to-everything (V2X) 3D target detection technology refers to the real-time detection and identification of 3D targets in the environment surrounding a vehicle through the combined action of the vehicle and the roadside beacon. Currently, single-vehicle autonomous driving systems still suffer from blind spots and weak long-distance perception capabilities. The roadside beacon's public knowledge layer, equipped with "five knowledge" functions, upgrades the independent cognition of all perception subjects in the V2X system to public knowledge, enhancing the overall system's perception capabilities. This comprehensive perception empowers the vehicle, providing crucial environmental perception information for single-vehicle autonomous driving systems, thereby ensuring the safety and reliability of autonomous driving.
[0067] In response to this situation, this application provides a target detection method based on the well-known low-bandwidth vehicle-road feature fusion for lighthouses. This method adopts a feature selection mechanism for target center point features, and selects low-dimensional sparse features at both the time and space levels, keeping the communication bandwidth at several times that of later collaboration, thus making deployment possible.
[0068] This application reduces communication bandwidth by transmitting feature information of sparse, low-dimensional target center points, enabling collaborative perception among multiple entities such as vehicles and infrastructure. It is a vehicle-road cooperative 3D target detection framework suitable for deployment.
[0069] like Figure 1 As shown, the roadside lighthouse sensing subject not only generates a top-down BEV feature map and 3D target detection results, but also receives M vehicle-side requests. In addition to its own detection results, it also uses the initial target results of other vehicle-side sensing subjects to supplement them, selects the spatial features and motion prediction feature maps of the target positions that the vehicle has not detected, and sends them to the vehicle-side sensing subject. This process is one round of vehicle-road communication.
[0070] The vehicle-side perception subject processes the perception data at the current moment to obtain the original BEV feature map of the vehicle at the current moment. It then acquires data sent by the nearest roadside beacon perception subject at the current moment. First, based on the motion prediction feature map, it calculates the spatial position of the geometric center point of the target detection box sent by the roadside at the current moment, thus obtaining the roadside feature of each detection box at the current moment. The roadside feature of each detection box at the current moment is processed based on a Gaussian function to obtain a heatmap corresponding to each detection box. The original BEV feature map of the vehicle-side perception subject and the heatmaps of all detection boxes are fused to obtain the BEV fused feature map of the current moment. The heatmaps of detection boxes belonging to the same category are processed to obtain several category feature maps. The BEV fused feature map of the current moment is superimposed with multiple category feature maps to obtain the final fused feature of the current moment. The detection head processes the final fused feature of the current moment to obtain the target detection result. This enables the roadside beacon perception subject to enable the detection of blind spots and uncertain areas on the vehicle-side, ultimately obtaining the target detection result. The training of this deep network system is achieved through supervision using the fused labels of all subjects.
[0071] The advantages of this application are:
[0072] 1. In the context of vehicle-road cooperative applications, cooperative vehicles can serve as an extension of the roadside lighthouse's sensing entity, enhancing the sensing capabilities of the roadside lighthouse entity.
[0073] 2. From the perspective of feature fusion, a large number of background pixels on the BEV feature map are redundant and invalid information. Only the location features of the target key point (center point) are transmitted, which has a bandwidth advantage on the basis of achieving high collaborative perception accuracy.
[0074] 3. From a temporal perspective, the location features of target key points in adjacent frames are similar / traceable. By interpolating target features in adjacent frames based on motion prediction, the communication transmission frequency can be reduced, thereby reducing the communication bandwidth.
[0075] After introducing the application scenarios and design concepts of the embodiments of this application, the technical solutions provided by the embodiments of this application will be described below.
[0076] like Figure 2 As shown, this application provides a target detection method based on the well-known low-bandwidth vehicle-road feature fusion for lighthouse applications, comprising the following steps:
[0077] Step 101: Obtain the current sensing data;
[0078] The perceived data is RGB images and / or 3D point cloud data.
[0079] Step 102: Process the perception data at the current moment to obtain the original BEV feature map of the vehicle at the current moment;
[0080] The vehicle-side sensing entity encodes the sensing data at the current time t, representing it as a top-down BEV feature map F∈R in a unified Cartesian coordinate system. W×H×C W×H represents the size of the top-view BEV feature map, and C represents the number of channels; this is to avoid complex coordinate system transformations and facilitate the collaborative sharing of vehicle-road features.
[0081] In addition, the vehicle-side sensing entity inputs the top-down BEV features into the detection head according to the preset communication frame rate to perform preliminary 3D target detection, and the resulting initial target detection box is packaged as a vehicle-side request and sent to the roadside lighthouse sensing entity.
[0082] Step 103: Obtain the roadside features of multiple detection boxes sent by the roadside lighthouse sensing subject at time t0. The roadside features include: detection box information, channel features, and motion features; t0≤t, where t is the current time.
[0083] For roadside lighthouse sensing entities, transmitting the entire feature map of the roadside would result in a large amount of redundant background information leading to ineffective communication. Therefore, it is necessary to filter effective sparse features. A common method is pixel-level filtering, filtering redundant background pixel areas to save valuable bandwidth. However, the bandwidth required for this feature selection method is far higher than the bandwidth limitations of on-site deployment.
[0084] Vehicle-to-infrastructure (V2I) communication typically occurs at a frequency of 10 Hz, meaning 10 V2I communications per second. Motion prediction feature maps are data representations used to predict the future motion of objects in a scene. Using prediction features can avoid high-frequency communication. By setting the communication frequency to 2 Hz, the frequency of communication is reduced from the time dimension, thereby reducing the total bandwidth of communication per second.
[0085] Specifically, the roadside features of the multiple detection boxes are generated by the roadside lighthouse sensing entity, and the generation steps include:
[0086] The roadside sensing data at time t0 is processed using a depth sensing feature extractor to obtain the top-view BEV feature map at time t0.
[0087] The top-view BEV feature map at time t0 is processed using the detection head to obtain the target detection box set object0; the initial target detection box sets sent by M vehicle-side perception subjects are obtained: object1, object2, ... and object0. M ;
[0088] Calculate the set of target detection boxes G sent to the nth vehicle-side sensing subject. n :
[0089]
[0090] Among them, for the object detection box set object m For any bounding box 'a' in the object, calculate IoU(a, b), where m ≠ n and b ∈ object. n If IoU(a, b) ≤ threshold, then a ∈ object m -object m ∩object n ; Object detection box set G n It includes P detection boxes, and the detection box information of the p-th detection box is represented by a seven-dimensional tensor Z. p =(x p y p , z p dx p dy p dz p , class p ) indicates that (x p y p , z p ) represents the spatial location of the geometric center point of the p-th detection box, (dx) p dy p dz p ) represents the length of the p-th detection box along the three directions, class p Let p be the target category of the detection box.
[0091] For the set of target detection boxes G n For the p-th detection box in the image, calculate the position of its location in the top-view BEV feature map based on the spatial position of its geometric center point, and use the feature at this position as the channel feature Q of the p-th detection box. p ∈R 1×C :
[0092] The following formula can be used to calculate (x) p y p , z p At the location (x′, y′) of the BEV feature map:
[0093]
[0094]
[0095] Where s is the real-world length corresponding to each pixel on the BEV plane, and x min and y min It represents the minimum value in the x and y directions in the real-world coordinate system;
[0096] The top-view BEV feature map at time t0 is processed using a motion prediction head to obtain a BEV motion prediction feature map, in which the feature value of each pixel includes the predicted velocity in the x and y directions.
[0097] For the set of target detection boxes G n For the p-th detection box in the BEV motion prediction feature map, the feature at that position is used as the motion feature M of the p-th detection box. p M p =(v p,x v p,y );v p,x v is the velocity in the x-direction; p,y The velocity is in the y-direction;
[0098] The detection box information Z of the p-th detection box p Channel characteristics Q p and motion characteristics M p By concatenating the data, the road-end features f of the p-th detection box are obtained. p :
[0099] f p =Z p concat Q p concat M p
[0100] Where 1≤p≤P.
[0101] The feature dimension of the transmission is f p ∈R 1×(7+C+2) Typically, C=80 is set, which is only about 12 times more than the later collaboration; at the same time, the communication frequency is reduced by 4 times, that is, the total transmission bandwidth per second is only 2.5 times that of the later collaboration, which is more than enough to meet the requirements for implementation.
[0102] Step 104: Using the motion features in the roadside features of each detection box at time t0, calculate the spatial position of the geometric center point of each detection box at the current time, thereby obtaining the roadside features of each detection box at the current time;
[0103] Since the vehicle-side target detection frequency is 10Hz, while the roadside information transmission frequency is only 2Hz, when there is no data frame with roadside information transmission at the current time t, the spatial position of the geometric center of the detection box can be predicted by using the motion vector and time difference in the features of the previous frame transmitted by the roadside at time t0, thus obtaining the position information of the center point of the detection box at the current time.
[0104] Specifically, the steps include:
[0105] Calculate the set of object detection boxes G nThe spatial position (x) of the geometric center point of the p-th bounding box at the current time. p,t y p,t , z p,t ):
[0106]
[0107] Then the detection box information of the p-th detection box at the current time is Z. p,t =(x p,t y p,t , z p,t dx p dy p dz p , class p );
[0108] The detection box information Z at the current time of the p-th detection box. p,t Channel characteristics Q p and motion characteristics M p By splicing the data, we obtain the roadside features of the p-th detection box at the current time.
[0109] Step 105: Process the roadside features of each detection box at the current time based on the Gaussian function to obtain the heatmap corresponding to each detection box; fuse the original BEV feature map at the vehicle end and the heatmaps of all detection boxes to obtain the BEV fused feature map at the current time.
[0110] In this embodiment, the point features are smoothed into heatmap features with the center point of the detection box as the origin and the size of the detection box as the radius δ, to prevent the single-pixel features from being overwhelmed in the deep neural network. Feature fusion is a key step in improving the perception performance of vehicle-road cooperative systems. Sparse feature selection strategies require appropriate feature fusion strategies to achieve effective information complementarity and realize the role of roadside empowerment of vehicle-side systems.
[0111] This step specifically includes:
[0112] Based on the target detection box set G n The geometric center point (x) of the p-th detection box at the current time. p,t y p,t , z p,t This allows us to obtain the BEV spatial position corresponding to the geometric center point of the p-th detection box at the current time.
[0113] Calculate the feature value of pixel (i, j) in the heatmap of the p-th detection box using the Gaussian function.
[0114]
[0115] Where δ is the area of the p-th detection box in the BEV space;
[0116] The feature value f1 of pixel (i, j) in the BEV fusion feature map f1 at the current time. ij for:
[0117]
[0118] Among them, f ij Let f be the feature value of pixel (i, j) in the original BEV feature map f at the vehicle end.
[0119] Step 106: Process the heatmaps of detection boxes belonging to the same category to obtain several category feature maps;
[0120] In this embodiment, the step includes:
[0121] Based on the categories in the detection box information, several detection boxes belonging to the c-th category are obtained; 1≤c≤NUM, where NUM is the total number of categories;
[0122] The feature value of each pixel in the category feature map of the c-th category is set to be the maximum feature value of the corresponding pixel in the heatmap of several detection boxes belonging to the c-th category.
[0123] Step 107: Overlay the BEV fusion feature map at the current time step with multiple category feature maps to obtain the final fusion feature at the current time step;
[0124] Among them, the final fusion features
[0125] Step 108: Use the detection head to process the final fusion features at the current moment to obtain the target detection result.
[0126] The final fused features at the current moment are input into the detection head to obtain the detection results. By using information from the roadside and other vehicles, the problems of blind spots and weak long-distance perception of the vehicle are greatly improved.
[0127] Compared to current vehicle-to-vehicle collaborative methods: In vehicle-to-road beacon communication, this application only transmits 3D detection bounding box information, with the same transmission bandwidth as later collaborative methods; In roadside beacon feedback to the vehicle, the feature selection method proposed in this application only increases the communication bandwidth by slightly more than 1 times, which is fully feasible for deployment. Moreover, the feature fusion method has higher robustness and accuracy in fusion of detection bounding box results compared to later collaborative methods; Furthermore, it does not require dense communication between vehicles and roads. Through the roadside beacon sensing subject, collaborative vehicles can act as extensions of the roadside beacon sensing subject, utilizing the sensing information of other vehicles and reducing communication between vehicles.
[0128] Based on the above embodiments, this application provides a target detection device for low-bandwidth vehicle-road feature fusion, as known in lighthouses, see [reference]. Figure 3 As shown, the target detection device 200 for low-bandwidth vehicle-road feature fusion based on lighthouses provided in this application embodiment includes at least:
[0129] The first acquisition unit 201 is used to acquire the perception data at the current moment;
[0130] The first processing unit 202 is used to process the perception data at the current moment to obtain the original BEV feature map of the vehicle at the current moment.
[0131] The second acquisition unit 203 is used to acquire roadside features of multiple detection boxes sent by the roadside lighthouse sensing subject at time t0. The roadside features include: detection box information, channel features and motion features; t0≤t, where t is the current time.
[0132] The motion prediction unit 204 is used to calculate the spatial position of the geometric center point of each detection box at the current time by using the motion features in the road end features of each detection box at time t0, thereby obtaining the road end features of each detection box at the current time.
[0133] The first fusion unit 205 is used to process the roadside features of each detection box at the current time based on the Gaussian function to obtain the heat map corresponding to each detection box; and to fuse the original BEV feature map at the vehicle end and the heat map of all detection boxes to obtain the BEV fusion feature map at the current time.
[0134] The second processing unit 206 is used to process the heatmaps of detection boxes belonging to the same category to obtain several category feature maps;
[0135] The second fusion unit 207 is used to superimpose the BEV fusion feature map at the current time and multiple category feature maps to obtain the final fusion feature at the current time.
[0136] The detection unit 208 is used to process the final fusion features at the current moment using the detection head to obtain the target detection result.
[0137] It should be noted that the principle of the target detection device 200 for low-bandwidth vehicle-road feature fusion for lighthouses provided in this application embodiment to solve the technical problem is similar to the method provided in this application embodiment. Therefore, the implementation of the target detection device 200 for low-bandwidth vehicle-road feature fusion for lighthouses provided in this application embodiment can refer to the implementation of the method provided in this application embodiment, and the repeated parts will not be described again.
[0138] Based on the above embodiments, this application also provides an electronic device, see below. Figure 4As shown, the electronic device 300 provided in this application embodiment includes at least: a processor 301, a memory 302, and a computer program stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program, it implements the target detection method for lighthouse-oriented low-bandwidth vehicle-road feature fusion provided in this application embodiment.
[0139] The electronic device 300 provided in this application embodiment may further include a bus 303 connecting different components (including processor 301 and memory 302). The bus 303 represents one or more types of bus structures, including memory bus, peripheral bus, local area bus, etc.
[0140] The memory 302 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 3021 and / or cache memory 3022, and may further include read-only memory (ROM) 3023.
[0141] The memory 302 may also include a program tool 3025 having a set (at least one) of program modules 3024, including but not limited to: an operating subsystem, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0142] Electronic device 300 can also communicate with one or more external devices 304 (e.g., keyboard, remote control, etc.), and with one or more devices that enable a user to interact with electronic device 300 (e.g., mobile phone, computer, etc.), and / or with any device that enables electronic device 300 to communicate with one or more other electronic devices 300 (e.g., router, modem, etc.). This communication can be performed through input / output (I / O) interface 305. Furthermore, electronic device 300 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) through network adapter 306. Figure 4 As shown, network adapter 306 communicates with other modules of electronic device 300 via bus 303. It should be understood that, although... Figure 4As not shown, other hardware and / or software modules may be used in conjunction with electronic device 300, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, Redundant Arrays of Independent Disks (RAID) subsystems, tape drives, and data backup storage subsystems.
[0143] It should be noted that, Figure 3 The electronic device 300 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0144] This application also provides a computer-readable storage medium storing computer instructions. When executed by a processor, these instructions implement the target detection method for low-bandwidth vehicle-road feature fusion based on lighthouse-specific knowledge provided in this application. Specifically, the executable program can be built into or installed in an electronic device 300, allowing the electronic device 300 to implement the target detection method for low-bandwidth vehicle-road feature fusion based on lighthouse-specific knowledge provided in this application by executing the built-in or installed executable program.
[0145] The method provided in this application embodiment can also be implemented as a program product, which includes program code. When the program product can be run on the electronic device 300, the program code is used to enable the electronic device 300 to execute the target detection method for lighthouse-oriented low-bandwidth vehicle-road feature fusion provided in this application embodiment.
[0146] The program product provided in this application embodiment can be any combination of one or more readable media, wherein the readable media can be a readable signal medium or a readable storage medium, and the readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. Specifically, more specific examples of readable storage media (a non-exhaustive list) include: electrical connections with one or more wires, portable disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0147] The program product provided in this application embodiment can be a CD-ROM and include program code, and can also run on a computing device. However, the program product provided in this application embodiment is not limited thereto. In this application embodiment, the readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0148] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.
[0149] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit it. Although this application has been described in detail with reference to the embodiments, those skilled in the art should understand that modifications or equivalent substitutions to the technical solutions of this application do not depart from the spirit and scope of the technical solutions of this application, and should all be covered within the scope of the claims of this application.
Claims
1. A target detection method based on well-known low-bandwidth vehicle-road feature fusion for lighthouse applications, applied to vehicle-side sensing entities; characterized in that, include: Obtain the perception data at the current moment; The perception data at the current moment is processed to obtain the original BEV feature map at the current moment. The roadside features of multiple detection boxes sent by the roadside lighthouse sensing subject at time t0 are obtained. The roadside features include: detection box information, channel features, and motion features; t0≤t, where t is the current time. Using the motion features in the roadside features of each detection box at time t0, the spatial position of the geometric center point of each detection box at the current time is calculated, thereby obtaining the roadside features of each detection box at the current time. The roadside features of each detection box at the current time are processed based on the Gaussian function to obtain the heatmap corresponding to each detection box; the original BEV feature map at the vehicle end and the heatmaps of all detection boxes are fused to obtain the BEV fused feature map at the current time. The heatmaps of detection boxes belonging to the same category are processed to obtain several category feature maps; The BEV fusion feature map at the current time step is superimposed with multiple category feature maps to obtain the final fusion feature at the current time step. The detection head is used to process the final fusion features at the current moment to obtain the target detection result.
2. The method according to claim 1, characterized in that, The perceived data is RGB images and / or 3D point cloud data.
3. The method according to claim 1, characterized in that, The method further includes: The original BEV feature map at the current time t is processed using the detection head to obtain multiple initial target detection boxes; Multiple initial target detection boxes are sent to the roadside lighthouse sensing subject.
4. The method according to claim 3, characterized in that, The roadside features of the multiple detection boxes are generated by the roadside lighthouse sensing entity, and the generation steps include: The roadside sensing data at time t0 is processed using a depth sensing feature extractor to obtain the top-view BEV feature map at time t0. The top-view BEV feature map at time t0 is processed using the detection head to obtain the target detection box set object0; the initial target detection box sets sent by M vehicle-side perception subjects are obtained: object1, object2, ... and object0. M ; Calculate the set of target detection boxes G sent to the nth vehicle-side sensing subject. n : Among them, for the object detection box set object m For any bounding box 'a' in the object, calculate IoU(a, b), where m ≠ n and b ∈ object. n If IoU(a, b) ≤ threshold, then a ∈ object m -object m ∩object n ; Object detection box set G n It includes P detection boxes, and the detection box information of the p-th detection box is represented by a seven-dimensional tensor Z. p =(x p y p , z p dx p dy p dz p , class p ) indicates that (x p y p , z p ) represents the spatial location of the geometric center point of the p-th detection box, (dx) p dy p dz p ) represents the length of the p-th detection box along the three directions, class p Let p be the target category of the detection box. For the set of target detection boxes G n For the p-th detection box in the image, calculate the position of its location in the top-view BEV feature map based on the spatial position of its geometric center point, and use the feature at this position as the channel feature Q of the p-th detection box. p ; The top-view BEV feature map at time t0 is processed using a motion prediction head to obtain a BEV motion prediction feature map, in which the feature value of each pixel includes the predicted velocity in the x and y directions. For the set of target detection boxes G n For the p-th detection box in the BEV motion prediction feature map, the feature at that position is used as the motion feature M of the p-th detection box. p M p =(v p,x v p,y );v p,x v is the velocity in the x-direction; p,y The velocity is in the y-direction; The detection box information Z of the p-th detection box p Channel characteristics Q p and motion characteristics M p By concatenating the data, the road-end features f of the p-th detection box are obtained. p Where 1≤p≤P.
5. The method according to claim 4, characterized in that, Using the motion features in the roadside features of each detection box at time t0, the spatial position of the geometric center point of each detection box at the current time is calculated, thereby obtaining the roadside features of each detection box at the current time; including: Calculate the set of object detection boxes G n The spatial position (x) of the geometric center point of the p-th bounding box at the current time. p,t y p,t , z p,t ): Then the detection box information of the p-th detection box at the current time is Z. p,t =(x p,t y p,t , z p,t dx p dy p dz p , class p ); The detection box information Z at the current time of the p-th detection box. p,t Channel characteristics Q p and motion characteristics M p By splicing the data, we obtain the roadside features of the p-th detection box at the current time.
6. The method according to claim 5, characterized in that, The road-end features of each detection box at the current time are processed based on the Gaussian function to obtain the heatmap corresponding to each detection box; The original BEV feature map from the vehicle end and the heatmaps of all detection boxes are fused to obtain the BEV fused feature map at the current time; including: Based on the target detection box set G n The spatial position (x) of the geometric center point of the p-th bounding box at the current time. p,t y p,t , z p,t This allows us to obtain the BEV spatial position corresponding to the geometric center point of the p-th detection box at the current time. Calculate the feature value of pixel (i, j) in the heatmap of the p-th detection box using the Gaussian function. Where δ is the area of the p-th detection box in the BEV space; The feature value f1 of pixel (i, j) in the BEV fusion feature map f1 at the current time. ij for: Among them, f ij Let f be the feature value of pixel (i, j) in the original BEV feature map f at the vehicle end.
7. The method according to claim 6, characterized in that, The heatmaps of detection boxes belonging to the same category are processed to obtain several category feature maps, including: Based on the categories in the detection box information, several detection boxes belonging to the c-th category are obtained; 1≤c≤NUM, where NUM is the total number of categories; The feature value of each pixel in the category feature map of the c-th category is set to be the maximum feature value of the corresponding pixel in the heatmap of several detection boxes belonging to the c-th category.
8. A target detection device based on low-bandwidth vehicle-road feature fusion, known for its application to lighthouse-based sensing entities; characterized in that, include: The first acquisition unit is used to acquire the sensing data at the current moment; The first processing unit is used to process the perception data at the current time t to obtain the original BEV feature map of the vehicle at the current time. The second acquisition unit is used to acquire the roadside features of multiple detection boxes sent by the roadside lighthouse sensing subject at time t0. The roadside features include: detection box information, channel features, and motion features; t0≤t, where t is the current time. The motion prediction unit is used to calculate the spatial position of the geometric center point of each detection box at the current time by utilizing the motion features in the roadside features of each detection box at time t0, thereby obtaining the roadside features of each detection box at the current time. The first fusion unit is used to process the roadside features of each detection box at the current time based on the Gaussian function to obtain the heatmap corresponding to each detection box; and to fuse the original BEV feature map at the vehicle end and the heatmaps of all detection boxes to obtain the BEV fusion feature map at the current time. The second processing unit is used to process the heatmaps of detection boxes belonging to the same category to obtain several category feature maps; The second fusion unit is used to superimpose the BEV fusion feature map at the current time and multiple category feature maps to obtain the final fusion feature at the current time. The detection unit is used to process the final fusion features at the current moment using the detection head to obtain the target detection result.
9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as claimed in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Vehicle environment information sensing method and device, electronic equipment and storage medium
CN112396043A
Vehicle-road cooperation target detection method based on multi-sensor fusion
CN115775378A