Multi-Agent Cooperative Perception Method Based on LiDAR in Complex Noise
Through the multi-agent collaborative perception method based on lidar, the detection range and robustness of traditional single-agent perception systems in complex noise environments are solved, and target detection with higher accuracy and longer distance is achieved.
Patent Information
- Application Number
- CN202410760827.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-13
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-06-13
AI Technical Summary
Traditional single agent perception systems have limited detection range in complex noise environments, poor long-distance detection effect, blind spots in field of view, and multi-agent collaborative perception methods lack robustness in complex noise environments, resulting in a degradation of detection performance.
The multi-agent collaborative perception method based on lidar is adopted. By converting and preprocessing the coordinate system of the collaborative agent lidar point cloud, feature maps are extracted and compressed and transmitted using deep learning, and the central vehicle is decompressed and deep learning is performed. Combining the attention mechanism and multi-agent window technology, multi-agent feature maps are integrated to enhance robustness.
The detection accuracy and perception range in complex noise environments are improved, the robustness of transmission time delay and relative pose errors is enhanced, and effective perception of longer distances and obscured targets is achieved.
Smart Images

Figure CN118731975B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving target detection, and particularly relates to a multi-agent collaborative perception method based on lidar under complex noise. Background Art
[0002] In a complex perception environment, accurate 3D target detection is crucial for ensuring road safety and traffic efficiency in autonomous driving, as it provides important information about the environment around the autonomous vehicle. In recent years, the perception capabilities of single-agent perception systems have been rapidly developed through deep learning technologies, showing great advantages in target detection tasks. Nevertheless, traditional single-agent perception systems still pose challenges.
[0003] The detection range of traditional single-agent perception systems is very limited. Due to the complexity of real roads and the heavy dependence of traditional single-agent 3D target detection on the perception data of in-vehicle sensors (such as lidar) of the central vehicle, its perception range has very obvious limitations. First, the detection range of a single in-vehicle lidar sensor is very limited, and the detection effect at long distances is poor; second, there are many blind spots in the field of vision during the driving process of autonomous vehicles. However, these problems are often the source of traffic accidents and affect traffic efficiency. In summary, traditional single-agent perception systems cannot solve these problems. Therefore, in order to solve the problems existing in the single-vehicle perception system, multi-agent collaborative perception systems have become one of the very effective methods. Collaborative perception enables the central vehicle to obtain more information to improve its perception performance by using the data of additional agent sensors (such as vehicles, infrastructure) through communication applications oriented to vehicle-to-everything (V2X).
[0004] Current multi-agent collaborative perception methods lack a certain degree of robustness against complex noise environments. In the actual application scenarios of autonomous driving collaborative perception, the communication system will inevitably be interfered by the surrounding environment, thus affecting the information sharing between collaborative agents. The resulting time asynchrony problem will lead to a decline in the detection performance of autonomous vehicles. At the same time, in existing in-vehicle positioning modules, the 6 DoF poses of the positioning estimates are not accurate, which will further lead to errors in the coordinate system conversion between the acquisition of agent sensor data and the acquired data. Inaccurate relative poses will cause misalignment of collaborative data, resulting in a decline in collaborative perception performance, and even being worse than single-agent perception performance.
[0005] In summary, designing an end-to-end multi-agent collaborative perception method with high detection accuracy and a certain degree of robustness against complex noise environments is an urgent problem to be solved. Summary of the Invention
[0006] The present invention provides a multi-agent collaborative perception method based on lidar, which solves the problem of performance degradation of the multi-agent perception system under transmission time delay and relative pose error.
[0007] The specific technical solution is as follows:
[0008] A multi-agent collaborative perception method based on lidar under complex noise includes the following steps:
[0009] Step S1: Perform coordinate transformation and preprocessing on the lidar point cloud of the cooperative agents;
[0010] Step S2: Use the method of deep learning to extract the feature maps of the lidar point clouds of all cooperative agents, and then compress the feature maps of the lidar point clouds of the cooperative agents and transmit them to the central vehicle;
[0011] Step S3: The central vehicle decompresses the received feature maps of the lidar point clouds of the cooperative agents, and at the same time uses a small amount of previously received feature maps of the lidar point clouds of the cooperative agents to perform corresponding deep learning processing on the transmission time delay and relative pose error;
[0012] Step S4: The central vehicle uses the attention mechanism to fuse the feature maps of the lidar point clouds of the cooperative agents into the feature map of the central vehicle to obtain the final feature map;
[0013] Step S5: Send the final feature map into the detection head to obtain the detection result.
[0014] Preferably, the coordinate transformation of the lidar point cloud of the cooperative agents in step S1 is specifically as follows: after receiving the metadata broadcast by the central vehicle, the cooperative agents project their local lidar point clouds into the coordinate system of the central vehicle;
[0015] The preprocessing includes removing the lidar point clouds near non-ground and collision detection.
[0016] Preferably, the backbone network for extracting the feature maps of the lidar point clouds of all cooperative agents by using the method of deep learning in step S2 is the PointPillar encoder; the compressor for feature map compression is the Conv-BN-ReLU block.
[0017] Preferably, the central vehicle decompresses the received feature maps of the lidar point clouds of the cooperative agents in step S3, specifically, a decompressor mainly composed of Conv-BN-ReLU blocks decompresses the received feature maps transmitted by the cooperative agents.
[0018] Preferably, the corresponding deep learning processing of the transmission time delay in step S3 includes the following steps:
[0019] Step S311: Combine a small amount of historical feature maps that are temporally continuous with the agent's transmitted feature map to form a temporally continuous feature map F t ∈R T×H×W×C ;
[0020] Step S312: Generate dynamic temporal kernels for the temporally continuous feature map of each agent, and learn the temporal context information in the continuous feature map in an adaptive manner to obtain the processed feature map Ht;
[0021] Step S313: To enhance the temporal information of the current transmitted feature map, fuse the current transmitted feature map t in the processed feature map H with the historical feature map :
[0022]
[0023]
[0024] where Mean and Max respectively represent taking the average and maximum value of the feature map in the channel dimension, represents the processed historical feature map, represents the fused current transmitted feature map, δ represents the Sigmoid function, || represents the concatenation operation, ⊙ represents the element-wise multiplication, and Conv represents the convolution.
[0025] Step S312 includes the following steps:
[0026] Step S3121: Use AveragePooling to compress the temporally continuous feature map F t in the spatial dimension to obtain the compressed feature map
[0027] Step S3122: The compressed feature map generates dynamic convolution kernels and weights sensitive to temporal positions through different convolution blocks respectively:
[0028]
[0029] where Y represents the feature map, Conv t represents the per-channel temporal convolution, B1 and B2 represent different convolution operations, and ⊙ represents the element-wise multiplication;
[0030] Step S3123: After performing temporal modeling on the temporally continuous feature map F t enrich all feature maps through receptive fields of different sizes.
[0031] Preferably, Step S3123 includes the following steps:
[0032] Step S31231: Divide the feature map Y along the channel dimension to obtain the divided feature map
[0033] Step S31232: According to the obtained divided feature map Use different convolutions to obtain spatial attention maps under different receptive fields respectively:
[0034]
[0035]
[0036] Among them, GN represents GroupNorm, δ represents the Sigmoid function, Conv1 and Conv3 represent convolutions of 1×1 and 3×3 respectively, || represents the concatenation operation, g h , g w represent global average pooling in different one-dimensional directions, Y g represents the spatial attention map under a larger receptive field, Y l represents the spatial attention map under a smaller receptive field, represents the divided feature map;
[0037] Step S31233: Aggregate the spatial attention maps within different channel dimensions, and then use the Sigmoid function and the feature map Y to obtain the processed feature map H t ∈R T×H×W×C :
[0038] H t =δ(g(Y l )⊙μ(Y g )+g(Y g )⊙μ(Y l ))⊙Y
[0039] Among them, g represents two-dimensional global average pooling, μ represents the Softmax function, Y g represents the spatial attention map under a larger receptive field, Y l represents the spatial attention map under a smaller receptive field, δ represents the Sigmoid function.
[0040] Preferably, the corresponding deep learning processing of the relative pose error in step S3 includes the following steps:
[0041] Step S321: Use multiple windows of different scale sizes and utilize multi-head self-attention respectively to capture feature information of different degrees. The large-scale window provides coarse-grained feature information to compensate for significant relative pose errors, while the small-scale window provides fine-grained feature information to compensate for more refined relative pose errors;
[0042] Step S322: Cross-learn between the same-scale window features of different agents;
[0043]
[0044]
[0045] Among them, Q, K, and V are the feature maps of all agents After being projected into different feature spaces, they form query vectors, key vectors, and value vectors. s ∈ [s1, s2, s3] represents different scale windows, m represents the number of attention heads, and represent the attention weights for cross-learning of different agents, d k represents the scaling factor, || represents the concatenation operation, represents the Q after completing window scale rearrangement, represents the matrix for window scale rearrangement of Q, represents the K after completing window scale rearrangement, represents the matrix for window scale rearrangement of K, represents the V after completing window scale rearrangement, represents the matrix for window scale rearrangement of V, h represents the maximum number of attention heads, μ represents the Softmax function, represents the scale window feature map;
[0046] Step S323: Aggregate adaptively the feature maps and processed by three scale windows to obtain the fused feature map H ms ∈ R L×H×W×C ;
[0047]
[0048] Among them, f(·) represents the convolution function; H ms represents the fused feature map.
[0049] Preferably, step S4 includes the following steps:
[0050] Step S41: The central vehicle feature H ego , the cooperative vehicle feature and the infrastructure feature H infraCombined into different feature sets;
[0051] Step S42: After completing the feature set grouping, perform feature fusion on the feature sets within each group using deformable cross-attention.
[0052] Preferably, step S42 includes the following steps:
[0053] Step S421: Use all the features within the group to obtain a reference center position map p, which represents the potential position of the detection target within the detection range;
[0054] Step S422: Use the reference center position map p to extract the features of the central vehicle at the reference point as the initial query embedding of the deformable cross-attention layer;
[0055] Step S423: Then use a linear layer to learn the two-dimensional offset Δp of the reference point through the reference center position map p. Through K key points at Δp + p, use bilinear sampling to extract its features as the participating features of the deformable cross-attention layer;
[0056] Step S424: Fill the output of the deformable cross-attention layer into the central vehicle features:
[0057]
[0058] H res = Fill(f defor (p), H ego )
[0059] where m represents the number of heads, W m / a represents W m learnable weights, l represents the number of cooperative agents, k represents the number of key points, μ represents the Softmax function, W a represents learnable weights, H ego (p) represents the features extracted by the central vehicle through the map p, represents the features extracted by the cooperative agent through the map p + Δp, f defor (p) represents the output of the deformable cross-attention layer, Fill(·) represents the filling operation, H ego represents the central vehicle features, and μ represents the Softmax function;
[0060] Preferably, step S5 includes the following steps:
[0061] Step S51: Feed the multiple groups of fused features H res into two different detection decoders to decode them into detection results, which are multiple groups of classification and regression outputs respectively;
[0062] Step S52: Calculate the mean of multiple sets of classification and regression output results to obtain the final detection result.
[0063] Compared with the prior art, the present invention has the following beneficial effects:
[0064] 1. The present invention uses multi-historical feature maps of multiple agents to simultaneously solve the problems of transmission time delay and relative pose error in cooperative perception; captures temporal context semantic information by simultaneously using the transmitted feature maps received by the central vehicle and a small number of consecutive historical feature maps; and finally enhances the robustness to time delay and relative pose error by capturing spatial information through multi-scale windows.
[0065] 2. The present invention gives full play to the respective advantages of cooperative vehicles, infrastructure, and the central vehicle to achieve efficient cooperative perception. All agents are equipped with lidar. The cooperative agents "cooperative vehicles and infrastructure" cooperate with the central vehicle to perceive farther distances and targets occluded by other objects; greatly improves the protection range of the comprehensive perception performance of the central vehicle. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] To more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0067] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention.
[0068] Figure 2 It is a schematic structural diagram of the PointPillars network for extracting high-dimensional feature vectors from 3D point cloud data according to an embodiment of the present invention.
[0069] Figure 3 It is a schematic structural diagram of the network framework for transmission time delay according to an embodiment of the present invention.
[0070] Figure 4 It is a schematic structural diagram of the network framework for relative pose error according to an embodiment of the present invention.
[0071] Figure 5 It is a schematic structural diagram of the network framework for grouping feature sets and feature fusion according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made. These all belong to the present invention.
[0073] A multi-agent collaborative perception method under complex noise based on lidar includes the following steps:
[0074] Step S1: Perform coordinate system transformation and preprocessing on the lidar point cloud of the collaborative agent;
[0075] Step S2: Use deep learning methods to extract the feature maps of the lidar point clouds of all collaborative agents, and then compress the feature maps of the lidar point clouds of the collaborative agents and transmit them to the central vehicle;
[0076] Step S3: The central vehicle decompresses the received feature maps of the lidar point clouds of the collaborative agents, and at the same time uses the feature maps of the lidar point clouds of a small number of previously received collaborative agents to perform corresponding deep learning processing on the transmission time delay and relative pose error;
[0077] Step S4: The central vehicle uses the attention mechanism to fuse the feature maps of the lidar point clouds of the collaborative agents into the feature map of the central vehicle to obtain the final feature map;
[0078] Step S5: Send the final feature map into the detection head to obtain the detection result.
[0079] In step S1 of this implementation scheme, the coordinate system transformation of the lidar point cloud of the collaborative agent is specifically as follows: After receiving the metadata broadcast by the central vehicle, the collaborative agents project their local lidar point clouds into the coordinate system of the central vehicle;
[0080] The preprocessing includes removing the lidar point clouds not near the ground and collision detection.
[0081] In step S2 of this implementation scheme, the backbone network for using deep learning methods to extract the feature maps of the lidar point clouds of all collaborative agents is the PointPillar encoder; the compressor for the feature map compression is the Conv-BN-ReLU block.
[0082] In step S3 of this implementation scheme, the central vehicle decompresses the received feature maps of the lidar point clouds of the collaborative agents specifically by using a decompressor mainly composed of Conv-BN-ReLU blocks to decompress the received feature maps transmitted by the collaborative agents.
[0083] The corresponding deep learning processing of the transmission time delay in step S3 of this implementation scheme includes the following steps:
[0084] Step S311: Combine a small amount of historical feature maps that are temporally continuous with the agent's transmitted feature map to form a temporally continuous feature map F t ∈R T×H×W×C ;
[0085] Step S312: Generate a dynamic temporal kernel for the temporally continuous feature map of each agent, and learn the temporal context information in the continuous feature map in an adaptive manner to obtain the processed feature map Ht;
[0086] Step S313: To enhance the temporal information of the current transmitted feature map, fuse the current transmitted feature map in the processed feature map H t with the historical feature map : :
[0087]
[0088]
[0089] where Mean and Max represent taking the average and maximum values of the feature map in the channel dimension respectively, represents the processed historical feature map, represents the fused current transmitted feature map, δ represents the Sigmoid function, || represents the concatenation operation, ⊙ represents the element-wise multiplication, and Conv represents the convolution.
[0090] Step S312 of this implementation scheme includes the following steps:
[0091] Step S3121: Use AveragePooling to compress the temporally continuous feature map F t in the spatial dimension to obtain the compressed feature map
[0092] Step S3122: The compressed feature map generates a dynamic convolution kernel and a weight sensitive to temporal position through different convolution blocks respectively:
[0093]
[0094] where Y represents the feature map, Conv t represents the per-channel temporal convolution, B1 and B2 represent different convolution operations, and ⊙ represents the element-wise multiplication;
[0095] Step S3123: After performing temporal modeling on the temporally continuous feature map F t , enrich all feature maps through receptive fields of different sizes;
[0096] Step S3123 includes the following steps:
[0097] Step S31231: Divide the feature map Y along the channel dimension to obtain the divided feature map
[0098] Step S31232: According to the obtained divided feature map Use different convolutions to obtain spatial attention maps under different receptive fields respectively:
[0099]
[0100]
[0101] Among them, GN represents GroupNorm, δ represents the Sigmoid function, Conv1 and Conv3 represent convolutions of 1×1 and 3×3 respectively, || represents the concatenation operation, g h 、g w represent global average pooling in one-dimensional different directions, Y g represents the spatial attention map under a larger receptive field, Y l represents the spatial attention map under a smaller receptive field, represents the divided feature map;
[0102] Step S31233: Aggregate the spatial attention maps within different channel dimensions, and then use the Sigmoid function and the feature map Y to obtain the processed feature map H t ∈R T×H×W×C :
[0103] H t =δ(g(Y l )⊙μ(Y g )+g(Y g )⊙μ(Y l ))⊙Y
[0104] Among them, g represents two-dimensional global average pooling, μ represents the Softmax function, Y g represents the spatial attention map under a larger receptive field, Y l represents the spatial attention map under a smaller receptive field, δ represents the Sigmoid function.
[0105] The corresponding deep learning processing of the relative pose error in step S3 of this implementation scheme includes the following steps:
[0106] Step S321: Use multiple windows of different scale sizes and respectively utilize multi-head self-attention to capture feature information at different levels. The large-scale window provides coarse-grained feature information to compensate for significant relative pose errors, while the small-scale window provides fine-grained feature information to compensate for more refined relative pose errors;
[0107] Step S322: Cross-learn between the same-scale window features of different agents;
[0108]
[0109]
[0110] where Q, K, and V are the feature maps of all agents The query vector, key vector, and value vector formed after being projected into different feature spaces. s ∈ [s1, s2, s3] represents different scale windows, m represents the number of attention heads, and represent the attention weights for cross-learning between different agents, d k represents the scaling factor, || represents the concatenation operation, represents Q after completing window scale rearrangement, represents the matrix for window scale rearrangement of Q, represents K after completing window scale rearrangement, represents the matrix for window scale rearrangement of K, represents V after completing window scale rearrangement, represents the matrix for window scale rearrangement of V, h represents the maximum number of attention heads, μ represents the Softmax function, represents the scale window feature map;
[0111] Step S323: Adaptively aggregate the feature maps and processed by three scale windows to obtain the fused feature map H ms ∈R L×H×W×C ;
[0112]
[0113] where f(·) represents the convolution function; H ms represents the fused feature map.
[0114] Step S4 of this implementation scheme includes the following steps:
[0115] Step S41: The central vehicle feature H ego , the cooperative vehicle feature and the infrastructure feature H infraCombine into different feature sets;
[0116] Step S42: After completing the feature set grouping, perform feature fusion on the feature sets within each group using deformable cross-attention.
[0117] Step S42 of this implementation scheme includes the following steps:
[0118] Step S421: Use all the features within the group to obtain a reference center position map p, which represents the potential position of the detection target within the detection range;
[0119] Step S422: Use the reference center position map p to extract the features of the central vehicle at the reference point as the initial query embedding of the deformable cross-attention layer;
[0120] Step S423: Then use a linear layer to learn the two-dimensional offset Δp of the reference point through the reference center position map p. Through K key points at Δp + p, use bilinear sampling to extract its features as the participating features of the deformable cross-attention layer;
[0121] Step S424: Fill the output of the deformable cross-attention layer into the central vehicle features:
[0122]
[0123] H res = Fill(f defor (p), H ego )
[0124] where m represents the number of heads, W m / a represents the learnable weight, l represents the number of collaborative agents, k represents the number of key points, μ represents the Softmax function, W m represents the learnable weight, H a represents the learnable weight, H ego (p) represents the features extracted by the central vehicle through the graph p, represents the features extracted by the collaborative agent through the graph p + Δp, f defor (p) represents the output of the deformable cross-attention layer, Fill(·) represents the filling operation, H ego represents the central vehicle features, and μ represents the Softmax function;
[0125] Step S5 of this implementation scheme includes the following steps:
[0126] Step S51: Send multiple groups of fused features H res into two different detection decoders to decode them into detection results, which are multiple groups of classification and regression outputs respectively;
[0127] Step S52: Calculate the mean of multiple sets of classification and regression output results to obtain the final detection result.
[0128] As Figure 1 shown, it is a schematic diagram of the method flow of this embodiment, and the steps include:
[0129] Step 1, during the autonomous driving process, the central vehicle collects the 3D point cloud data collected by the on-vehicle lidar and broadcasts its metadata (such as position, etc.). The cooperative agent collects the 3D point cloud data collected by the lidar and performs coordinate system conversion.
[0130] Step 2, extract high-dimensional feature vectors from the 3D point cloud data. As Figure 2 shown, the PointPillar encoder first divides the point cloud data into multiple columnar units for processing, and then generates pseudo-image features through a scatter operator. Finally, the cooperative agent compresses the feature map with Conv-BN-ReLU blocks and transmits it to the central vehicle.
[0131] Step 3, after decompressing the feature map transmitted by the cooperative agent received by the decompressor mainly composed of Conv-BN-ReLU blocks. Corresponding processing is performed on the transmission time delay and relative pose error respectively.
[0132] As Figure 3 shown, for the transmission time delay, temporal modeling and fusion are performed on the feature vectors, specifically as follows:
[0133] First, a small number of historical feature maps that are temporally continuous with the feature map transmitted by the agent and the transmitted feature map are combined to form a temporally continuous feature map F t ∈R T×H×W×C . Then, a dynamic temporal kernel is generated for the temporally continuous feature map of each agent to adaptively learn the temporal context information in the continuous feature map for temporal modeling. Specifically, first use AveragePooling to compress the spatial dimension of F t to obtain Then respectively generate a dynamic convolution kernel and a weight sensitive to the temporal position through different convolution blocks:
[0134]
[0135] Among them, Conv t represents per-channel temporal convolution, B1, B2 represent different convolution operations, and ⊙ represents element-wise multiplication. After temporal modeling, enrich all feature maps through receptive fields of different sizes. Specifically, first divide the feature map Y along the channel dimension to obtain Then use different convolutions to obtain spatial attention maps under different receptive fields:
[0136]
[0137]
[0138] Among them, GN represents GroupNorm, δ represents the Sigmoid function, Conv1 and Conv3 represent convolutions of 1×1 and 3×3 respectively, || represents the concatenation operation, and g h and g w represent global average pooling in different one-dimensional directions. Then, the spatial attention maps within different channel dimensions are aggregated, and the processed feature map H t ∈ R T×H×W×C :
[0139] H t = δ(g(Y l ) ⊙ μ(Y g ) + g(Y g ) ⊙ μ(Y l )) ⊙ Y(1 - 4)
[0140] Among them, g represents two-dimensional global average pooling, and μ represents the Softmax function.
[0141] Secondly, to enhance the temporal information of the obtained current transmission feature map, the current transmission feature map in the processed feature map H t is fused with the historical feature map : As
[0142]
[0143]
[0144] shown, for the relative pose error, attention learning is performed on the feature vectors as follows: Figure 4 First, multiple windows with different scale sizes are used to capture feature information at different levels using multi-head self-attention. The large-scale window provides coarse-grained feature information to compensate for significant relative pose errors, while the small-scale window provides fine-grained feature information to compensate for more refined relative pose errors. Then, cross-learning is performed between the feature maps of the same scale windows of different agents (e.g., vehicle-vehicle, vehicle-infrastructure, infrastructure-infrastructure):
[0145]
[0146]
[0147]
[0148] Among them, Q, K, and V are the feature maps of all agents The query vector (Query), key vector (Key), and value vector (Value) formed after projecting into different feature spaces. s ∈ [s1, s2, s3] represents different scale windows, such as 4×4, 8×8, 16×16, m represents the number of attention heads, and represents the attention weights for cross-learning of different agents, d k represents the scaling factor.
[0149] Then, the feature maps processed by the three scale windows and are adaptively aggregated. Specifically, the corresponding confidence maps C1, C2, and C3 are learned from the feature maps processed by the three scale windows (convolution function f(·)). Finally, the feature maps are fused according to the confidence to obtain H ms ∈R L×H×W×C :
[0150]
[0151] Step 4, in order to more effectively fuse the features of multiple cooperative agents, as Figure 5 shown, the central vehicle groups the feature vectors participating in the aggregation and then aggregates them, specifically as follows:
[0152] First, the central vehicle feature H ego , the cooperative vehicle feature and the infrastructure feature H infra are combined into different feature sets. Specifically, the central vehicle feature, the infrastructure feature, and one cooperative vehicle are grouped into a set of features, that is
[0153] Secondly, after completing the feature grouping, the feature sets within each group are fused using deformable cross-attention. First, a reference center position map p is obtained using all the features within the group, which represents the potential position of the detection target within the detection range. Then, the feature of the central vehicle at the reference point is extracted using the map p as the initial query embedding of the deformable cross-attention layer. Secondly, a linear layer is used to learn the two-dimensional offset Δp of the reference point through the map p. Through K key points at Δp + p, bilinear sampling is used to extract its features as the participating features of the deformable cross-attention layer. Finally, the output of the deformable cross-attention layer is filled into the central vehicle feature:
[0154]
[0155] Hres = Fill(f defor (p), H ego ) (1 - 11)
[0156] Among them, m represents the number of heads, W m / a represents W m learnable weights, l represents the number of collaborative agents, k represents the number of key points, and μ represents the Softmax function.
[0157] Step 5: The central vehicle sends multiple groups of fused features H res to two different detection decoders to decode them into detection results, namely classification and regression outputs. Among them, the classification output is the confidence score of the object or background for each anchor box, and the regression output is the bounding box (x, y, z, w, l, h, θ) for each position, where x, y, z are the positions of the bounding box, w, l, h are the sizes of the bounding box, and θ is the heading angle of the bounding box. Finally, the final detection result is obtained by calculating the mean of multiple groups of classification and regression output results. Among them, smooth l1 is used as the regression loss function, and focal is used as the classification loss function.
Claims
1. A multi-agent collaborative perception method based on lidar in complex noise, characterized in that, It includes the following steps: Step S1: Perform coordinate system transformation and preprocessing on the collaborative agent lidar point cloud; Step S2: Use the method of deep learning to extract the feature maps of all collaborative agent lidar point clouds, and then compress the feature maps of the collaborative agent lidar point clouds and transmit them to the central vehicle; Step S3: The central vehicle decompresses the received feature maps of the collaborative agent lidar point clouds, and at the same time uses the feature maps of a small amount of previously received collaborative agent lidar point clouds to perform corresponding deep learning processing on the transmission time delay and relative pose error; Step S4: The central vehicle uses the attention mechanism to fuse the feature maps of the collaborative agent lidar point clouds into the feature maps of the central vehicle to obtain the final feature maps; Step S5: Send the final feature maps into the detection head to obtain the detection results; The corresponding deep learning processing of the transmission time delay in step S3 includes the following steps: Step S311: Combine a small amount of historical feature maps that are temporally continuous with the agent's transmitted feature maps and the transmitted feature maps to form a temporally continuous feature map F t ∈R T×H×W×C ; Step S312: Generate dynamic temporal kernels for the temporally continuous feature maps of each agent, and learn the temporal context information in the continuous feature maps in an adaptive manner to obtain the processed feature map Ht; Step S313: To enhance the temporal information of the obtained current transmission feature map, fuse the current transmission feature map in the processed feature map H t with the historical feature map : Among them, Mean and Max respectively represent taking the average value and the maximum value of the feature map in the channel dimension. represents the processed historical feature map. represents the current transmitted feature map after fusion, δ represents the Sigmoid function, || represents the concatenation operation, ⊙ represents the element-wise multiplication, and Conv represents the convolution.
2. The multi-agent collaborative perception method based on lidar under complex noise according to claim 1, wherein The specific coordinate system transformation of the collaborative agent lidar point cloud in step S1 is as follows: After receiving the metadata broadcast by the central vehicle, the collaborative agents project their local lidar point clouds into the coordinate system of the central vehicle; The preprocessing includes removing the lidar point clouds not near the ground and collision detection.
3. The multi-agent collaborative perception method based on lidar under complex noise according to claim 1, wherein, The backbone network for using the method of deep learning to extract the feature maps of all collaborative agent lidar point clouds in step S2 is the PointPillar encoder; the compressor for the feature map compression is the Conv-BN-ReLU block.
4. The multi-agent collaborative perception method based on lidar under complex noise according to claim 1, characterized in that, The central vehicle in step S3 decompresses the received feature maps of the collaborative agent lidar point clouds specifically by using a decompressor mainly composed of Conv-BN-ReLU blocks to decompress the received feature maps transmitted by the collaborative agents.
5. The multi-agent collaborative perception method based on lidar under complex noise according to claim 1, wherein Step S312 includes the following steps: Step S3121: Use AveragePooling to compress the time - continuous feature map F t in the spatial dimension to obtain the compressed feature map Step S3122: Compressed feature map Generate dynamic convolution kernels and weights sensitive to temporal positions respectively through different convolutional blocks: Among them, Y represents the feature map, Conv t represents per-channel temporal convolution, B1 and B2 represent different convolution operations, and ⊙ represents element-wise multiplication; Step S3123: After performing temporal modeling on the temporally continuous feature map F t all feature maps are enriched by receptive fields of different sizes; Step S3123 includes the following steps: Step S31231: Divide the feature map Y along the channel dimension to obtain the divided feature map Step S31232: According to the obtained partitioned feature map Use different convolutions to obtain spatial attention maps under different receptive fields respectively: Among them, GN represents GroupNorm, δ represents the Sigmoid function, Conv1 and Conv3 respectively represent convolutions of 1×1 and 3×3, || represents the concatenation operation, g h and g w represent global average pooling in different one-dimensional directions, Y g represents the spatial attention map under a larger receptive field, Y l represents the spatial attention map under a smaller receptive field, represents the divided feature map; Step S31233: Aggregate the spatial attention maps within different channel dimensions, and then use the Sigmoid function and the feature map Y to obtain the processed feature map H t ∈R T×H×W×C : H t = δ(g(Y l ) ⊙ μ(Y g ) + g(Y g ) ⊙ μ(Y l )) ⊙ Y Among them, g represents two-dimensional global average pooling, μ represents the Softmax function, and Y g represents the spatial attention map under a larger receptive field, and Y l represents the spatial attention map under a smaller receptive field, and δ represents the Sigmoid function.
6. The multi-agent collaborative perception method based on lidar under complex noise according to claim 1, characterized in that The corresponding deep learning processing of the relative pose error in step S3 includes the following steps: Step S321: Use multiple windows of different scale sizes and respectively use multi-head self-attention to capture feature information of different degrees. The large-scale windows provide coarse-grained feature information to compensate for significant relative pose errors, while the small-scale windows provide fine-grained feature information to compensate for more refined relative pose errors; Step S322: Cross-learn between the same-scale window features of different agents; Among them, Q, K, and V are the feature maps of all agents The query vector, key vector, and value vector formed after projecting into different feature spaces. s ∈ [s1, s2, s3] represents different scale windows, and m represents the number of attention heads. and represents the attention weights for cross-learning among different agents, and d k represents the scaling factor, and || represents the concatenation operation. represents Q after completing window scale rearrangement. represents the matrix for window scale rearrangement of Q. represents K after completing window scale rearrangement. represents the matrix for window scale rearrangement of K. represents V after completing window scale rearrangement. represents the matrix for window scale rearrangement of V. h represents the maximum number of attention heads, and μ represents the Softmax function. represents the scale window feature map; Step S323: Aggregate adaptively the feature maps and processed by three scale windows to obtain a fused feature map H ms ∈ R L×H×W×C ; where f(·) represents the convolution function; H ms represents the fused feature map.
7. The multi-agent collaborative perception method based on lidar under complex noise according to claim 1, characterized in that Step S4 includes the following steps: Step S41: Combine the central vehicle feature H ego , the cooperative vehicle feature and the infrastructure feature H infra into different feature sets; Step S42: After completing the feature set grouping, perform feature fusion on the feature sets within each group using deformable cross-attention.
8. The multi-agent collaborative perception method based on lidar under complex noise according to claim 7, characterized in that Step S42 includes the following steps: Step S421: Use all the features within the group to obtain a reference center position map p, which represents the potential position of the detection target within the detection range; Step S422: Use the reference center position map p to extract the features of the central vehicle at the reference points as the initial query embedding of the deformable cross-attention layer; Step S423: Then use a linear layer to learn the two-dimensional offset Δp of the reference points through the reference center position map p. Through the K key points at Δp + p, use bilinear sampling to extract their features as the participating features of the deformable cross-attention layer; Step S424: Fill the output of the deformable cross-attention layer into the center vehicle features: H res = Fill(f defor (p), H ego ) where m represents the number of heads, W m / a represents W m learnable weights, l represents the number of cooperative agents, k represents the number of key points, μ represents the Softmax function, and W a represents learnable weights, H ego (p) represents the features extracted by the central vehicle from graph p, represents the features extracted by the cooperative agent from graph p + Δp, and f defor (p) represents the output of the deformable cross-attention layer, Fill(·) represents the filling operation, and H ego represents the central vehicle features, and μ represents the Softmax function.
9. The multi-agent collaborative perception method based on lidar under complex noise according to claim 1, wherein, The said Step S5 includes the following steps: Step S51: Send multiple sets of fused features H res into two different detection decoders to decode them into detection results, namely multiple sets of classification and regression outputs respectively; Step S52: Obtain the final detection result by calculating the mean of multiple groups of classification and regression output results.
Citation Information
Patent Citations
Roadside laser radar beyond visual range sensing method
CN114565901A
Mapping positioning method fusing laser radar and depth camera point cloud
CN115330866A
Cited By
Multi-agent collaborative relative positioning error estimation method and system
CN121685918A