Space-time error tolerant multi-agent collaborative perception method, apparatus and electronic device
By performing feature encoding and importance selection in a bird's-eye view space, combined with redundancy enhancement and complementary enhancement, the spatiotemporal alignment error problem in multi-agent collaborative perception is solved, improving the accuracy and stability of perception.
Patent Information
- Application Number
- CN202310560996.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-17
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2043-05-17
AI Technical Summary
In the process of multi-agent collaborative perception, there are spatial misalignments caused by errors in agent pose and sensor correction parameters, as well as motion misalignments caused by asynchronous sampling times, resulting in spatiotemporal alignment errors that affect the accuracy and stability of perception.
Feature encoding is performed in the bird's-eye view space, important features are selected and communication data is packaged, shared information is aggregated, redundancy enhancement and complementary enhancement are performed, and spatiotemporal alignment errors are handled by redundancy enhancement and complementary enhancement to achieve collaborative perception.
It improves the stability and performance of the collaborative perception algorithm, reduces the impact of spatiotemporal alignment errors on perception results, and enhances the accuracy and stability of scene perception.
Smart Images

Figure CN116740514B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of collaborative perception, and in particular to a multi-agent collaborative perception method and device tolerant to spatial-temporal error and an electronic device. BACKGROUND
[0002] In recent years, deep learning technology has achieved good performance in various scene perception tasks and has been widely used in the fields of automatic driving and intelligent monitoring. However, single-agent perception has inherent defects such as limited sensor field of view, mutual occlusion between targets, and sparse data at long distances, and the perception capability in real complex scenes still needs to be further improved, which limits the large-scale popularization of automatic driving cars and autonomous mobile robots with strict safety requirements.
[0003] In the prior art, multi-agent collaborative perception exchanges perception information within a certain range to achieve feature enhancement and information complementation of scene targets, thereby improving the accuracy and stability of scene perception.
[0004] However, in the process of multi-agent collaborative perception, the pose of the agent and the correction parameters of the sensor usually have errors, causing spatial misplacement of the perceived scene. In addition, to improve the timeliness of the response of the collaborative perception system, the latest cached collaborative agent perception information is usually used, which has an obvious sampling time synchronization problem with the self-agent perception information, causing motion misplacement of the perceived object. The above factors cause the local misplacement of the collaborative agent features and the self-agent features, that is, there is a spatial-temporal alignment error. SUMMARY
[0005] The present application provides a multi-agent collaborative perception method and device tolerant to spatial-temporal error and an electronic device to solve the defect that the collaborative agent features and the self-agent features are consistent in the spatial-temporal whole but misaligned locally, that is, there is a spatial-temporal alignment error in the prior art.
[0006] The present application provides a multi-agent collaborative perception method tolerant to spatial-temporal error, comprising:
[0007] obtaining point cloud data of a laser radar of a self-agent and pose data of the laser radar;
[0008] performing feature encoding on the point cloud data to obtain self-agent features in an aerial view space;
[0009] performing importance selection on the self-agent features, and performing communication data packaging based on the importance selected self-agent features and the pose data to obtain shared information;
[0010] performing shared information aggregation on the shared information to obtain collaborative agent features;
[0011] Based on the self-agent feature and the coordination agent feature, redundancy enhancement and complementary enhancement are performed to obtain a fusion feature;
[0012] Based on the fusion feature, cooperative perception is performed to obtain a cooperative perception result.
[0013] According to the multi-agent cooperative perception method provided by the application, the redundancy enhancement and the complementary enhancement are performed based on the self-agent feature and the coordination agent feature to obtain a fusion feature, including:
[0014] Based on the self-agent feature and the coordination agent feature, a candidate error set required for adaptive accurate alignment of each feature space position and a candidate confidence set corresponding to the candidate error set are obtained;
[0015] Based on the self-agent feature, the coordination agent feature, the candidate error set and the candidate confidence set, a redundancy enhancement feature is obtained;
[0016] Based on the redundancy enhancement feature, the coordination agent feature and the self-agent perception blind area graph, complementary enhancement is performed to obtain a fusion feature.
[0017] According to the multi-agent cooperative perception method provided by the application, the determination of the self-agent perception blind area graph includes:
[0018] The spatial probability graph is smoothed, binarized and inverted to obtain a spatial demand graph;
[0019] Based on the intensity value of the coordination agent feature, an effective spatial graph is obtained;
[0020] Based on the spatial demand graph and the effective spatial graph, a self-agent perception blind area graph is obtained.
[0021] According to the multi-agent cooperative perception method provided by the application, the communication data packaging based on the self-agent feature selected according to importance and the pose data obtains shared information, including:
[0022] The self-agent feature selected according to importance is subjected to target probability estimation to obtain a spatial probability graph;
[0023] Based on the self-agent feature and the spatial probability graph, and the pose data, communication data packaging is performed to obtain shared information.
[0024] According to the multi-agent cooperative perception method provided by the application, the communication data packaging based on the self-agent feature and the spatial probability graph, and the pose data obtains shared information, including:
[0025] Based on the features of the autonomous agent and the spatial probability map, thresholding is used to select features and obtain shared features.
[0026] Based on the shared features and the pose data, communication data is packaged to obtain shared information.
[0027] According to the spatiotemporal error-tolerant multi-agent cooperative perception method provided by the present invention, the step of aggregating the shared information to obtain cooperative agent features includes:
[0028] Based on multiple shared features and multiple pose data in the shared information, position reorganization is performed to restore the multiple shared features to multiple feature maps in the bird's-eye view space;
[0029] Based on the effective perception area of the intelligent agent, the feature regions of the multiple feature maps are cropped to obtain multiple coarsely aligned feature maps.
[0030] Information is aggregated from the multiple coarsely aligned feature maps to obtain the cooperative agent features.
[0031] According to a spatiotemporal error-tolerant multi-agent cooperative sensing method provided by the present invention, the step of performing cooperative sensing based on the fused features to obtain cooperative sensing results includes:
[0032] The fused features are enhanced to obtain enhanced features;
[0033] Based on the enhanced features, three-dimensional target detection is performed to obtain the target's position and size information;
[0034] Based on the enhanced features, scene segmentation is performed to obtain typical attribute information for each spatial location in the scene;
[0035] Based on the location and size information and the typical attribute information, the collaborative perception result is obtained.
[0036] According to the spatiotemporal error-tolerant multi-agent cooperative perception method provided by the present invention, the step of performing feature encoding on the point cloud data in a bird's-eye view space to obtain the features of the agent includes:
[0037] The point cloud data is projected onto the bird's-eye view space and resampled to obtain normalized point columns;
[0038] Based on the dot-column coding network, dot-column abstract feature extraction is performed on the normalized dot columns to obtain the self-intelligent agent features;
[0039] The dot-column encoding network is trained based on sample normalized dot columns, 3D object detection labels, and BEV semantic segmentation labels, and is combined with a 3D object detection model and a scene segmentation model. The 3D object detection model is used to perform 3D object detection based on the features of the autonomous agent, and the scene segmentation model is used to perform semantic segmentation based on the features of the autonomous agent.
[0040] The present invention also provides a spatiotemporally error-tolerant multi-agent cooperative sensing device, comprising:
[0041] The acquisition unit is used to acquire point cloud data of the lidar from the intelligent agent and pose data of the lidar.
[0042] The feature encoding unit is used to encode the point cloud data in a bird's-eye view space to obtain the features of the autonomous agent;
[0043] A shared information unit is determined, which is used to select the importance of the features of the autonomous agent, and to package communication data based on the selected features of the autonomous agent and the pose data to obtain shared information;
[0044] An information aggregation unit is used to aggregate the shared information to obtain cooperative agent features.
[0045] An enhancement unit is used to perform redundancy enhancement and complementary enhancement based on the features of the autonomous agent and the features of the cooperative agent to obtain fused features;
[0046] The collaborative sensing unit is used to perform collaborative sensing based on the fused features to obtain collaborative sensing results.
[0047] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the spatiotemporal error-tolerant multi-agent cooperative sensing method as described above.
[0048] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a spatiotemporally error-tolerant multi-agent cooperative sensing method as described above.
[0049] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements a spatiotemporal error-tolerant multi-agent cooperative perception method as described above.
[0050] The present invention provides a spatiotemporally error-tolerant multi-agent cooperative sensing method, apparatus, and electronic device that acquires point cloud data and pose data of lidar from the autonomous agents. In a bird's-eye view space, feature encoding is performed on the point cloud data to obtain autonomous agent features. Importance selection is performed on the autonomous agent features, and communication data is packaged based on the importance-selected autonomous agent features and pose data to obtain shared information. Shared information is aggregated to obtain co-agent features. Based on the autonomous agent features and co-agent features, redundancy enhancement and complementary enhancement are performed to obtain fused features. Based on the fused features, cooperative sensing is performed to obtain the cooperative sensing result. The cooperative agent features are obtained by aggregating shared information from a variable number of cooperative agents. This facilitates the acquisition of cooperative agent features with fixed formats and rich information, further enhancing the stability of the cooperative perception algorithm. Furthermore, the fused features are obtained by performing redundancy enhancement and complementary enhancement on the self-agent features and cooperative agent features. In the process of redundancy enhancement and complementary enhancement, the position misalignment problem is handled, which can reduce the impact of the spatiotemporal alignment error between cooperative agent features and self-agent features on the cooperative perception effect, further improving the performance of cooperative perception under the condition of spatiotemporal alignment error. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0052] Figure 1 This is one of the flowcharts of the spatiotemporal error-tolerant multi-agent cooperative sensing method provided by the present invention;
[0053] Figure 2 This is the second flowchart of the spatiotemporal error-tolerant multi-agent cooperative perception method provided by the present invention;
[0054] Figure 3 This is a flowchart illustrating the process of determining the characteristics of a cooperative agent provided by the present invention;
[0055] Figure 4 This is a flowchart illustrating step 150 in the spatiotemporal error-tolerant multi-agent cooperative sensing method provided by the present invention.
[0056] Figure 5 This is a flowchart illustrating step 120 in the spatiotemporal error-tolerant multi-agent cooperative sensing method provided by the present invention.
[0057] Figure 6This is a schematic diagram of the spatiotemporal error-tolerant multi-agent collaborative sensing device provided by the present invention.
[0058] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0059] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0060] In related technologies, multi-agent collaborative perception enhances the features and complements information of scene targets by exchanging perception information within a certain range, thereby improving the accuracy and stability of scene perception. Furthermore, collaborative perception can also enhance the effective perception range of agents, alleviate reliance on long-distance, high-precision perception data, and reduce sensor costs.
[0061] Common collaborative perception tasks include 3D target detection and BEV (Bird's Eye View) semantic segmentation, which have important application value in fields such as vehicle-road cooperative autonomous vehicles, vehicle-to-vehicle cooperative autonomous vehicles, multi-robot warehouse automation systems, UAV swarm collaborative search and rescue, underwater robot swarm collaborative target detection, UAV-quadruped robot collaborative transportation of important materials, and near-ground space three-dimensional monitoring.
[0062] In multi-agent collaborative perception, errors often exist in the pose of agents and the calibration parameters of sensors, causing spatial misalignment of the perceived scene. Furthermore, to improve the timeliness of the collaborative perception system's response, the latest cached perception information of the co-agents is typically used. This cached information has a significant sampling time asynchrony with the perception information of the self-agents, causing motion misalignment of the perceived objects. These factors result in the co-agent features and self-agent features being spatiotemporally consistent overall but locally misaligned, i.e., a spatiotemporal alignment error exists.
[0063] To address the aforementioned problems, this invention provides a spatiotemporally error-tolerant multi-agent cooperative sensing method. Figure 1 This is one of the flowcharts of the spatiotemporal error-tolerant multi-agent cooperative sensing method provided by the present invention. Figure 2 This is the second flowchart of the spatiotemporal error-tolerant multi-agent cooperative sensing method provided by the present invention, as shown below. Figure 1 , Figure 2 As shown, this method can be applied to a server, and the method includes:
[0064] Step 110: Obtain the point cloud data of the lidar from the intelligent agent and the pose data of the lidar.
[0065] Specifically, point cloud data and pose data of the LiDAR sensors on the autonomous agent can be acquired. Here, the autonomous agent refers to the central agent performing the task, such as a central autonomous vehicle in motion. The LiDAR point cloud data refers to the perceived point cloud data scanned and acquired by the LiDAR sensors installed on the autonomous agent, and can be presented in the form of a sparse 3D point cloud. LiDAR point cloud data can typically be represented as... Where, N LiDAR The number of measurements acquired by the lidar sensor in one scan cycle, (x i ,y i ,z i ) and α i Let be the three-dimensional coordinates of the reflection point and its reflectivity for the i-th measurement value.
[0066] LiDAR pose data refers to the position and attitude estimates of the LiDAR sensor. These estimates typically have a certain error, which can be represented as O. e = {x0,y0,z0,α0,β0,γ0}, where (x0,y0,z0) are the three-dimensional coordinates of the center point of the lidar sensor, and (α0,β0,γ0) are the angles between the lidar sensor and the three coordinate planes.
[0067] Step 120: In the bird's-eye view space, the point cloud data is feature-encoded to obtain the features of the autonomous agent.
[0068] Specifically, after acquiring the point cloud data of the LiDAR of the autonomous agent, feature encoding can be performed on the point cloud data in the Bird's Eye View (BEV) space to obtain the autonomous agent's features. The BEV space refers to a two-dimensional planar coordinate system with the three-dimensional coordinates of the center point of the LiDAR sensor as the origin, the horizontal plane as the coordinate plane, and the geographic location as the coordinate axis. The perception range is usually represented as a discrete grid of size w×h in the BEV space.
[0069] Here, the BEV spatial feature extraction module f can be used. BEV For point cloud data P e Feature encoding is performed to obtain the self-agent feature F. e The characteristics of an autonomous agent can be represented as: F e =f BEV (P e ).
[0070] BEV spatial features refer to the abstract features of LiDAR point clouds acquired in the BEV space, i.e., the features of the autonomous agent refer to the abstract features of LiDAR point clouds acquired in the BEV space, usually represented as tensors F∈R. w×h×c , where c is the number of feature channels. The BEV spatial feature extraction module is a network model that maps the point cloud data of the LiDAR into BEV spatial features.
[0071] Step 130: Select the importance of the features of the autonomous agent, and package the communication data based on the selected features of the autonomous agent and the pose data to obtain shared information.
[0072] Specifically, importance selection is performed on the features of the autonomous agent, that is, the importance of the features of the autonomous agent at each location in the BEV space is evaluated, and only the spatially sparse but perceptually critical BEV space features are retained to reduce the amount of communication data.
[0073] Then, communication data can be packaged based on the characteristics and pose data of the selected intelligent agents according to their importance, thus obtaining shared information. Communication data packaging refers to lossless compression and data splitting of the data to be shared to achieve efficient communication requirements.
[0074] That is, the server obtains the self-agent characteristics F after importance selection. e and pose data O e Then, the shared information selection and transmission module f can be used. share Obtain spatially sparse but perceptually critical shared information S for use by other intelligent agents. e It can be represented as: S e =f share (F e O e ).
[0075] The shared information selection and transmission module is a mathematical model that maps the features and pose information of the autonomous agent into shared data packets.
[0076] Step 140: Aggregate the shared information to obtain the cooperative agent features.
[0077] Specifically, the autonomous agent can acquire shared information from a variable number of co-agents and aggregate the shared information to obtain co-agent features.
[0078] Figure 3 This is a flowchart illustrating the process of determining the characteristics of a cooperative agent provided by the present invention, as shown below. Figure 3As shown, the shared information aggregation here includes feature restoration, spatial alignment, and information aggregation. Feature restoration refers to decoding and recombining the shared information composed of shared data packets into a sparse co-agent BEV feature map. Spatial alignment refers to transforming the co-agent BEV feature map into the self-agent BEV spatial coordinate system to obtain a co-agent BEV feature map that is roughly spatially aligned with the self-agent. Information aggregation refers to integrating the effective features of the roughly spatially aligned co-agent BEV feature maps obtained from multiple co-agents (of varying numbers) to obtain a single co-agent feature to be fused.
[0079] That is, the server obtains shared information from K cooperative agents. (K is a non-fixed value and can change dynamically) After that, the co-agent feature restoration, spatial alignment and information aggregation modules f can be used. merge Obtain a single spatiotemporally consistent but locally misaligned co-agent feature F. s , can be represented as:
[0080] The co-agent feature restoration, spatial alignment and information aggregation module here is a mathematical model that maps the shared information of a variable number of co-agents into a single co-agent feature to be fused.
[0081] Step 150: Based on the features of the self-intelligent agent and the features of the cooperative agent, perform redundancy enhancement and complementary enhancement to obtain fused features.
[0082] Specifically, spatiotemporal error refers to the spatial misalignment of the perceived scene caused by the pose of the intelligent agent and the error of the sensor correction parameters, as well as the motion misalignment of the perceived object caused by the asynchronous sampling time of multiple intelligent agents' sensors. This results in the overall spatiotemporal consistency of the features of the cooperative intelligent agent and the features of the self intelligent agent, but local misalignment, i.e., there is a spatiotemporal alignment error.
[0083] Considering that the features of the cooperative agent and the features of the autonomous agent are consistent in space and time but misaligned locally, i.e., there is a space-time alignment error.
[0084] Therefore, after obtaining the features of the self-agent and the co-agent, redundancy enhancement and complementary enhancement can be performed based on these features to obtain fused features. Here, fused features refer to the features obtained after redundancy enhancement and complementary enhancement of the self-agent and co-agent features.
[0085] Redundancy enhancement refers to using multiple observations of the same target object to enhance information, aiming to improve the perception of uncertain information in a scene. Complementary enhancement refers to using the observables of the co-agent to compensate for the blind spots of the self-agent, aiming to improve the perception of occluded targets in a scene.
[0086] The server obtains the characteristics F from the intelligent agent.e Features of Harmonious Intelligent Agent F s Then, the feature fusion module f, which is tolerant to spatiotemporal errors, can be used. fuse The fusion feature F obtained after redundancy enhancement and complementary enhancement is obtained. fusion It can be represented as: F fusion =f fusion (F e F s ).
[0087] The spatiotemporal error-tolerant feature fusion module performs spatiotemporal error-tolerant redundancy enhancement and complementary enhancement on the features of the self-agent and the co-agent to obtain a network model with fused features.
[0088] Step 160: Based on the fusion features, perform collaborative perception to obtain collaborative perception results.
[0089] Specifically, after obtaining the fusion features, collaborative perception can be performed based on the fusion features to obtain collaborative perception results.
[0090] That is, the server obtains the fusion feature F fusion Then, the convolutional neural network module f can be used. enhance This yields enhanced features F with stronger abstraction and discriminative abilities. enhance It can be represented as: F enhance =f enhance (F fusion ).
[0091] In obtaining enhanced feature F enhance Then, the three-dimensional target detection module f can be used. det Obtain the three-dimensional target detection result O det =f det (F enhance Alternatively, the BEV semantic segmentation module f can be used. seg Obtain the BEV semantic segmentation result O seg =f seg (F enhance Finally, based on the 3D target detection results O det and BEV semantic segmentation results O seg Together they constitute the results of collaborative perception.
[0092] The convolutional neural network module, 3D object detection module, and BEV semantic segmentation module mentioned here all refer to network models composed of multiple layers of convolutional neural networks. The convolutional neural network module aims to enhance the semantic abstraction and spatial discrimination capabilities of features, such as the frequently used "ResNet" model or "ResNet+FPN (Feature Pyramid Networks)" model.
[0093] The 3D object detection module aims to obtain the predicted probability values of typical object categories and the bounding box regression values, and based on these, obtain the 3D object detection results. For example, it often uses a 3×3 convolution layer plus a 1×1 convolution layer. The BEV semantic segmentation module aims to obtain the predicted probability values of semantic categories for each spatial location, and based on these, obtain the semantic segmentation results. For example, it often uses a 3×3 convolution layer plus a 1×1 convolution layer.
[0094] Collaborative perception refers to a technology that involves multiple sensors working together and processing and integrating the data collected by the sensors to produce more accurate and complete perception results.
[0095] The method provided in this invention acquires point cloud data and pose data of a LiDAR from an autonomous agent. In a bird's-eye view space, feature encoding is performed on the point cloud data to obtain autonomous agent features. Importance selection is performed on these autonomous agent features, and communication data is packaged based on the selected features and pose data to obtain shared information. This shared information is then aggregated to obtain co-agent features. Redundancy enhancement and complementary enhancement are performed based on the autonomous and co-agent features to obtain fused features. Based on these fused features, collaborative perception is performed to obtain the collaborative perception result. The co-agent features are obtained by aggregating shared information from multiple co-agents of varying numbers. This facilitates the acquisition of co-agent features with fixed formats and rich information, further enhancing the stability of the collaborative perception algorithm. Furthermore, the fused features are obtained by redundancy enhancement and complementary enhancement based on the autonomous and co-agent features. During the redundancy enhancement and complementary enhancement process, positional misalignment issues are addressed, reducing the impact of spatiotemporal alignment errors between co-agent and autonomous agent features on the collaborative perception effect, further improving the performance of collaborative perception under conditions of spatiotemporal alignment errors.
[0096] Based on the above embodiments, Figure 4 This is a flowchart illustrating step 150 in the spatiotemporal error-tolerant multi-agent cooperative sensing method provided by the present invention, as shown below. Figure 4 As shown, step 150 includes:
[0097] Step 151: Based on the features of the self-intelligent agent and the features of the co-intelligent agent, obtain the set of candidate error quantities required for adaptive and precise alignment of each feature spatial location and the set of candidate confidence values corresponding to the set of candidate error quantities;
[0098] Step 152: Based on the self-agent features, the co-agent features, the candidate error set, and the candidate confidence set, the redundancy enhancement features are obtained;
[0099] Step 153: Based on the redundant enhancement features, the cooperative agent features, and the self-agent perception blind spot map, complementary enhancements are performed to obtain fused features.
[0100] Specifically, the server obtains the characteristics F of the autonomous agent. e Features of Harmonious Intelligent Agent F s Then, the self-agent characteristics F can be... e Features of Harmonious Intelligent Agent F s Stacking them along the feature dimension yields the concatenated feature F. c ∈R w×h×(2c) .
[0101] Obtain splicing features F c Then, the bias weight estimation module f can be used. align To obtain the set of candidate error quantities required for adaptive and accurate alignment of each feature space location (x,y) and the set of candidate confidence scores corresponding to the set of candidate error values. The candidate error set and the candidate confidence set can be represented as: Where M is the number of output estimation points, to ensure that the correctly aligned points are in the candidate set, and the deviation weight estimation module f align One 3×3 convolution layer plus one 1×1 convolution layer can be used. The bias weight estimation module f... align The network parameters to be trained can be obtained through end-to-end training using a collaborative sensing algorithm.
[0102] Obtain the set of candidate error quantities and candidate confidence set Subsequently, based on the features of the self-agent agent, the features of the co-agent agent, the candidate error set, and the candidate confidence set, redundant enhancement features F can be obtained. ef The formula is as follows:
[0103]
[0104] Among them, f bilinear This is a bilinear interpolation function used to obtain candidate redundant features after adding the offset. M is the number of output estimated points to ensure that the correctly aligned points are in the candidate set. This represents the characteristics of the intelligent agent. The characteristics of the cooperative agent are represented by η, which is an adjustment coefficient used to balance the characteristics of the autonomous agent. Characteristics of Harmonious Intelligent Agents The relative importance of.
[0105] Here, the method for obtaining redundant enhancement features based on self-agent features, co-agent features, candidate error sets, and candidate confidence sets is normalized weighted fusion.
[0106] After obtaining the redundancy enhancement features, complementary enhancements can be performed based on the redundancy enhancement features, cooperative agent features, and the self-agent perception blind spot map to obtain fused features.
[0107] That is, the server obtains the redundancy enhancement feature F ef Features of cooperative agents F s Blind spot diagram of the self-intelligent agent perception blind Then, the complementary enhanced fusion feature F can be obtained by using blind zone probability weighted fusion. fusion , fusion feature F fusion The formula is as follows:
[0108] F fusion =(1-B blind )F ef +B blind F s
[0109] Among them, F fusion B represents the fusion feature. blind F represents the blind spot map of the self-agent perception. ef F represents the redundancy enhancement feature. s This represents the characteristics of the cooperative agent.
[0110] Based on the above embodiments, the steps for determining the blind spot map of the autonomous agent include:
[0111] Step 310: Smooth the spatial probability map, binarize it, and invert it to obtain the spatial demand map;
[0112] Step 311: Based on the intensity values of the cooperative agent features, obtain the effective spatial map;
[0113] Step 312: Based on the spatial demand map and the effective spatial map, obtain the self-aware agent perception blind spot map.
[0114] Specifically, the server obtains the characteristics F of the autonomous agent. e Then, the confidence estimation module f can be used. evd Obtain spatial probability map B spatial =f evd (F e )∈R w×h×1 The confidence estimation module f evd The goal is to obtain a target probability estimate for each spatial location. The confidence estimation module can use two 3×3 convolutional layers plus one 1×1 convolutional layer.
[0115] Obtain spatial probability map B spatialSubsequently, a ground truth map of the target coverage area, independent of category, can be generated based on manually labeled data, and the network can be trained end-to-end using cross-entropy loss. After the network converges and stabilizes, the network parameters of the confidence estimation module are retained and solidified, and used to synchronously generate spatial probability maps of BEV features of different agents.
[0116] Obtain spatial probability map B spatial Then, the error-tolerant collaborative demand module can be used to smooth, binarize, and invert the spatial probability map to obtain the spatial demand map B perceived by the autonomous agent. require Space requirement diagram B require The formula is as follows:
[0117] B require =1-f binary (f smooth (B spatial ),T r )
[0118] Among them, f smooth (B spatial () represents a smoothing function, such as using 5×5 mean smoothing, which expands the target confidence level to the corresponding smoothed region, thereby increasing the tolerance for local spatiotemporal alignment errors; f binary (B x,y ,T r ) as T r As a differentiable approximate binarization function for the threshold, f binary (B x,y ,T r The formula is as follows:
[0119]
[0120] Here, γ is a threshold controlling the degree of binarization, typically set to 20. Next, based on the co-agent feature F... s Using the effective region estimation module f valid Obtain the effective spatial graph B perceived by the co-agent valid =f valid (F s )∈R w×h×1 The effective region estimation module f valid It can be obtained by calculating the intensity value of the feature.
[0121] Obtain Space Requirements Map B require And effective space diagram B valid Then, the error-tolerant self-aware agent perception blind zone map B can be obtained by using spatial point-by-point multiplication. blind Similarly, the network parameters to be trained involved in the embodiments of the present invention can be obtained by end-to-end training using a collaborative sensing algorithm.
[0122] Based on the above embodiments, step 130 includes:
[0123] Step 131: Estimate the target probability of the features of the autonomous agent after importance selection to obtain a spatial probability map;
[0124] Step 132: Based on the features of the autonomous agent, the spatial probability map, and the pose data, the communication data is packaged to obtain shared information.
[0125] Specifically, the server's self-agent characteristics F after obtaining the importance selection. e Then, the confidence estimation module f can be used. evd Target probability estimation is performed on the features of the autonomous agent after importance selection to obtain spatial probability map B. spatial =f evd (F e )∈R w×h×1 The confidence estimation module f evd The goal is to obtain a target probability estimate for each spatial location. The confidence estimation module can use two 3×3 convolutional layers plus one 1×1 convolutional layer.
[0126] After obtaining the spatial probability map, communication data can be packaged based on the characteristics of the autonomous agent, the spatial probability map, and pose data to obtain shared information. Here, shared information refers to spatially sparse but perceptually crucial information that can be shared and used by other agents.
[0127] Here, feature selection can be performed based on the characteristics of the autonomous agent and the spatial probability map to obtain shared features. Then, communication data can be packaged based on the shared features and pose data to obtain shared information.
[0128] Based on the above embodiments, step 132 includes:
[0129] Step 1321: Based on the features of the autonomous agent and the spatial probability map, thresholding is used to select features and obtain shared features;
[0130] Step 1322: Based on the shared features and the pose data, the communication data is packaged to obtain shared information.
[0131] Specifically, the server obtains the characteristics F of the autonomous agent. e And spatial probability graph F evd Afterwards, only retain and transmit the target confidence level that is not lower than a certain threshold T. 1 The characteristics of the autonomous agent are used to obtain spatially sparse but perceptually crucial shared features. Where δ(F) evd -T 1 )and These are the binary function and the pointwise product, respectively.
[0132] After obtaining the shared feature F share and pose data O e Then, general lossless data compression algorithms such as Huffman coding, arithmetic coding, or run-length coding can be used to encode the data to be transmitted (shared features and pose data), and the resulting compressed data can be split and packaged to obtain shared information S for communication and use by other intelligent agents. e .
[0133] Based on the above embodiments, step 140 includes:
[0134] Step 141: Based on multiple shared features and multiple pose data in the shared information, perform position reorganization to restore the multiple shared features to multiple feature maps in the bird's-eye view space;
[0135] Step 142: Based on the effective perception area of the autonomous agent, the feature regions of the multiple feature maps are cropped to obtain multiple roughly aligned feature maps.
[0136] Step 143: Aggregate information from the multiple coarsely aligned feature maps to obtain the cooperative agent features.
[0137] Specifically, the server obtains shared information from K cooperative agents. Where K is a non-fixed value and changes dynamically.
[0138] After obtaining the shared information of each cooperative agent Afterwards, the data can be decompressed to obtain its corresponding shared features. pose data
[0139] Obtaining shared features of the cooperative agent Then, the feature map of the BEV space is restored by repositioning. Then based on the pose data of the cooperative agent And pose data of the autonomous agent O e The feature map of the cooperative agent The coordinate origin is moved to the feature map F of the autonomous agent. e The origin of the coordinate system is determined, and feature regions are cropped using the effective perception area of the autonomous agent, retaining only the collaborative features useful for the agent's perception. This results in a roughly aligned feature map that is spatiotemporally consistent with the autonomous agent but locally misaligned.
[0140] After obtaining a rough alignment feature map for each co-agent Subsequently, max pooling can be used at the agent level to aggregate information, resulting in a single spatiotemporally consistent but locally misaligned co-agent feature F. s .
[0141] Due to the coarse alignment feature map of each co-agent It only covers a portion of the effective perception area of the autonomous agent, and the features of each co-agent are also spatially sparse. Max pooling is beneficial for obtaining a single co-agent feature containing rich feature information.
[0142] Based on the above embodiments, step 160 includes:
[0143] Step 161: Perform feature enhancement on the fused features to obtain enhanced features;
[0144] Step 162: Based on the enhanced features, perform three-dimensional target detection to obtain the target's position and size information;
[0145] Step 163: Based on the enhanced features, perform scene segmentation to obtain typical attribute information for each spatial location in the scene;
[0146] Step 164: Based on the location and size information and the typical attribute information, obtain the collaborative perception result.
[0147] Specifically, the server obtains the fusion feature F fusion Then, the convolutional neural network module f can be used. enhance This yields enhanced features F with stronger abstraction and discriminative abilities. enhance It can be represented as: F enhance =f enhance (F fusion ).
[0148] In obtaining enhanced feature F enhance Then, based on the enhanced features, 3D target detection can be performed to obtain the target's position and size information. det =f det (F enhance That is, the three-dimensional target detection module f can be used. det The target's location and size information can also be obtained using the BEV semantic segmentation module. seg Scene segmentation is performed on the enhanced features to obtain typical attribute information O for each spatial location in the scene. seg =f seg (F enhance Finally, based on the target's position and size information O det Typical attribute information O for each spatial location in the scene seg Together they constitute the results of collaborative perception.
[0149] Here, scene segmentation refers to classifying each pixel to determine its class. Instance segmentation, a subtype of scene semantic segmentation, performs both localization and semantic segmentation on each target, with each target being an instance. The task is ultimately evaluated based on the segmentation accuracy of each instance.
[0150] The convolutional neural network module, 3D object detection module, and BEV semantic segmentation module mentioned here all refer to network models composed of multiple layers of convolutional neural networks. The convolutional neural network module aims to enhance the semantic abstraction and spatial discrimination capabilities of features, such as the frequently used "ResNet" model or "ResNet+FPN" model.
[0151] Among them, the collaborative perception results obtained based on the fusion features can be achieved through a collaborative perception model, which can be trained based on the sample point cloud data, sample pose data and label collaborative perception results of the self-intelligent agent's LiDAR.
[0152] Sample point cloud data, sample pose data, and label-based collaborative perception results from the LiDAR can be collected in advance from the intelligent agent. An initial collaborative perception model can also be built in advance. Here, the label-based collaborative perception results include 3D target detection labels and BEV semantic segmentation labels.
[0153] The initial cooperative perception model can functionally include two parts: 3D object detection and scene segmentation. In this process, the initial 3D object detection model and the initial scene segmentation model can be used as the initial cooperative perception model.
[0154] After obtaining the initial cooperative perception model, which includes the initial 3D target detection model and the initial scene segmentation model, the initial cooperative perception model can be trained using pre-collected sample point cloud data, sample pose data, and labeled cooperative perception results from the LiDAR of the autonomous agent.
[0155] First, the sample point cloud data and sample pose data of the LiDAR of the autonomous agent are input into the initial cooperative perception model, and the initial cooperative perception model outputs the 3D target detection results and BEV semantic segmentation results.
[0156] After obtaining the 3D object detection results and BEV semantic segmentation results based on the initial collaborative perception model, the 3D object detection labels and 3D object detection results can be compared, and the 3D object detection loss L can be calculated based on the degree of difference between the two. det The BEV semantic segmentation results are compared with the BEV semantic segmentation labels, and the BEV semantic segmentation loss L is calculated based on the degree of difference between the two. segThen, the total loss is determined based on the 3D target detection loss and the BEV semantic segmentation loss. Finally, the parameters of the initial collaborative perception model are iterated as a whole based on the total loss. The initial collaborative perception model after parameter iteration is denoted as the collaborative perception model.
[0157] It is understandable that the greater the difference between the pre-collected 3D target detection labels and the 3D target detection results, the greater the 3D target detection loss; conversely, the smaller the difference between the pre-collected 3D target detection labels and the 3D target detection results, the smaller the 3D target detection loss.
[0158] It is understandable that the greater the difference between the BEV semantic segmentation result and the pre-collected BEV semantic segmentation labels, the greater the BEV semantic segmentation loss; conversely, the smaller the difference between the BEV semantic segmentation result and the pre-collected BEV semantic segmentation labels, the smaller the BEV semantic segmentation loss.
[0159] The initial collaborative perception model after parameter iteration has the same structure as the initial collaborative perception model. Therefore, the collaborative perception model can be divided into two parts: 3D object detection and BEV semantic segmentation.
[0160] Here, the Cross Entropy Loss Function, the Mean Squared Error (MSE) Loss Function, or the Stochastic Gradient Descent Method can be used to update the parameters of the initial co-sensing model. This embodiment of the invention does not impose specific limitations on these methods.
[0161] Wherein, the loss for 3D target detection is expressed as L det The BEV semantic segmentation loss is represented as L seg Thus, the total loss function L = L det +η·L seg , where η is the weighting adjustment coefficient.
[0162] Based on the above embodiments, Figure 5 This is a flowchart illustrating step 120 in the spatiotemporal error-tolerant multi-agent cooperative sensing method provided by the present invention, as shown below. Figure 5 As shown, step 120 includes:
[0163] Step 121: Project the point cloud data onto the bird's-eye view space and resample it to obtain normalized point columns;
[0164] Step 122: Based on the dot-column coding network, extract the dot-column abstract features from the normalized dot columns to obtain the self-intelligent agent features;
[0165] The dot-column encoding network is trained based on sample normalized dot columns, 3D object detection labels, and BEV semantic segmentation labels, and is combined with a 3D object detection model and a scene segmentation model. The 3D object detection model is used to perform 3D object detection based on the features of the autonomous agent, and the scene segmentation model is used to perform semantic segmentation based on the features of the autonomous agent.
[0166] Specifically, after acquiring the point cloud data, it can be projected onto the bird's-eye view space and resampled to obtain normalized point columns. That is, each observation is projected onto a corresponding discrete grid in the BEV space. Each discrete grid contains a different number of point clouds, commonly referred to as a point column. Let n be the standard number of observations in a point column. pillar Each observation contains 4 values. Then, for observations with a number greater than n... pillar Random sampling is performed on the point pillars, and the number of observations is less than n. pillar The point columns are randomly copied to obtain rasterized, normalized point columns with a fixed number of observations.
[0167] The server obtains the normalized point column F. pillar Then, the dot-column coding network f can be used. pillar The geometric structural features of the point pillars are obtained, i.e., the features of the autonomous agent in the BEV space. The features of the autonomous agent can be represented as:
[0168] Among them, the point-column coding network f pillar PointNet can be used, or the data can be further divided into voxels along the z-axis. The geometric features of each voxel can be calculated and concatenated, and then 1×1 convolution can be used for feature aggregation and dimension adjustment.
[0169] In order to better extract features from the agent, the dot-column coding network needs to be obtained through the following steps before step 122:
[0170] Normalized point pillars of samples can be collected in advance, and an initial point pillar encoding network, an initial 3D target detection model, and an initial scene segmentation model can be constructed in advance. The initial 3D target detection model is used to perform 3D target detection based on the features of the autonomous agent to obtain the 3D target detection result. The initial scene segmentation model is used to perform BEV semantic segmentation based on the features of the autonomous agent to obtain the BEV semantic segmentation result.
[0171] After obtaining the initial point-pillar encoding network, it can be trained based on the sample normalized point pillars, in conjunction with the initial 3D object detection model and the initial scene segmentation model, and the trained initial point-pillar encoding network can be used as the point-pillar encoding network.
[0172] In this process, the initial point-pillar encoding network, the initial 3D object detection model, and the initial scene segmentation model can be used as the initial detection model, which is the initial model used to train the initial point-pillar encoding network. The initial point-pillar encoding network here can be a PointNet network, etc., and this embodiment of the invention does not specifically limit it.
[0173] After obtaining the initial detection model, the pre-collected normalized point columns of samples, 3D object detection labels, and BEV semantic segmentation labels can be used to train the initial detection model:
[0174] First, the normalized point pillars of the samples are input into the initial point pillar encoding network. The initial point pillar encoding network extracts abstract features from the normalized point pillars of the samples to obtain the initial agent features. It can be understood that the initial point pillar encoding network is the initial model before the initial detection model is trained. To distinguish it from the agent features output by the initial detection model, the agent features output by the initial point pillar encoding network are denoted as the initial agent features.
[0175] Secondly, the initial autonomous agent features are input into the initial convolutional neural network, which obtains and outputs enhanced features. The enhanced features are then input into the initial 3D object detection model and the initial scene segmentation model, respectively, which obtain and output the 3D object detection results and the BEV semantic segmentation results.
[0176] After obtaining the 3D object detection results and BEV semantic segmentation results based on the initial detection model, the 3D object detection labels and 3D object detection results can be compared, and the 3D object detection loss L can be calculated based on the degree of difference between the two. det The BEV semantic segmentation results are compared with the BEV semantic segmentation labels, and the BEV semantic segmentation loss L is calculated based on the degree of difference between the two. seg Then, the total loss is determined based on the 3D object detection loss and the BEV semantic segmentation loss. Finally, based on the total loss, the parameters of the initial detection model are iterated as a whole. The initial point-column coding network in the initial detection model after parameter iteration can be directly used as the point-column coding network.
[0177] It is understandable that the greater the difference between the pre-collected 3D target detection labels and the 3D target detection results, the greater the 3D target detection loss; conversely, the smaller the difference between the pre-collected 3D target detection labels and the 3D target detection results, the smaller the 3D target detection loss.
[0178] It is understandable that the greater the difference between the BEV semantic segmentation result and the pre-collected BEV semantic segmentation labels, the greater the BEV semantic segmentation loss; conversely, the smaller the difference between the BEV semantic segmentation result and the pre-collected BEV semantic segmentation labels, the smaller the BEV semantic segmentation loss.
[0179] After training the dot-column encoding network, the dot-column abstract feature extraction can be performed on the normalized dot columns based on the dot-column encoding network to obtain the features of the autonomous agent.
[0180] Based on any of the above embodiments, a spatiotemporally error-tolerant multi-agent cooperative sensing method comprises the following steps:
[0181] The first step is to acquire point cloud data and pose data of the lidar from the intelligent agent.
[0182] The second step involves projecting the point cloud data onto the bird's-eye view space and resampling it to obtain normalized point pillars. Based on the point pillar coding network, the normalized point pillars are used to extract point pillar abstract features to obtain the features of the autonomous agent.
[0183] The dot-column encoding network here is trained based on sample normalized dot columns, 3D object detection labels, and BEV semantic segmentation labels, and is jointly trained with a 3D object detection model and a scene segmentation model. The 3D object detection model is used to perform 3D object detection based on the features of the autonomous agent, and the scene segmentation model is used to perform semantic segmentation based on the features of the autonomous agent.
[0184] The third step is to estimate the target probability of the features of the autonomous agent after importance selection, and obtain the spatial probability map.
[0185] The fourth step involves using thresholding to select shared features based on the features of the autonomous agent and the spatial probability map.
[0186] The fifth step is to package the communication data based on the shared features and pose data to obtain the shared information.
[0187] The sixth step involves reorganizing the positions of multiple shared features and multiple pose data from the shared information, restoring the multiple shared features to multiple feature maps in the bird's-eye view space.
[0188] Step 7: Based on the effective perception area of the autonomous agent, perform feature region cropping on multiple feature maps to obtain multiple roughly aligned feature maps.
[0189] The eighth step is to aggregate information from multiple roughly aligned feature maps to obtain the cooperative agent features.
[0190] Step 9: Based on the features of the self-intelligent agent and the features of the cooperative agent, obtain the set of candidate error quantities and the set of candidate confidence values corresponding to the set of candidate error quantities required for adaptive and accurate alignment of each feature spatial location.
[0191] Step 10: Based on the features of the self-agent agent, the features of the co-agent agent, the set of candidate error quantities, and the set of candidate confidence scores, the redundancy enhancement features are obtained.
[0192] In the eleventh step, complementary enhancements are performed based on the redundancy enhancement features, the cooperative agent features, and the self-agent perception blind spot map to obtain the fused features.
[0193] The steps for determining the blind spot map of the autonomous agent here include:
[0194] The spatial probability map is smoothed, binarized, and inverted to obtain the spatial demand map.
[0195] Based on the intensity values of the cooperative agent features, an effective spatial graph is obtained;
[0196] Based on the spatial demand map and the effective space map, the perception blind spot map of the autonomous agent is obtained.
[0197] Step 12: Enhance the fused features to obtain enhanced features.
[0198] Based on enhanced features, 3D target detection is performed to obtain the target's position and size information;
[0199] Based on the enhanced features, scene segmentation is performed to obtain typical attribute information for each spatial location in the scene;
[0200] Based on location and size information and typical attribute information, collaborative perception results are obtained.
[0201] The spatiotemporal error-tolerant multi-agent cooperative sensing device provided by the present invention is described below. The spatiotemporal error-tolerant multi-agent cooperative sensing device described below and the spatiotemporal error-tolerant multi-agent cooperative sensing method described above can be referred to in correspondence.
[0202] Based on any of the above embodiments, the present invention provides a spatiotemporal error-tolerant multi-agent cooperative sensing device. Figure 6 This is a schematic diagram of the spatiotemporal error-tolerant multi-agent cooperative sensing device provided by the present invention, as shown below. Figure 6 As shown, the device includes:
[0203] The acquisition unit 610 is used to acquire the point cloud data of the lidar of the intelligent agent and the pose data of the lidar.
[0204] The feature encoding unit 620 is used to perform feature encoding on the point cloud data in a bird's-eye view space to obtain the features of the autonomous agent;
[0205] The shared information unit 630 is used to select the importance of the features of the autonomous agent, and to package communication data based on the selected features of the autonomous agent and the pose data to obtain shared information;
[0206] Information aggregation unit 640 is used to aggregate the shared information to obtain cooperative agent features;
[0207] The enhancement unit 650 is used to perform redundancy enhancement and complementary enhancement based on the features of the self-intelligent agent and the features of the cooperative agent to obtain fused features;
[0208] The collaborative sensing unit 660 is used to perform collaborative sensing based on the fused features to obtain collaborative sensing results.
[0209] The apparatus provided in this invention acquires point cloud data and pose data of a LiDAR from an autonomous agent. In a bird's-eye view space, it performs feature encoding on the point cloud data to obtain autonomous agent features. Importance selection is performed on these autonomous agent features, and communication data is packaged based on the selected features and pose data to obtain shared information. This shared information is then aggregated to obtain co-agent features. Redundancy enhancement and complementary enhancement are performed on the autonomous agent features and co-agent features to obtain fused features. Based on these fused features, collaborative perception is performed to obtain the collaborative perception result. The co-agent features are obtained by aggregating shared information from multiple co-agents of varying numbers. This facilitates the acquisition of co-agent features with fixed formats and rich information, further enhancing the stability of the collaborative perception algorithm. Furthermore, the fused features are obtained through redundancy enhancement and complementary enhancement based on the autonomous agent features and co-agent features. The redundancy enhancement and complementary enhancement processes address positional misalignment issues, reducing the impact of spatiotemporal alignment errors between co-agent and autonomous agent features on the collaborative perception effect, further improving the performance of collaborative perception under conditions of spatiotemporal alignment errors.
[0210] Based on any of the above embodiments, the enhancement unit 650 is specifically used for:
[0211] An error determination unit is used to obtain, based on the features of the autonomous agent and the features of the cooperative agent, a set of candidate error quantities required for adaptive and precise alignment of each feature spatial position and a set of candidate confidence values corresponding to the set of candidate error quantities;
[0212] A redundancy enhancement feature unit is determined, which is used to obtain redundancy enhancement features based on the self-intelligent agent features, the co-intelligent agent features, the candidate error quantity set, and the candidate confidence set;
[0213] The fusion unit is used to perform complementary enhancements based on the redundant enhancement features, the cooperative agent features, and the self-agent perception blind spot map to obtain fused features.
[0214] Based on any of the above embodiments, the step of determining the blind spot map of the autonomous agent includes:
[0215] The spatial probability map is smoothed, binarized, and inverted to obtain the spatial demand map.
[0216] Based on the intensity values of the features of the cooperative agent, an effective spatial map is obtained;
[0217] Based on the spatial demand map and the effective spatial map, a perception blind spot map of the autonomous agent is obtained.
[0218] Based on any of the above embodiments, the shared information unit 630 is specifically used for:
[0219] Determine spatial probability map units to perform target probability estimation on the features of the autonomous agent after importance selection, and obtain the spatial probability map;
[0220] A shared information subunit is determined, which is used to package communication data based on the features of the autonomous agent, the spatial probability map, and the pose data to obtain shared information.
[0221] Based on any of the above embodiments, the shared information subunit is specifically used for:
[0222] Based on the features of the autonomous agent and the spatial probability map, thresholding is used to select features and obtain shared features.
[0223] Based on the shared features and the pose data, communication data is packaged to obtain shared information.
[0224] Based on any of the above embodiments, the information aggregation unit 640 is specifically used for:
[0225] Based on multiple shared features and multiple pose data in the shared information, position reorganization is performed to restore the multiple shared features to multiple feature maps in the bird's-eye view space;
[0226] Based on the effective perception area of the intelligent agent, the feature regions of the multiple feature maps are cropped to obtain multiple coarsely aligned feature maps.
[0227] Information is aggregated from the multiple coarsely aligned feature maps to obtain the cooperative agent features.
[0228] Based on any of the above embodiments, the collaborative sensing unit 660 is specifically used for:
[0229] The fused features are enhanced to obtain enhanced features;
[0230] Based on the enhanced features, three-dimensional target detection is performed to obtain the target's position and size information;
[0231] Based on the enhanced features, scene segmentation is performed to obtain typical attribute information for each spatial location in the scene;
[0232] Based on the location and size information and the typical attribute information, the collaborative perception result is obtained.
[0233] Based on any of the above embodiments, the feature encoding unit 620 is specifically used for:
[0234] The point cloud data is projected onto the bird's-eye view space and resampled to obtain normalized point columns;
[0235] Based on the dot-column coding network, dot-column abstract feature extraction is performed on the normalized dot columns to obtain the self-intelligent agent features;
[0236] The dot-column encoding network is trained based on sample normalized dot columns, 3D object detection labels, and BEV semantic segmentation labels, and is combined with a 3D object detection model and a scene segmentation model. The 3D object detection model is used to perform 3D object detection based on the features of the autonomous agent, and the scene segmentation model is used to perform semantic segmentation based on the features of the autonomous agent.
[0237] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include a processor 710, a communication interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute a spatiotemporally error-tolerant multi-agent cooperative perception method. This method includes: acquiring point cloud data and pose data of the lidar of the agent; performing feature encoding on the point cloud data in a bird's-eye view space to obtain agent features; performing importance selection on the agent features, and packaging communication data based on the importance-selected agent features and the pose data to obtain shared information; aggregating the shared information to obtain co-agent features; performing redundancy enhancement and complementary enhancement based on the agent features and the co-agent features to obtain fused features; and performing cooperative perception based on the fused features to obtain cooperative perception results.
[0238] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0239] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the spatiotemporal error-tolerant multi-agent cooperative perception method provided by the above methods. The method includes: acquiring point cloud data and pose data of the lidar of the agent; performing feature encoding on the point cloud data in a bird's-eye view space to obtain agent features; performing importance selection on the agent features, and packaging communication data based on the importance-selected agent features and the pose data to obtain shared information; aggregating the shared information to obtain co-agent features; performing redundancy enhancement and complementary enhancement based on the agent features and the co-agent features to obtain fused features; and performing cooperative perception based on the fused features to obtain cooperative perception results.
[0240] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements a spatiotemporally error-tolerant multi-agent cooperative perception method provided by the methods described above. This method includes: acquiring point cloud data and pose data of the lidar of the agent; performing feature encoding on the point cloud data in a bird's-eye view space to obtain agent features; performing importance selection on the agent features, and packaging communication data based on the importance-selected agent features and the pose data to obtain shared information; aggregating the shared information to obtain cooperative agent features; performing redundancy enhancement and complementary enhancement based on the agent features and the cooperative agent features to obtain fused features; and performing cooperative perception based on the fused features to obtain a cooperative perception result.
[0241] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0242] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0243] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A spatiotemporally error-tolerant multi-agent cooperative sensing method, characterized in that, include: The point cloud data and pose data of the lidar are obtained from the intelligent agent. In the bird's-eye view space, the point cloud data is feature-encoded to obtain the features of the autonomous agent; The importance of the agent features is selected, and communication data is packaged based on the selected agent features and the pose data to obtain shared information. The shared information is aggregated to obtain the characteristics of the cooperative agent; Based on the features of the self-intelligent agent and the features of the cooperative agent, redundancy enhancement and complementary enhancement are performed to obtain fused features; Based on the fusion features, collaborative perception is performed to obtain collaborative perception results; The process of performing redundancy enhancement and complementary enhancement based on the features of the self-intelligent agent and the features of the co-intelligent agent to obtain fused features includes: Based on the features of the self-intelligent agent and the features of the co-intelligent agent, a set of candidate error quantities required for adaptive and precise alignment of each feature spatial location and a set of candidate confidence values corresponding to the set of candidate error quantities are obtained. Based on the self-agent features, the co-agent features, the candidate error set, and the candidate confidence set, redundancy enhancement features are obtained. Based on the redundancy enhancement features, the cooperative agent features, and the self-agent perception blind spot map, complementary enhancements are performed to obtain fused features.
2. The spatiotemporal error-tolerant multi-agent cooperative sensing method according to claim 1, characterized in that, The steps for determining the blind spot map of the autonomous agent include: The spatial probability map is smoothed, binarized, and inverted to obtain the spatial demand map. Based on the intensity values of the features of the cooperative agent, an effective spatial map is obtained; Based on the spatial demand map and the effective spatial map, a perception blind spot map of the autonomous agent is obtained.
3. The spatiotemporal error-tolerant multi-agent cooperative sensing method according to claim 1, characterized in that, The communication data is packaged based on the self-intelligent agent features selected according to importance and the pose data to obtain shared information, including: Target probability estimation is performed on the features of the autonomous agent after importance selection to obtain a spatial probability map; Based on the features of the autonomous agent, the spatial probability map, and the pose data, communication data is packaged to obtain shared information.
4. The spatiotemporal error-tolerant multi-agent cooperative sensing method according to claim 3, characterized in that, The process of packaging communication data based on the autonomous agent's features, the spatial probability map, and the pose data to obtain shared information includes: Based on the features of the autonomous agent and the spatial probability map, thresholding is used to select features and obtain shared features. Based on the shared features and the pose data, communication data is packaged to obtain shared information.
5. The spatiotemporal error-tolerant multi-agent cooperative sensing method according to claim 1, characterized in that, The process of aggregating the shared information to obtain cooperative agent features includes: Based on multiple shared features and multiple pose data in the shared information, position reorganization is performed to restore the multiple shared features to multiple feature maps in the bird's-eye view space; Based on the effective perception area of the intelligent agent, the feature regions of the multiple feature maps are cropped to obtain multiple coarsely aligned feature maps. Information is aggregated from the multiple coarsely aligned feature maps to obtain the cooperative agent features.
6. The spatiotemporal error-tolerant multi-agent cooperative sensing method according to any one of claims 1 to 5, characterized in that, The process of performing collaborative perception based on the fused features to obtain collaborative perception results includes: The fused features are enhanced to obtain enhanced features; Based on the enhanced features, three-dimensional target detection is performed to obtain the target's position and size information; Based on the enhanced features, scene segmentation is performed to obtain typical attribute information for each spatial location in the scene; Based on the location and size information and the typical attribute information, the collaborative perception result is obtained.
7. The spatiotemporal error-tolerant multi-agent cooperative sensing method according to any one of claims 1 to 5, characterized in that, The feature encoding of the point cloud data in the bird's-eye view space to obtain the features of the autonomous agent includes: The point cloud data is projected onto the bird's-eye view space and resampled to obtain normalized point columns; Based on the dot-column coding network, dot-column abstract feature extraction is performed on the normalized dot columns to obtain the self-intelligent agent features; The dot-column encoding network is trained based on sample normalized dot columns, 3D object detection labels, and BEV semantic segmentation labels, and is combined with a 3D object detection model and a scene segmentation model. The 3D object detection model is used to perform 3D object detection based on the features of the autonomous agent, and the scene segmentation model is used to perform semantic segmentation based on the features of the autonomous agent.
8. A spatiotemporally error-tolerant multi-agent cooperative sensing device, characterized in that, include: The acquisition unit is used to acquire point cloud data of the lidar from the intelligent agent and pose data of the lidar. The feature encoding unit is used to encode the point cloud data in a bird's-eye view space to obtain the features of the autonomous agent; A shared information unit is determined, which is used to select the importance of the features of the autonomous agent, and to package communication data based on the selected features of the autonomous agent and the pose data to obtain shared information; An information aggregation unit is used to aggregate the shared information to obtain cooperative agent features. An enhancement unit is used to perform redundancy enhancement and complementary enhancement based on the features of the autonomous agent and the features of the cooperative agent to obtain fused features; A collaborative sensing unit is used to perform collaborative sensing based on the fused features to obtain collaborative sensing results; The enhancement unit is specifically used for: Based on the features of the self-intelligent agent and the features of the co-intelligent agent, a set of candidate error quantities required for adaptive and precise alignment of each feature spatial location and a set of candidate confidence values corresponding to the set of candidate error quantities are obtained. Based on the self-agent features, the co-agent features, the candidate error set, and the candidate confidence set, redundancy enhancement features are obtained. Based on the redundancy enhancement features, the cooperative agent features, and the self-agent perception blind spot map, complementary enhancements are performed to obtain fused features.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the spatiotemporal error-tolerant multi-agent cooperative perception method as described in any one of claims 1 to 7.