Multi-agent collaborative sensing system and method based on evidence regression

Through the multi-agent collaborative perception system of evidence regression, the point cloud features are processed using convolutional neural network and sparse graph generation module, which solves the problem of insufficient perception accuracy and real-time performance of the multi-agent system in a dynamic environment, and achieves more efficient information fusion and robustness enhancement.

CN120354916APending Publication Date: 2025-07-22HARBIN INST OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411475406.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-22
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Multi-agent systems cannot effectively deal with noise, delay and sensor errors when dealing with dynamic environments, resulting in insufficient perceptual accuracy and real-time performance.

Method used

A multi-agent collaborative perception system based on evidence regression is adopted to extract point cloud features through convolutional neural networks, generate sparse graphs and broadcast transmissions, and use sparse graph generation modules, feature selection modules and evidence regression modules for information fusion and uncertainty estimation to optimize perceived information.

Benefits of technology

It improves the perceived accuracy and robustness of multi-agent systems in complex environments, enhances independent decision-making capabilities, and can effectively deal with sensor heterogeneity and noise interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354916A_ABST
    Figure CN120354916A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent cooperative sensing system based on an evidence neural network, and relates to a multi-agent cooperative sensing implementation method and implementation system. The invention aims to solve the problems that an intelligent agent cannot deal with noise, delay and sensor errors when processing a dynamic environment, and the sensing precision and the real-time performance cannot be ensured. The method comprises the following steps: step 1, extracting point cloud features by using a convolutional neural network; 2, obtaining a sparse graph from the point cloud features through a sparse graph generation module, and transmitting the sparse graph to other agents in a broadcast form; step 3, the adjacent intelligent agent determines whether to transmit the features to the own intelligent agent or not through a feature selection module; step 4, the self intelligent agent receives the characteristics transmitted by the adjacent intelligent agent and fuses the characteristics; and step 5, the fusion features pass through a detection module and an evidence regression module to obtain perception information and uncertainty estimation, and the perception information is optimized according to the uncertainty estimation module to obtain a reliable detection result. The invention relates to the field of multi-agent collaborative awareness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-agent collaborative perception system and method, belonging to the technical field of multi-agent collaborative environment perception. Background Art

[0002] Multi-Agent Collaborative Perception is an important technology that has emerged in recent years with the development of intelligent systems and sensor technologies. Due to the limitations of the sensor's perspective, the finiteness of the perception range, and the constraints of computing resources, traditional single-agent perception systems often cannot provide comprehensive and accurate environmental perception capabilities in complex environments. With the gradually rich application scenarios of multi-agent systems such as autonomous driving, intelligent robots, and drone swarms, the perception capabilities of single agents are no longer sufficient to meet the requirements of these complex tasks. In order to improve the perception capabilities of multi-agent systems in dynamic environments, collaborative perception technology has emerged;

[0003] In multi-agent collaborative perception, multiple agents share their perception data with each other and combine information collected from different perspectives, sensor types, and different time points to achieve comprehensive perception and understanding of the target area. Collaborative perception can not only improve the perception accuracy of agents for the environment but also reduce perception blind spots through cooperation among multiple agents, enhancing the robustness and fault tolerance of the system;

[0004] In the scenario of autonomous driving, collaborative perception technology can significantly improve the safety and driving efficiency of vehicles. For example, in platoon driving or intersections, the front and rear vehicles can share sensor data to timely perceive potential obstacles or dangerous situations, and thus make early responses to avoid accidents. Similarly, collaborative perception can also be used for drone formation operations or robot swarm collaborations to complete tasks such as environmental monitoring and object recognition in large-scale areas;

[0005] However, multi-agent collaborative perception faces many technical challenges in practical applications. First, the transmission of perception information between agents is limited by bandwidth, communication latency, and data synchronization problems. Second, the perception data of different agents may be heterogeneous, including differences in aspects such as timestamps, resolutions, coordinate systems, and sensor noise, which brings difficulties to data fusion. In addition, perception uncertainty is also an issue that needs to be addressed. How agents handle noise, latency, and sensor errors when dealing with dynamic environments and ensure the accuracy and real-time nature of perception is a key direction for technological development. Summary of the Invention

[0006] To solve the problem that agents cannot handle noise, latency, and sensor errors when dealing with dynamic environments, and cannot guarantee perception accuracy and real-time performance, the present invention proposes a multi-agent collaborative perception system and method based on evidence regression;

[0007] The technical solution adopted by the present invention to solve the above problems is as follows: The steps of the method of the present invention include:

[0008] Step 1: Use a convolutional neural network to extract point cloud features;

[0009] Step 2: The point cloud features pass through a sparse graph generation module to obtain a sparse graph;

[0010] Step 3: The self-agent transmits the request graph to neighboring agents in a broadcast form, and the neighboring agents determine whether to transmit the features to the self-agent through a feature selection module;

[0011] Step 4: The self-agent receives the features transmitted by neighboring agents and fuses them;

[0012] Step 5: The fused features pass through a detection module and an evidence regression module respectively to obtain perception information and uncertainty estimation, and optimize the perception information according to the latter to obtain a reliable detection result.

[0013] Further, in Step 1, the PointPillar network is used to extract point cloud features from the input point cloud data; the specific process is as follows:

[0014] Step 101: Each agent uses a lidar or a depth camera to capture the point cloud data P of the environment i ;

[0015] Step 102: The point cloud data P in Step 101 i is used as the input of the PointPillar network to obtain the corresponding point cloud features f i ;

[0016] Further, in Step 102, the point cloud features use a sparse graph generation module to generate a point cloud sparse graph M i , and obtain an information request graph P i = 1 - M i ;

[0017] Further, the request graph R i is transmitted to other agents in a broadcast communication form to construct a communication group, and the point cloud features representing environmental perception are mutually transmitted; the specific process is as follows:

[0018] Step 201: Taking the self-agent as an example, the information request graph R ego is transmitted to neighboring agent j in a broadcast form;

[0019] Step 202. The neighboring agent matches the point cloud sparse graph M of the ontology j with the received information request graph R ego for matching S(M j ΘR ego );

[0020] Step 203. Determine the communication group centered on its own agent according to the matching result;

[0021] Step 204. The neighboring agents within the communication group transmit the point cloud features to its own agent f j→i .

[0022] Furthermore, the specific process of the self-agent receiving and fusing the features transmitted by the neighboring agents in Step 4 is as follows:

[0023] Step 401. The self-agent receives the point cloud features f of the neighboring agents with complementary information j→i ;

[0024] Step 402. Use the multi-head attention mechanism to fuse the features obtained by the ontology with the received features to obtain the final self-agent feature F ego = ATTN(f j→ego , f ego ).

[0025] Furthermore, in Step 5, the fused features pass through the detection module and the evidence regression module respectively. The detection module includes a regression module and a classification module, and the evidence regression module includes a detection module and an uncertainty estimation module; the specific process is as follows:

[0026] Step 501. The classification module that the fused features pass through obtains the class probability of the detection target, and at the same time the regression module obtains the bounding box B-(x, y, z, h, w, l, a) of the detection target, where (x, y, z) represents the center point coordinates of the bounding box, (l, w, h) represents the length, width and height of the bounding box, and α represents the offset angle of the bounding box;

[0027] Step 502. The fused features pass through the uncertainty estimation module to model the probability distribution of the parameters (x, y, z, h, w, l, a) of the bounding box;

[0028] Step 503. Generate a confidence estimate for each parameter of the bounding box, and its parameters follow a Gaussian distribution B-N(μ, σ 2 with an unknown mean μ and an unknown variance δ 2 );

[0029] Step 504. In order to estimate the unknown mean μ and unknown variance δ 2 , Gaussian priors μ-N(γ, σ 2 υ-1) and the inverse gamma prior σ 2 -Γ -1 (α, β) is used to model it;

[0030] Step 505: Using the Gaussian conjugate prior, estimate the posterior distributions q(μ, σ 2 ), that is, obtain the normal

[0031] inverse gamma distribution:

[0032]

[0033] Step 506: By marginalizing the likelihood parameter m, calculate the probability distribution of the observed values, that is:

[0034]

[0035] Step 507: Apply the normal inverse gamma distribution to the Gaussian likelihood function to obtain an analytical solution:

[0036]

[0037] where St(y∣μ, σ 2 , ν) represents the Student t-distribution evaluated at a specific value y, and its three parameters are the location parameter μ, the scale parameter δ 2 , and the degrees of freedom parameter v.

[0038] Furthermore, the loss function is:

[0039] L = L Det + L Unc = η cls L Cls + η reg L Reg + L NLL + η R L R + η U L U

[0040] where η cls , η reg , η R , η U are hyperparameters, and L Det is the loss function of the detection module, including:

[0041] The classification loss L Cls based on evidence perception:

[0042] L Cls = (ν + 2α - 1)L cls - log(ν + 2α - 1)

[0043] Evidence-aware regression loss L Reg ;

[0044] L Reg =(ν + 2α - 1)L reg -log(ν + 2α - 1)

[0045] L Unc is the loss function of the uncertainty estimation module, including:

[0046] Negative log-likelihood loss L NLL :

[0047]

[0048] Evidence misleading regularization term L R :

[0049] L R =|y - γ|(2ν + α)

[0050] Non-saturation regularization term L U :

[0051]

[0052] The perception system described in the present invention includes a feature extraction module, a sparse graph generation module, a selection and exchange module, a feature fusion module, a detection module, an evidence regression module, a terminal, and a computer-readable storage medium;

[0053] The feature extraction module uses a convolutional neural network to extract point cloud features;

[0054] The sparse graph generation module generates a feature sparse graph using the data features;

[0055] The selection and exchange module packages the request graph, constructs a communication group, and transmits the information packet among the agents through the communication group;

[0056] The feature fusion module is used to fuse the received information with the features local to each agent;

[0057] The detection module is used to decode the fused features to obtain the detection result;

[0058] The evidence regression module is used to model the probability distribution of the bounding box parameters and combine with the detection module to obtain the final detection result;

[0059] The terminal includes a memory, a processor, and a computer program stored on the memory and executable on the processor;

[0060] The computer-readable storage medium is used to store the computer program.

[0061] The beneficial effects of the present invention are as follows: By introducing a mechanism for uncertainty estimation, the multi-agent system can more effectively integrate data from different sensors. In practical applications, each agent may generate errors due to factors such as sensor accuracy and environmental interference. Evidential deep learning can model these uncertainties, adopt the framework of evidential theory to quantify the uncertainty of information, convert sensor data into probability distributions, and quantify the confidence and credibility of the data. This process can not only identify which data is reliable but also help the system understand the correlations between data. By introducing the concepts of trust and confidence, agents can more effectively handle information ambiguity and ensure that all possible situations are considered in the decision-making process. This method can significantly improve the robustness of the system in harsh environments.

[0062] The significance of the present invention lies in enhancing the collaborative perception performance of the multi-agent system, thereby enhancing the autonomous decision-making ability of agents in complex environments Brief Description of the Drawings

[0063] Figure 1 is the system structure diagram of the present invention;

[0064] Figure 2 is the schematic diagram of the evidential regression module. Detailed Embodiments

[0065] Detailed Embodiment 1: As Figure 1 shown, a multi-agent collaborative perception method based on evidential regression, the specific steps include:

[0066] Step 1: Use a convolutional neural network to extract point cloud features; use the PointPillar network to extract point cloud features from the input point cloud data; the specific process is as follows:

[0067] Step 101: Each agent uses a lidar or a depth camera to capture the point cloud data P of the environment i ;

[0068] Step 102: Take the point cloud data P in Step 101 i as the input of the PointPillar network to obtain the corresponding point cloud feature f i ; The point cloud feature uses a sparse graph generation module to generate a point cloud sparse graph M i , and obtain an information request graph P i= 1 - M i ; The request graph R i is transmitted to other agents in the form of broadcast communication to construct a communication group and mutually transmit the point cloud features representing environmental perception; the specific process is as follows:

[0069] Step 201: Taking the agent itself as an example, take the information request graph R egoTransmit it to neighboring agent j in the form of broadcast;

[0070] Step 202, the neighboring agent sparsifies the point cloud graph M of the ontology j And the received information request graph R ego Perform matching S(M j ΘR ego );

[0071] Step 203, determine the communication group centered on its own agent according to the matching result;

[0072] Step 204, the neighboring agents within the communication group transmit the point cloud features to its own agent f j→i ;

[0073] Step 2, the point cloud features pass through the sparse graph generation module to obtain a sparse graph and;

[0074] Step 3, the agent itself transmits the request graph to neighboring agents in the form of broadcast, and the neighboring agents determine whether to transmit the features to the agent itself through the feature selection module;

[0075] Step 4, the agent itself receives the features transmitted by neighboring agents and fuses them; the specific process of the agent itself receiving the features transmitted by neighboring agents and fusing them is as follows:

[0076] Step 401, the agent itself receives the point cloud features f of neighboring agents with complementary information j→i ;

[0077] Step 402, use the multi-head attention mechanism to fuse the features obtained by the ontology with the received features to obtain the final agent feature F ego = ATTN(f j→ego , f ego );

[0078] Step 5, the fused features pass through the detection module and the evidence regression module respectively to obtain the perception information and uncertainty estimation, and optimize the perception information according to the latter to obtain a reliable detection result; the fused features pass through the detection module and the evidence regression module respectively, the detection module includes a regression module and a classification module, and the evidence regression module includes a detection module and an uncertainty estimation module; the specific process is as follows:

[0079] Step 501, the classification module that the fused features pass through obtains the class probability of the detection target, and at the same time the regression module obtains the bounding box B-(x, y, z, h, w, l, a) of the detection target, where (x, y, z) represents the center point coordinates of the bounding box, (l, w, h) represents the length, width and height of the bounding box, and α represents the offset angle of the bounding box;

[0080] Step 502: The fused features are used by the uncertainty estimation module to model the probability distribution of the parameters (x, y, z, h, w, l, a) of the bounding box;

[0081] Step 503: Generate a confidence estimate for each parameter of the bounding box, and its parameters follow a Gaussian distribution B-N(μ,σ 2 with an unknown mean μ and an unknown variance δ 2 );

[0082] Step 504: To estimate the unknown mean μ and the unknown variance δ 2 , introduce a Gaussian prior μ-N(γ,σ 2 υ -1) and an inverse gamma prior σ 2 -Γ -1 (α, β) to model them;

[0083] Step 505: Adopt a Gaussian conjugate prior to estimate the posterior distribution q(μ, σ 2 ), that is, obtain a normal

[0084] inverse gamma distribution:

[0085]

[0086] Step 506: Calculate the probability distribution of the observed values by marginalizing the likelihood parameter m, that is:

[0087]

[0088] Step 507: Apply a normal inverse gamma distribution to the Gaussian likelihood function to obtain an analytical solution:

[0089]

[0090] where St(y∣μ,σ 2 ,ν) represents the Student's t-distribution evaluated at a specific value y, and its three parameters are the location parameter μμ, the scale parameter δ 2 , and the degrees of freedom parameter v.

[0091] Specific Embodiment 2: As Figure 1 shown, the loss function is:

[0092] L = L Det + L Unc = η cls L Cls + η reg L Reg + L NLL + η R L R + η U LU

[0093] where η cls 、η reg 、η R 、η U are hyperparameters, and L Det is the loss function of the detection module, including:

[0094] Evidence-aware classification loss L Cls :

[0095] L Cls = (ν + 2α - 1)L cls - log(ν + 2α - 1)

[0096] Evidence-aware regression loss L Reg ;

[0097] L Reg = (ν + 2α - 1)L reg - log(ν + 2α - 1)

[0098] L Unc is the loss function of the uncertainty estimation module, including:

[0099] Negative log-likelihood loss L NLL :

[0100]

[0101] Evidence misleading regularization term L R :

[0102] L R = |y - γ|(2ν + α)

[0103] Non-saturation regularization term L U :

[0104]

[0105] Specific implementation method three: As Figure 1 shown, a multi-agent collaborative perception system based on evidence regression includes a feature extraction module, a sparse graph generation module, a selection and exchange module, a feature fusion module, a detection module, an evidence regression module, a terminal, and a computer-readable storage medium;

[0106] The feature extraction module uses a convolutional neural network to extract point cloud features;

[0107] The sparse graph generation module generates a feature sparse graph using the data features;

[0108] Select and exchange module to pack request diagrams, construct communication groups, and transmit the information packets among agents through the communication groups;

[0109] The feature fusion module is used to fuse the received information with the features local to each agent;

[0110] The detection module is used to decode the fused features to obtain detection results;

[0111] The evidence regression module is used to model the probability distribution of the bounding box parameters and combine with the detection module to obtain the final detection results;

[0112] The terminal includes a memory, a processor, and a computer program stored on the memory and executable on the processor;

[0113] The computer-readable storage medium is used to store the computer program.

[0114] Symbol description table

[0115]

[0116] As mentioned above, it is only the preferred embodiment of the present invention, and it does not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any simple modification, equivalent replacement, and improvement made to the above embodiments according to the technical essence of the present invention within the spirit and principle of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A multi-agent collaborative perception method based on evidence regression, characterized in that, The specific steps include: Step 1: Extract point cloud features using a convolutional neural network; Step 2: The point cloud features pass through a sparse graph generation module to obtain a sparse graph; Step 3: The self-agent transmits the request graph in a broadcast form to neighboring agents, and the neighboring agents determine whether to transmit the features to the self-agent through a feature selection module; Step 4: The self-agent receives the features transmitted by the neighboring agents and fuses them; Step 5: The fused features pass through a detection module and an evidence regression module respectively to obtain perception information and uncertainty estimation, and optimize the perception information according to the latter to obtain a reliable detection result.

2. The multi-agent collaborative perception method based on evidence regression according to claim 1, wherein, Step 1 uses the PointPillar network to extract point cloud features from the input point cloud data; The specific process is as follows: Step 101: Each agent captures the point cloud data P of the environment using a lidar or a depth camera i ; Step 102: Use the point cloud data P from Step 101 i as the input of the PointPillar network to obtain the corresponding point cloud feature f i .

3. A multi-agent collaborative perception method based on evidence regression according to claim 2, characterized in that, In step 102, the point cloud features are used by the sparse graph generation module to generate a point cloud sparse graph M i , and an information request graph R i = 1 - M i .

4. A multi-agent collaborative perception method based on evidence regression according to claim 3, characterized in that, The requested graph R i is transmitted to other agents in a broadcast communication form to construct a communication group and mutually transmit the point cloud features representing environmental perception; the specific process is as follows: Step 201: Taking its own agent as an example, transmit the information request graph R ego to neighboring agent j in the form of broadcast; Step 202, the neighboring agent matches the point cloud sparse graph M of the ontology j with the received information request graph R ego to perform a match S(M j ΘR ego ); Step 203: Determine a communication group centered on the self-agent according to the matching result; Step 204: The neighboring agents within the communication group transmit the point cloud features to their own agent f j→i .

5. A multi-agent collaborative perception method based on evidence regression according to claim 1, characterized in that, The specific process of the self-agent receiving the features transmitted by the neighboring agents and fusing them in Step 4 is as follows: Step 401, the self-agent receives the point cloud feature f of neighboring agents with complementary information j→i ; Step 402: Use the multi-head attention mechanism to fuse the features obtained from the ontology with the received features to obtain the final self-agent feature F ego = ATTN(f j→ego , f ego ).

6. A multi-agent collaborative perception method based on evidence regression according to claim 1, characterized in that, In Step 5, the fused features pass through a detection module and an evidence regression module respectively. The detection module includes a regression module and a classification module, and the evidence regression module includes a detection module and an uncertainty estimation module; the specific process is as follows: Step 501: The fused features pass through the classification module to obtain the class probability of the detection target, and at the same time pass through the regression module to obtain the bounding box B~(x, y, z, h, w, l, a) of the detection target, where (x, y, z) represents the center point coordinates of the bounding box, (l, w, h) represents the length, width and height of the bounding box, and α represents the offset angle of the bounding box; Step 502: The fused features pass through the uncertainty estimation module to model the probability distribution of the parameters (x, y, z, h, w, l, a) of the bounding box; Step 503, generate a confidence estimate for each parameter of the bounding box, where the parameters follow a Gaussian distribution B~N(μ,σ 2 ) with an unknown mean μ and an unknown variance δ 2 ); Step 504: To estimate the unknown mean μ and unknown variance δ 2 respectively, Gaussian prior μ ~ N(γ, σ 2 υ -1) and inverse gamma prior σ 2 ~ Γ -1 (α, β) are introduced to model them; Step 505, using the Gaussian conjugate prior, estimate the posterior distributions q(μ, σ 2 ), that is, obtain the normal inverse gamma distribution: Step 506: Calculate the probability distribution of the observed value by marginalizing the likelihood parameter m, that is: Step 507: Apply a normal inverse gamma distribution to the Gaussian likelihood function to obtain an analytical solution: where, St(y∣μ,σ 2 ,ν) represents the Student's t-distribution evaluated at a specific value y, with its three parameters being the location parameter μ, the scale parameter δ 2 , and the degrees of freedom parameter v.

7. A multi-agent collaborative perception method based on evidence regression according to claim 1, characterized in that The loss function is: L = L Det + L Unc = η cls L Cls + η reg L Reg + L NLL + η R L R + η U L U where η cls , η reg , η R , η U are hyperparameters, and L Det is the loss function of the detection module, including: Evidence-aware Classification Loss L Cls : L Cls =(ν + 2α - 1)L cls -log(ν + 2α - 1) Evidence-aware regression loss L Reg ; L Reg =(ν + 2α - 1)L reg -log(ν + 2α - 1) L Unc It is the loss function of the uncertainty estimation module, including: Negative log-likelihood loss L NLL : Evidence misleading regularization term L R : L R = |y - γ|(2ν + α) Non-saturated regularization term L U :

8. A multi-agent collaborative perception system based on evidence regression, characterized in that, It includes a feature extraction module, a sparse graph generation module, a selection and exchange module, a feature fusion module, a detection module, an evidence regression module, a terminal and a computer-readable storage medium; The feature extraction module uses a convolutional neural network to extract point cloud features; The sparse graph generation module generates a feature sparse graph using the data features; The selection and exchange module packs the request graph, constructs a communication group, and transmits the information packet among agents through the communication group; The feature fusion module is used to fuse the received information with the local features of each agent; The detection module is used to decode the fused features to obtain a detection result; The evidence regression module is used to model the probability distribution of the bounding box parameters and combine with the detection module to obtain the final detection result; The terminal includes a memory, a processor and a computer program stored in the memory and executable on the processor; The computer-readable storage medium is used to store the computer program.

Citation Information

Cited By

  • Multi-agent cooperative sensing method and system for Internet of Vehicles

    CN121837878A