Multi-agent collaborative perception method and system based on double attention, terminal and storage medium
By introducing a dual attention weighting module in the multi-agent system, point cloud feature extraction and sparse graph generation is performed, the problem of insufficient adaptability of traditional systems in complex environments is solved, and a more efficient and reliable multi-agent collaborative perception effect is achieved.
Patent Information
- Application Number
- CN202510076283.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional multi-agent perception systems are not adaptable enough in complex and uncertain environments, making it difficult to build an efficient and reliable multi-agent collaborative perception system.
The multi-agent collaborative perception method based on dual attention is adopted to extract point cloud features through convolutional neural networks, generate sparse graphs, and feature fusion is performed through weighted dual-channel attention modules to improve the accuracy and real-timeness of perceived information.
The perceived performance and autonomous decision-making capabilities of multi-agent systems in complex environments are improved, the robustness and fault tolerance of the system are enhanced, and the perceived blind spots are effectively reduced.
Smart Images

Figure CN119992270A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multi-agent collaborative environmental perception based on deep learning, and specifically relates to a multi-agent collaborative perception method, system, terminal and storage medium based on dual attention. Background Art
[0002] As a product of the rapid development of intelligent systems and sensor technology, multi-agent collaborative perception has gradually become a research hotspot in recent years. Traditional single-agent perception systems often have difficulty achieving comprehensive and accurate environmental perception in complex environments due to the limitations of sensor perspectives, limited perception range, and constraints on computing resources. With the increasing complexity of multi-agent system application scenarios such as autonomous driving, intelligent robots, and drone swarms, the perception capabilities of a single agent can no longer meet the requirements of high-complexity tasks. Therefore, in order to improve the perception capabilities of multi-agent systems in dynamic environments, collaborative perception technology has emerged as an important means to solve this problem.
[0003] In the framework of multi-agent collaborative perception, multiple agents share their own perception data and integrate information obtained from different perspectives, sensor types and different time points, thereby achieving comprehensive perception and understanding of the target area. This collaborative approach can not only improve the accuracy of overall perception, but also effectively reduce perception blind spots through collaboration between multiple agents, and enhance the robustness and fault tolerance of the system. In addition, collaborative perception can optimize resource allocation and improve perception efficiency, especially in complex and dynamically changing environments, reflecting its unique advantages. Taking autonomous driving as an example, collaborative perception technology has significantly improved vehicle safety and driving efficiency. In scenarios such as convoy driving or intersections, vehicles can identify potential obstacles or dangerous situations in a timely manner by sharing sensor data, so as to take countermeasures in advance and reduce the risk of accidents. Similarly, collaborative perception also shows strong application potential in drone formations and robot group collaboration, which can support large-scale environmental monitoring, accurate object recognition and efficient completion of complex tasks.
[0004] However, multi-agent collaborative perception still faces many technical challenges in practical applications. First, data transmission between agents is constrained by bandwidth limitations, communication delays, and data synchronization issues, which affect collaborative efficiency. Second, the perception data of different agents are heterogeneous, including differences in timestamps, resolutions, coordinate systems, and sensor noise, which brings significant difficulties to data fusion. In a dynamic environment, agents need to effectively deal with noise and sensor errors to ensure the accuracy and real-time performance of perception. Solving these problems is a key direction for future technological development. Summary of the invention
[0005] The purpose of the present invention is to solve the problem that traditional multi-agent perception systems usually rely on predefined role allocation and information fusion mechanisms, but these methods have insufficient adaptability in complex and highly uncertain environments, to build an efficient and reliable multi-agent collaborative perception system, to improve the autonomy and intelligence level of multi-intelligent systems in complex environments, and then propose a multi-agent collaborative perception method, system, terminal and storage medium based on dual attention weighting.
[0006] The technical solution adopted by the present invention to solve the above problems is:
[0007] In a first aspect, the present invention provides a multi-agent collaborative perception method based on dual attention weighting, comprising the following steps:
[0008] Step 1: Extract point cloud features using convolutional neural network;
[0009] Step 2: The point cloud features are passed through a sparse graph generation module to obtain a sparse graph;
[0010] Step 3: The agent transmits the request graph to the neighboring agent in the form of broadcast, and the neighboring agent determines whether to transmit the feature to the agent through the feature selection module;
[0011] Step 4: The agent receives the features transmitted by the neighboring agents and performs feature fusion through the weighted dual-channel attention module;
[0012] Step 5: The fusion features are passed through the detection module to obtain perception information and obtain reliable detection results.
[0013] Furthermore, in step 1, the PointPillar network is used to extract point cloud features from the input point cloud data; the specific process is:
[0014] Step 1.1 Each agent uses a lidar or depth camera to capture point cloud data P of the environment;
[0015] Step 1.2: Use the point cloud data P in step 1 as the input of the PointPillar network to obtain the corresponding point cloud features f.
[0016] Furthermore, the point cloud features in step 1.2 use a sparse graph generation module to generate a point cloud sparse graph M, and obtain an information request graph R=1-M.
[0017] Furthermore, in step 2, the request graph R is transmitted to other agents in the form of broadcast communication to build a communication group, and point cloud features representing environmental perception are transmitted to each other; the specific process is:
[0018] Step 2.1 Taking the self-agent ego as an example, the information request graph R egoTransmitted to neighboring agent j in the form of broadcast;
[0019] Step 2.2 The neighboring agent j takes the point cloud sparse graph M of the entity j Figure R with received information request ego Match S(M j ⊙R ego );
[0020] Step 2.3: Determine the communication group centered on the self-agent ego based on the matching results;
[0021] Step 2.4 The neighboring agent j in the communication group transmits the point cloud features to its own agent.
[0022] Furthermore, the features in step 4 are fused, and the specific process includes:
[0023] Step 4.1 The self-agent ego receives the point cloud features f of the neighboring agents with complementary information j→ego ;
[0024] Step 4.2 Use the self-attention mechanism to concatenate the features acquired by the entity with the received features to obtain the fusion feature F ego =ATTN(fj → ego,f ego ).
[0025] Furthermore, the fused features in step 5 are passed through a weighted dual attention module; the specific process is:
[0026] Step 5.1 directly generates a channel-specific attention map that captures the interactions between channels by reshaping the fused features, performing matrix multiplication, and applying softmax normalization:
[0027]
[0028] where a ji Represents the influence of the i-th channel on the j-th channel, capturing the dependency between channels;
[0029] Step 5.2 applies it to the feature map through matrix operations, reshaping and scaling parameters to generate the final output C containing the dependencies between channels. j :
[0030]
[0031] Among them C j is the weighted sum of all channel features, and γ is a learnable parameter;
[0032] Step 5.3 Weight the channel features and C jMax pooling and average pooling are applied along the channel dimension of the input feature map, producing two feature maps:
[0033] X = concat(f m ,f a );
[0034] Step 5.4 applies three convolution kernels with different dilation rates to process the feature map X and generate the feature map f r=1 ,f r=2 ,f r=4 ;
[0035] Step 5.5 introduces a local convolutional layer to generate local features f local :
[0036] Step 5.6 performs weighted fusion of feature maps with different expansion rates and local convolution feature maps to obtain a spatial attention map:
[0037] S = σ(concat(f r=1 ,f r=2 ,f r=4 ,f local ))
[0038] Where σ represents the Sigmoid function;
[0039] Step 5.7 performs weighted fusion on the above two attention maps to obtain the final fusion feature F′:
[0040] F′=σ(ω c )·C+σ(ω s )·S
[0041] where ω c and ω s is a learnable weight parameter.
[0042] Furthermore, the loss function is:
[0043] L=η cls L Cls +η reg L Reg
[0044] where η cls , η reg is a hyperparameter, L Cls , L Reg is the loss function of the detection module.
[0045] In a second aspect, the present invention provides a multi-agent collaborative perception system based on dual attention, the system comprising:
[0046] Feature extraction module, which uses convolutional neural network to extract point cloud features;
[0047] A sparse graph generation module, which generates a feature sparse graph using the data features;
[0048] A selection and exchange module that packages a request graph, builds a communication group, and transmits the information packet between agents through the communication group;
[0049] The feature fusion module fuses the received information with the local features of each agent;
[0050] The detection module decodes the fused features to obtain the detection results.
[0051] In a third aspect, the present invention provides a terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, can be used to execute the steps of the method described in the first aspect, or to run the system described in the second aspect.
[0052] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can be used to execute the steps of the method described in the first aspect, or to execute the system described in the second aspect.
[0053] The beneficial effects of the present invention are:
[0054] The present invention introduces a dual attention weighting module to enable the multi-agent system to more effectively integrate data from different sensors. In practical applications, each agent may produce errors due to factors such as sensor accuracy and environmental interference. Channel attention highlights channels containing key semantic information by weighting feature maps across channels, while spatial attention focuses on the spatial position within the feature map to enhance the model's ability to focus on specific areas. By weighted fusion of these two types of attention at the same level, the model's attention to key channels and important spatial areas can be enhanced at the same time, making the feature information more detailed and refined. This fusion not only enhances the effectiveness of the channel dimension, but also captures multi-scale spatial details through multi-scale spatial attention, further improving the model's global perception and local detail processing capabilities. This method enables the model to focus on key areas more effectively when processing complex scenes, while suppressing irrelevant or noise information, thereby improving the performance of perception tasks.
[0055] The significance of the present invention is to enhance the collaborative perception performance of multi-agent systems, thereby enhancing the autonomous decision-making ability of agents in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a system structure diagram of the present invention;
[0057] Figure 2It is a channel attention structure diagram of the present invention;
[0058] Figure 3 is a diagram of the spatial attention structure of the present invention;
[0059] Figure 4 Schematic diagram of the dual attention weighted mechanism of the present invention. DETAILED DESCRIPTION
[0060] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0061] Specific implementation method 1: The specific process of a multi-agent collaborative perception method based on dual attention described in this implementation method is as follows:
[0062] Step 1: Use convolutional neural network to extract point cloud features;
[0063] The steps are as Figure 1 The PointPillar network is used to extract point cloud features from the input point cloud data; the specific process is:
[0064] Each agent uses a lidar or depth camera to capture point cloud data P of the environment;
[0065] The point cloud data P is used as the input of the PointPillar network to obtain the corresponding point cloud features f.
[0066] The point cloud feature uses a sparse graph generation module to generate a point cloud sparse graph M and obtain an information request graph R = 1-M.
[0067] Step 2: The point cloud features are processed by the sparse graph generation module to obtain a sparse graph;
[0068] The request graph R is transmitted to other agents in the form of broadcast communication to build a communication group and transmit point cloud features representing environmental perception to each other. The specific process is as follows:
[0069] Taking our own agent as an example, we will request information R ego Transmitted to neighboring agent j in the form of broadcast;
[0070] The neighboring agent will be the point cloud sparse graph M of the main body j Figure R with received information request ego Match S(M j ⊙R ego );
[0071] According to the matching results, determine the communication group centered on the self-agent ego;
[0072] The neighboring agent j in the communication group transmits the point cloud features to its own agent f j→ego .
[0073] Step 3: The agent transmits the request graph to the neighboring agents in the form of broadcast, and the neighboring agents determine whether to transmit the features to the agent through the feature selection module;
[0074] Step 4: The agent receives the features transmitted by the neighboring agents and performs feature fusion through the weighted dual-channel attention module. The specific process is as follows:
[0075] The point cloud features f of the neighboring agents that receive complementary information from their own agents j→ego ;
[0076] The multi-head attention mechanism is used to fuse the features acquired by the entity with the received features to obtain the fused feature F ego =ATTN(fj → ego,f ego ).
[0077] Step 5: The fusion features are passed through the detection module to obtain perception information and obtain reliable detection results.
[0078] The fusion features in step 5 are passed through a weighted dual attention module, including a channel attention module and a spatial attention module; the specific process is as follows:
[0079] By reshaping the fused features, performing matrix multiplication, and applying softmax normalization, it is straightforward to generate channel-specific attention maps that capture the interactions between channels:
[0080]
[0081] where a ji Represents the influence of the i-th channel on the j-th channel, capturing the dependency between channels;
[0082] Through matrix operations, reshaping and scaling parameters, it is applied to the feature map to generate the final output C containing the dependencies between channels j :
[0083]
[0084] Among them C j is the weighted sum of all channel features, and γ is a learnable parameter;
[0085] Weight the channel features and C jMax pooling and average pooling are applied along the channel dimension of the input feature map, producing two feature maps:
[0086] X = concat(f m ,f a );
[0087] Apply three convolution kernels with different expansion rates to process the feature map X and generate the feature map f r=1 ,f r=2 ,f r=4 ;
[0088] At the same time, a local convolutional layer is introduced to generate local features f local ;
[0089] The feature maps with different expansion rates are weighted fused with the local convolution feature map to obtain the spatial attention map:
[0090] S = σ(concat(f r=1 ,f r=2 ,f r=4 ,f local ))
[0091] Where σ represents the Sigmoid function;
[0092] The above two attention maps are weighted fused to obtain the final fusion feature F′:
[0093] F′=σ(ω c )·C+σ(ω s )·S
[0094] where ω c and ω s is a learnable weight parameter.
[0095] Specific implementation method 2: This implementation method is different from the specific implementation method 1 in that the loss function is:
[0096] L=η cls L Cls +η reg L Reg
[0097] where η cls , η reg is a hyperparameter, L Cls , L Reg Loss function for the detection module:
[0098] Based on the same inventive concept, the present invention also provides a multi-agent collaborative perception system based on dual attention, which implements the method described in the first specific implementation mode, including:
[0099] Feature extraction module, which uses convolutional neural network to extract point cloud features;
[0100] A sparse graph generation module, which generates a feature sparse graph using the data features;
[0101] A selection and exchange module that packages a request graph, builds a communication group, and transmits the information packet between agents through the communication group;
[0102] The feature fusion module fuses the received information with the local features of each agent;
[0103] The detection module decodes the fused features to obtain the detection results.
[0104] The specific implementation techniques of each module / unit in the above example of the present invention may refer to the corresponding steps of the multi-agent collaborative perception method based on dual attention weighting in the above embodiment, which will not be repeated here.
[0105] Specific implementation method three: This implementation method is different from specific implementation method one or two in that, based on the same inventive concept, a terminal is provided in other embodiments of the present invention, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it can be used to execute the method described in specific implementation method one, or to run the system described in specific implementation method two.
[0106] Specific implementation method four: This implementation method is different from specific implementation methods one, two or three in that, based on the same inventive concept, a computer-readable storage medium is provided in other embodiments of the present invention, on which a computer program is stored, which, when executed by a processor, can be used for the method described in specific implementation method one, or to execute the system described in specific implementation method two.
[0107] It should be noted that the steps in the method provided by the present invention can be implemented by using corresponding modules, devices, units, etc. in the system. Those skilled in the art can refer to the technical solution of the system to implement the step flow of the method, that is, the embodiments in the system can be understood as preferred examples for implementing the method, which will not be elaborated here.
[0108] Those skilled in the art know that, in addition to implementing the system and its various devices provided by the present invention in a purely computer-readable program code, the system and its various devices provided by the present invention can be made to implement the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices provided by the present invention can be considered as a hardware component, and the devices for implementing various functions included therein can also be considered as structures within the hardware component; the devices for implementing various functions can also be considered as both software modules for implementing the method and structures within the hardware component.
[0109] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various modifications or variations within the scope of the claims, which does not affect the essence of the present invention. The above preferred features can be used in any combination without conflicting with each other.
[0110] Table 1 Symbols
[0111]
[0112]
[0113] The above is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technician familiar with this profession can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent replacement and improvement made to the above embodiments without departing from the content of the technical solution of the present invention, based on the technical essence of the present invention, within the spirit and principles of the present invention, still fall within the protection scope of the technical solution of the present invention.
Claims
1. A multi-agent collaborative perception method based on dual attention, characterized in that: The method comprises: Step 1: Extract point cloud features using convolutional neural network; Step 2: The point cloud features are passed through a sparse graph generation module to obtain a sparse graph; Step 3: The agent transmits the request graph to the neighboring agent in the form of broadcast, and the neighboring agent determines whether to transmit the feature to the agent through the feature selection module; Step 4: The agent receives the features transmitted by the neighboring agents and performs feature fusion through the weighted dual-channel attention module; Step 5: The fusion features are passed through the detection module to obtain perception information and obtain reliable detection results.
2. According to the multi-agent collaborative perception method based on dual attention in claim 1, it is characterized in that: In step 1, the PointPillar network is used to extract point cloud features from the input point cloud data; The specific process is: Step 1.1 Each agent uses a lidar or depth camera to capture point cloud data P of the environment; Step 1.2 uses the point cloud data P in step 1.1 as the input of the PointPillar network to obtain the corresponding point cloud features f.
3. According to the multi-agent collaborative perception method based on dual attention in claim 2, it is characterized in that: The point cloud features in step 1.2 use a sparse graph generation module to generate a point cloud sparse graph M, and obtain an information request graph R=1-M.
4. According to the multi-agent collaborative perception method based on dual attention in claim 1, it is characterized in that: In step 2, the request graph R is transmitted to other agents in the form of broadcast communication to build a communication group and transmit point cloud features representing environmental perception to each other; the specific process is as follows: Step 2.1 Taking the self-agent ego as an example, the information request graph R ego Transmitted to neighboring agent j in the form of broadcast; Step 2.2 The neighboring agent j takes the point cloud sparse graph M of the entity j Figure R with received information request ego Match S(M j ⊙R ego ); Step 2.3: Determine the communication group centered on the self-agent ego based on the matching results; Step 2.4 The neighboring agent j in the communication group transmits the point cloud features to its own agent.
5. According to the multi-agent collaborative perception method based on dual attention in claim 1, it is characterized in that: The features in step 4 are fused, and the specific process includes: Step 4.1 The self-agent ego receives the point cloud features f of the neighboring agents with complementary information j→ego ; Step 4.2 Use the self-attention mechanism to concatenate the features acquired by the entity with the received features to obtain the fusion feature F ego =ATTN(fj → ego,f ego ).
6. A multi-agent collaborative perception method based on dual attention according to claim 5, characterized in that: The fused features in step 5 are passed through a weighted dual attention module; The specific process is: Step 5.1 directly generates a channel-specific attention map that captures the interactions between channels by reshaping the fused features, performing matrix multiplication, and applying softmax normalization: where a ji Represents the influence of the i-th channel on the j-th channel, capturing the dependency between channels; Step 5.2 applies it to the feature map through matrix operations, reshaping and scaling parameters to generate the final output C containing the dependencies between channels. j : Among them C j is the weighted sum of all channel features, and γ is a learnable parameter; Step 5.3 Weight the channel features and C j Max pooling and average pooling are applied along the channel dimension of the input feature map, producing two feature maps: X=concat(f m ,f a ); Step 5.4 applies three convolution kernels with different dilation rates to process the feature map X and generate the feature map f r=1 ,f r=2 ,f r=4 ; Step 5.5 introduces a local convolutional layer to generate local features f local : Step 5.6 performs weighted fusion of feature maps with different expansion rates and local convolution feature maps to obtain a spatial attention map: S=σ(concat(f r=1 ,f r=2 ,f r=4 ,f local )) Where σ represents the Sigmoid function; Step 5.7 performs weighted fusion on the above two attention maps to obtain the final fusion feature F′: F′=σ(ω c )·C+σ(ω s )·S where ω c and ω s is a learnable weight parameter.
7. The multi-agent collaborative perception method based on dual attention according to claim 1 is characterized in that: The loss function is: L=η cls L Cls +n reg L Reg where η cls , η reg is a hyperparameter, L Cls , L Reg is the loss function of the detection module.
8. A multi-agent collaborative perception system based on dual attention, which implements the method described in any one of claims 1 to 7, characterized in that: The system comprises: Feature extraction module, which uses convolutional neural network to extract point cloud features; A sparse graph generation module, which generates a feature sparse graph using the data features; A selection and exchange module that packages a request graph, builds a communication group, and transmits the information packet between agents through the communication group; The feature fusion module fuses the received information with the local features of each agent; The detection module decodes the fused features to obtain the detection results.
9. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it can be used to execute the method described in any one of claims 1 to 7, or run the system described in claim 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it can be used to execute the method described in any one of claims 1 to 7, or execute the system described in claim 8.
Citation Information
Cited By
Multi-agent cooperative sensing method and system for Internet of Vehicles
CN121837878A
An internet of vehicles multi-agent cooperative perception method and system
CN121837878B