Multi-agent collaborative sensing system and method

Through the coordinated work of positioning, environment perception, feature compression and fusion perception modules, the transmission bandwidth and feature coherence problems in multi-agent collaborative perception are solved, and efficient and accurate global perception results are achieved.

CN120543809APending Publication Date: 2025-08-26HONG KONG UNIV OF SCI & TECH (GUANGZHOU)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510592473.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Existing multiagent collaborative perception methods cannot maintain high accuracy and feature coherence while reducing transmission bandwidth requirements.

Method used

The positioning module is used to obtain positioning information, the environment perception module extracts a bird's-eye feature map, the feature compression module compresses data in the spatial dimension and channel dimension, the communication module transmits the compressed feature map, and the fused perception module is used to fusion attention to obtain global perception results.

Benefits of technology

Through data compression, the transmission bandwidth requirements are reduced, the data transmission efficiency is improved, and the accuracy and consistency of global perceptual results are ensured, which enhances the perception and decision-making capabilities of multi-agent systems in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543809A_ABST
    Figure CN120543809A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent collaborative sensing system and method, and belongs to the field of collaborative sensing. The system comprises at least two intelligent agents. Randomly selecting an agent as a main agent; the positioning module is used for transmitting the obtained positioning information to the communication module and the environment sensing module; the environment sensing module obtains a bird's-eye view feature map according to the obtained sensing data; wherein the main intelligent body transmits current positioning information to the cooperative intelligent body through the communication module, so that the cooperative intelligent body obtains a bird's-eye view feature map at a corresponding view angle according to the obtained bird's-eye view feature map; the feature compression module is used for performing spatial dimension and channel dimension compression according to the aerial view feature map and a preset spatial block to obtain a compressed feature map; the compressed feature map of the cooperative agent is transmitted to the main agent through the corresponding communication module; and the fusion sensing module is used for performing fusion according to all the received compressed feature maps to obtain a global sensing result. By implementing the method and the device, the intelligent agent sensing result can be kept at high accuracy at low bandwidth.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of collaborative perception technology, and in particular to a multi-agent collaborative perception system and method. Background Art

[0002] Environmental perception, key to the interaction between mobile agents and the external environment, aims to efficiently and accurately understand the external environment and provide reliable security for subsequent decision-making and operations. With the promotion of concepts such as swarm intelligence and social intelligence, and considering the visual blind spots caused by occlusion and limited perception distance of single-agent perception, multi-agent collaborative perception has rapidly developed.

[0003] Existing multi-agent collaborative perception methods can be broadly categorized as pre-fusion and post-fusion. Pre-fusion methods provide the most complete environmental information by sharing raw observation data, but this also introduces the highest transmission bandwidth requirements. Post-fusion methods only share perception results between different agents, such as an object's position, posture, and boundaries. This reduces transmission bandwidth requirements, but the fused perception results are highly dependent on the perception results of individual agents. Furthermore, some unit-point-based compression methods are sensitive to positioning errors and can easily lead to feature incoherence between agents.

[0004] Therefore, existing technologies cannot reduce transmission bandwidth requirements while maintaining a high level of collaborative sensing results. Summary of the Invention

[0005] The embodiments of the present invention provide a multi-agent collaborative perception system and method, which can effectively solve the problem that the existing technology cannot reduce the transmission bandwidth requirement while maintaining a high level of collaborative perception results.

[0006] An embodiment of the present invention provides a multi-agent collaborative perception system, comprising: at least two agents; each agent comprises: a positioning module, an environment perception module, a feature compression module, a communication module, and a fusion perception module; one agent is randomly selected as a master agent; and the remaining agents except the master agent are used as collaborative agents;

[0007] The positioning module is used to obtain positioning information and transmit the positioning information to the communication module and the environment perception module;

[0008] The environmental perception module is used to obtain perception data, perform feature extraction based on the perception data to obtain a bird's-eye view feature map, and transmit the bird's-eye view feature map to the feature compression module; wherein the master agent transmits current positioning information to the collaborative agent through the communication module, so that the collaborative agent performs coordinate system conversion and feature extraction based on the current positioning information of the master agent, its own corresponding positioning information and the perception data, and obtains the corresponding bird's-eye view feature map from the perspective of the master agent;

[0009] The feature compression module is configured to perform data compression in the spatial dimension and the channel dimension based on the bird's-eye view feature map and the preset spatial block to obtain a compressed feature map, and transmit the compressed feature map to the communication module; wherein the collaborative intelligent agent performs data compression in the spatial dimension and the channel dimension based on its corresponding bird's-eye view feature map and the preset spatial block to obtain a corresponding compressed feature map;

[0010] The communication module transmits the compressed feature map to the fusion perception module; wherein the compressed feature map of the collaborative agent is transmitted to the communication module of the master agent through the corresponding communication module;

[0011] The fusion perception module is used to perform attention fusion based on all compressed feature maps to obtain a global perception result; and transmit the global perception result to the communication module, so that the communication module transmits the global perception result to the collaborative intelligent agent.

[0012] Furthermore, the perception data includes radar point cloud data; the collaborative agent performs coordinate system conversion and feature extraction based on the current positioning information of the master agent, its own corresponding positioning information, and the perception data to obtain a bird's-eye view feature map from the perspective of the corresponding master agent, including:

[0013] Calculate the transformation matrix based on the current positioning information of the main agent and its corresponding positioning information;

[0014] Decomposing the transformation matrix into a rotation matrix and a translation matrix;

[0015] According to the rotation matrix and the translation matrix, the radar point cloud data is projected into the intelligent body coordinate system for coordinate system conversion to obtain target radar point cloud data;

[0016] The target radar point cloud data is encoded and processed and features are extracted to obtain a bird's-eye view feature map from the perspective of the corresponding main intelligent agent.

[0017] Furthermore, determining the preset space block includes:

[0018] Divide the global bird's-eye view space corresponding to each agent's bird's-eye view feature map into several spatial blocks;

[0019] Divide the bird's-eye view feature map corresponding to each intelligent agent according to a number of spatial blocks to obtain a divided bird's-eye view feature map;

[0020] Reconstruct and flatten the divided bird's-eye view feature map to obtain a block feature sequence;

[0021] Embed the position information of several spatial blocks into the block feature sequence to obtain token information;

[0022] The token information is input into a preset encoder and encoded into a high-dimensional feature space to obtain a preset space block for characterizing the high-dimensional information features.

[0023] Furthermore, the preset encoder includes several multi-head attention networks and convolutional neural networks; the preset encoder is used to encode the token information into a high-dimensional feature space to obtain a preset space block for representing the high-dimensional information features, including:

[0024] Map the token information to obtain a query vector, a key vector, and a value vector;

[0025] Performing a nonlinear affine transformation in the multi-head attention network according to the query vector, the key vector, the value vector, and a preset scaling factor, obtaining a sub-attention feature map of the multi-head attention network;

[0026] The sub-attention feature maps are integrated according to the convolutional neural network to obtain preset spatial blocks for characterizing high-dimensional information features.

[0027] Furthermore, the feature compression module is used to perform data compression in the spatial dimension and the channel dimension according to the bird's-eye view feature map and the preset spatial block to obtain a compressed feature map, including:

[0028] Quantify spatial complementarity and perceptual importance according to a preset spatial block and a preset activation function to obtain a quantized value;

[0029] Reordering the quantized values ​​and taking the quantized values ​​greater than a preset threshold as the target space block of the master agent;

[0030] Compressing the bird's-eye view feature map in spatial dimension according to the index corresponding to the target spatial block to obtain a first feature map after spatial compression;

[0031] The first feature map is compressed in the channel dimension according to a preset compression rate to obtain a compressed feature map.

[0032] Furthermore, the fusion perception module is used to perform attention fusion based on all compressed feature maps to obtain a global perception result, including:

[0033] Perform global maximum pooling based on the compressed feature map to obtain a preliminary fused global feature map;

[0034] According to the preset perception window, the corresponding eigenvalue size of the preliminary fused global feature map is adjusted based on the sliding window attention mechanism to make the preliminary fused global feature map spatially coherent and obtain the final global feature map;

[0035] Environmental perception is performed based on the final global feature map and the preset environmental perception task head to obtain the global perception result.

[0036] Furthermore, the feature compression module of the collaborative intelligent agent is used to compress data in the spatial dimension and channel dimension based on its own corresponding bird's-eye view feature map and the preset spatial block, obtain the corresponding compressed feature map, and transmit each compressed feature map to the communication module of the main intelligent agent through its own communication module.

[0037] Furthermore, the environmental perception module obtains perception data through environmental perception equipment; the environmental perception equipment includes: a lidar sensor, a surround-view camera, and a millimeter-wave radar sensor.

[0038] Furthermore, each agent can move.

[0039] As an improvement to the above solution, another embodiment of the present invention provides a multi-agent collaborative perception method applicable to a master agent of a multi-agent collaborative perception system; the multi-agent collaborative perception system includes at least two agents; one agent is randomly selected as the master agent; the remaining agents except the master agent are used as collaborative agents; each agent includes: a positioning module, an environment perception module, a feature compression module, a communication module, and a fusion perception module;

[0040] The multi-agent collaborative perception method comprises:

[0041] Acquire positioning information through the positioning module, and transmit the positioning information to the communication module and the environment perception module;

[0042] Acquire perception data through the environmental perception module, perform feature extraction based on the perception data to obtain a bird's-eye view feature map, and transmit the bird's-eye view feature map to the feature compression module; wherein, the current positioning information of the master agent is transmitted to the collaborative agent through the communication module, so that the collaborative agent performs coordinate system conversion and feature extraction based on the current positioning information of the master agent, its own corresponding positioning information, and the perception data, to obtain the corresponding bird's-eye view feature map from the perspective of the master agent;

[0043] The feature compression module performs data compression in the spatial dimension and the channel dimension based on the bird's-eye view feature map and the preset spatial block to obtain a compressed feature map, and transmits the compressed feature map to the communication module; wherein the collaborative intelligent agent performs data compression in the spatial dimension and the channel dimension based on its corresponding bird's-eye view feature map and the preset spatial block to obtain a corresponding compressed feature map;

[0044] The feature map compressed by the collaborative agent is transmitted to the communication module of the master agent through the corresponding communication module;

[0045] Through the fusion perception module, attention fusion is performed based on all compressed feature maps to obtain a global perception result; and the global perception result is transmitted to the communication module, so that the communication module transmits the global perception result to its collaborative intelligent agent.

[0046] By implementing the present invention, at least the following beneficial effects are achieved:

[0047] The present invention provides a multi-agent collaborative perception system and method, the system includes: at least two agents; each agent includes: a positioning module, an environmental perception module, a feature compression module, a communication module and a fusion perception module; an agent is randomly selected as the main agent; the remaining agents except the main agent are used as collaborative agents; the positioning module is used to obtain positioning information and transmit the positioning information to the communication module and the environmental perception module; the environmental perception module is used to obtain perception data, perform feature extraction based on the perception data to obtain a bird's-eye view feature map, and transmit the bird's-eye view feature map to the feature compression module; wherein the main agent transmits the current positioning information to the collaborative agent through the communication module, so that the collaborative agent performs feature extraction based on the current positioning information of the main agent, its own corresponding positioning information and perception data, and obtains the corresponding main agent. A bird's-eye view feature map under the perspective; the feature compression module is used to perform data compression of the spatial dimension and channel dimension according to the bird's-eye view feature map and the preset spatial block to obtain the compressed feature map, and transmit the compressed feature map to the communication module; wherein the collaborative intelligent agent performs data compression of the spatial dimension and channel dimension according to its own corresponding bird's-eye view feature map and the preset spatial block to obtain the corresponding compressed feature map; the communication module is used to transmit the compressed feature map to the fusion perception module; wherein the compressed feature map of the collaborative intelligent agent is transmitted to the communication module of the main intelligent agent through the corresponding communication module; the fusion perception module is used to perform attention fusion according to all compressed feature maps to obtain a global perception result; and transmit the global perception result to the communication module, so that the communication module transmits the global perception result to its collaborative intelligent agent.

[0048] The feature compression module compresses the bird's-eye view feature map in both spatial and channel dimensions. The communication module transmits only the compressed feature map, not the original perception data or uncompressed bird's-eye view feature map. Since the compressed feature map data size is significantly reduced, the bandwidth required for data transmission between agents is also reduced. This reduces the amount of data transmitted, thereby reducing bandwidth requirements and improving the system's data transmission efficiency. By sharing positioning information, using a unified feature extraction and compression method, and employing an attention fusion mechanism, each agent can extract features in a unified coordinate system. This ensures spatial consistency and coherence between features extracted by different agents, resulting in a more accurate and reliable global perception result. This helps improve the perception and decision-making capabilities of the multi-agent collaborative perception system in complex environments. The fused perception module performs attention fusion on all compressed feature maps to generate a global perception result. This organically integrates the features of different agents, not only fully utilizing the local perception information of each agent but also making the global perception result more coherent and accurate, avoiding feature conflicts and incoherence. Therefore, the characteristics of each agent are made coherent while reducing the transmission bandwidth requirement, and accurate global perception results are obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 This is a schematic diagram of the structure of a multi-agent collaborative perception system provided by one embodiment of the present invention;

[0050] Figure 2 This is a diagram showing the structure of a multi-agent collaborative perception neural network based on attention selection provided by one embodiment of the present invention;

[0051] Figure 3 This is a structural framework diagram of an attention selection network algorithm module provided by an embodiment of the present invention;

[0052] Figure 4 This is a structural diagram of a Transformer encoder provided by one embodiment of the present invention;

[0053] Figure 5 It is a flow chart of a multi-agent collaborative perception method provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0055] See also Figure 1 To address the problem in existing technologies that it is impossible to reduce transmission bandwidth requirements while maintaining high accuracy and spatial feature consistency in multi-agent collaborative perception, an embodiment of the present invention provides a schematic structural diagram of a multi-agent collaborative perception system, comprising: at least two agents; each agent includes: a positioning module, an environment perception module, a feature compression module, a communication module, and a fusion perception module; one agent is randomly selected as the master agent; and the remaining agents except the master agent are used as collaborative agents;

[0056] The positioning module is used to obtain positioning information and transmit the positioning information to the communication module and the environment perception module;

[0057] The environmental perception module is used to obtain perception data, perform feature extraction based on the perception data to obtain a bird's-eye view feature map, and transmit the bird's-eye view feature map to the feature compression module; wherein the master agent transmits current positioning information to the collaborative agent through the communication module, so that the collaborative agent performs coordinate system conversion and feature extraction based on the current positioning information of the master agent, its own corresponding positioning information and the perception data, and obtains the corresponding bird's-eye view feature map from the perspective of the master agent;

[0058] The feature compression module is configured to perform data compression in the spatial dimension and the channel dimension based on the bird's-eye view feature map and the preset spatial block to obtain a compressed feature map, and transmit the compressed feature map to the communication module; wherein the collaborative intelligent agent performs data compression in the spatial dimension and the channel dimension based on its corresponding bird's-eye view feature map and the preset spatial block to obtain a corresponding compressed feature map;

[0059] The communication module is used to transmit the compressed feature map to the fusion perception module; wherein the compressed feature map of the collaborative agent is transmitted to the communication module of the master agent through the corresponding communication module;

[0060] The fusion perception module is used to perform attention fusion based on all compressed feature maps to obtain a global perception result; and transmit the global perception result to the communication module, so that the communication module transmits the global perception result to the collaborative intelligent agent.

[0061] In a preferred embodiment of the present invention, the positioning module of the main intelligent body is used to obtain the current positioning information of the main intelligent body and transmit the current positioning information to the communication module of the main intelligent body and the environmental perception module of the main intelligent body; the positioning module of the collaborative intelligent body is used to obtain the respective positioning information of the collaborative intelligent body and transmit the respective positioning information to the respective communication module and environmental perception module. The environmental perception module of the main intelligent body is used to obtain perception data, perform feature extraction based on the perception data of the main intelligent body to obtain a bird's-eye view feature map, and transmit the bird's-eye view feature map to the feature compression module; wherein, the main intelligent body transmits the current positioning information to the collaborative intelligent body through the communication module, so that the collaborative intelligent body performs coordinate system conversion and feature extraction based on the current positioning information of the main intelligent body, its own corresponding positioning information and perception data, and obtains the corresponding bird's-eye view feature map from the perspective of the main intelligent body. The feature compression module of the master agent is used to perform data compression in the spatial and channel dimensions based on the bird's-eye view feature map and the preset spatial block of the master agent to obtain a compressed feature map, and transmit the compressed feature map to the communication module of the master agent; wherein the collaborative agent performs data compression in the spatial and channel dimensions based on its corresponding bird's-eye view feature map and the preset spatial block to obtain the corresponding compressed feature map. The communication module of the master agent is used to transmit the compressed feature map to the fusion perception module; wherein the compressed feature map of the collaborative agent is transmitted to the communication module of the master agent via the corresponding communication module; and before feature compression, the communication module of the master agent is also used to transmit the current positioning information to the communication module of the collaborative agent. The fusion perception module of the master agent is used to perform attention fusion based on all compressed feature maps to obtain a global perception result; and transmit the global perception result to the communication module, so that the communication module transmits the global perception result to the collaborative agent.

[0062] Specifically, the positioning module is used to obtain positioning information and transmit the positioning information to the communication module and the environment perception module.

[0063] Each agent is equipped with motion sensors such as the Global Navigation Satellite System (GNSS) and an inertial measurement unit (IMU). The positioning module uses the sensor data from each agent to determine its position and outputs an estimate of its self-positioning result, known as positioning information. Each agent is mobile.

[0064] In a preferred embodiment of the present invention, Figure 2As shown in the figure, two intelligent connected cars and one connected road test unit are used as working agents, and the laser radar data input and 3D target detection tasks in environmental perception are used as examples. Each movable agent obtains an independent positioning result of its own position through its own motion sensor data in the positioning module. Assume that the autonomous positioning result of the agent numbered a is (x a ,y a ,z a ,r a ,θ a ,p a ), where (x a ,y a ,z a ) is the spatial position information of agent a in the world coordinate system, (r a ,θ a ,p a ) is the deflection angle information of the agent a in the space under the world coordinate system. The two together constitute the posture information of the agent, that is, the positioning information.

[0065] Similarly, the autonomous positioning result of the agent numbered b is (x b ,y b ,z b ,r b ,θ b ,p b ), where (x b ,y b ,z b ) is the spatial position information of agent b in the world coordinate system, (r b ,θ b ,p b ) is the deflection angle information of agent b in the space under the world coordinate system. The two together constitute the pose information of the agent, that is, the positioning information. The autonomous positioning result of the agent numbered c is (x c ,y c ,z c ,r c ,θ c ,p c ), where (x c ,y c ,z c ) is the spatial position information of the agent c in the world coordinate system, (r c ,θ c ,p c) is the deflection angle information of agent c in space within the world coordinate system. Furthermore, during the collaborative process, an agent must be selected as the carrier for data fusion and perception output, referred to here as the Ego agent (master agent). This selection process can randomly select an agent or select the agent with the most computing resources. Each agent must be equipped with sufficient computing equipment to deploy a multi-agent collaborative environmental perception neural network. The master agent must also complete the multi-source perception information fusion and final environmental perception tasks.

[0066] Specifically, the environmental perception module is used to obtain perception data, perform feature extraction based on the perception data to obtain a bird's-eye view feature map, and transmit the bird's-eye view feature map to the feature compression module; wherein the main intelligent agent transmits the current positioning information to the collaborative intelligent agent through the communication module, so that the collaborative intelligent agent performs coordinate system conversion and feature extraction based on the current positioning information of the main intelligent agent, its own corresponding positioning information and perception data, and obtains the corresponding bird's-eye view feature map from the perspective of the main intelligent agent.

[0067] Preferably, the environmental perception module obtains perception data through an environmental perception device; the environmental perception device includes: a lidar sensor, a surround-view camera, and a millimeter-wave radar sensor.

[0068] Each agent is equipped with one or more environmental perception sensors, including at least a lidar. The environmental perception module obtains the corresponding perception data from each agent's own sensors. This perception data is extracted based on the pose information of other agents in the communication module and transferred to the BEV space (bird's-eye view space) of the corresponding agent's coordinate system, outputting a BEV feature map (bird's-eye view feature map) from the agent's own perspective.

[0069] Specifically, the perception data includes radar point cloud data; the collaborative agent performs coordinate system conversion and feature extraction based on the current positioning information of the master agent, its own corresponding positioning information, and the perception data to obtain a bird's-eye view feature map from the perspective of the corresponding master agent, including:

[0070] Calculate the transformation matrix based on the current positioning information of the main agent and its corresponding positioning information;

[0071] Decomposing the transformation matrix into a rotation matrix and a translation matrix;

[0072] According to the rotation matrix and the translation matrix, the radar point cloud data is projected into the intelligent body coordinate system for coordinate system conversion to obtain target radar point cloud data;

[0073] The target radar point cloud data is encoded and processed and features are extracted to obtain a bird's-eye view feature map from the perspective of the corresponding main intelligent agent.

[0074] Specifically, agents a, b, and c obtain radar point cloud data P through their own lidar sensors in the environment perception module. a ,P b ,P c , where P∈R N×4 Represents a set of radar point clouds, P represents a lidar point cloud set, R represents a real number field, and N represents the number of points in the lidar point cloud set. Each set contains N lidar point data, and each lidar point contains coordinate information and reflection intensity information.

[0075] In a preferred embodiment of the present invention, based on the posture information (positioning information) obtained in step S1, the posture information is first shared as metadata with the help of the communication module, and the amount of data transmission here is extremely small. After the collaborative agent obtains the posture information of the Ego agent b, it can calculate the transformation matrix T of the collaborative agent a from its own coordinate system to the Ego coordinate system based on its corresponding positioning information. a→b and the transformation matrix T of the collaborative agent c c→b , transformation matrix T a→b Can be decomposed into the rotation matrix R a→b and the translation matrix t a→b , transformation matrix T c→b Can be decomposed into the rotation matrix R c→b and the translation matrix t c→b Based on the rotation matrix and translation matrix, the following expression can be used to project the radar point cloud data into the Ego agent coordinate system for coordinate system transformation to obtain the target radar point cloud data for subsequent multi-agent data fusion:

[0076] P a→b / c→b =R a→b / c→b ·P a / c T +t a→b / c→b ;

[0077] Next, with the help of existing mature solutions, such as PointPillar network or VoxelNet network, the target radar point cloud data can be encoded and processed and features extracted to obtain a bird's-eye view feature map from a bird's-eye view perspective. This represents a bird's-eye view feature map of agents a, b, and c. This bird's-eye view feature map represents the local information each agent perceives about its environment. Furthermore, the process of acquiring positioning information and perception data follows parallel computation.

[0078] Specifically, the feature compression module is used to perform data compression in the spatial dimension and channel dimension according to the bird's-eye view feature map and the preset spatial block to obtain a compressed feature map, and transmit the compressed feature map to the communication module; wherein the collaborative intelligent agent performs data compression in the spatial dimension and channel dimension according to its corresponding bird's-eye view feature map and the preset spatial block to obtain the corresponding compressed feature map.

[0079] The feature compression module simultaneously compresses the BEV feature graph of each agent in both spatial and channel dimensions, significantly reducing the amount of feature data. The compressed feature graph is then shared via the communication module with the Ego agent, which performs the final calculations. The Ego agent serves as the master agent.

[0080] Preferably, the degree of complementarity of spatial information between different intelligent agents and the importance of the perception level in a specific area are first considered simultaneously, and the bird's-eye view feature map obtained by each intelligent agent is compressed from the spatial dimension, and then the feature data volume is further compressed in the channel dimension, and the compressed feature map is shared.

[0081] In a preferred embodiment of the present invention, a spatial compression idea based on the spatial block level is first adopted to divide the entire global BEV space (global bird's-eye view space) of size (H, W) into spatial blocks of specific sizes (h, w). The bird's-eye view feature map corresponding to each intelligent agent is divided according to several spatial blocks. The divided BEV feature map (divided bird's-eye view feature map) is expressed as Where (H′, W′) = (H / h, W / w). (H, W) is the size of the original BEV feature map (length and width); (h, w) is the size of the spatial block used to divide the BEV feature map; (H', W') is the size of the BEV feature map after division (the number of spatial blocks in the length and width directions respectively).

[0082] Specifically, determining the preset space block includes:

[0083] Divide the global bird's-eye view space corresponding to each agent's bird's-eye view feature map into several spatial blocks;

[0084] Divide the bird's-eye view feature map corresponding to each intelligent agent according to a number of spatial blocks to obtain a divided bird's-eye view feature map;

[0085] Reconstruct and flatten the divided bird's-eye view feature map to obtain a block feature sequence;

[0086] Embed the position information of several spatial blocks into the block feature sequence to obtain token information;

[0087] The token information is input into a preset encoder and encoded into a high-dimensional feature space to obtain a preset space block for characterizing the high-dimensional information features.

[0088] Specifically, based on the degree of complementarity of spatial information between different agents and the perceived importance of a specific spatial patch, a reordering method is used to define the specific spatial patches that need to be shared. This step is mainly implemented by the Attentive Patch Selection Module (APS) based on the attention network. For details on the spatial patch selection neural network module, refer to Figure 3 The divided BEV feature map is reconstructed based on the spatial block level and flattened into a block feature sequence Where N = H′·W′, D = C·h·w, and then the block feature sequence is encoded into a token form that can be recognized by the Transformer network through the linear network layer. In addition, the position information of the spatial block is embedded in the block feature sequence. Next, the APS module uses the Transformer encoder to encode the Token information into the high-dimensional feature space, and the same depth dimension is reduced to d to obtain the preset spatial block for characterizing the high-dimensional information features.

[0089] Specifically, the preset encoder includes several multi-head attention networks and convolutional neural networks; the preset encoder is used to encode the token information into a high-dimensional feature space to obtain a preset space block for representing the high-dimensional information features, including:

[0090] Map the token information to obtain a query vector, a key vector, and a value vector;

[0091] Performing a nonlinear affine transformation in the multi-head attention network according to the query vector, the key vector, the value vector, and a preset scaling factor, obtaining a sub-attention feature map of the multi-head attention network;

[0092] The sub-attention feature maps are integrated according to the convolutional neural network to obtain preset spatial blocks for characterizing high-dimensional information features.

[0093] In a preferred embodiment of the present invention, after the BEV feature map is divided, the spatial block selection neural network module explores how to efficiently and accurately select complementary and perceptually important spatial blocks: First, the preset VisionTransformer encoder is used to select F′ a / b / c Encode the high-dimensional information within each spatial block. The reason for choosing this encoder is that it is highly consistent with the spatial block division process. The specific structure of the Vision Transformer encoder is referenced Figure 4 In each multi-head attention network, the linear neural network layer is first used to map the token information into the query vector Q, the key vector K, and the value vector V, and the sub-attention feature map α is obtained based on the nonlinear affine transformation. The calculation process is as follows:

[0094] Where Softmax() is the activation function and δ is the preset scaling factor. After each multi-head attention network completes parallel calculations, it completes splicing along the depth dimension and uses a convolutional neural network to complete the integration of multiple sub-attention feature maps to obtain high-dimensional spatial block features, that is, the preset spatial blocks used to represent high-dimensional information features. The process is as shown in the formula: T′ a / b / c =σ(|| h∈[1,L] (α(h)·T a / b / c )), where L is the number of attention heads in the multi-head attention network, and the symbols || and σ represent the concatenation operation and the convolutional neural network respectively.

[0095] Next, with the help of the communication module, different agents share the high-dimensional spatial block features (about 34.375Kb), and quantify the spatial complementarity and perceptual importance of a specific feature block based on the L2 norm and SoftMax activation function. The expression is as follows: ij =Softmax(‖T′ i -T′ j ‖2). The quantized values ​​are reordered based on the Top k method, defining the top k spatial blocks with the highest spatial complementarity and perceptual importance within each agent as the target spatial blocks. Different spatial compression rates can be achieved by adjusting the value of k. In experimental scenarios, this embodiment further reduces data transmission by approximately 40% compared to traditional methods that use only channel compression. The specific partitioning resolution for the spatial blocks is (h, w) = (3.2m, 3.2m).

[0096] Preferably, the feature compression module is configured to perform data compression in the spatial dimension and the channel dimension according to the bird's-eye view feature map and the preset spatial block to obtain a compressed feature map, including:

[0097] Quantify spatial complementarity and perceptual importance according to a preset spatial block and a preset activation function to obtain a quantized value;

[0098] Reordering the quantized values ​​and taking the quantized values ​​greater than a preset threshold as the target space block of the master agent;

[0099] Compressing the bird's-eye view feature map in spatial dimension according to the index corresponding to the target spatial block to obtain a first feature map after spatial compression;

[0100] The first feature map is compressed in the channel dimension according to a preset compression rate to obtain a compressed feature map.

[0101] Preferably, the feature compression module of the collaborative intelligent agent is used to perform data compression in the spatial dimension and channel dimension according to its own corresponding bird's-eye view feature map and the preset spatial block, obtain the corresponding compressed feature map, and transmit the corresponding compressed feature map to the communication module of the main intelligent agent through its own communication module.

[0102] In a preferred embodiment of the present invention, based on the selected k spatial blocks, the spatial dimension compression of the BEV feature map is completed according to the index corresponding to the target spatial block, and the first feature map after spatial compression is obtained as Next, similar to the existing technology, a 1×1 convolutional neural network (CNN) is used to further compress the data transmission along the channel dimension, and the preset compression rate is set to R c , and finally the size of the compressed feature map shared in the communication module is Where C′=C / R c .

[0103] Specifically, the communication module is used to transmit the compressed feature map to the fusion perception module; wherein the compressed feature map of its collaborative intelligent body is transmitted to the communication module of the main intelligent body through the corresponding communication module.

[0104] The communication module of each agent can receive and send data to ensure that data can be transmitted between each agent.

[0105] In a preferred embodiment of the present invention, the communication module mainly includes two submodules: data sending and data receiving. The Ego agent needs to share its preliminary autonomous positioning results (x b ,y b ,z b ,r b ,θ b ,p b ) and high-dimensional spatial block feature data T′ (preset spatial block), also need to pass through the data receiving submodule to obtain the feature map F′ after spatial compression and compression of other collaborative agents by the spatial block selection neural network module APS a / c→b On the contrary, other collaborative agents need to share their compressed feature maps F′ through the data sending submodule. a / c→b , it is also necessary to obtain the autonomous positioning results shared by the Ego agent through the data receiving submodule (x b ,y b ,z b ,r b ,θ b ,p b) and high-dimensional space block feature data T′.

[0106] Specifically, the fusion perception module is used to perform attention fusion based on all compressed feature maps to obtain a global perception result; and transmit the global perception result to the communication module, so that the communication module transmits the global perception result to the collaborative intelligent agent.

[0107] When the Ego agent selected to complete the final calculation obtains the compressed feature map shared by other collaborative agents, it fuses the compressed feature maps provided by all agents based on the fusion perception module, repairs the spatial incoherence generated in the fusion, and finally outputs the fused global BEV feature map; based on the global BEV feature map, it further uses the environmental perception task head to output specific global perception results.

[0108] Specifically, the fusion perception module is used to perform attention fusion based on the compressed feature map to obtain a global perception result, including:

[0109] Performing global maximum pooling on the compressed feature map to obtain a preliminary fused global feature map;

[0110] According to the preset perception window, the corresponding eigenvalue size of the preliminary fused global feature map is adjusted based on the sliding window attention mechanism to make the preliminary fused global feature map spatially coherent and obtain the final global feature map;

[0111] Environmental perception is performed based on the final global feature map and the preset environmental perception task head to obtain the global perception result.

[0112] In a preferred embodiment of the present invention, the master agent has obtained the compressed feature graph F′ of other collaborative agents through the data receiving submodule. a / c→b , the subsequent calculation process will be completed on the Ego agent, mainly including the compressed feature graph F′ of multiple agents a→b ,F′ c→b ,F′ b Complete effective fusion and complete the corresponding perception task based on the specific perception task head.

[0113] The main agent has the complementary and important perception data features F′ provided by each agent, and the compressed perception data features a→b ,F′ c→b ,F′ b This embodiment adopts a simple but efficient Max Pooling method to retain the most obvious perceptual features in each spatial block, and finally obtains the global feature map after preliminary fusion. Compared with the Summation and Mean Pooling methods used in traditional methods, this method can, to a certain extent, avoid the problem that some objects perceived repeatedly by multiple agents occupy too much attention, while some objects perceived only once by a certain agent are ignored by the network. Global feature map after preliminary fusion There is still a certain spatial incoherence problem, which is introduced by using isolated spatial blocks as the basic unit of data feature sharing and fusion. For example, if two adjacent spatial blocks come from different agents, and one agent has a certain positioning error, it may cause significant feature changes at the boundary between the adjacent spatial blocks. In order to fix this spatial incoherence problem, the Local Attention mechanism is used in this framework. Through a preset perception window that is constantly translated, the sliding window attention mechanism is used in the preset perception window to readjust the feature value size, thereby achieving the effect of repairing spatial incoherence and obtaining the final global feature map. The perception window size used by the local attention network to complete spatial coherence restoration is (4.8m, 4.8m). Based on the final global feature map A specific environmental perception task head is used to complete the corresponding perception task and obtain a global perception result. After obtaining the global feature map, other collaborative intelligent agents can also obtain complete and comprehensive environmental perception information to understand the surrounding environment and complete downstream tasks after perception, such as path planning and motion decision-making.

[0114] In a preferred embodiment of the present invention, a 3D target detection task head is used as an embodiment. Based on the preset anchor, two convolutional networks are used to output the final classification result and regression result of each preset object frame; wherein the classification result represents the confidence that the object frame is detected as an object, and the regression result is presented as (x, y, z, w, l, h, θ), which respectively represent the (x, y, z) object frame center position, (w, h, l) object frame size, and deflection angle θ. During the training process of the multi-agent collaborative perception neural network, Composite Focal Loss is used as the loss function of the classification result, and Smooth L1 Loss is used to calculate the loss of the regression result. The loss function of the entire network is expressed as the weighted value of the two, such as the formula: L = β1·L cls +β2·L regAs shown in the figure, different weight values ​​β1 and β2 are used to balance the scales of the two losses. During the training of the multi-agent collaborative environmental perception neural network, the weight values ​​of the different components of the loss function are β1 = 1 and β2 = 2. In a training sample, the calculated loss value is backpropagated through the network, the gradient of each parameter is calculated using the chain rule, and the network parameters are updated based on the Adam optimizer to reduce the loss, ultimately obtaining a fitted model. Based on this, the multi-agent collaborative perception neural network can be used, taking the perception data and autonomous positioning data of multiple agents as input, and the Ego agent outputs the final fused global perception result. This result can be further distributed to the networked agents and networked non-agent agents in the scene through the communication module to help understand the surrounding environment.

[0115] By utilizing the data from motion sensors and environmental perception sensors from different intelligent agent perspectives and adopting a data feature compression method based on the spatial block level in data sharing, the fusion results in beyond-visual-range, accurate, and robust global environmental perception information. Compared with the traditional method of using only channel dimension compression, the data feature transmission volume is further compressed by about 60%, without affecting the original layout of the system and without additional hardware requirements.

[0116] This embodiment utilizes data from motion sensors and environmental perception sensors from different agent perspectives to obtain global environmental perception information beyond visual range, overcoming the problems of limited perception range and visual blind spots from a single agent's perspective. It also constructs a spatial compression method based on the spatial block level, further compressing the data feature transmission by approximately 60% compared to traditional methods that only use channel dimension compression, significantly alleviating bandwidth pressure on the data transmission module and, to a certain extent, promoting the practical application of multi-agent collaborative perception. Furthermore, compared to existing unit point-level compression, the local semantic information in the spatial block of this spatial compression method can reduce the system's sensitivity to positioning errors and improve system robustness. The motion sensors, environmental perception devices, and data transmission modules are all installed by the mobile networked agents in the target scenario. This embodiment, without requiring additional hardware, can achieve accurate and robust global environmental perception information beyond visual range, which has application value for multi-agent systems and meets practical requirements such as simple deployment, flexible and reliable operation, and no additional cost.

[0117] The proposed spatial block-based attention selection module can effectively select highly complementary and perceptually important spatial regions, completing feature map compression in the spatial dimension. Using only approximately 40% of the data transmission, it surpasses the accuracy of existing feature fusion methods for 3D object detection. While point-level spatial compression methods are sensitive to agent positioning errors, spatial block-level compression methods benefit from a larger perceptual domain and local semantic information, making them less sensitive to positioning errors and more robust. Multi-agent collaborative perception methods can accomplish a variety of downstream tasks based on specific needs and offer the potential for sharing perception results with connected agents that are not equipped with environmental perception sensors.

[0118] By implementing this embodiment, the feature compression module compresses the bird's-eye view feature map in both spatial and channel dimensions. The communication module is responsible for transmitting only the compressed feature map, not the original perception data or uncompressed bird's-eye view feature map. Since the data volume of the compressed feature map is significantly reduced, the bandwidth required for data transmission between agents is also reduced accordingly, reducing the amount of data transmitted, thereby reducing the transmission bandwidth requirement and improving the system's data transmission efficiency. Through the sharing of positioning information, unified feature extraction and compression methods, and attention fusion mechanisms, the sharing of positioning information enables each agent to extract features in a unified coordinate system, ensuring spatial consistency and coherence of the features extracted by different agents. This makes the final global perception result more accurate and reliable, helping to improve the perception and decision-making capabilities of the multi-agent collaborative perception system in complex environments. The fusion perception module performs attention fusion based on all compressed feature maps to obtain a global perception result, organically integrating the features of different agents. This not only fully utilizes the local perception information of each agent, but also makes the global perception result more coherent and accurate, avoiding conflicts and incoherence between features. Therefore, the characteristics of each intelligent agent are made coherent while reducing the transmission bandwidth requirement, thereby obtaining a coherent global perception result.

[0119] See also Figure 5 , is a flow chart of a multi-agent collaborative perception method provided by one embodiment of the present invention, applicable to the master agent of a multi-agent collaborative perception system; the multi-agent collaborative perception system includes at least two agents; one agent is randomly selected as the master agent; the remaining agents except the master agent are used as collaborative agents; each agent includes: a positioning module, an environment perception module, a feature compression module, a communication module, and a fusion perception module;

[0120] The multi-agent collaborative perception method comprises:

[0121] S1. Obtaining positioning information through the positioning module and transmitting the positioning information to the communication module and the environment perception module;

[0122] S2. Acquire perception data through the environmental perception module, perform feature extraction based on the perception data to obtain a bird's-eye view feature map, and transmit the bird's-eye view feature map to the feature compression module; wherein the current positioning information of the master agent is transmitted to the collaborative agent through the communication module, so that the collaborative agent performs coordinate system conversion and feature extraction based on the current positioning information of the master agent, its own corresponding positioning information, and the perception data, to obtain a bird's-eye view feature map from the perspective of the corresponding master agent;

[0123] S3. The feature compression module compresses data in the spatial and channel dimensions according to the bird's-eye view feature map and the preset spatial block to obtain a compressed feature map, and transmits the compressed feature map to the communication module; wherein the collaborative agent compresses data in the spatial and channel dimensions according to its corresponding bird's-eye view feature map and the preset spatial block to obtain a corresponding compressed feature map;

[0124] S4. Transmitting the compressed feature map to the fusion perception module through the communication module; wherein the feature map compressed by the collaborative agent is transmitted to the communication module of the master agent through the corresponding communication module;

[0125] S5. Perform attention fusion based on all compressed feature maps through the fusion perception module to obtain a global perception result; and transmit the global perception result to the communication module, so that the communication module transmits the global perception result to the collaborative intelligent agent.

[0126] The present invention provides a multi-agent collaborative perception method. A feature compression module compresses bird's-eye view feature maps in both spatial and channel dimensions. The communication module transmits only the compressed feature maps, not the original perception data or uncompressed bird's-eye view feature maps. Because the compressed feature map data volume is significantly reduced, the bandwidth required for data transmission between agents is also reduced, reducing the amount of data transmitted and thus the bandwidth requirement, thereby improving the system's data transmission efficiency. By sharing positioning information, employing a unified feature extraction and compression method, and employing an attention fusion mechanism, the shared positioning information enables each agent to extract features within a unified coordinate system, ensuring spatial consistency and coherence of the features extracted by different agents. This makes the final global perception result more accurate and reliable, helping to improve the perception and decision-making capabilities of the multi-agent collaborative perception system in complex environments. The fused perception module performs attention fusion based on all compressed feature maps to obtain a global perception result, organically integrating the features of different agents. This not only fully utilizes the local perception information of each agent, but also makes the global perception result more coherent and accurate, avoiding conflicts and incoherence between features. Therefore, the characteristics of each intelligent agent are made coherent while reducing the transmission bandwidth requirement, thereby obtaining a coherent global perception result.

[0127] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A multi-agent collaborative perception system, characterized in that: include: At least two intelligent agents; each intelligent agent includes: a positioning module, an environment perception module, a feature compression module, a communication module, and a fusion perception module; one intelligent agent is randomly selected as the main intelligent agent; the remaining intelligent agents except the main intelligent agent are used as collaborative intelligent agents; The positioning module is used to obtain positioning information and transmit the positioning information to the communication module and the environment perception module; The environmental perception module is used to obtain perception data, perform feature extraction based on the perception data to obtain a bird's-eye view feature map, and transmit the bird's-eye view feature map to the feature compression module; wherein the master agent transmits current positioning information to the collaborative agent through the communication module, so that the collaborative agent performs coordinate system conversion and feature extraction based on the current positioning information of the master agent, its own corresponding positioning information and the perception data, and obtains the corresponding bird's-eye view feature map from the perspective of the master agent; The feature compression module is configured to perform data compression in the spatial dimension and the channel dimension based on the bird's-eye view feature map and the preset spatial block to obtain a compressed feature map, and transmit the compressed feature map to the communication module; wherein the collaborative intelligent agent performs data compression in the spatial dimension and the channel dimension based on its corresponding bird's-eye view feature map and the preset spatial block to obtain a corresponding compressed feature map; The communication module is used to transmit the compressed feature map to the fusion perception module; wherein the compressed feature map of the collaborative agent is transmitted to the communication module of the master agent through the corresponding communication module; The fusion perception module is used to perform attention fusion based on all compressed feature maps to obtain a global perception result; and transmit the global perception result to the communication module, so that the communication module transmits the global perception result to the collaborative intelligent agent.

2. A multi-agent collaborative perception system according to claim 1, characterized in that: The perception data includes radar point cloud data; the collaborative agent performs coordinate system conversion and feature extraction based on the current positioning information of the master agent, its own corresponding positioning information, and the perception data to obtain a bird's-eye view feature map from the perspective of the corresponding master agent, including: Calculate the transformation matrix based on the current positioning information of the main agent and its corresponding positioning information; Decomposing the transformation matrix into a rotation matrix and a translation matrix; According to the rotation matrix and the translation matrix, the radar point cloud data is projected into the main intelligent body coordinate system for coordinate system conversion to obtain target radar point cloud data; The target radar point cloud data is encoded and processed and features are extracted to obtain a bird's-eye view feature map from the perspective of the corresponding main intelligent agent.

3. The multi-agent collaborative perception system according to claim 1, characterized in that: Determination of preset space blocks includes: Divide the global bird's-eye view space corresponding to each agent's bird's-eye view feature map into several spatial blocks; Divide the bird's-eye view feature map corresponding to each intelligent agent according to a number of spatial blocks to obtain a divided bird's-eye view feature map; Reconstruct and flatten the divided bird's-eye view feature map to obtain a block feature sequence; Embed the position information of several spatial blocks into the block feature sequence to obtain token information; The token information is input into a preset encoder and encoded into a high-dimensional feature space to obtain a preset space block for characterizing the high-dimensional information features.

4. A multi-agent collaborative perception system according to claim 3, characterized in that: The preset encoder includes several multi-head attention networks and convolutional neural networks; A preset encoder is used to encode the token information into a high-dimensional feature space to obtain a preset spatial block for characterizing the high-dimensional information features, including: Map the token information to obtain a query vector, a key vector, and a value vector; Performing a nonlinear affine transformation in the multi-head attention network according to the query vector, the key vector, the value vector, and a preset scaling factor, obtaining a sub-attention feature map of the multi-head attention network; The sub-attention feature maps are integrated according to the convolutional neural network to obtain preset spatial blocks for characterizing high-dimensional information features.

5. The multi-agent collaborative perception system according to claim 1, characterized in that: The feature compression module is used to perform data compression in spatial dimensions and channel dimensions based on the bird's-eye view feature map and the preset spatial block to obtain a compressed feature map, including: Quantify spatial complementarity and perceptual importance according to a preset spatial block and a preset activation function to obtain a quantized value; Reordering the quantized values ​​and taking the quantized values ​​greater than a preset threshold as the target space block of the master agent; Compressing the bird's-eye view feature map in spatial dimension according to the index corresponding to the target spatial block to obtain a first feature map after spatial compression; The first feature map is compressed in the channel dimension according to a preset compression rate to obtain a compressed feature map.

6. The multi-agent collaborative perception system according to claim 1, characterized in that: The fusion perception module is used to perform attention fusion based on all compressed feature maps to obtain a global perception result, including: Perform global maximum pooling based on the compressed feature map to obtain a preliminary fused global feature map; According to the preset perception window, the corresponding eigenvalue size of the preliminary fused global feature map is adjusted based on the sliding window attention mechanism to make the preliminary fused global feature map spatially coherent and obtain the final global feature map; Environmental perception is performed based on the final global feature map and the preset environmental perception task head to obtain the global perception result.

7. The multi-agent collaborative perception system according to claim 1, characterized in that: The feature compression module of the collaborative intelligent agent is used to compress data in the spatial dimension and channel dimension according to its own corresponding bird's-eye view feature map and the preset spatial block, obtain the corresponding compressed feature map, and transmit each compressed feature map to the communication module of the main intelligent agent through its own communication module.

8. The multi-agent collaborative perception system according to claim 1, characterized in that: The environmental perception module obtains perception data through environmental perception equipment; the environmental perception equipment includes: a lidar sensor, a surround-view camera, and a millimeter-wave radar sensor.

9. The multi-agent collaborative perception system according to claim 1, characterized in that: Each agent can move.

10. A multi-agent collaborative perception method, characterized in that: A master agent for a multi-agent collaborative perception system; the multi-agent collaborative perception system includes at least two agents; one agent is randomly selected as the master agent; the remaining agents except the master agent are used as collaborative agents; each agent includes: a positioning module, an environment perception module, a feature compression module, a communication module, and a fusion perception module; The multi-agent collaborative perception method comprises: Acquire positioning information through the positioning module, and transmit the positioning information to the communication module and the environment perception module; Acquire perception data through the environmental perception module, perform feature extraction based on the perception data to obtain a bird's-eye view feature map, and transmit the bird's-eye view feature map to the feature compression module; wherein, the current positioning information of the master agent is transmitted to the collaborative agent through the communication module, so that the collaborative agent performs coordinate system conversion and feature extraction based on the current positioning information of the master agent, its own corresponding positioning information, and the perception data, to obtain the corresponding bird's-eye view feature map from the perspective of the master agent; The feature compression module performs data compression in the spatial dimension and the channel dimension based on the bird's-eye view feature map and the preset spatial block to obtain a compressed feature map, and transmits the compressed feature map to the communication module; wherein the collaborative intelligent agent performs data compression in the spatial dimension and the channel dimension based on its own corresponding bird's-eye view feature map and the preset spatial block to obtain a corresponding compressed feature map; The feature map compressed by the collaborative agent is transmitted to the communication module of the master agent through the corresponding communication module; Through the fusion perception module, attention fusion is performed according to all compressed feature maps to obtain a global perception result; and the global perception result is transmitted to the communication module, so that the communication module transmits the global perception result to the collaborative intelligent agent.

Citation Information

Cited By

  • A multi-intelligent machine coordination method and system based on relative pose

    CN122550689A