Multi-agent system collaborative perception method and system based on query mechanism
By transmitting query information related to target objects between agents, efficient information sharing and fusion of multi-agent systems is achieved, and the problems of information redundancy and high communication costs in the prior art are solved, and perception accuracy and real-timeness are improved.
Patent Information
- Application Number
- CN202411808909.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-05-16
AI Technical Summary
The information redundancy and communication cost between existing agents is high, making it difficult to promote in practical applications, and the system is complex and lacks real-time, which cannot meet the demand for immediate response in high-speed dynamic environments.
The multi-agent system collaborative perception method based on the query mechanism is adopted to realize efficient information sharing and fusion by transmitting query information related to the target object between the agents. The specific steps include obtaining sensor data, encoding the target object into a query through a preset perception model, generating a single agent query, and interacting and fusion with other agents through the network to form an object query diagram and a fusion query.
It significantly reduces communication costs, improves perception accuracy and system real-time performance, and can achieve instant response in a high-speed dynamic environment to meet the perception needs of multi-agent systems.
Smart Images

Figure CN120011585A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information collaboration technology, and in particular to a multi-agent system collaborative perception method and system based on a query mechanism. Background Art
[0002] With the rapid development of autonomous driving technology, the perception capabilities of single agents (such as vehicles, drones, etc.) face the problems of limited perception range and insufficient accuracy in complex traffic environments. Collaborative perception methods improve the overall perception capabilities by sharing information between multiple agents. However, existing collaborative perception methods usually require the transmission of a large amount of raw data or feature data between agents, which not only occupies a large amount of communication bandwidth, but is also difficult to implement in practical applications.
[0003] In the existing technologies, the communication bandwidth consumption is high. Many existing technologies rely on the transmission of a large amount of sensor data or feature data, especially in the scenarios of multimodal perception and multi-sensor fusion, which will lead to extremely high communication costs and are difficult to promote in practical applications with limited communication resources. The system complexity is high and the real-time performance is insufficient. Some solutions propose complex feature alignment and fusion mechanisms, such as depth alignment, feature completion, and perspective conversion. These operations greatly increase the computational complexity of the system and may affect the real-time performance of the system, and cannot meet the needs of instant response in high-speed dynamic environments. Strong dependence on the roadside or a single modality, strong dependence on the perception results of roadside equipment or a single sensor (such as LiDAR or camera), lack of information sharing and collaboration between multi-agent systems, resulting in poor perception effects in the case of incomplete coverage or missing information. Information redundancy and limited perception accuracy: Although the existing mid-term fusion reduces some redundant data, the features it transmits are still redundant, and it is unable to flexibly and efficiently process important target information in the scene, affecting the perception accuracy. At the same time, some methods rely on the perception results of a single agent and cannot fully utilize the advantages of the multi-agent system to improve the accuracy of perception.
[0004] In order to find a better balance between perception performance and communication cost, the present invention proposes a collaborative perception method based on query mechanism, which realizes efficient information sharing and fusion by transmitting query information related to the target object between intelligent agents. Summary of the invention
[0005] The present invention provides a multi-agent system collaborative perception method based on a query mechanism, which is used to solve the problems of high information redundancy and high communication cost in information interaction between existing agents.
[0006] The present invention provides a multi-agent system collaborative perception method based on a query mechanism, comprising: Get sensor data; Encoding each target object in the sensor data as a query through a preset perception model to achieve single-agent query generation; The query of a single intelligent agent interacts with other intelligent agents through the network and is integrated with the queries of other intelligent agents to realize the information query of the target object and complete the collaborative perception of the target object.
[0007] According to a multi-agent system collaborative perception method based on a query mechanism provided by the present invention, the obtaining of sensor data specifically includes: The LiDAR point cloud and camera images of multiple target objects in the surrounding environment are acquired through various types of sensors to generate initial sensor data.
[0008] According to a query mechanism-based multi-agent system collaborative perception method provided by the present invention, before encoding each target object in the sensor data as a query through a preset perception model to realize single-agent query generation, the method comprises: Processing the sensor data through a preset feature encoder; The feature encoder converts sensor data into multi-scale bird's-eye view features through a 3D backbone network and a feature pyramid network.
[0009] According to a multi-agent system collaborative perception method based on a query mechanism provided by the present invention, each target object in the sensor data is encoded as a query through a preset perception model to realize single-agent query generation, which specifically includes: Initialize a single agent as a query set, sample from multi-scale bird's-eye view features through a preset decoder to generate sampled features; The sampled features are interacted with the initial query in the query set to complete the dynamic update of the query and generate a single-agent query.
[0010] According to a multi-agent system collaborative perception method based on a query mechanism provided by the present invention, the query of a single agent is interacted with other agents through a network and merged with other agent queries to realize the information query of the target object and complete the collaborative perception of the target object, which specifically includes: Single-agent queries receive other agent queries through the network; Connect multiple agent queries for the same target object to form an object query graph; Based on the object query graph, query aggregation is performed through an attention mechanism, query information of different agents is integrated into a unified target representation, and a fusion query is generated; The target object information query is completed through fusion query, and the collaborative perception of the target object is completed.
[0011] According to a multi-agent system collaborative perception method based on a query mechanism provided by the present invention, the multiple agent queries of the same target object are connected to form an object query graph, specifically including: Obtain multiple agent queries of the target object, and calculate the position encoding of the reference points based on the query through a preset linear neural network; Summing the position code and the query to obtain an enhanced query; The enhanced queries are used to calculate the similarity of queries between different agents through the inner product. Queries with similarity exceeding the set threshold are associated with the same target object to generate an object query graph.
[0012] The present invention also provides a multi-agent system collaborative perception system based on a query mechanism, the system comprising: A data acquisition module, used to acquire sensor data; A single-agent query module, used to encode each target object in the sensor data into a query through a preset perception model to realize single-agent query generation; The multi-agent fusion module is used to interact the single-agent query with other agents through the network and fuse it with other agent queries to realize the target object information query and complete the collaborative perception of the target object.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements any of the above-described multi-agent system collaborative perception methods based on the query mechanism.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the query mechanism-based collaborative perception methods for a multi-agent system as described above.
[0015] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the query mechanism-based multi-agent system collaborative perception methods described above.
[0016] The present invention provides a multi-agent system collaborative perception method and system based on a query mechanism, which encodes the sensor data collected by each agent into queries and shares them in the multi-agent system; the query mechanism focuses on the feature information of the target object, avoids the redundancy problem of full-scene feature transmission, and significantly reduces communication costs; each agent generates a set of queries from the collected original point cloud data through a perception model, and these queries can be transmitted between agents with low communication overhead through an efficient compression and encoding process; after receiving query information from other agents, the system uses spatial query matching to match queries for the same object, constructs an object query graph through the query matching results, and further uses the attention mechanism for fusion in the target query aggregation to generate a more accurate object representation, which can integrate diverse information from different agents and improve perception accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is a flow chart of the collaborative perception method of a multi-agent system based on a query mechanism provided by the present invention.
[0019] Figure 2 It is a structural diagram of the collaborative perception method based on the query mechanism provided by the present invention.
[0020] Figure 3 It is a single-agent query generation module based on a perception model provided by the present invention.
[0021] Figure 4 It is the cross-agent query fusion module provided by the present invention.
[0022] Figure 5 It is a schematic diagram of a collaborative perception case of a multi-agent system using a query mechanism provided by the present invention.
[0023] Figure 6 It is a schematic diagram of communication cost comparison provided by the present invention.
[0024] Figure 7 Schematic diagram of posture error provided by the present invention.
[0025] Figure 8 It is a schematic diagram of module connection of a multi-agent system collaborative perception system based on a query mechanism provided by the present invention.
[0026] Fig. 9 It is a structural schematic diagram of the electronic device provided by the present invention.
[0027] Figure numerals: 110: data acquisition module; 120: single-agent query module; 130: multi-agent fusion module; 910: processor; 920: communication interface; 930: memory; 940: communication bus. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0029] Some professional names that appear in this invention are explained as follows: mAP: an indicator that comprehensively reflects the precision and recall rate of the target detection model in multiple categories. For a specific category, AP is the area under the Precision-Recall curve of the category. When there are multiple categories to be detected, mAP is the average of APs of all categories.
[0030] DETR: DEtection Transformer, applies the Transformer architecture to the field of target detection, designs the Hungarian Algorithm to achieve optimal matching, and can realize end-to-end training of the overall model. In this invention, PointDETR is a Transformer-based perception model.
[0031] Query: Query is an abstract representation for object detection in DETR. They interact with image features through learning and are ultimately used to generate object detection results. Each Query represents a potential target and converts image features into detection outputs through a decoder.
[0032] FPN: FPN (Feature Pyramid Network) solves the problem of object detection at different scales by constructing a feature pyramid. It combines the bottom-up feature extraction path (using a convolutional network to generate multi-layer feature maps) and the top-down feature fusion path (upsampling high-level features and fusing them with low-level features), enhancing the expressiveness of multi-scale features. FPN is commonly used in target detection and instance segmentation tasks, and can effectively improve the detection performance of small and large objects.
[0033] Combine the following Figure 1The present invention describes a multi-agent system collaborative perception method based on a query mechanism, including: step 100, obtaining sensor data.
[0034] Specifically, each agent collects sensor data in the environment (such as LiDAR point clouds, camera images, etc.), and then processes this data through the PointDETR perception model to encode each target object in the scene as a query. The query contains key information such as the object's category, location, shape, etc. Compared with traditional regional features or raw data transmission, the object query generation process greatly reduces information redundancy while retaining key feature information related to the object.
[0035] refer to Figure 2 In this invention, the sensor data in the multi-agent system is encoded into an object query (ObjectQuery), and query information related to the target object is transmitted between agents instead of the entire scene features. The system improves the accuracy of target detection and reduces communication costs through query matching and fusion across agents.
[0036] Step 200: Encode each target object in the sensor data as a query through a preset perception model to achieve single-agent query generation.
[0037] Specifically, the sensor data is processed through a preset feature encoder; The feature encoder converts sensor data into multi-scale bird's-eye view features through a 3D backbone network and a feature pyramid network.
[0038] Initialize a single agent as a query set, sample from multi-scale bird's-eye view features through a preset decoder to generate sampled features; The sampled features are interacted with the initial query in the query set to complete the dynamic update of the query and generate a single-agent query.
[0039] In the present invention, each agent first uses the PointPillar model as a feature encoder to process point cloud data. The point cloud data is processed by a 3D backbone network and a feature pyramid network (FPN). The feature encoder converts the point cloud data into multi-scale Bird's-Eye View (BEV) features, which are expressed as the following formula.
[0040] .
[0041] Among them, L represents the number of feature layers, C is the number of channels, and are the height and width of the l-th layer features respectively.
[0042] Each agent initializes a query set, represented as ,in is the number of queries, is the channel dimension of the query. Each query is dynamically updated through the PointDETR module, which samples from multi-scale BEV features based on the Transformer decoder and interacts the sampled features with the query. The specific structure is as follows Figure 3 shown.
[0043] In the Deformable Cross-Attention module of PointDETR, the input is the initialized target query and multi-scale BEV features, and the output is the updated target query. In this process, each query generates a corresponding reference point through a linear neural network. , which represents the center position of the target object. The Multi-scale Deformable Attention mechanism samples features from the BEV features to interact with the target query, and the coordinates of the reference point are used to generate the sampling position and attention weight.
[0044] The specific sampling process is as follows: Sampling location By reference point Add offset Calculated, the offset and attention weights All are generated through a linear network. Finally, the sampled features are obtained from the BEV features through a bilinear interpolation method.
[0045] .
[0046] Among them, k represents the index of the sampling point, l represents the index of the feature scale, and m represents the index of the attention head. Finally, we sample from the multi-scale features. points and aggregate them through the attention mechanism: in, is the normalized attention weight, ensuring that the sum of the weights of sampling points at different scales is 1.
[0047] Step 300: The single agent query interacts with other agents through the network and is integrated with other agent queries to realize the target object information query and complete the collaborative perception of the target object.
[0048] Specifically, a single agent query receives other agent queries through the network; Connect multiple agent queries for the same target object to form an object query graph; Based on the object query graph, query aggregation is performed through an attention mechanism, query information of different agents is integrated into a unified target representation, and a fusion query is generated; The target object information query is completed through fusion query, and the collaborative perception of the target object is completed.
[0049] In the present invention, each agent shares the generated object query through a network (such as V2X or 5G). After each agent (called the ego vehicle or ego agent) receives the query information of other agents, it performs cross-agent query fusion. This module contains two key processes: Spatial Query Matching (SQM): The system uses a query matching algorithm to match multiple agents' queries about the same target object to form an object query graph (Object QueryGraph), which provides a basis for subsequent fusion. Object Query Aggregation (OQA): Queries from multiple agents are aggregated through an attention mechanism to generate a more complete target representation, which is ultimately used to predict object categories, positions, and shapes.
[0050] Among them, in the multi-agent system, the queries generated by each agent are shared with other agents. Assume that after an agent (called Ego Agent) receives queries from other agents, it needs to associate different queries for the same target object through spatial query matching. If the reference point is directly used for matching, the matching may be inaccurate due to positioning errors. Therefore, spatial query matching combines the context information of the query and the position encoding to improve the matching accuracy. Its working principle is as follows: Figure 4 shown.
[0051] The position encoding of the reference point based on the query is calculated through a linear neural network, and the (position) encoding is summed with the query to obtain the enhanced query: in, represents a linear neural network, It is the position encoding function in Transformer.
[0052] The enhanced query will calculate the similarity of queries between different agents i and j through the inner product: Similarity exceeds threshold The queries will be associated with the same target object, generating a target query graph.
[0053] Object Query Aggregation After completing spatial query matching, CoopDETR aggregates all queries in the same object query graph. Aggregation is achieved through a multi-head attention mechanism. Take a query graph as an example, the query of Ego Agent is q and the queries of other agents are K and V. The aggregated query The calculation is as follows: in, Represents a mask operation. The attention weight between the masked K and q will be set to 0, ensuring that queries with low relevance to the Egoagent query will not affect the aggregation process. The final aggregated query will be used for target classification and bounding box prediction through MLP.
[0054] In a specific embodiment, referring to Figure 5 A typical case of collaborative perception of a multi-agent system using a query mechanism is shown. In this scenario, there are three agents (1 to 3) that can communicate with each other and five targets (A to E) that need to be detected. Each agent processes its own point cloud data and generates queries using a DETR-based model. Therefore, different agents will generate different queries for the same target. These queries can be connected to form an object query graph. For example, target A will contain two queries from agents 1 and 3.
[0055] Existing technologies often require the transmission of large amounts of raw sensor data or scene-level features, which consumes a lot of bandwidth. CoopDETR significantly reduces the amount of data transmission by transmitting only the query information of the target object instead of all the feature data of the scene. Figure 6 ,Experiments show that the communication cost of CoopDETR is hundreds of times lower ,than that of traditional mid-term fusion methods (e.g., reduced to 1 / 782), making it suitable ,for practical application scenarios with limited bandwidth.
[0056] Existing methods often have a lot of redundant information in feature transmission, resulting in low information utilization. CoopDETR directly shares and fuses information about the target object of interest through object query, avoiding the transmission of scene-level redundant data. Referring to Table 1, through spatial query matching and target query aggregation, the different perspective information of multiple agents on the same target can be effectively fused to form a more accurate target representation.
[0057] Table 1 .
[0058] refer to Figure 7The method of the present invention has high robustness in complex environments. In particular, by using the deformable attention mechanism, it can effectively alleviate the feature alignment problem caused by posture errors (such as positioning and orientation errors) between multiple agents, further improving the stability and perception performance of the system.
[0059] In addition, in this application, query fusion based on geometric consistency matching replaces the existing SQM mechanism, and a matching method based on geometric consistency can be used. By detecting the geometric features in the perception data of multiple agents, the observation consistency of different agents on the same object is found. Geometric consistency can be achieved through triangulation or point cloud registration. Geometric consistency matching can directly perform cross-agent fusion based on the spatial features of sensor data, which is suitable for complex scenarios.
[0060] Fusion based on visual feature alignment can replace the existing query matching when there is multimodal perception (such as camera images and LiDAR point clouds). By aligning the visual features perceived by different agents (such as camera image features, depth maps, etc.), objects are fused based on the aligned features. This solution can effectively handle multimodal perception, is suitable for multimodal fusion scenarios, and can reduce errors between multiple modalities.
[0061] Query matching based on spatiotemporal alignment introduces a spatiotemporal alignment mechanism to replace spatial query matching. By analyzing the perception data of multiple agents at different times and combining continuous observation information in the time dimension, the targets of different agents are matched. The tracking and fusion of target objects are enhanced by using historical information. The spatiotemporal alignment mechanism can use historical data for more reliable object tracking and perception, which is particularly suitable for dynamic targets.
[0062] The present invention can significantly reduce the communication bandwidth requirement by encoding sensor data into object queries for transmission, compared with traditional raw data or feature data transmission. Experiments show that the communication cost of CoopDETR is only 1 / 782 of the existing mid-term fusion method, which greatly improves the communication efficiency of the system and is suitable for scenarios with limited bandwidth.
[0063] A simplified cross-agent query matching and fusion mechanism is adopted. Through the spatial query matching (SQM) and object query aggregation (OQA) modules, the complex feature alignment and fusion process is reduced while retaining the perception accuracy, thereby improving the system's computational efficiency and real-time performance and adapting to high-speed dynamic environments.
[0064] Flexible collaboration among multiple agents is achieved through cross-agent information sharing and object-level query fusion. Object queries generated by different agents can be flexibly matched and fused in the system, making full use of the perception results of each agent, improving the accuracy and comprehensiveness of perception, and solving the problem of limited perception of a single agent.
[0065] Different from traditional region-level feature transmission, CoopDETR focuses on target-level feature transmission, avoiding redundant information of full-scene features. Through the object query mechanism, CoopDETR can more efficiently extract and transmit key information related to the target, further improving perception accuracy and transmission efficiency.
[0066] refer to Figure 8 The present invention also discloses a multi-agent system collaborative perception system based on a query mechanism, the system comprising: A data acquisition module 110, used to acquire sensor data; A single-agent query module 120, configured to encode each target object in the sensor data into a query through a preset perception model, thereby realizing single-agent query generation; The multi-agent fusion module 130 is used to interact the single-agent query with other agents through the network and fuse it with other agent queries to realize the target object information query and complete the collaborative perception of the target object.
[0067] The acquiring of sensor data specifically includes: The LiDAR point cloud and camera images of multiple target objects in the surrounding environment are acquired through various types of sensors to generate initial sensor data.
[0068] Before encoding each target object in the sensor data into a query through a preset perception model to realize single-agent query generation, the method includes: Processing the sensor data through a preset feature encoder; The feature encoder converts sensor data into multi-scale bird's-eye view features through a 3D backbone network and a feature pyramid network.
[0069] The method of encoding each target object in the sensor data into a query through a preset perception model to realize single-agent query generation specifically includes: Initialize a single agent as a query set, sample from multi-scale bird's-eye view features through a preset decoder to generate sampled features; The sampled features are interacted with the initial query in the query set to complete the dynamic update of the query and generate a single-agent query.
[0070] The query of a single agent interacts with other agents through the network and merges with other agent queries to realize the information query of the target object and complete the collaborative perception of the target object, including: Single-agent queries receive other agent queries through the network; Connect multiple agent queries for the same target object to form an object query graph; Based on the object query graph, query aggregation is performed through an attention mechanism, query information of different agents is integrated into a unified target representation, and a fusion query is generated; The target object information query is completed through fusion query, and the collaborative perception of the target object is completed.
[0071] Connect multiple agent queries for the same target object to form an object query graph, which includes: Obtain multiple agent queries of the target object, and calculate the position encoding of the reference points based on the query through a preset linear neural network; Summing the position code and the query to obtain an enhanced query; The enhanced queries are used to calculate the similarity of queries between different agents through the inner product. Queries with similarity exceeding the set threshold are associated with the same target object to generate an object query graph.
[0072] A multi-agent system collaborative perception system based on a query mechanism provided by the present invention encodes the sensor data collected by each agent into queries and shares them in the multi-agent system; the query mechanism focuses on the feature information of the target object, avoids the redundancy problem of full-scene feature transmission, and significantly reduces the communication cost; each agent generates a set of queries from the collected original point cloud data through a perception model, and these queries can be transmitted between agents with low communication overhead through an efficient compression and encoding process; after receiving query information from other agents, the system uses spatial query matching to match queries of the same object, builds an object query graph through the query matching results, and further uses the attention mechanism for fusion in the target query aggregation to generate a more accurate object representation, which can integrate diverse information from different agents and improve perception accuracy.
[0073] Fig. 9 An example of a physical structure diagram of an electronic device is shown in FIG. Fig. 9As shown, the electronic device may include: a processor 910, a communications interface 920, a memory 930 and a communication bus 940, wherein the processor 910, the communications interface 920 and the memory 930 communicate with each other through the communication bus 940. The processor 910 may call the logic instructions in the memory 930 to execute a multi-agent system collaborative perception method based on a query mechanism, the method comprising: acquiring sensor data; encoding each target object in the sensor data into a query through a preset perception model to realize single-agent query generation; interacting the single-agent query with other agents through a network and fusing it with other agent queries to realize target object information query and complete collaborative perception of the target object.
[0074] In addition, the logic instructions in the above-mentioned memory 930 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0075] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a multi-agent system collaborative perception method based on a query mechanism provided by the above methods. The method includes: acquiring sensor data; encoding each target object in the sensor data as a query through a preset perception model to achieve single-agent query generation; interacting the single-agent query with other agents through a network and fusing it with other agent queries to achieve target object information query and complete collaborative perception of the target object.
[0076] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is implemented to execute a multi-agent system collaborative perception method based on a query mechanism provided by the above-mentioned methods, the method comprising: acquiring sensor data; encoding each target object in the sensor data as a query through a preset perception model to realize single-agent query generation; interacting the single-agent query with other agents through a network and fusing it with other agent queries to realize target object information query and complete collaborative perception of the target object.
[0077] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0078] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A collaborative perception method for a multi-agent system based on a query mechanism, characterized in that: include: Get sensor data; Encoding each target object in the sensor data as a query through a preset perception model to achieve single-agent query generation; The query of a single intelligent agent interacts with other intelligent agents through the network and is integrated with the queries of other intelligent agents to realize the information query of the target object and complete the collaborative perception of the target object.
2. The multi-agent system collaborative perception method based on query mechanism according to claim 1 is characterized in that: The obtaining of sensor data specifically includes: The LiDAR point cloud and camera images of multiple target objects in the surrounding environment are acquired through various types of sensors to generate initial sensor data.
3. The multi-agent system collaborative perception method based on query mechanism according to claim 1 is characterized in that: Before encoding each target object in the sensor data into a query through a preset perception model to realize single-agent query generation, the method includes: Processing the sensor data through a preset feature encoder; The feature encoder converts sensor data into multi-scale bird's-eye view features through a 3D backbone network and a feature pyramid network.
4. The multi-agent system collaborative perception method based on query mechanism according to claim 1 is characterized in that: The method of encoding each target object in the sensor data into a query through a preset perception model to realize single-agent query generation specifically includes: Initialize a single agent as a query set, sample from multi-scale bird's-eye view features through a preset decoder to generate sampled features; The sampled features are interacted with the initial query in the query set to complete the dynamic update of the query and generate a single-agent query.
5. The multi-agent system collaborative perception method based on query mechanism according to claim 1 is characterized in that: The single agent query is interacted with other agents through the network and integrated with other agent queries to realize the target object information query and complete the collaborative perception of the target object, specifically including: Single-agent queries receive other agent queries through the network; Connect multiple agent queries for the same target object to form an object query graph; Based on the object query graph, query aggregation is performed through an attention mechanism, query information of different agents is integrated into a unified target representation, and a fusion query is generated; The target object information query is completed through fusion query, and the collaborative perception of the target object is completed.
6. The multi-agent system collaborative perception method based on query mechanism according to claim 5 is characterized in that: The step of connecting multiple agent queries for the same target object to form an object query graph specifically includes: Obtain multiple agent queries of the target object, and calculate the position encoding of the reference points based on the query through a preset linear neural network; Summing the position code and the query to obtain an enhanced query; The enhanced queries are used to calculate the similarity of queries between different agents through the inner product. Queries with similarity exceeding the set threshold are associated with the same target object to generate an object query graph.
7. A multi-agent system collaborative perception system based on query mechanism, characterized in that: The system comprises: A data acquisition module, used to acquire sensor data; A single-agent query module, used to encode each target object in the sensor data into a query through a preset perception model to realize single-agent query generation; The multi-agent fusion module is used to interact the single-agent query with other agents through the network and fuse it with other agent queries to realize the target object information query and complete the collaborative perception of the target object.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, it implements the collaborative perception method of a multi-agent system based on a query mechanism as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the collaborative perception method of a multi-agent system based on a query mechanism as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the collaborative perception method of a multi-agent system based on a query mechanism as described in any one of claims 1 to 6.