Multi-agent system dynamic heterogeneous data collaborative perception training method and device

Through a two-stage collaborative training method, the problem of insufficient perception ability under dynamic heterogeneous data in multi-agent systems is solved, the independent perception model development of each agent and the convenient integration of new agents are realized, and the perception performance and system stability are improved.

CN120687832APending Publication Date: 2025-09-23HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510786201.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In dynamic heterogeneous data scenarios, the existing technology has insufficient perception capabilities of multi-agent systems and it is difficult to maintain the perception capabilities of each agent unaffected. In particular, when the agent types change dynamically, the perception model cannot be further optimized.

Method used

A two-stage collaborative training method is adopted. First, the same type of intelligent agents are trained in the first stage to solidify the feature extraction and fusion methods. Then, different types are trained in the second stage. A unified perception model of heterogeneous data is realized through the forward projection module to ensure that each intelligent agent independently develops a perception model.

Benefits of technology

It achieves the maintenance of the perception capabilities of each intelligent agent in dynamic heterogeneous data scenarios and allows new intelligent agents to easily join the system, improving perception performance and stability and reducing storage overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687832A_ABST
    Figure CN120687832A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent system dynamic heterogeneous data collaborative perception training method and device, and belongs to the technical field of model training. The method comprises the following steps: classifying all agents in the multi-agent system according to types; selecting a type which is not subjected to one-stage cooperative training from all types, and performing one-stage cooperative training on all agents belonging to the currently selected type until all types complete one-stage cooperative training; selecting a type which is not subjected to two-stage cooperative training from all types, taking the currently selected type as a main type, taking other types as auxiliary types, and performing two-stage cooperative training on all agents in the multi-agent system until all types are taken as the main type to complete the two-stage cooperative training; and obtaining the trained multi-agent system. According to the method, two stages of cooperative training are adopted, and the self-sensing ability of each agent can be maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of model training technology, and in particular to a method and device for training dynamic heterogeneous data collaborative perception of a multi-agent system. Background Art

[0002] Multi-agent collaborative perception is an effective solution to address the limitations of individual perception and an important means of enhancing self-perception capabilities. Agents communicate within the system, exchanging and sharing their own data and information, achieving multi-source information fusion and generating more comprehensive and reliable high-level analysis and decision-making capabilities.

[0003] In order to balance performance and bandwidth, it is necessary to select perceptual features as the information exchanged between intelligent agents. This is usually achieved using a mid-term collaborative solution: the intelligent agent analyzes the perceptual information in the original data, extracts perceptual features, and shares the perceptual features and 6DoF posture with other intelligent agents in the system through a broadcast mechanism. The perceptual features received from other intelligent agents are first spatially transformed according to the 6DoF posture, and then feature fusion is performed. Finally, the fused features are used to achieve environmental perception and understanding.

[0004] Because different agents come from different sources and are equipped with different sensor models, perception data is often heterogeneous. Furthermore, this heterogeneity is dynamic. First, agent configurations are constantly being iterated and updated; second, various agent collaboration scenarios exist, including system memory disparity, partial collaboration, and full collaboration; and third, some agents may be added to multiple systems due to demand, or new agent types may be added to the current system due to new functions or tasks. Therefore, an effective collaborative perception training method for dynamic heterogeneous data is needed to maintain the perception capabilities of each agent. Summary of the Invention

[0005] The present invention provides a method and apparatus for training dynamic heterogeneous data collaborative perception in a multi-agent system. The technical solution is as follows:

[0006] In one aspect, a method for training dynamic heterogeneous data collaborative perception of a multi-agent system is provided, the method comprising:

[0007] Determine the feature extraction and fusion methods for each agent in the multi-agent system, and classify all agents in the multi-agent system by type; the feature extraction and fusion methods for agents of the same type are the same;

[0008] Select a type from all types that has not undergone one-stage collaborative training. For the currently selected type, conduct one-stage collaborative training on all agents of this type until all types have completed one-stage collaborative training. The one-stage collaborative training is the collaborative training between agents of the same type. After one-stage collaborative training, each agent of this type obtains its own first perception model.

[0009] Select one type from all types that has not undergone the second-stage collaborative training, use the currently selected type as the main type, and the other types as secondary types, and conduct the second-stage collaborative training for all agents in the multi-agent system until all types have completed the second-stage collaborative training as the main type; the second-stage collaborative training is collaborative training between different types; after the second-stage collaborative training is completed, each type obtains its own second perception model;

[0010] A trained multi-agent system is obtained, in which each agent in the multi-agent system uses the second perception model of its type to perceive the environment.

[0011] In another aspect, a multi-agent system dynamic heterogeneous data collaborative perception training device is provided, the device comprising:

[0012] The classification unit is used to determine the feature extraction and feature fusion methods for each agent in the multi-agent system, and classify all agents in the multi-agent system according to type; the feature extraction and feature fusion methods for agents of the same type are the same;

[0013] The first-stage training unit is used to select a type from all types that has not undergone the first-stage collaborative training. For the currently selected type, all agents belonging to this type undergo the first-stage collaborative training until all types have completed the first-stage collaborative training. The first-stage collaborative training is the collaborative training between agents within the same type. After the first-stage collaborative training, each agent in this type obtains its own first perception model.

[0014] The two-stage training unit is used to select a type that has not undergone two-stage collaborative training from all types, use the currently selected type as the main type, and other types as sub-types, and conduct two-stage collaborative training on all agents in the multi-agent system until all types complete the two-stage collaborative training as the main type; the two-stage collaborative training is collaborative training between different types; after the two-stage collaborative training is completed, each type obtains the second perception model of its own type; a trained multi-agent system is obtained, and each agent in the multi-agent system uses the second perception model of its type to perceive the environment.

[0015] On the other hand, a computer device is provided, which includes a memory and a processor, wherein the memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory to implement the steps of the above-mentioned dynamic heterogeneous data collaborative perception training method for multi-agent systems.

[0016] On the other hand, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned dynamic heterogeneous data collaborative perception training method for a multi-agent system are implemented.

[0017] On the other hand, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the above-mentioned method for training dynamic heterogeneous data collaborative perception of a multi-agent system.

[0018] The technical solution provided by the present invention can at least bring the following beneficial effects:

[0019] First, a first-stage collaborative training is used to train agents with the same feature extraction and fusion methods, solidifying the same type of feature extraction and fusion. Then, a second-stage collaborative training is used to train agents of different types, enabling each type of agent to perceive the environment using its own independent perception model without affecting the perception capabilities of other agents. This two-stage collaborative training approach maintains the inherent perception capabilities of each agent. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 This is a flow chart of a method for training dynamic heterogeneous data collaborative perception of a multi-agent system provided by one embodiment of the present invention;

[0022] Figure 2 is a schematic structural diagram of a first perception model provided by an embodiment of the present invention;

[0023] Figure 3 is a schematic structural diagram of a second perception model provided by an embodiment of the present invention;

[0024] Figure 4 1 is a schematic structural diagram of a forward projection module provided by an embodiment of the present invention;

[0025] Figure 5 This is a schematic diagram of qualitative comparative analysis provided by an embodiment of the present invention;

[0026] Figure 6 This is a structural diagram of a multi-agent system dynamic heterogeneous data collaborative perception training device provided by one embodiment of the present invention;

[0027] Figure 7 This is a hardware architecture diagram of a computer device provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0029] The following two papers provide different solutions for collaborative perception training of heterogeneous data:

[0030] Reference 1 (Hao Xiang, Runsheng Xu, and Jiaqi Ma. Hm-vit: Hetero-modal vehicle-to-vehicle cooperative perception with vision transformer. In Proceedings of the IEEE / CVF International Conference on Computer Vision, pages 284-295, 2023) discusses scenarios where agents use a single type of modal data, and different agents use different modal sensors. Building on a mid-term collaborative approach, it requires a unified feature extraction method for the same modality and proposes a heterogeneous 3D graph attention mechanism (H3GAT) for feature fusion to capture the heterogeneity of features from different sensor modalities. This mechanism encodes agent interactions from both local and global perspectives, capturing 3D ambiguity in the feature space. Local attention helps preserve feature details, while global attention allows for a better understanding of environmental context, such as road and topology.

[0031] Reference 2 (Yifan Lu, Yue Hu, Yiqi Zhong, Dequan Wang, Yanfeng Wang, and Siheng Chen. An extensible framework for open heterogeneous collaborative perception. In The Twelfth International Conference on Learning Representations, 2024) discusses application scenarios in which different agents use different single-modality sensors while allowing new heterogeneous agents to join the system. Based on the mid-term collaborative approach, it allows the use of arbitrary feature extraction methods for modal data and proposes a reverse alignment training mechanism. This mechanism first selects one modality as the base modality for perception model training, creating a unified feature space. The perception model trained for the base modality is then fixed, and one of the remaining modalities is selected each time to train the corresponding feature extraction model. This aligns the feature space with the unified feature space, solving the heterogeneity problem of different modalities through feature homogenization.

[0032] However, Reference 1 is based on an application scenario where each agent uses only a single-modality sensor and requires that the same-modality data must adopt a unified feature extraction method, resulting in insufficient perception capabilities of this method in dynamic heterogeneous data scenarios. Reference 2 is also based on an application scenario where each agent uses a single-modality sensor. Although it allows the same modality to use any feature extraction method and allows new heterogeneous agents to join the system, once training is completed, the perception model is fixed and its perception capabilities cannot be further improved and optimized, especially in the application scenario of dynamic heterogeneous data collaboration.

[0033] The specific implementation of the concept of the present invention is described below.

[0034] Please refer to Figure 1 , an embodiment of the present invention provides a multi-agent system dynamic heterogeneous data collaborative perception training method, the method comprising:

[0035] Step 100: Determine the feature extraction method and feature fusion method for each agent in the multi-agent system, and classify all agents in the multi-agent system by type; agents of the same type have the same feature extraction method and feature fusion method;

[0036] Step 102: Select a type from all types that has not undergone the first-stage collaborative training. For the currently selected type, perform the first-stage collaborative training on all agents of the type until all types have completed the first-stage collaborative training. The first-stage collaborative training is the collaborative training between agents of the same type. After the first-stage collaborative training, each agent of the type obtains its own first perception model.

[0037] Step 104: Select a type from all types that has not undergone the second-stage collaborative training, use the currently selected type as the main type, and the other types as secondary types. Perform the second-stage collaborative training on all agents in the multi-agent system until all types have completed the second-stage collaborative training as the main type. The second-stage collaborative training is collaborative training between different types. After the second-stage collaborative training is completed, each type obtains its own second perception model.

[0038] Step 106: A trained multi-agent system is obtained, in which each agent in the multi-agent system uses the second perception model of its type to perceive the environment.

[0039] In this embodiment of the present invention, a first-stage collaborative training is used to train agents with the same feature extraction and fusion methods. This solidifies the same type of feature extraction and fusion. Then, a second-stage collaborative training is used to train agents of different types. This allows each agent to use its own independent perception model for environmental perception without affecting the perception capabilities of other agents. This two-stage collaborative training approach maintains the inherent perception capabilities of each agent.

[0040] Described below Figure 1 How to perform the steps shown.

[0041] First, for step 100, the feature extraction method and feature fusion method for each agent in the multi-agent system are determined, and all agents in the multi-agent system are classified according to type.

[0042] In the embodiments of the present invention, an agent is an agent that can perceive its environment and take actions to achieve specific goals. It can be software, hardware, or a system, and possesses autonomy, adaptability, and interaction capabilities. For example, when the agent is a car, a multi-agent system is a system composed of multiple cars.

[0043] Since different agents may have different feature extraction methods or feature fusion methods, in order to effectively perform collaborative training of heterogeneous data, all agents in the multi-agent system can be classified according to type.

[0044] Among them, the feature extraction method and feature fusion method of agents of the same type are the same. In other words, as long as there is at least one difference between the feature extraction method and the feature fusion method between two agents, it means that the two agents belong to different types.

[0045] Then, for step 102, a type that has not undergone one-stage collaborative training is selected from all types, and for the currently selected type, all agents belonging to this type are subjected to one-stage collaborative training until all types have completed one-stage collaborative training.

[0046] In this embodiment of the present invention, the first phase of collaborative training involves collaborative training between agents within the same class. After classification in step 100, multiple classes are obtained, each containing at least one agent. If a class contains only one agent, the first perception model is trained directly using that single agent. If a class contains multiple agents, each agent is collaboratively trained using its own characteristics and those of other agents within the same class to obtain its own first perception model. In other words, after the first phase of collaborative training, each agent within that class obtains its own first perception model.

[0047] For one implementation, see Figure 2 , is a schematic diagram of the structure of the first perception model, the first perception model includes: a feature extraction and fusion module, a first spatial transformation module, a first feature aggregation module and a first environment perception module;

[0048] The one-stage collaborative training method includes:

[0049] For each agent of this type, execute:

[0050] S1: using the feature extraction and fusion module to obtain the fusion features obtained by the current agent according to its own feature extraction method and feature fusion method;

[0051] The agent extracts features from the raw modal data, including but not limited to images (from cameras) and point clouds (from radar / lidar). After extracting features from each modality, it fuses the features of each modality using feature fusion to generate fused features. If the raw modal data only includes data from a single modality, feature fusion is not required.

[0052] S2: Using the first spatial transformation module, the fusion features sent by other agents of the same type are spatially transformed into the coordinate space of the current agent;

[0053] After obtaining its own fused features, each agent packages the fused features and 6DoF pose into a message and broadcasts this packaged message to other agents in the system, while also receiving packaged messages from other agents. In this embodiment of the present invention, an agent processes packaged messages from other agents of the same type as itself by spatially transforming the fused features of other agents based on their own and other agents' 6DoF poses to achieve spatial feature alignment.

[0054] S3: Using the first feature aggregation module to perform feature aggregation on the fusion features after spatial transformation and the fusion features of the current agent;

[0055] S4: Utilize the first environment perception module to perform environment perception based on the aggregated features.

[0056] Taking type A as an example, type A includes agent A1, agent A2 and agent A3. A one-stage collaborative training is performed on agent A1. Then, agent A can obtain its own fusion feature a1 by using the feature extraction and fusion module, and receive the fusion feature a2 and fusion feature a3 sent by agent A2 and agent A3 respectively. In agent A1, the fusion feature a2 and fusion feature a3 are spatially transformed respectively, and the fusion feature a1 and the fusion feature a2 and fusion feature a3 after spatial transformation are feature aggregated to use the aggregated features for environmental perception.

[0057] In this way, after all types complete one stage of collaborative training, each agent obtains its own first perception model.

[0058] Next, for step 104, a type that has not undergone the second-stage collaborative training is selected from all types, the currently selected type is used as the main type, and the other types are used as secondary types, and all agents in the multi-agent system are subjected to the second-stage collaborative training until all types have completed the second-stage collaborative training as the main type.

[0059] In an embodiment of the present invention, one-stage collaborative training is collaborative training between multiple agents of the same type. Its main purpose is to fix / freeze the feature extraction and fusion modules. At this time, the fixed / frozen feature extraction and fusion modules not only combine the fusion features obtained by other agents using the same feature extraction method and feature fusion method, but also focus on feature extraction and fusion from their own perspective, providing fusion features that are more suitable for their own type for subsequent collaborative training with other types.

[0060] In the embodiment of the present invention, the two-stage collaborative training is collaborative training between different types to learn the perception process of the second perception model.

[0061] For one implementation, see Figure 3 , is a schematic diagram of the structure of the second perception model, which includes a feature extraction and fusion module after the first stage of collaborative training, a forward projection module, a second spatial transformation module, a second feature aggregation module, and a second environment perception module;

[0062] The two-stage collaborative training method includes:

[0063] For each main type, the following steps are performed: using the forward projection module to forward project the fusion features of the main type and each intelligent agent in each sub-type; using the second spatial transformation module to spatially transform the forward projection features of each intelligent agent in the sub-type into the coordinate space of the main type; using the second feature aggregation module to aggregate the forward projection features after spatial transformation with the forward projection features of each intelligent agent in the main type to obtain the aggregated features of the main type, so as to use the aggregated features of the main type for environmental perception.

[0064] The feature extraction and fusion module in the second perception model is the feature extraction and fusion module in the first perception model after the completion of the first stage of collaborative training. In the second perception model, the feature extraction and fusion module is a fixed / frozen module ( Figure 3 The background color is grayed out), and the other modules are learnable modules.

[0065] In order to aggregate the fusion features of other types of intelligent agents with the fusion features of the intelligent agents of the own type, it is necessary to introduce an additional forward projection operation for each type of intelligent agent in the second perception model of the own type to perform forward projection operation on all the fusion features.

[0066] In one implementation, the forward projection module may include a plurality of key-value pairs; the keys in the key-value pairs correspond one-to-one to all types in the multi-agent system, the keys are represented using binary strings, and the values ​​are represented using lightweight neural networks;

[0067] When using the forward projection module to forward project the fusion features of each intelligent agent in the main type and each sub-type, it includes: for each type, performing: determining the target key corresponding to the type, and determining the target value corresponding to the target key based on the key-value pair, and using the lightweight neural network of the target value to forward project the fusion features of each intelligent agent of the type.

[0068] Among them, the length of the binary string can be determined in accordance with the following principle: the upper limit of the representable decimal number (that is, the decimal number corresponding to when all bits are 1, such as 1111 represents the 16th) is not less than the number of all types in the current system. In this way, if a new agent joins the system, it is convenient to quickly obtain a new key. Each value is composed of a lightweight neural network to solidify the network parameters through the training process and realize the forward projection of the fused features. In one implementation, the parameter level of the lightweight neural network is less than 1M.

[0069] Please refer to Figure 4 , which are multiple key-value pairs included in the forward projection module. Figure 4 In [1], the key is represented by a 4-digit binary number, for example, type A is 0001, type B is 0010, type C is 0011, and so on, and the value is a lightweight neural network. Using key-value pairs, it is possible to quickly match the fused features received from different types of agents, and then use the matched lightweight neural network to perform forward projection operations on the corresponding fused features.

[0070] The forward projection module of the embodiment of the present invention can project heterogeneous data into a unified feature space, thereby effectively alleviating the heterogeneity of heterogeneous data, and the use of lightweight neural networks can ensure that when the number of new heterogeneous intelligent agents increases, the entire system can still maintain a low storage overhead growth.

[0071] During the two-stage collaborative training process, assuming type A is the primary type and types B and C are secondary types, the system receives the fused feature b from each agent in type B and the fused feature c from each agent in type C. It then forward-projects the corresponding lightweight neural network values ​​of its own fused features a, b, and c to obtain forward-projected features a', b', and c'. It then spatially transforms b' and c' to align them with a'. The three aligned forward-projected features are then aggregated to obtain an aggregated feature, which is then used for environmental perception. Similarly, when type B is the primary type, types A and C are secondary types; when type C is the primary type, types A and B are secondary types. The environmental perception process is the same as when type A is the primary type.

[0072] After the two-stage collaborative training is completed, each type obtains its own second perception model.

[0073] It should be noted that, in the embodiment of the present invention, no matter in the first stage or the second stage, the fusion features are obtained by extracting features according to each agent's own feature extraction method and fusing features according to its own feature fusion method.

[0074] In the embodiment of the present invention, intelligent agents interact with each other using features rather than raw data, which can effectively protect data privacy and model privacy.

[0075] Finally, for step 106, a trained multi-agent system is obtained, in which each agent in the multi-agent system uses the second perception model of its type to perceive the environment.

[0076] In the embodiment of the present invention, through two stages of collaborative training, each intelligent agent can independently develop its own model, freely choose the type and number of modal sensors and the feature extraction and feature fusion methods suitable for itself according to needs, without affecting the perception capabilities of other intelligent agents.

[0077] In an embodiment of the present invention, there may be a scenario in a multi-agent system where a new agent is added to a trained multi-agent system. For the trained multi-agent system, how to integrate the new agent is also a technical problem that needs to be solved urgently.

[0078] Based on this, one embodiment of the present invention can handle this scenario in the following manner:

[0079] When a new agent joins the trained multi-agent system, determining whether the new agent has a trained third perception model;

[0080] If so, the new agent is used as a new type, and the two-stage collaborative training is performed on the new type and all types in the multi-agent system as main types respectively;

[0081] If not, determine whether the target type of the new agent is included in all types of the multi-agent system. If included, use the second perception model corresponding to the target type in the trained multi-agent system as the perception model of the new agent; if not included, perform the first-stage collaborative training for the new agent, and for all types and the target type in the multi-agent system, use each type as the main type to perform the second-stage collaborative training to obtain a new system that has been trained.

[0082] Among them, the third perception model can have the same structure and training process as the first perception model.

[0083] If the new agent has a trained third perception model, there is no need to conduct the first-stage collaborative training for the new agent. Therefore, it can be directly used as a new type to enter the second-stage collaborative training. Since new types are added to the system, all types in the new system need to be trained in the second stage to ensure that each type in the new system can aggregate the fusion features of every other type.

[0084] If the new agent does not have a trained third perception model, it is first necessary to check whether the existing types in the system include the type of the new agent. If it does, the second perception model of the existing type is directly used without the need for retraining. If it does not, it indicates that the newly added agent is of a new type and requires re-training of the first and second stages.

[0085] In the embodiment of the present invention, the agents in the system can be easily integrated into a new multi-agent system, and the new agents can also be conveniently added to a multi-agent system that has already completed training.

[0086] The following experiments illustrate the effect of the dynamic heterogeneous data collaborative perception training method for a multi-agent system according to an embodiment of the present invention.

[0087] Experimental conditions:

[0088] Datasets: OPV2V dataset and DAIR-V2X dataset;

[0089] Heterogeneous agent types: Lpillar (PointPillar for lidar point cloud data), Lsec (SECOND for lidar point cloud data), Cls (LiftSplat for camera RGB image data), and Cunproj (Unprojection for camera RGB image data).

[0090] Perception tasks: object detection (D), semantic segmentation (S)

[0091] Performance indicators: AP (target detection), IoU (semantic segmentation)

[0092] Experimental methods:

[0093] Quantitative and qualitative comparative analysis of the embodiments of the present invention and the prior art.

[0094] Experimental results:

[0095] First, quantitative comparative analysis

[0096] Table 1: Quantitative comparison of fully collaborative and non-collaborative results (AP@IoU 0.7 / IoU) in a heterogeneous agent collaborative perception scenario. Systems I and II were tested on the OPV2V dataset, while Systems III, IV, and V were tested on the DAIR-V2X dataset. Agents in the primary modality combination are marked with a gray background.

[0097] The number of model parameters involved in training is in millions (M).

[0098]

[0099]

[0100] Table 2: Quantitative comparison of fully collaborative, partially collaborative, and non-collaborative performance on the OPV2V dataset (AP@IoU0.7 / IoU). The detection task (D) uses AP@IoU0.7, while the segmentation task (S) uses IoU. The background of the agent in the main modality combination is gray. The number of model parameters involved in training is in millions (M).

[0101]

[0102]

[0103] Second, qualitative comparative analysis

[0104] Please refer to Figure 5 Visual comparison of the heterogeneous agent scenario in the OPV2V dataset Lsec+Lpillar with existing technologies: HM-ViT (left), HEAL (center), and this embodiment (right). The object detection results of each technology are in red, and the standard perception results are in green.

[0105] The experimental results show that compared with existing technologies, the embodiments of the present invention achieve superior perception performance in dynamic heterogeneous data collaborative perception scenarios (including full, partial, and uncoordinated perception), mitigating the impact of dynamic heterogeneous data on the perception capabilities of individual agents. Furthermore, the embodiments of the present invention offer more stable training parameters, requiring only a minimal increase in storage space to record new agent projections as the number of agents increases.

[0106] Please refer to Figure 6 The embodiment of the present invention provides a multi-agent system dynamic heterogeneous data collaborative perception training device, the device comprising:

[0107] The classification unit 600 is used to determine the feature extraction method and feature fusion method for each agent in the multi-agent system, and classify all agents in the multi-agent system according to type; agents of the same type have the same feature extraction method and feature fusion method;

[0108] The first-stage training unit 602 is configured to select a type from all types that has not undergone the first-stage collaborative training, and perform the first-stage collaborative training on all agents of the currently selected type until all types have completed the first-stage collaborative training. The first-stage collaborative training is the collaborative training between agents of the same type. After the first-stage collaborative training, each agent of the type obtains its own first perception model.

[0109] The two-stage training unit 604 is used to select a type that has not undergone two-stage collaborative training from all types, use the currently selected type as the main type, and other types as sub-types, and perform two-stage collaborative training on all agents in the multi-agent system until all types have completed the two-stage collaborative training as the main types; the two-stage collaborative training is collaborative training between different types; after the two-stage collaborative training is completed, each type obtains the second perception model of its own type; a trained multi-agent system is obtained, and each agent in the multi-agent system uses the second perception model of its type to perceive the environment.

[0110] In one embodiment of the present invention, the first perception model includes: a feature extraction and fusion module, a first spatial transformation module, a first feature aggregation module and a first environment perception module;

[0111] The one-stage collaborative training method includes:

[0112] For each intelligent agent of this type, the following are executed: using the feature extraction and fusion module to obtain the fusion features obtained by the current intelligent agent according to its own feature extraction method and feature fusion method, and using the first spatial transformation module to spatially transform the fusion features sent by other intelligent agents of this type into the coordinate space of the current intelligent agent, using the first feature aggregation module to feature aggregate the fusion features after spatial transformation with the fusion features of the current intelligent agent, and using the first environmental perception module to perform environmental perception based on the aggregated features.

[0113] In one embodiment of the present invention, the second perception model includes a feature extraction and fusion module after the completion of the first stage of collaborative training, and further includes a forward projection module, a second spatial transformation module, a second feature aggregation module and a second environment perception module;

[0114] The two-stage collaborative training method includes:

[0115] For each main type, the following steps are performed: using the forward projection module to forward project the fusion features of the main type and each intelligent agent in each sub-type; using the second spatial transformation module to spatially transform the forward projection features of each intelligent agent in the sub-type into the coordinate space of the main type; using the second feature aggregation module to aggregate the forward projection features after spatial transformation with the forward projection features of each intelligent agent in the main type to obtain the aggregated features of the main type, so as to use the aggregated features of the main type for environmental perception.

[0116] In one embodiment of the present invention, the forward projection module includes a plurality of key-value pairs; the keys in the key-value pairs correspond one-to-one to all types in the multi-agent system, the keys are represented by binary strings, and the values ​​are represented by lightweight neural networks;

[0117] When using the forward projection module to forward project the fusion features of each intelligent agent in the main type and each sub-type, it includes: for each type, performing: determining the target key corresponding to the type, and determining the target value corresponding to the target key based on the key-value pair, and using the lightweight neural network of the target value to forward project the fusion features of each intelligent agent of the type.

[0118] In one embodiment of the present invention, the fusion features are obtained by extracting features according to each agent's own feature extraction method and fusing features according to its own feature fusion method.

[0119] In one embodiment of the present invention, the apparatus may further include: a new processing unit, specifically configured to:

[0120] When a new agent joins the trained multi-agent system, determining whether the new agent has a trained third perception model;

[0121] If so, the new agent is used as a new type, and the two-stage collaborative training is performed on the new type and all types in the multi-agent system as main types respectively;

[0122] If not, determine whether the target type of the new agent is included in all types of the multi-agent system. If included, use the second perception model corresponding to the target type in the trained multi-agent system as the perception model of the new agent; if not included, perform the first-stage collaborative training for the new agent, and for all types and the target type in the multi-agent system, use each type as the main type to perform the second-stage collaborative training to obtain a new system that has been trained.

[0123] It should be noted that the multi-agent system dynamic heterogeneous data collaborative perception training device provided in the above embodiment is only illustrated by the division of the above-mentioned functional modules. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the multi-agent system dynamic heterogeneous data collaborative perception training device provided in the above embodiment and the multi-agent system dynamic heterogeneous data collaborative perception training method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0124] The embodiment of the present application also provides a computer device, please refer to Figure 7 The computer device includes a processor and a memory, in which at least one instruction, at least one program, code set or instruction set is stored. The at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the dynamic heterogeneous data collaborative perception training method for a multi-agent system provided by the above-mentioned method embodiments.

[0125] An embodiment of the present application also provides a computer-readable storage medium, which stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the dynamic heterogeneous data collaborative perception training method for a multi-agent system provided by the above-mentioned method embodiments.

[0126] An embodiment of the present application also provides a computer program product, which includes a computer program. The processor of a computer device reads the computer program from a computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the dynamic heterogeneous data collaborative perception training method for a multi-agent system described in any of the above embodiments.

[0127] For the convenience of description, the above systems or devices are described as being divided into various modules or units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0128] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.

[0129] Finally, it should be noted that, in this document, relational terms such as first, second, third, and fourth are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.

[0130] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for training dynamic heterogeneous data collaborative perception of a multi-agent system, characterized by: The method comprises: Determine the feature extraction and fusion methods for each agent in the multi-agent system, and classify all agents in the multi-agent system by type; the feature extraction and fusion methods for agents of the same type are the same; Select a type from all types that has not undergone one-stage collaborative training. For the currently selected type, conduct one-stage collaborative training on all agents of this type until all types have completed one-stage collaborative training. The one-stage collaborative training is the collaborative training between agents of the same type. After one-stage collaborative training, each agent of this type obtains its own first perception model. Select one type from all types that has not undergone the second-stage collaborative training, use the currently selected type as the main type, and the other types as secondary types, and conduct the second-stage collaborative training for all agents in the multi-agent system until all types have completed the second-stage collaborative training as the main type; the second-stage collaborative training is collaborative training between different types; after the second-stage collaborative training is completed, each type obtains its own second perception model; A trained multi-agent system is obtained, in which each agent in the multi-agent system uses the second perception model of its type to perceive the environment.

2. The method according to claim 1, characterized in that The first perception model includes: a feature extraction and fusion module, a first spatial transformation module, a first feature aggregation module and a first environment perception module; The one-stage collaborative training method includes: For each intelligent agent of this type, the following are executed: using the feature extraction and fusion module to obtain the fusion features obtained by the current intelligent agent according to its own feature extraction method and feature fusion method, and using the first spatial transformation module to spatially transform the fusion features sent by other intelligent agents of this type into the coordinate space of the current intelligent agent, using the first feature aggregation module to feature aggregate the fusion features after spatial transformation with the fusion features of the current intelligent agent, and using the first environmental perception module to perform environmental perception based on the aggregated features.

3. The method according to claim 1, characterized in that The second perception model includes a feature extraction and fusion module after the first stage of collaborative training, and also includes a forward projection module, a second spatial transformation module, a second feature aggregation module and a second environment perception module; The two-stage collaborative training method includes: For each main type, the following steps are performed: using the forward projection module to forward project the fusion features of the main type and each intelligent agent in each sub-type; using the second spatial transformation module to spatially transform the forward projection features of each intelligent agent in the sub-type into the coordinate space of the main type; using the second feature aggregation module to aggregate the forward projection features after spatial transformation with the forward projection features of each intelligent agent in the main type to obtain the aggregated features of the main type, so as to use the aggregated features of the main type for environmental perception.

4. The method according to claim 3, characterized in that The forward projection module includes a plurality of key-value pairs; the keys in the key-value pairs correspond one-to-one to all types in the multi-agent system, the keys are represented by binary strings, and the values ​​are represented by lightweight neural networks; When using the forward projection module to forward project the fusion features of each intelligent agent in the main type and each sub-type, it includes: for each type, performing: determining the target key corresponding to the type, and determining the target value corresponding to the target key based on the key-value pair, and using the lightweight neural network of the target value to forward project the fusion features of each intelligent agent of the type.

5. The method according to any one of claims 2 to 4, characterized in that: The fusion features are all obtained by extracting features according to each agent's own feature extraction method and fusing features according to its own feature fusion method.

6. The method according to claim 1, characterized in that Also includes: When a new agent joins the trained multi-agent system, determining whether the new agent has a trained third perception model; If so, the new agent is used as a new type, and the two-stage collaborative training is performed on the new type and all types in the multi-agent system as main types respectively; If not, determine whether the target type of the new agent is included in all types of the multi-agent system. If included, use the second perception model corresponding to the target type in the trained multi-agent system as the perception model of the new agent; if not included, perform the first-stage collaborative training for the new agent, and for all types and the target type in the multi-agent system, use each type as the main type to perform the second-stage collaborative training to obtain a new system that has been trained.

7. A multi-agent system dynamic heterogeneous data collaborative perception training device, characterized by: The device comprises: The classification unit is used to determine the feature extraction and feature fusion methods for each agent in the multi-agent system, and classify all agents in the multi-agent system according to type; the feature extraction and feature fusion methods for agents of the same type are the same; The first-stage training unit is used to select a type from all types that has not undergone the first-stage collaborative training. For the currently selected type, all agents belonging to this type undergo the first-stage collaborative training until all types have completed the first-stage collaborative training. The first-stage collaborative training is the collaborative training between agents within the same type. After the first-stage collaborative training, each agent in this type obtains its own first perception model. The two-stage training unit is used to select a type that has not undergone two-stage collaborative training from all types, use the currently selected type as the main type, and other types as sub-types, and conduct two-stage collaborative training on all agents in the multi-agent system until all types complete the two-stage collaborative training as the main type; the two-stage collaborative training is collaborative training between different types; after the two-stage collaborative training is completed, each type obtains the second perception model of its own type; a trained multi-agent system is obtained, and each agent in the multi-agent system uses the second perception model of its type to perceive the environment.

8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to implement the steps of any one of the methods described in claims 1-6.

9. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, characterized in that The method comprises a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.