Multi-agent control method and device based on orthogonal decomposition representation
By introducing orthogonal decomposition representation and dynamic task adaptation hypernetwork in multi-agent reinforcement learning, the problem of redundancy and generalization capabilities of strategy expression in multi-task scenarios is solved, and efficient multi-agent collaborative control and cross-task adaptation are achieved.
Patent Information
- Application Number
- CN202510349999.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-13
AI Technical Summary
The existing multi-agent reinforcement learning technology has problems such as redundancy in strategy expression, insufficient generalization capabilities, and gradient interference in multi-task scenarios, making it difficult to quickly adapt to new tasks and achieve cross-task collaboration.
Using a multi-agent control method based on orthogonal decomposition representation, a decoupled task-specific representation is generated through orthogonal representation learning and multi-agent reinforcement learning algorithm, and a hyper-network architecture is adapted to the hyper-network architecture through dynamic tasks, and the strategic representation is optimized to adapt to multi-task scenarios.
It significantly improves the generalization of collaborative generalization of multi-agents, the efficiency of strategy expression and the robustness of complex tasks, solves problems such as policy conflict, low data efficiency and rigid tactical collaboration, and realizes reinforcement learning solutions with high generalization and low latency.
Smart Images

Figure CN120146144A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-agent reinforcement learning, and particularly relates to a multi-agent control method and device based on orthogonal decomposition representation. Background Art
[0002] In recent years, the technology of multi-agent reinforcement learning (MARL) has developed rapidly and achieved breakthrough applications in fields such as multi-robot systems, multi-manipulator decision-making, and autonomous driving fleet cooperation. However, in practical scenarios, multi-agent systems often need to cope with dynamic multi-task requirements (for example, a robot cluster needs to switch between tasks such as material handling and environmental exploration), which poses a severe challenge to the cross-task generalization ability of algorithms. The defects of existing technologies in multi-task expression ability have become the core bottleneck restricting their large-scale implementation.
[0003] Existing online methods rely on a fixed network structure (such as the Transformer architecture) to handle the changes in the input and output dimensions of multi-tasks, but their parameter sharing mechanism lacks explicit modeling of orthogonal features between tasks, resulting in rigid expression and policy conflicts of the algorithm. For example, the network weights are dynamically adjusted through the self-attention mechanism, but the independence of the representations corresponding to different tasks is not constrained, leading to gradient interference (Task Interference) between tasks during policy update. In addition, synchronously training a predefined task set (such as policy distillation) forces all tasks to share the same representation space, resulting in feature collapse (that is, the feature vectors of different tasks are highly similar), and the task specificity (such as the morphological differences of robots and the physical rules of the environment) cannot be effectively distinguished, limiting the generalization ability of the policy in unseen tasks.
[0004] Offline methods (such as ODIS) extract skills from static data as the primitives of task expression, but the construction of their skill libraries is limited by the coverage and optimality of the data set, and the data quality will severely restrict the representation diversity and algorithm performance. When the data does not contain cross-task common features (such as the spatio-temporal patterns of cooperation strategies), the skill representation lacks orthogonal constraints, resulting in high redundancy and poor combination flexibility between skills (such as the random factorization method). Experiments show that in a medium-quality data set, the dimensionality of the representation space of the skill library is significantly lower than the actual requirements of the tasks, causing the policy to be unable to interpolate the feature combinations suitable for new tasks.
[0005] Although the hybrid RL framework improves the representation ability through offline-online joint optimization in single-agent scenarios, its design does not consider the collaborative expression requirements of multi-agent tasks. For example, offline skill reuse is used to accelerate online training, but the conflict problem of skill combination among multi-agents is not solved (e.g., skill A is effective for agent 1, but is incompatible with skill B of agent 2). In addition, the existing hybrid buffer pool design (such as balanced replay) does not introduce a mechanism for orthogonalization of multi-task representations, resulting in interference between different task samples during joint training, and it is difficult for the policy network to converge to a representation space that takes into account both diversity and consistency.
[0006] The above problems can be attributed to the structural defects of the representation space:
[0007] 1) Low orthogonality: The similarity between task feature vectors is high, and decoupled basis vectors cannot be formed, reducing the expression ability of cross-task feature combinations;
[0008] 2) Low coverage: The dimension of the representation space is insufficient or the manifold structure (such as the Stiefel manifold) is not embedded, and the non-linear relationship between complex multi-tasks cannot be encoded;
[0009] 3) Low collaboration: When multi-agent policies are updated, there is a lack of a cross-agent representation alignment mechanism, resulting in local optimization destroying the global representation consistency.
[0010] The above defects make it difficult for existing methods to quickly generate strategies adapted to new tasks when facing multi-task groups with large differences in input-output dimensions and heterogeneous physical rules (such as changes in object shape / mass in multi-robot collaborative handling), seriously restricting the scalability of practical applications.
[0011] The existing multi-robot cooperative control technologies mainly focus on distributed reinforcement learning architectures. Typical solutions include the following four types of technical routes:
[0012] ① The multi-agent robot cooperative control scheme based on distributed reinforcement learning adopts a distributed training framework and realizes experience sharing among agents through independent policy networks. However, in multi-task scenarios, the policy gradients of each agent are coupled and interfered, resulting in conflicts in policy update directions and affecting the efficiency of multi-task collaboration.
[0013] ② The training scheme for multi-manipulator multi-target search constructs a memory sample through an LSTM network to improve the policy network. Although it can improve the training stability, the fixed architecture is difficult to dynamically adapt to changes in the task environment. Especially when the new target space distribution changes suddenly, the network parameters need to be retrained.
[0014] ③The dynamic hierarchical multi-agent control scheme based on reinforcement learning uses a hierarchical communication structure to reduce bandwidth requirements. However, the static hierarchical division cannot adapt to sudden task changes, and the fixed selection mechanism of representative nodes is prone to decision-making delays in dynamic environments.
[0015] ④The target acquisition scheme for multi-manipulator multi-tasks. Although the autonomous exploration mechanism based on the SAC algorithm can complete multi-target trajectory planning, it lacks the design of policy decoupling between tasks, and there is a risk of implicit conflicts in multi-task action decisions.
[0016] The above methods have systematic bottlenecks in realizing multi-robot collaborative control: (1) The coupling of multi-task policy gradients leads to low learning efficiency; (2) The fixed architecture is difficult to achieve cross-task dynamic adaptation; (3) There is a lack of strict consistency guarantee between global collaboration and individual policies; (4) Adapting to new tasks requires re-training network parameters. These defects are particularly significant in open dynamic environments, resulting in the difficulty of existing systems to achieve a balance among policy decoupling, zero-shot generalization, and multi-task collaboration. Summary of the Invention
[0017] To address the above technical bottlenecks, it is urgent to solve the following core problems: In an open dynamic environment, how to achieve policy decoupling and dynamic adaptation of a multi-robot system, realize cross-task generalization control through dynamic parameter combination rather than re-training, and at the same time eliminate the impact of multi-task gradient interference on collaborative efficiency.
[0018] In view of this, the present invention proposes a multi-agent control method and device based on orthogonal decomposition representation, which solves the problems of redundant multi-agent policy expression and insufficient generalization ability in multi-task scenarios through orthogonal representation learning and multi-agent reinforcement learning algorithms.
[0019] According to one aspect of the present invention, a multi-agent control method based on orthogonal decomposition representation is proposed, including: receiving the local observations and task contexts of agents provided by the upper-layer scheduling system in real time; generating a corresponding original feature matrix according to the local observations of the agents, and performing orthogonalization on the original feature matrix to obtain a corresponding orthogonal basis matrix; generating weights according to the task contexts; calculating task-specific representations according to the orthogonal basis matrix and the weights; inputting the task-specific representations into a fully connected layer, outputting an agent action value function to measure the value of observation-action, and selecting the action corresponding to the maximum value estimation.
[0020] Further, it also includes: if a part of the local observations is missing, using historical observations for interpolation and completion; when multiple task contexts are observed to be activated simultaneously, adopting a weighted voting mechanism to fuse the weights of each task.
[0021] Further, the task context is received by the task-aware dynamic weight generator, and non-negative weights are output; the non-negative weights are normalized, and the normalized weights are multiplied by the orthogonal basis matrix to obtain the task-specific representation.
[0022] Further, each agent receives its local observation through a separate orthogonal mixture of experts Q-network to generate the original feature matrix; each agent receives its task context through a separate task-aware dynamic weight generator to generate the weights; the value estimate is calculated by a hypernetwork-driven monotonic mixture network.
[0023] Further, it also includes a multi-agent training process:
[0024] 1) Collect data: Randomly initialize a neural network as the control strategy for the agent, interact with the environment, collect the reward feedback r, and store the data at each step in the multi-task experience pool.
[0025] 2) Offline experience replay: Sample data (o i , u i , r, s', o i ', c) from the multi-task experience pool, which are the local observation, joint action, reward feedback, state at the next moment, local observation at the next moment, and task context respectively.
[0026] 3) Forward calculation: Input the data (o i , u i , r, s', o i ', c) into the orthogonal mixture of experts Q-network of agent i to calculate the orthogonal basis matrix V i and the action value function Q i , then generate the mixture network parameters W mix , b mix through the hypernetwork-driven monotonic mixture network, and calculate the value estimate Q tot ;
[0027] 4) Calculate the loss function and backpropagation:
[0028] The loss function is the temporal difference error
[0029]
[0030] where γ is the discount factor constant;
[0031] Jointly optimize the parameters of the orthogonal mixture of experts Q-network, the parameters of the task-aware dynamic weight generator, and the parameters of the hypernetwork-driven monotonic mixture network through backpropagation;
[0032] 5) Update the strategy:
[0033] Update the policy of the agent to the optimized artificial neural network, and return to step 1) until convergence.
[0034] According to another aspect of the present invention, a multi-agent control device based on orthogonal decomposition representation is proposed, including: an orthogonal hybrid expert Q network, which is used for each agent to independently generate an orthogonal basis matrix to achieve decoupled representation between tasks; a task-aware dynamic weight generator, which is used to output corresponding weights based on the task context of the agent, and generate task-specific representations according to the weights and the orthogonal basis matrix to achieve cross-task knowledge transfer; a fully connected layer, which takes the task-specific representation as input and outputs the agent action value function to measure the value of the observation-action; a hypernetwork-driven monotonic hybrid network, which takes the global state as input, generates hybrid network parameters through a multi-layer perceptron, and calculates the value estimate according to the hybrid network parameters and the action value function.
[0035] Further, each agent corresponds to an orthogonal hybrid expert Q network, and each orthogonal hybrid expert Q network includes k independent expert networks; the orthogonal hybrid expert Q network receives the local observation at the current moment of the corresponding agent, generates an original feature matrix through its k independent expert networks, and performs Schmidt orthogonalization on the original feature matrix to obtain the orthogonal basis matrix.
[0036] Further, for agent i, its orthogonal hybrid expert Q network receives the local observation o at the current moment i and then generates an original feature matrix U through k independent expert networks i
[0037]
[0038] where d represents the output dimension of the independent expert network;
[0039] Perform Schmidt orthogonalization on the original feature matrix U i to obtain an orthogonal basis matrix satisfying the unit orthogonality constraint where I k is a k×k identity matrix; in order to satisfy the unit orthogonality constraint, the steps of using Schmidt orthogonalization are:
[0040]
[0041] where, u ij represents the j-th row of the original feature matrix U i and v ij represents the j-th row of the orthogonal basis matrix V i of.
[0042] Furthermore, the task-aware dynamic weight generator receives the task context c and outputs non-negative weights and performs Softmax normalization to obtain the normalized weights:
[0043]
[0044] Then, based on the normalized weights and the orthogonal basis matrix, the task-specific representation is generated:
[0045]
[0046] Furthermore, the hypernetwork-driven monotonic mixing network takes the global state s as input and generates mixing network parameters through a multi-layer perceptron:
[0047] W mix ,b mix =f hyper (s;θ hyper )
[0048] The mixing network parameters are processed by the absolute value activation function to ensure Thus:
[0049]
[0050] Among them, the value estimate Q tot is calculated from the action value function Q i and the mixing network parameters W mix ,b mix as follows:
[0051] Q tot =W mix Q+b mix
[0052] Among them, Q=[Q 1 ,Q 2 ,…,Q i ,…,Q n T , and the action value function Q i is:
[0053]
[0054] Among them, represents the action of agent i, and f θ (·) is a fully connected layer function.
[0055] The beneficial effects of the technical solution of the present invention are reflected in that: the multi-agent control method and device of the present invention are significantly superior to the prior art in terms of multi-agent collaborative generalization, policy expression efficiency, and complex task robustness by introducing an explicit orthogonalization representation learning mechanism and a dynamic task adaptation hypernetwork architecture. Among them, for the decoupled orthogonal representation, the orthogonality of expert features is explicitly constrained by Gram-Schmidt, eliminating feature interference between multi-task policies; through dynamic hypernetwork adaptation, the hypernetwork-driven monotonic mixture network generates mixture network parameters according to the global state, making the weight allocation of the orthogonal basis matrix match the optimal collaborative policy of the current task; for the strict monotonicity guarantee: the non-negative weight constraint ensures that the global Q value (value estimation) is monotonically consistent with the single-agent Q value, avoiding the destruction of the multi-agent cooperation stability caused by orthogonalization. In short, the present invention overcomes the core problems such as agent policy conflicts, low data efficiency, and rigid tactical collaboration in multi-task scenarios through orthogonalization representation decoupling, dynamic task weight adaptation, and strict multi-agent collaborative constraints, providing a high-generalization and low-latency reinforcement learning solution for scenarios such as multi-robot systems and UAV swarms. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 FIG. is an architecture diagram of a multi-agent control scheme based on orthogonal decomposition representation according to an embodiment of the present invention.
[0057] Figure 2 FIG. is a schematic flowchart of multi-agent training according to an embodiment of the present invention.
[0058] Figure 3 FIG. is a schematic diagram of an experimental scenario according to an embodiment of the present invention.
[0059] Figure 4 FIG. is a schematic diagram of deploying the multi-agent control scheme of an embodiment of the present invention in an e-commerce warehouse center. DETAILED DESCRIPTION OF THE INVENTION
[0060] The present invention will be further described below in conjunction with the accompanying drawings, specific embodiments, and examples. It should be understood that the purpose of providing the embodiments is only for illustration, not for limitation.
[0061] An embodiment of the present invention provides a multi-agent control method and device based on orthogonal decomposition representation. The architecture of the method and device is as Figure 1As shown, the core components of the device architecture include: an orthogonal mixture of experts Q-network 10, a task-aware dynamic weight generator 20, and a hypernetwork-driven monotonic mixture network 30 (hereinafter referred to as "hypernetwork" for short); n agents, each agent is configured with an independent orthogonal mixture of experts Q-network 10 and a task-aware dynamic weight generator 20. The orthogonal mixture of experts Q-network 10 independently generates an orthogonal basis matrix for each agent to achieve decoupled representation between tasks; and the task-aware dynamic weight generator 20 generates a combined weight based on the task context to achieve cross-task knowledge transfer. The hypernetwork 30 dynamically generates the parameters of the mixture network to ensure the monotonicity of the global Q-value (the value estimate of the action).
[0062] Figure 1 In represents the local observation and task context of agent 1 at time t, represents the local observation and task context of agent n at time t. The orthogonal mixture of experts Q-network receives the local observation o of agent i (i = 1, 2,..., n) provided by the upper-level scheduling system in real time i , and generates the original feature matrix U through its k independent expert networks i :
[0063]
[0064] where d represents the output dimension of the independent expert network;
[0065] Next, perform Gram-Schmidt orthogonalization (Schmidt orthogonalization) on the original feature matrix U i to obtain the orthogonal basis matrix satisfying:
[0066] (unit orthogonality constraint)
[0067] where, I k is a k×k identity matrix. To satisfy the unit orthogonality constraint, the steps of using Schmidt orthogonalization are:
[0068]
[0069] where, u ij represents the j-th row of the original feature matrix U i , and v ij represents the j-th row of the orthogonal basis matrix V i .
[0070] The task-aware dynamic weight generator g ψ receives the task context c (task ID or feature vector) and outputs a non-negative weight vector through the "task encoder" The task weights are obtained through Softmax normalization:
[0071]
[0072] Through the "task status representation", according to the task weights obtained by normalization and the orthogonal basis matrix, a task-specific representation is generated:
[0073]
[0074] The task-specific representation is input into the fully connected layer f θ , and the action value function Q of agent i is output i to measure the value of this observation-action:
[0075]
[0076] where represents the action of agent i.
[0077] Figure 1 In , represents the independent value estimation of the "state sequence-action pair" of agent i at time t under the task context c, that is, the action value function, and its value is used to measure the quality of this action, where tot represents the historical observation-action sequence of agent i, and i = 1, 2,..., n represents the agent number; Q
[0078] For example Figure 1 , the hypernetwork 30 takes the global state s at time t t and the task context c as inputs, and generates the hybrid network parameters through a multi-layer perceptron:
[0079] W mix , b mix = f hyper (s; θ hyper )
[0080] The hybrid network parameters are processed by the absolute value activation function to ensure so that:
[0081]
[0082] Among them, the value estimation Q tot is calculated from the action value function Q i and the hybrid network parameters W mix , b mix :
[0083] Q tot = Wmix Q + b mix
[0084] where Q = [Q 1 , Q 2 ,..., Q i ,..., Q n T .
[0085] The algorithm adopts a centralized training and decentralized execution model. That is, during training, for the training data, the value estimation network of each agent calculates the corresponding value estimation, and then the overall value (value estimation Q tot ) is calculated through the monotonic mixing network driven by the hypernetwork. Then, the value estimation network is optimized using the gradient descent method, and the value estimation network of each agent consists of an orthogonal mixing expert Q-network and a task-aware dynamic weight generator. The orthogonal mixing expert Q-network is used to represent the local observations of the agent, and the task-aware dynamic weight generator is used to generate task weights. The local observations are multiplied by the task weights to obtain a representation of a specific task. Then, the value estimation is obtained using the state-action value function network (Q-network) and the mixing network parameters. During execution (deployment), only the orthogonal mixing expert Q-network and the task-aware dynamic weight generator are used as the policy to control the actions of the agent.
[0086] For the specific training process, please refer to Figure 2 , including:
[0087] 1) Data collection: Randomly initialize the neural network as the control policy for the agent, interact with the environment, collect the reward feedback r, and store the data at each step in the multi-task experience pool;
[0088] 2) Offline experience replay: Sample data (o i , u i , r, s', o i ', c) from the multi-task experience pool, which are the local observation, joint action, reward feedback, state at the next moment, local observation at the next moment, and task context, respectively;
[0089] 3) Forward calculation: Input the data (o i , u i , r, s', o i ', c) into the orthogonal mixing expert Q-network of agent i to calculate the orthogonal basis matrix V i and the action value function Q i , and then generate the mixing network parameters W mix , b mix through the monotonic mixing network driven by the hypernetwork, and calculate the value estimation Q tot ;
[0090] 4) Calculate the loss function and backpropagation:
[0091] The loss function is the temporal difference error
[0092]
[0093] where γ is the discount factor constant, usually set to 0.99;
[0094] Jointly optimize the parameters φ of the orthogonal hybrid expert Q-network, the parameters ψ of the task-aware dynamic weight generator, and the parameters θ of the hypernetwork-driven monotonic hybrid network through backpropagation hyper ;
[0095] 5) Update the policy:
[0096] Update the agent's policy to the optimized artificial neural network and return to step 1) until convergence.
[0097] In summary, in the actual application of the algorithms of the method and device of the present invention, the deployment process is divided into two stages: ① model solidification and ② real-time inference, ensuring that the trained policy can be efficiently adapted to the dynamic multi-task scenario.
[0098] ① Model solidification and lightweighting
[0099] Fix the parameters φ of the orthogonal hybrid expert Q-network, the parameters ψ of the task-aware dynamic weight generator, and the hypernetwork parameters θ after training convergence, and export them as a lightweight model file (such as ONNX or TensorRT engine). At the same time, solidify the orthogonal hybrid expert Q-network, retain the Gram-Schmidt computation graph to ensure real-time orthogonality during deployment; and compile the hypernetwork into a static computation graph to avoid the delay of dynamic parameter generation. Use GPU / TensorCore to perform parallel acceleration on the orthogonal matrix operation for hardware acceleration; perform FP16 or INT8 quantization on the hybrid network parameters W hyper and the orthogonal basis matrix V mix to reduce memory occupancy and inference latency. i
[0100] ② Real-time multi-task inference
[0101] For the current local observation o of the agent i and the task context c input, the processing flow is as follows:
[0102] 1. Orthogonal basis generation: Each agent's orthogonal hybrid expert Q-network generates the original feature matrix U i according to o i and performs Gram-Schmidt orthogonalization to obtain the orthogonal basis matrix V i ;
[0103] 2. Task weight interpolation: The task-aware dynamic weight generator generates a weight w based on the task context c, and after normalization, the task weight is obtained. c , and the task weight is obtained after normalization. Calculate the task-specific representation
[0104] 3. Action decision-making: Calculate Select the action corresponding to the maximum Q value
[0105] 4. Exception handling: If the local observation o i is partially missing, use the historical observation value for interpolation and completion; when multiple task contexts c are observed to be activated simultaneously, adopt a weighted voting mechanism to fuse the task weights.
[0106] To verify the effectiveness of the present invention, in the embodiments of the present invention, a multi-task scenario of multi-vehicle cooperation is built using a parallel multi-agent simulator (Vectorized MultiAgent Simulator) simulation environment. For different tasks, we designed different numbers of vehicles and different road maps, and used the overall circulation efficiency as the environmental reward. When a vehicle collision occurs, the task ends. If the vehicle does not collide within 1000 steps and the average vehicle passing efficiency is higher than the threshold, it is considered that the task is completed. The multi-robot system needs to improve the circulation efficiency as much as possible without collision. We designed a total of 5 tasks: 10-vehicle scenario, 20-vehicle scenario, 50-vehicle scenario, 20-vehicle scenario with fault disturbance, and 50-vehicle scenario with fault disturbance. The task visualization is shown in Figure 3 . The specific technical effects are as follows in 3.1 - 3.3:
[0107] 3.1 The multi-task generalization performance is significantly improved
[0108] We used the 10-vehicle scenario and the 20-vehicle scenario for training, and tested in the 50-vehicle scenario after training convergence. The cross-task average completion rate of the method of the present invention is ≥28% higher than that of QMIX (the task completion rate reaches 43%, and the baseline algorithm is 33%). Based on the converged policy, training is carried out on new tasks until convergence. Compared with the baseline algorithm QMIX, the number of interaction steps required to adapt to new tasks is reduced by 32% (the convergence steps are about 500,000 steps, and the convergence steps of the baseline algorithm are about 750,000 steps). The performance improvement is first due to the orthogonal decoupling representation used in this method. By the Gram-Schmidt process, the expert network of the agent is forced to generate orthogonal basis vectors, so that the basis vectors encode state features and task roles in different dimensions, avoiding policy conflicts caused by feature redundancy in traditional methods.
[0109] 3.2 Data efficiency and training stability in complex scenarios
[0110] In the 20-vehicle scenario and 50-vehicle scenario with disturbances in high-difficulty tasks, having a fault disturbance means that the vehicle will randomly have a fault and cannot move (the fault probability is 0.1). The method of the present invention has a task completion rate 15% higher than QMIX (reaching 76%, while QMIX is 66%) under the same number of training steps (3 million steps), and the variance in the training process is reduced by 23%. The improvement in the training effect benefits from the orthogonalization of the algorithm, which alleviates the gradient conflict in training. Based on the calculation of the gradient cosine similarity, the gradient conflict between tasks is reduced by 83%.
[0111] 3.3 Algorithm Compatibility and Real-Time Deployment Capability
[0112] The algorithm of the present invention can be seamlessly integrated into the standard training framework of multi-agent reinforcement learning algorithms (based on PyMARL, PyMARL2, or EPyMARL), with a code modification amount <15%, and the deployment inference latency ≤ 8 ms (NVIDIA Jetson AGX Xavier). The high compatibility of the algorithm benefits from the modular lightweight design of the algorithm, and the hypernetwork and Q-value hybrid calculation maintain the original design. Such modular processing enables the algorithm to be seamlessly migrated to other multi-agent reinforcement learning methods (such as MAPPO, MADDPG, etc.). The efficient deployment benefits from the fact that the algorithm can reduce the calculation latency and memory occupancy through hardware-level acceleration and TensorRT quantization.
[0113] In summary, through orthogonal representation decoupling, dynamic task weight adaptation, and strict multi-agent collaboration constraints, the present invention has overcome the core problems such as agent policy conflicts, low data efficiency, and rigid tactical collaboration in multi-task scenarios, providing a highly generalized and low-latency reinforcement learning solution for scenarios such as multi-robot systems and unmanned aerial vehicle clusters.
[0114] The following provides a real-world deployment case of the multi-agent control solution of the present invention - taking an e-commerce warehouse center as an example.
[0115] The difficulties in the e-commerce warehouse center scenario are as follows: Emergency orders need to be processed out of turn, and the task allocation of robots needs to be adjusted in real time; Multiple automatic guided vehicles cross each other frequently in narrow channels; The grasping speed of the robotic arm needs to be accurately matched with the transportation rhythm of the automatic guided vehicle to avoid congestion in the packing area. The entire deployment plan is as Figure 4 shown, where:[[]]END]]
[0116] Perception layer: ① The automatic guided vehicle is equipped with a camera, lidar, and inertial navigation module to obtain environmental information; ② The robotic arm is equipped with a visual grasping system.
[0117] Decision-making layer: ① Control center: Runs the algorithm of the multi-agent control method / device of the present invention, and dynamically generates multi-robot cooperation strategies; ② Task scheduling interface: Receives the order flow of the WMS (Warehouse Management System), and converts it into a reinforcement learning task context (such as order priority, target shelf ID).
[0118] Execution layer: Each robot executes the strategy generated by the neural network in real time through the edge computing unit (NVIDIA Jetson AGX Xavier).
[0119] The decision-making process is as follows:
[0120] First, input the task context c (including order type, target shelf location, goods type, goods volume) and the local observation of each robot (such as its own location, power, sensor data) to generate an orthogonal basis matrix. Then use the task-aware dynamic weight generator to generate task weights and calculate the task-specific representation. Next, pass the task representation through the value network to calculate the action value function, and finally select the appropriate action according to the state-action value estimation (i.e., the action corresponding to the maximum state-action value estimation).
[0121] The core of the technical solution of the present invention lies in enhancing the multi-agent multi-task expression ability through orthogonalized representation. Based on this principle, the following variant implementation manners can be derived. Although these variant solutions differ in implementation details from the best embodiment, they all share the same innovative principle and are applicable to different application scenarios or technical constraint conditions.
[0122] Transition solution: Approximate orthogonalization method based on feature decoupling
[0123] In the initial stage of development, an attempt was made to approximately achieve feature orthogonality through sparse regularization and decoupled loss functions instead of explicit Gram-Schmidt orthogonalization. The specific implementation is as follows:
[0124] The agent's orthogonal mixture of experts Q-network contains k experts and outputs the original feature matrix Add an orthogonal loss function:
[0125]
[0126] Forcibly reduce the correlation between the features of different experts. When calculating the task weights, add a sparsity constraint (L1 regularization) to make the weight vector activate only a few experts.
[0127] The advantage of this transition solution is that the computational complexity is reduced, which can slightly improve the training speed and deployment efficiency. The disadvantage is that the orthogonality of the expert features is weak, which will reduce the performance on unseen tasks. It is applicable to scenarios with extremely high real-time requirements but low task complexity (such as simple collaborative transportation tasks).
[0128] Regardless of which embodiment, the core lies in:
[0129] ① Orthogonal Mixture of Experts Q-Network: Each agent deploys an independent multi-expert network to generate the original feature matrix, and enforces the generation of an orthogonal basis matrix through Gram-Schmidt orthogonalization. The orthogonal basis matrix is dynamically fused with the task-aware dynamic weights generated by the task-aware dynamic weight generator to generate task-specific representations. It supports the decoupling of internal policies of the agent and eliminates multi-task gradient interference.
[0130] ② Task-Aware Dynamic Weight Generation Mechanism: The task-aware dynamic weight generator generates a non-negative weight vector based on the task context description, and dynamically combines the orthogonal bases after Softmax normalization. It realizes the zero-shot generalization ability across tasks (adapts to new tasks through weight interpolation).
[0131] ③ Hypernetwork-Driven Monotonic Mixture Network: The hypernetwork takes the global state as input, generates the parameters of the mixture network, and ensures non-negative weights through absolute value constraints. It ensures that the global Q-value and the single-agent Q-value are strictly monotonically consistent.
[0132] ④ Multi-Task Cooperative Training Framework: Design of the loss function for jointly optimizing the parameters of orthogonal experts, task encoders, and hypernetworks.
[0133] ⑤ Associated storage and sampling mechanism for task context in the offline experience replay pool.
[0134] The key algorithm implementations include:
[0135] 1) Explicit Orthogonalization Constraint Method: Embed a Gram-Schmidt orthogonalization layer in the agent's Q-network to generate the orthogonal basis matrix in real time; support the orthogonalization computation graph solidification technology for hardware acceleration (such as CUDA kernel optimization).
[0136] 2) Dynamic Task Adaptation and Exception Handling: Historical observation interpolation and completion strategy when the task context is missing; weighted voting fusion mechanism in multi-task conflict scenarios (such as weight superposition when multiple tasks c are activated simultaneously).
[0137] 3) Approximate Orthogonalization Alternative: Implicit feature decoupling method based on sparse regularization (L decouple =∑(<u ij , u im > 2 )); hierarchical orthogonalization design (perform global feature orthogonalization in the mixture network layer).
[0138] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several equivalent substitutions or obvious modifications can be made, and as long as the performance or use is the same, they should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A multi-agent control method based on orthogonal decomposition representation, characterized in that: include: Receive local observations and task contexts of the agent provided by the upper scheduling system in real time; Generate a corresponding original feature matrix according to the local observation of the agent, and perform orthogonalization on the original feature matrix to obtain a corresponding orthogonal basis matrix; generating a weight according to the task context; Compute a task-specific representation based on the orthogonal basis matrix and the weights; The task-specific representation is input into the fully connected layer, and the agent action value function is output to measure the value of the observation-action, and the action corresponding to the maximum value estimate is selected.
2. The multi-agent control method according to claim 1, characterized in that: Also includes: If the local observation is partially missing, historical observations are used to interpolate and complete it; When multiple task contexts are observed to be activated simultaneously, a weighted voting mechanism is used to fuse the weights of each task.
3. The multi-agent control method according to claim 1 or 2, characterized in that: The task context is received by a task-aware dynamic weight generator, and a non-negative weight is output; the non-negative weight is normalized, and the normalized weight is multiplied by the orthogonal basis matrix to obtain the task-specific representation.
4. The multi-agent control method according to any one of claims 1 to 3, characterized in that: Each agent receives its local observations through a separate orthogonal mixture of experts Q network to generate the original feature matrix; each agent receives its task context through a separate task-aware dynamic weight generator to generate the weight; The value estimates are computed via a monotonic mixing network driven by a hypernetwork.
5. The multi-agent control method according to claim 4, characterized in that: It also includes the multi-agent training process: 1) Collect data: randomly initialize a neural network as a control strategy for the agent, interact with the environment, collect reward feedback r, and store the data of each step in a multi-task experience pool; 2) Offline experience replay: sampling data from the multi-task experience pool (o i ,u i ,r,s',o i ',c), respectively, local observation, joint action, reward feedback, state at the next moment, local observation at the next moment, and task context; 3) Forward calculation: input data (o i ,u i ,r,s',o i ',c) to the orthogonal hybrid expert Q network of agent i to calculate the orthogonal basis matrix V i And the action value function Q i , and then generate the hybrid network parameters W through the monotone hybrid network driven by the hypernetwork mix ,b mix , and calculate the value estimate Q tot ; 4) Calculate loss function and back propagation: The loss function is the temporal difference error Where γ is the discount factor constant; Jointly optimize the orthogonal hybrid expert Q-network parameters, the task-aware dynamic weight generator parameters, and the hypernetwork-driven monotonic hybrid network parameters via back-propagation. 5) Update strategy: Update the agent's strategy to the optimized artificial neural network and return to step 1) until convergence.
6. A multi-agent control device based on orthogonal decomposition representation, characterized in that: include: Orthogonal hybrid expert Q network, which is used for each agent to independently generate an orthogonal basis matrix to achieve decoupled representation between tasks; A task-aware dynamic weight generator, configured to output corresponding weights based on the agent's task context, and generate task-specific representations based on the weights and the orthogonal basis matrix to achieve cross-task knowledge transfer; A fully connected layer, which takes the task-specific representation as input and outputs an agent action-value function to measure the value of the observation-action; The hypernetwork-driven monotone hybrid network takes the global state as input, generates hybrid network parameters through a multi-layer perceptron, and calculates value estimates based on the hybrid network parameters and the action-value function.
7. The multi-agent control device according to claim 6, characterized in that: Each intelligent agent corresponds to an orthogonal hybrid expert Q network, and each orthogonal hybrid expert Q network contains k independent expert networks; the orthogonal hybrid expert Q network receives the local observation of the corresponding intelligent agent at the current moment, generates an original feature matrix through its k independent expert networks, and performs Schmidt orthogonalization on the original feature matrix to obtain the orthogonal basis matrix.
8. The multi-agent control device according to claim 7, characterized in that: For agent i, its orthogonal hybrid expert Q network receives the local observation o at the current moment i Then the original feature matrix U is generated through k independent expert networks i Wherein, d represents the output dimension of the independent expert network; For the original feature matrix U i Perform Schmidt orthogonalization to obtain the orthogonal basis matrix Satisfy the unit orthogonality constraint Among them I k is a k×k identity matrix; in order to satisfy the unit orthogonal constraint, the steps of using Schmidt orthogonalization are: Among them, u ij Represents the original feature matrix U i The jth row of v ij Denotes the orthogonal basis matrix V i The jth row of .
9. The multi-agent control device according to claim 8, characterized in that: The task-aware dynamic weight generator receives the task context c and outputs a non-negative weight And perform Softmax normalization to obtain the normalized weights: Then, the task-specific representation is generated based on the normalized weights and the orthogonal basis matrix:
10. The multi-agent control device according to claim 9, characterized in that: The monotone hybrid network driven by the hypernetwork takes the global state s as input and generates the hybrid network parameters through a multi-layer perceptron: W mix ,b mix =f hyper (s;θ hyper ) The hybrid network parameters are processed by the absolute value activation function to ensure thereby: Among them, the value estimate Q tot The action value function Q i and the hybrid network parameter W mix ,b mix The calculation results are: Q tot =W mix Q+b mix Where Q = [Q1, Q2, ..., Q i , …, Q n ] T , the action value function Q i for: in, represents the action of agent i, f θ (·) is the fully connected layer function.