Robot motion control method and system applicable to multiple scenes
By introducing distributed multi-agent reinforcement learning algorithms and experience sharing mechanisms in multi-robot systems, the problem of low collaboration efficiency of multi-robot systems in the existing technology in complex environments is solved, and efficient path planning and collaboration performance is achieved.
Patent Information
- Application Number
- CN202510027929.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-13
AI Technical Summary
Existing multi-robot systems are difficult to achieve efficient collaboration and path planning in complex environments. Centralized control algorithms are complex in calculations and are susceptible to single point failures, while distributed control algorithms lack globality and limited collaboration efficiency.
The experience sharing mechanism between agents is adopted, and efficient collaboration and path planning between robots are achieved through distributed multi-agent reinforcement learning algorithm (MARL) and state-action-reward function design. Specific steps include task priority allocation, local policy generation, path planning, experience sharing, and conflict detection and optimization.
It significantly improves the efficiency of strategy optimization, improves learning speed and collaboration performance, and can efficiently divide labor and flexibly adjust strategies in a dynamic environment, which is suitable for multiple application scenarios.
Smart Images

Figure CN119990179A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent control technology, and in particular to a robot motion control method and system applicable to multiple scenarios. Background Art
[0002] In recent years, with the rapid development of artificial intelligence and robotics, the application of multi-robot systems in industrial production, logistics and transportation, drone formations, disaster relief and other scenarios has gradually become popular. However, achieving efficient multi-robot collaboration and path planning in complex environments is still a technical difficulty. At present, the multi-robot system in the industry usually adopts the following two technical methods:
[0003] (1) Centralized control algorithm
[0004] The global information of the multi-robot system is collected, analyzed and decided through a single control center. This type of method relies on the complete perception of the environment and tasks, and the unified allocation of resources and path planning. For example, the traditional ant colony optimization algorithm (ACO) and particle swarm optimization algorithm (PSO) are used for path planning and task allocation under a centralized architecture. However, since the centralized system needs to process a large amount of environmental data, its computational complexity and dependence on communication are extremely high, resulting in poor performance in dynamic or partially observable environments. The uncertainty and dynamics in the actual environment make it difficult for centralized control to cope with. It is vulnerable to single point failures and lacks reliability.
[0005] (2) Distributed control algorithm
[0006] By equipping each robot with an independent control unit, the robot can make decisions based on local perception and interaction information between neighboring robots. For example, distributed collaboration models based on game theory (such as Nash Equilibrium) have been widely used in drone formation path planning. However, existing distributed methods usually lack a global grasp of the environment, and the efficiency of collaboration between agents is limited, especially in high-density, multi-task environments, and cannot efficiently solve resource allocation and path conflict problems. The lack of unified optimization of global strategies may lead to local optimal solutions. The learning efficiency is low, and a large number of iterations are usually required to achieve a relatively ideal collaboration effect. Summary of the invention
[0007] In order to avoid the above-mentioned problems existing in the prior art, the purpose of the present invention is to provide a robot motion control method and system applicable to multiple scenarios that can improve the efficiency of strategy optimization and accelerate the learning process through an experience sharing mechanism among intelligent agents, and can collaborate efficiently in a dynamic environment, have a clear division of labor and flexibly adjust strategies.
[0008] To achieve the above object, the present invention provides the following technical solution: a robot motion control method applicable to multiple scenarios, comprising the following steps:
[0009] S1: Model the task scenario and initialize the robot parameters;
[0010] S2: Dynamically assign tasks to each robot according to the task priority table;
[0011] S3: Generate a local strategy for each robot based on the goal of the task assigned in step S2 using a distributed multi-agent reinforcement learning algorithm (MARL);
[0012] S4: Each robot uses the SAC algorithm for path planning;
[0013] S5: The robot extracts the experience optimization strategy shared by other robots from the global shared pool, and stores the key experience in the global shared pool after completing a single-step action;
[0014] S6: Detect path conflicts and optimize the path through game theory algorithms. The robot executes according to the optimized path.
[0015] S7: Update the task priority table in real time based on task completion status or environmental changes;
[0016] S8: Loop steps S4-S6 until the task is completed or the system stops running.
[0017] The present invention is further configured such that step S1 specifically models the task scenario of multi-robot work as a multi-task scheduling problem with dynamic changes, defines the state S, action A and reward function R in the task scenario; and initializes an independent strategy network and value network for each robot.
[0018] The state S includes a collection of global information such as the robot position, the target task point position, and the current task state;
[0019] The action A includes a set of parameters such as the movement direction and speed adjustment of the robot at a specific time step;
[0020] Reward function R: It is designed according to the actual situation, taking into account factors such as task completion efficiency and path safety, and is used to optimize system performance.
[0021] The present invention is further configured that step S3 specifically comprises: each robot based on its own observed local state s i and the current task goal, using the distributed multi-agent reinforcement learning algorithm (MARL) to generate the corresponding action a i ; where i represents the robot number, i=1, 2,…, n; n is the number of robots.
[0022] It should be noted that s i It refers to the local state that robot i can observe, such as its own position, the distribution of surrounding obstacles provided by sensors, and the state information of surrounding robots obtained through communication. It is part of the global state observed by robot i. i is the robot i in the local state s i The corresponding actions generated below.
[0023] s i 、a i For distributed control, each robot only needs to be based on the local state s i Perform path planning or decision making without the need for the full global state s t , reducing communication requirements and computational complexity.
[0024] The present invention is further configured such that the robots exchange status information via a local area network to optimize path planning and avoid task conflicts.
[0025] The present invention is further configured that, in step S4, each robot updates its strategy network parameters and value network parameters using the SAC algorithm;
[0026] The policy network parameter update formula is:
[0027]
[0028] Among them, J πθ is the objective function of the policy update, which represents the goal that the policy network needs to optimize. πθ , improve the quality of the strategy. It refers to the expectation, which means that under the strategy π, by sampling state s t and action a t The expected value of the objective function is calculated from the distribution of . α is the entropy regularization coefficient, which is used to balance exploration and utilization. logπθ(a t |s t ) is the logarithm of the action probability, indicating that the strategy π is in state s t Next select action a t The probability of is used to calculate the entropy of the strategy. π (s t ,a t ) is the Q value evaluated by the value network, indicating that in state s t Take action a t After that, the expectation of future accumulated rewards.
[0029] The value network update formula is:
[0030]
[0031] is the objective function of the value network, which represents the goal that the value network needs to optimize. Make the predicted Q value closer to the target value. is the expected symbol, indicating the sampled state s t 、Action a t 、Instant Rewards t , the next state s t+1 Calculate the expectation of the objective function; is the output of the current value network, indicating that in state s t Take action a t The Q value of y is the target value, which is used to measure the accuracy of the current value network prediction.
[0032] It should be noted that s t It refers to the global state of the system at time step t, including the complete information of all robots and the environment, and is the total state of all local states s i Collection of t Used for global strategy optimization, training value network, and recording s in the experience pool t For subsequent experience sharing.
[0033] The present invention is further configured such that the target value y is calculated by the following formula:
[0034]
[0035] where r t is the immediate reward of the current time step; γ is the discount factor, which is used to balance the weights of long-term rewards and short-term rewards; is the expectation operator, indicating that from the strategy π θ a sampled from (a|s) t+1 Calculate the weighted average of all possible values of ; is the Q value prediction of the value network for the next time step, indicating that in state s t+1 Next select action a t+1 The future cumulative reward of αlogπθ(a t+1 |s t+1 ) is a regularization term of action entropy, which is used to encourage the strategy to explore new actions.
[0036] The present invention is further configured such that step S5 specifically comprises the following steps: the robot extracts data from the global shared pool through a priority sampling algorithm, and each robot stores its key experience in the global shared pool after completing a single-step action; the key experience includes the state s at the previous single-step action, the executed action a, the reward r obtained, and the state s′ after completing the single-step action.
[0037] The priority sampling algorithm can choose to extract different valuable data according to the current situation. Simple tasks can be selected based on immediate rewards. The larger the immediate reward, the more significant the positive impact of the action on the system. In complex or dynamic environments, rare or under-explored states can be selected based on state scarcity, which may contain more learning value, strengthen exploration capabilities and avoid local optimality. It can also be based on task relevance to determine whether the experience data is related to the current task goal, and prioritize sampling of experience related to the target point or high-priority tasks, which is more in line with the actual task. Or consider it comprehensively.
[0038] The purpose of the above scheme is to enrich training data, accelerate learning, and improve the speed of strategy convergence.
[0039] The present invention is further configured that step S6 includes the following steps:
[0040] S61: discretize the path planning result of each robot to generate a path sequence (x, y, t), where t represents the time step, and x and y are two-dimensional coordinate values;
[0041] S62: Perform conflict detection to check at each time step t whether there are any two robots R i and R j satisfy:
[0042] (x i ,y i , t) = (x j ,y j , t);
[0043] Among them, x i ,y i For robot R i The coordinates at time step t, x j ,y j For robot R j The coordinates at time step t, i.e., the two robots arrive at the same position at the same time;
[0044] S63: Recording conflict points that satisfy conflict detection;
[0045] S64: assigning a priority level P to the robot according to the task priority table;
[0046] S65: traverse all recorded conflict points, handle conflicts in order of priority, and modify the path of the robot with low priority;
[0047] S66: Repeat steps S62-S65 for the adjusted path until there are no conflict points on the new path.
[0048] The present invention is further configured such that the path modification in step S65 adopts time delay or spatial detour.
[0049] The present invention also relates to a robot motion control system applicable to multiple scenarios, which is used in conjunction with the above-mentioned robot motion control method applicable to multiple scenarios, and comprises:
[0050] A global shared pool, used to store and manage key experiences generated by multiple robots, and to replace obsolete data using a first-in-first-out (FIFO) mechanism; the global shared pool is a distributed shared pool;
[0051] The multi-agent communication module is used for information exchange between robots through the local area network; the multi-agent communication module broadcasts global task information within the specified time step and allocates tasks and paths through a distributed negotiation mechanism.
[0052] An algorithm execution module is used to execute the algorithm flow of the robot motion control method applicable to multiple scenarios.
[0053] The present invention accelerates strategy convergence through experience sharing. Compared with traditional distributed reinforcement learning, the shared pool mechanism can significantly improve the efficiency of strategy optimization. Combined with priority allocation and conflict resolution algorithms, efficient multi-robot collaboration is achieved. The system has strong adaptability to dynamic environments and is suitable for multiple application scenarios such as logistics, agriculture, and rescue.
[0054] In summary, the beneficial effects of the above technical solution of the present invention are as follows:
[0055] (1) The present invention introduces a distributed multi-agent reinforcement learning (MARL) algorithm, which enables robots to make independent decisions based on local information and achieve collaboration through a communication mechanism, thus avoiding the inefficiency problems caused by information transmission delays and computing bottlenecks in traditional centralized methods.
[0056] (2) The experience sharing mechanism is used to store the key experiences learned by individual robots in a global shared pool, and other robots are allowed to sample and use them during training. This design overcomes the problems of low sample efficiency and slow convergence in traditional multi-agent reinforcement learning.
[0057] (3) By combining the advantages of the SAC algorithm, the single-robot strategy optimization process is made more accurate. At the same time, the path conflict detection and optimization algorithm based on game theory is adopted to effectively solve the path conflict problem that may occur between multiple robots.
[0058] (4) A dynamic task priority allocation mechanism is adopted to adjust the task allocation order and strategy in real time according to the urgency, complexity and resource availability of the task, overcoming the problem of poor adaptability of the traditional static task allocation method in environmental changes. The dynamic task priority mechanism is combined with the multi-robot collaboration mechanism to make resource allocation more reasonable and avoid task duplication and resource waste.
[0059] (5) Through modular architecture design, the multi-robot control system of this patent can be flexibly expanded to more robots and mission scenarios, including but not limited to disaster relief, agricultural spraying, warehouse management and other fields. By adopting a distributed control architecture, robots make independent decisions and collaborate through local communication, avoiding the reliance of traditional centralized control on global communication, while reducing the risk of system failure due to single point failure. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0061] Figure 1 The figure is a flow chart of a robot motion control method applicable to multiple scenarios. DETAILED DESCRIPTION
[0062] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention is clearly and completely described below in conjunction with the accompanying drawings of the present invention. Based on the embodiments of the present invention, other similar embodiments obtained by ordinary technicians in the field without making any creative work should all fall within the scope of protection of the present invention.
[0063] The present invention will be further described below in conjunction with the accompanying drawings and preferred embodiments.
[0064] Example:
[0065] like Figure 1 As shown in the figure, a preferred embodiment of the present invention is a robot motion control method applicable to multiple scenarios, comprising the following steps:
[0066] S1: Model the task scenario and initialize the robot parameters;
[0067] In this embodiment, the multi-UAV collaborative task in logistics transportation is used as the simulation scenario, and efficient path planning and task allocation are achieved by constructing a distributed multi-agent system and a dynamic task scheduling solution.
[0068] The hardware equipment includes:
[0069] Number of drones: 10 multi-rotor drones, with a payload of 3kg and a flight time of 30 minutes.
[0070] Communication equipment: Wi-Fi communication module.
[0071] Mission scenario: 50 randomly distributed logistics distribution points, with dynamic changes in the locations of distribution points and mission requirements.
[0072] The task scenario of multi-robot work is modeled as a multi-task scheduling problem with dynamic changes. The state S, action A and reward function R in the task scenario are defined; and an independent policy network and value network are initialized for each robot.
[0073] The state S includes the drone location, the delivery point location, and the task status (completed / uncompleted);
[0074] The action A includes adjusting the movement direction and speed of the drone;
[0075] The reward function R is designed to be the negative value of the task completion time.
[0076] S2: Dynamically assign tasks to each drone according to the task priority table; assign priorities based on task urgency and distance, and dynamically update task priorities.
[0077] S3: Use the distributed multi-agent reinforcement learning algorithm (MARL) based on its own observed local state s i and the goal of the task assigned in step S2, using the distributed multi-agent reinforcement learning algorithm (MARL) to generate the corresponding action a i ; where i represents the robot number, i=1, 2,…, n; n is the number of robots.
[0078] S4: Each UAV uses the SAC algorithm to plan the path and form a preliminary path;
[0079] Each robot uses the SAC algorithm to update its policy network parameters πθ and value network parameters Qφ;
[0080] The policy network parameter update formula is:
[0081]
[0082] Among them, J πθ is the objective function of the policy update, which represents the goal that the policy network needs to optimize. πθ , improve the quality of the strategy. It refers to the expectation, which means that under the strategy π, by sampling state s t and action a tThe expected value of the objective function is calculated from the distribution of . α is the entropy regularization coefficient, which is used to balance exploration and utilization. logπθ(a t |s t ) is the logarithm of the action probability, indicating that the strategy π is in state s t Next select action a t The probability of is used to calculate the entropy of the strategy. π (s t ,a t ) is the Q value evaluated by the value network, indicating that in state s t Take action a t After that, the expectation of future accumulated rewards.
[0083]
[0084] is the objective function of the value network, which represents the goal that the value network needs to optimize. Make the predicted Q value closer to the target value. is the expected symbol, indicating the sampled state s t 、Action a t 、Instant Rewards t , the next state s t+1 Calculate the expectation of the objective function; is the output of the current value network, indicating that in state s t Take action a t The Q value of y is the target value, which is used to measure the accuracy of the current value network prediction.
[0085] The present invention is further configured such that the target value y is calculated by the following formula:
[0086]
[0087] where r t is the immediate reward of the current time step; γ is the discount factor, which is used to balance the weights of long-term rewards and short-term rewards; is the expectation operator, indicating that from the strategy π θ a sampled from (a|s) t+1 Calculate the weighted average of all possible values of ; is the Q value prediction of the value network for the next time step, indicating that in state s t+1 Next select action a t+1 The future cumulative reward of αlogπθ(a t+1 |s t+1 ) is a regularization term of action entropy, which is used to encourage the strategy to explore new actions.
[0088] S5: The drone uploads key experience data to the global shared pool in real time, and extracts the experience of other drones from the global shared pool through the priority sampling algorithm for strategy optimization, so as to enrich training data, accelerate learning, and improve the speed of strategy convergence.
[0089] The key experience includes the state s at the last single-step action, the executed action a, the reward r obtained, and the state s′ after completing the single-step action.
[0090] S6: The drone optimizes the path according to the path conflict detection algorithm to avoid conflicts with other drones, and the drone executes according to the optimized path;
[0091] Step S6 includes the following steps:
[0092] S61: discretize the path planning result of each robot to generate a path sequence (x, y, t), where t represents the time step, and x and y are two-dimensional coordinate values;
[0093] S62: Perform conflict detection to check at each time step t whether there are any two robots R i and R j satisfy:
[0094] (x i ,y i , t) = (x j ,y j , t);
[0095] Among them, x i ,y i For robot R i The coordinates at time step t, x j ,y j For robot R j The coordinates at time step t, i.e., the two robots arrive at the same position at the same time;
[0096] S63: Recording conflict points that satisfy conflict detection;
[0097] S64: assigning a priority level P to the robot according to the task priority table;
[0098] S65: traverse all recorded conflict points, handle conflicts in order of priority, and modify the path of the robot with a low priority; the path modification adopts time delay or spatial detour;
[0099] The time delay is to modify the path of the low-priority robot at the conflict time step t, insert a waiting time step t+1 near the conflict point, and wait for the high-priority robot to pass;
[0100] The spatial detour is to re-plan a local path to bypass the conflict point;
[0101] S66: Repeat steps S62-S65 for the adjusted path until there are no conflict points on the new path.
[0102] S7: When encountering task changes, such as adding a new delivery point or a blocked route, the system updates the task priority table in real time and reallocates tasks.
[0103] The system performance of the above method of this embodiment and the traditional centralized control algorithm was run 30 times in the same task scenario. The average task completion time, average number of path conflicts and system convergence time obtained by this embodiment and the traditional centralized control algorithm were recorded respectively.
[0104] The results show that the system task completion time of this patented technology is 30% shorter than that of traditional methods. The number of path conflicts is reduced to 1 / 3 of that of traditional methods. The system convergence time is greatly shortened, which improves learning efficiency.
[0105] Embodiment 2:
[0106] The present invention also relates to a robot motion control system applicable to multiple scenarios, which is used in conjunction with the above-mentioned robot motion control method applicable to multiple scenarios, and comprises:
[0107] A global shared pool, which is a distributed shared pool used to store and manage key experiences generated by multiple robots, and uses a first-in-first-out (FIFO) mechanism to replace obsolete data;
[0108] The multi-agent communication module is used for information exchange between robots through the local area network; the multi-agent communication module broadcasts global task information within the specified time step and allocates tasks and paths through a distributed negotiation mechanism.
[0109] An algorithm execution module is used to execute the algorithm flow of the robot motion control method applicable to multiple scenarios.
[0110] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A robot motion control method applicable to multiple scenarios, characterized in that: The following steps are involved: S1: Model the task scenario and initialize the robot parameters; S2: Dynamically assign tasks to each robot according to the task priority table; S3: Generate a local strategy for each robot based on the goal of the task assigned in step S2 using a distributed multi-agent reinforcement learning algorithm; S4: Each robot uses the SAC algorithm for path planning; S5: The robot extracts the experience optimization strategy shared by other robots from the global shared pool, and stores the key experience in the global shared pool after completing a single-step action; S6: Detect path conflicts and optimize the path through game theory algorithms. The robot executes according to the optimized path. S7: Update the task priority table in real time based on task completion status or environmental changes; S8: Loop steps S4-S6 until the task is completed or the system stops running.
2. The robot motion control method applicable to multiple scenarios according to claim 1 is characterized in that: Specifically, step S1 models the task scenario of multi-robot work as a multi-task scheduling problem with dynamic changes, defines the state S, action A and reward function R in the task scenario; and initializes an independent strategy network and value network for each robot.
3. The robot motion control method applicable to multiple scenarios according to claim 1 is characterized in that: Step S3 specifically includes: each robot based on its own observed local state s i and the current task goal, using the distributed multi-agent reinforcement learning algorithm to generate the corresponding action a i ; where i represents the robot number, i=1, 2,…, n; n is the number of robots.
4. The robot motion control method applicable to multiple scenarios according to claim 1 is characterized in that: The robots exchange status information via a local area network.
5. The robot motion control method applicable to multiple scenarios according to claim 1, characterized in that: In step S4, each robot uses the SAC algorithm to update its policy network parameters πθ and value network parameters The policy network parameter update formula is: Among them, J πθ is the objective function of the policy update, which indicates the goal that the policy network needs to optimize; is the expectation, which means that under the strategy π, by sampling state s t and action a t The expected value of the objective function is calculated by the distribution; α is the entropy regularization coefficient; logπθ(a t |s t ) is the logarithm of the action probability, indicating that the strategy π is in state s t Next select action a t The probability of π (s t , a t ) is the Q value evaluated by the value network, indicating that in state s t Take action a t After that, the expectation of future cumulative rewards; The value network update formula is: is the objective function of the value network, which indicates the objective that the value network needs to optimize; is the expected symbol, indicating the sampled state s t 、Action a t 、Instant Rewards t , the next state s t+1 Calculate the expectation of the objective function; is the output of the current value network, indicating that in state s t Take action a t The Q value of , y is the target value.
6. The robot motion control method applicable to multiple scenarios according to claim 5, characterized in that: The target value is calculated by the following formula: where r t is the instant reward of the current time step; γ is the discount factor; is the expectation operator, indicating that from the strategy π θ a sampled from (a|s) t+1 Calculate the weighted average of all possible values of ; is the Q value prediction of the value network for the next time step, indicating that in state s t+1 Next select action a t+1 The future cumulative reward of αlogπθ(a t+1 |s t+1 ) is the regularization term of action entropy.
7. The robot motion control method applicable to multiple scenarios according to claim 1, characterized in that: Specifically, step S5 is that the robot extracts data from the global shared pool through a priority sampling algorithm. After completing a single-step action, each robot stores its key experience in the global shared pool; the key experience includes the state s at the previous single-step action, the executed action a, the reward r obtained, and the state s′ after completing the single-step action.
8. The robot motion control method applicable to multiple scenarios according to claim 1, characterized in that: Step S6 includes the following steps: S61: discretize the path planning result of each robot to generate a path sequence (x, y, t), where t represents the time step, and x and y are two-dimensional coordinate values; S62: Perform conflict detection to check at each time step t whether there are any two robots R i and R j satisfy: (x i ,y i ,t)=(x j ,y j ,t); Among them, x i ,y i For robot R i The coordinates at time step t, x j ,y j For robot R j Coordinates at time step t; S63: Recording conflict points that satisfy conflict detection; S64: assigning a priority level P to the robot according to the task priority table; S65: traverse all recorded conflict points, handle conflicts in order of priority, and modify the path of the robot with low priority; S66: Repeat steps S62-S65 for the adjusted path until there is no conflict point on the new path.
9. The robot motion control method applicable to multiple scenarios according to claim 8, characterized in that: The path modification in step S65 adopts time delay or spatial detour.
10. A robot motion control system applicable to multiple scenarios, used in conjunction with a robot motion control method applicable to multiple scenarios as claimed in any one of claims 1 to 9, characterized in that: include: A global shared pool for storing and managing key experiences generated by multiple robots, using a first-in-first-out mechanism to replace obsolete data; The global shared pool is a distributed shared pool; The multi-agent communication module is used for information exchange between robots through the local area network. The multi-agent communication module broadcasts global task information within a specified time step and allocates tasks and paths through a distributed negotiation mechanism. An algorithm execution module is used to execute the algorithm flow of the robot motion control method applicable to multiple scenarios.
Citation Information
Patent Citations
Design method of distributed autonomous robot traffic coordination mechanism
CN111638717A
Task allocation and route planning optimization method for multiple unmanned aerial vehicles in dynamic environment
CN117035435A
Full-automatic storage bin and control method thereof
CN119045587A
Cited By
Urban end distribution scheduling method and system based on multi-agent reinforcement learning
CN121414086A