Robot group real-time control method based on edge calculation
Through the distributed architecture of edge computing and deep reinforcement learning, data transmission delay, network pressure and single point of failure in traditional robot population control are solved, and efficient, reliable and real-time collaborative control of robot population is achieved.
Patent Information
- Application Number
- CN202510446986.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional robot group control methods rely on central controllers, resulting in data transmission delay, high network bandwidth pressure, and high risk of single point failure, which affects real-time, synergy efficiency and system reliability, making it difficult to adapt to the growth of robot group size and dynamic environmental changes.
Adopting a distributed architecture based on edge computing, combined with deep reinforcement learning collaborative control algorithms and data compression technology, localized data processing and control instruction generation are realized, and collaborative control strategies of robot populations are optimized through multi-agent reinforcement learning and distributed training, and deep learning models are used for environmental modeling and prediction.
Significantly reduce data transmission delay, reduce network bandwidth pressure, improve system reliability and environmental adaptability, and enhance real-time performance of robot groups and task completion efficiency.
Smart Images

Figure BDA0005352889660000071 
Figure BDA0005352889660000072 
Figure BDA0005352889660000081
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robots, and specifically to a real-time control method for a group of robots based on edge computing. Background Art
[0002] With the rapid development of robot technology, the collaborative operation of a group of robots has shown great potential in fields such as disaster rescue, environmental monitoring, and military reconnaissance. However, traditional control methods for a group of robots usually rely on a central controller for centralized processing, and there are the following problems:
[0003] 1. Data transmission delay: In a centralized control architecture, the sensor data of all robots needs to be transmitted to the central controller for processing, and after generating control instructions, they are sent to each robot. This data transmission process will introduce significant delays. Especially when the number of robots is large, the amount of sensor data is large, or the network environment is poor, a large amount of sensor data needs to be transmitted to the central controller for processing, resulting in delayed control instructions and affecting the real-time performance and response speed of the group of robots. Specifically manifested as:
[0004] A) Reduced real-time performance: The delay of control instructions will cause the response speed of the group of robots to slow down and they cannot respond to environmental changes or task requirements in a timely manner.
[0005] B) Decreased collaborative efficiency: The collaborative operation of a group of robots requires a high degree of synchronization and coordination, and data transmission delay will disrupt this synchronization and reduce the collaborative efficiency.
[0006] C) Increased safety hazards: In a dynamic environment, data transmission delay may cause robots to be unable to avoid obstacles or make correct decisions in a timely manner, increasing safety hazards.
[0007] For example, in a disaster rescue scenario, a group of robots need to collaborate to search for survivors. If the data transmission delay is high, the robots may not be able to obtain the position information and environmental information of other robots in a timely manner, resulting in low search efficiency and even missing the rescue opportunity.
[0008] 2. Network bandwidth pressure: Simultaneous data transmission by multiple robots will cause huge pressure on the network bandwidth. Especially when high-precision sensors (such as lidar, cameras) generate a large amount of data, it may lead to data loss or network congestion.
[0009] Specifically manifested as:
[0010] A) Data loss: Insufficient network bandwidth may cause data packets to be lost, affecting the stability and reliability of the control system.
[0011] B) Network congestion: A large amount of data transmission may cause network congestion, further exacerbating the problem of data transmission delay.
[0012] C) Increased cost: To meet high - bandwidth requirements, high - performance network devices need to be deployed, increasing the system cost.
[0013] For example, in an environmental monitoring scenario, a swarm of robots needs to cooperate to collect environmental data (such as temperature, humidity, air quality, etc.). If the network bandwidth is insufficient, it may lead to incomplete or delayed data collection, affecting the accuracy and timeliness of monitoring results.
[0014] 3. Single - point failure risk: The centralized control architecture relies on a central controller for data processing and decision - making. Once the central controller fails, the entire swarm of robots will be paralyzed, and the system reliability is low.
[0015] Specifically manifested as:
[0016] A) Reduced system reliability: The single - point failure risk will reduce the reliability of the entire system, affecting the smooth completion of tasks.
[0017] B) Increased maintenance cost: Additional measures need to be taken to ensure the reliability of the central controller, such as redundant backup, fault diagnosis, etc., increasing the maintenance cost.
[0018] C) Limited scalability: The centralized control architecture is difficult to adapt to the growth of the swarm size of robots, with limited scalability.
[0019] For example, in a military reconnaissance scenario, a swarm of robots needs to cooperate to perform reconnaissance tasks. If the central controller fails, the entire reconnaissance task will be forced to interrupt, possibly resulting in serious consequences.
[0020] Moreover, the existing centralized control architecture is difficult to adapt to the growth of the swarm size of robots, with poor scalability. The centralized control architecture is difficult to adapt to dynamic environments and task requirements, lacking flexibility. All sensor data needs to be transmitted to the central controller, presenting risks of data leakage and privacy security.
[0021] Therefore, those skilled in the art have provided a real - time control method for a swarm of robots based on edge computing to solve the problems raised in the above - mentioned background technology. Summary of the Invention
[0022] The purpose of the present invention is to provide a real - time control method for a swarm of robots based on edge computing to solve the problems raised in the above - mentioned background technology.
[0023] To achieve the above - mentioned purpose, the present invention provides the following technical solutions:
[0024] A real - time control method for a swarm of robots based on edge computing includes the following steps:
[0025] Step 1): Deploy edge computing nodes on each robot using a distributed edge computing architecture, which is responsible for real-time processing of local sensor data and generation of control instructions. Communication and cooperation are carried out between the edge computing nodes through a wireless network to achieve distributed decision-making and control;
[0026] Step 2): Use a cooperative control algorithm based on deep reinforcement learning to enable each robot to autonomously learn a cooperative control strategy according to its own sensor data and the status information of surrounding robots.
[0027] This algorithm can adapt to changes in the dynamic environment and optimize the overall performance of the robot swarm, such as task completion efficiency, resource utilization, etc.;
[0028] Step 3): Adopt data compression and transmission optimization technologies to reduce the amount of sensor data transmitted, relieve the network bandwidth pressure, and dynamically adjust the data transmission priority according to the importance and real-time requirements of the data to ensure the timely transmission of critical data.
[0029] As a further solution of the present invention: The cooperative control algorithm based on deep reinforcement learning in the step 2) is designed and optimized for the cooperative control problem of the robot swarm, and its main features include:
[0030] a) Multi-agent reinforcement learning: Each robot acts as an agent and learns a cooperative control strategy through interaction with other agents;
[0031] b) Distributed training: Each edge computing node independently trains a local model and regularly shares model parameters with other nodes to achieve distributed learning and knowledge sharing;
[0032] c) Environment modeling and prediction: Use a deep learning model to model and predict the dynamic environment to improve the environmental adaptability and decision-making efficiency of the robot swarm.
[0033] As a further solution of the present invention: The specific formula of the cooperative control algorithm based on deep reinforcement learning in the step 2) is:
[0034] 1) Multi-agent reinforcement learning
[0035] a) State Space
[0036] The state s_i of each robot i includes its own sensor data (such as position, speed, attitude, etc.) and the status information of surrounding robots (such as relative position, speed, etc.);
[0037] The overall state space S = {s_1, s_2,..., s_N}, where N is the number of robots;
[0038] b) Action Space
[0039] The action \(a_i\) of each robot \(i\) includes control instructions (such as speed, direction, etc.);
[0040] The overall action space \(A=\{a_1,a_2,\cdots,a_N\}\);
[0041] c) Reward Function
[0042] The reward function \(R(s,a)\) is used to evaluate the overall performance of the robot swarm, such as task completion progress, resource utilization rate, collision avoidance, etc.;
[0043] The design of the reward function needs to be adjusted according to the specific application scenario;
[0044] d) Value Function
[0045] The state value function \(V(s)\) represents the long-term cumulative reward expectation of the robot swarm in state \(s\);
[0046] The action value function \(Q(s,a)\) represents the long-term cumulative reward expectation of the robot swarm after executing action \(a\) in state \(s\);
[0047] e) Policy
[0048] The policy \(\pi(a|s)\) represents the action selection probability distribution of the robot swarm in state \(s\);
[0049] The goal is to find the optimal policy \(\pi^*\) to maximize the long-term cumulative reward of the robot swarm;
[0050] 2) Distributed Training
[0051] a) Local Model Update
[0052] Each edge computing node \(i\) updates the local Q function \(Q_i(s,a)\) using local data \((s_i,a_i,r_i,s_i')\);
[0053] The update formula can adopt algorithms such as Q-learning or Deep Q-Network (DQN);
[0054] b) Model Parameter Sharing
[0055] Periodically upload the local model parameters \(\theta_i\) to the cloud or other edge computing nodes;
[0056] Adopt algorithms such as Federated Learning to achieve distributed aggregation and update of model parameters;
[0057] c) Global model update
[0058] Update the global Q - function \(Q_{global}(s,a)\) using the aggregated model parameter \(\theta_{global}\);
[0059] Each edge computing node regularly downloads the latest global model parameter \(\theta_{global}\) from the cloud or other edge computing nodes and updates the local model;
[0060] 3) Environment modeling and prediction
[0061] a) Environment model
[0062] Use a deep - learning model (such as LSTM, Transformer, etc.) to model the environment dynamics and predict the future state \(s'\);
[0063] The input of the environment model includes the current state \(s\) and the action \(a\), and the output is the predicted future state \(s'\);
[0064] b) Prediction error
[0065] Calculate the error between the predicted state \(s'\) and the actual state \(s''\) for optimizing the environment model;
[0066] The prediction error can use loss functions such as mean - squared error (MSE) or cross - entropy (Cross - Entropy).
[0067] As a further solution of the present invention: The algorithm flow of the collaborative control algorithm based on deep reinforcement learning in step 2) is as follows:
[0068] 1) Initialize the state \(s\), select an action \(a\) to execute;
[0069] 2) Observe the reward \(r\) and the next state \(s'\);
[0070] 3) Update the local Q - function \(Q_i(s,a)\);
[0071] 4) Regularly upload the local model parameter \(\theta_i\) and download the latest global model parameter \(\theta_{global}\);
[0072] 6) Use the environment model to predict the future state \(s'\) and calculate the prediction error;
[0073] 6) Repeat steps 1 - 5 until convergence.
[0074] As a further solution of the present invention: The specific formula of the collaborative control algorithm based on deep reinforcement learning in step 2) further includes:
[0075] 1) Multi - agent reinforcement learning
[0076] a) Q - learning Update Formula
[0077] For each robot i, the update formula for its local Q - function is as follows:
[0078] Qi(s,a)←Qi(s,a)+α[r + γmaxa′Qi(s′,a′)-Qi(s,a)]
[0079] Where:
[0080] Qi(s,a) is the action - value function for robot i to execute action a in state s;
[0081] α is the learning rate, which controls the update step size;
[0082] r is the current reward;
[0083] γ is the discount factor, which measures the importance of future rewards;
[0084] maxa′Qi(s′,a′) is the maximum Q - value in the next state s′;
[0085] b) Deep Q - Network (DQN) Loss Function
[0086] If a deep Q - network is used, the loss function is:
[0087] L(θi)=E(s,a,r,s′)~D[(r + γmaxa′Qi(s′,a′;θi - )-Qi(s,a;θi)) 2
[0088] Where:
[0089] θi are the local Q - network parameters of robot i;
[0090] θi - are the parameters of the target network (copied from θi regularly);
[0091] D is the experience replay buffer, which stores historical data (s,a,r,s′);
[0092] c) Multi - Agent Policy Gradient (MAPG)
[0093] For multi - agent cooperative control, the policy gradient method can be used. The policy gradient update formula for each robot i is:
[0094]
[0095] Where:
[0096] πi(ai∣si) is the policy function of robot i;
[0097] Qπ(s,a) is the joint action value function, representing the expected cumulative reward for executing the joint action a in state s;
[0098] 2) Distributed training
[0099] a) Local model update
[0100] The update formula for the local Q function of each robot i is:
[0101] Qi(s,a;θi)←Qi(s,a;θi)+α[r+γmaxa′Qi(s′,a′;θi-)-Qi(s,a;θi)]
[0102] b) Federated learning parameter aggregation
[0103] In distributed training, federated learning is used to aggregate model parameters. The update formula for the global model parameter θglobal is:
[0104]
[0105] Where:
[0106] N is the number of robots;
[0107] ni is the local data volume of robot i;
[0108] ni is the total data volume;
[0109] c) Local model synchronization
[0110] Each robot i periodically downloads the parameters from the global model:
[0111] θi←θglobal
[0112] 3) Environment modeling and prediction
[0113] a) Environment model prediction
[0114] Use a deep learning model (such as LSTM or Transformer) to model the environment dynamics. The prediction formula of the environment model is:
[0115]
[0116] Where:
[0117] is the predicted next state;
[0118] fenv is the environment model with parameters φ;
[0119] s is the current state;
[0120] a is the current action;
[0121] b) Environment model loss function
[0122] The training objective of the environment model is to minimize the prediction error, and the loss function is:
[0123]
[0124] where:
[0125] s′ is the actual next state;
[0126] is the predicted next state;
[0127] c) Q-function update incorporating the environment model
[0128] In the Q-function update, the environment model is used to generate simulated data:
[0129]
[0130] where is the next state predicted by the environment model.
[0131] As a further solution of the present invention: The algorithm flow of the collaborative control algorithm based on deep reinforcement learning in step 2) further includes:
[0132] 1) Initialization:
[0133] Initialize the local Q-function Q i (s,a;θ i ) and the environment model f env (s,a;φ) for each robot i;
[0134] Initialize the global model parameter θ global ;
[0135] 2) Local training:
[0136] Each robot i collects data (s,a,r,s′);
[0137] Update the local Q-function:
[0138] Q i (s,a;θ i )←Q i (s,a;θ i )+α[r+γmax a′ Q i (s′,a′;θ i - )-Q i (s,a;θ i )]
[0139] Update the environmental model:
[0140]
[0141] 3) Parameter aggregation:
[0142] Regularly upload the local parameter θ i to the cloud;
[0143] Update the global model parameters:
[0144]
[0145] 4) Model synchronization:
[0146] Each robot i downloads the latest global model parameters:
[0147] θ i ← θ global
[0148] 5) Repeat:
[0149] Repeat steps 2 - 4 until convergence.
[0150] Compared with the prior art, the beneficial effects of the present invention are:
[0151] 1. Reduce data transmission latency: Localized data processing and control instruction generation significantly reduce data transmission latency, improving the real - time performance and response speed of the swarm of robots.
[0152] 2. Alleviate network bandwidth pressure: Data compression and transmission optimization technologies effectively reduce network bandwidth occupancy, improving network resource utilization.
[0153] 3. Improve system reliability: The distributed architecture avoids the risk of single - point failure. Even if some nodes fail, other nodes can still work normally, improving the overall reliability of the system.
[0154] 4. Enhance environmental adaptability: The deep reinforcement learning algorithm enables the swarm of robots to autonomously learn to adapt to changes in the dynamic environment, improving the task completion efficiency and success rate. Specific implementation manner
[0155] In the embodiment of the present invention, the real - time control method for a swarm of robots based on edge computing includes the following steps:
[0156] Step 1), Deploy edge computing nodes on each robot using a distributed edge computing architecture, which is responsible for real - time processing of local sensor data and generation of control instructions, and communicate and cooperate between edge computing nodes through a wireless network to achieve distributed decision - making and control;
[0157] Step 2): Use the cooperative control algorithm based on deep reinforcement learning to enable each robot to autonomously learn the cooperative control strategy according to its own sensor data and the state information of surrounding robots.
[0158] This algorithm can adapt to the changes in the dynamic environment and optimize the overall performance of the robot swarm, such as task completion efficiency, resource utilization rate, etc.
[0159] Step 3): Adopt data compression and transmission optimization technology to reduce the amount of sensor data transmitted, relieve the network bandwidth pressure, dynamically adjust the data transmission priority according to the importance and real-time requirements of the data, and ensure the timely transmission of key data.
[0160] By adopting the above technical solutions, through the distributed edge computing architecture, deep reinforcement learning algorithm and data compression and transmission optimization technology, the problems of poor real-time performance, large network bandwidth pressure, low system reliability, etc. existing in the traditional methods are effectively solved, and it has important application value and development prospects.
[0161] Among them, the cooperative control algorithm based on deep reinforcement learning in the step 2) is designed and optimized for the cooperative control problem of the robot swarm, and its main features include:
[0162] a) Multi-agent reinforcement learning: Each robot acts as an agent and learns the cooperative control strategy through interaction with other agents.
[0163] b) Distributed training: Each edge computing node independently trains the local model and regularly shares the model parameters with other nodes to achieve distributed learning and knowledge sharing.
[0164] c) Environment modeling and prediction: Use the deep learning model to model and predict the dynamic environment to improve the environmental adaptability and decision-making efficiency of the robot swarm.
[0165] Among them, the specific formula of the cooperative control algorithm based on deep reinforcement learning in the step 2) is:
[0166] 1) Multi-agent reinforcement learning
[0167] a) State Space
[0168] The state s_i of each robot i includes its own sensor data (such as position, speed, attitude, etc.) and the state information of surrounding robots (such as relative position, speed, etc.).
[0169] The overall state space S = {s_1, s_2,..., s_N}, where N is the number of robots.
[0170] b) Action Space
[0171] The action ai of each robot i includes control instructions (such as speed, direction, etc.);
[0172] The overall action space A = {a1, a2,..., aN};
[0173] c) Reward Function
[0174] The reward function R(s, a) is used to evaluate the overall performance of the swarm of robots, such as task completion progress, resource utilization rate, collision avoidance, etc.;
[0175] The design of the reward function needs to be adjusted according to the specific application scenario;
[0176] d) Value Function
[0177] The state value function V(s) represents the long-term cumulative reward expectation of the swarm of robots in state s;
[0178] The action value function Q(s, a) represents the long-term cumulative reward expectation of the swarm of robots after executing action a in state s;
[0179] e) Policy
[0180] The policy π(a|s) represents the action selection probability distribution of the swarm of robots in state s;
[0181] The goal is to find the optimal policy π* to maximize the long-term cumulative reward of the swarm of robots;
[0182] 2) Distributed Training
[0183] a) Local Model Update
[0184] Each edge computing node i updates the local Q function Qi(s, a) using local data (si, ai, ri, si');
[0185] The update formula can adopt algorithms such as Q-learning or Deep Q-Network (DQN);
[0186] b) Model Parameter Sharing
[0187] Periodically upload the local model parameters θi to the cloud or other edge computing nodes;
[0188] Adopt algorithms such as Federated Learning to achieve distributed aggregation and update of model parameters;
[0189] c) Global model update
[0190] Update the global Q - function \(Q_{global}(s,a)\) using the aggregated model parameter \(\theta_{global}\);
[0191] Each edge computing node regularly downloads the latest global model parameter \(\theta_{global}\) from the cloud or other edge computing nodes and updates the local model;
[0192] 3) Environment modeling and prediction
[0193] a) Environment model
[0194] Use deep learning models (such as LSTM, Transformer, etc.) to model the environment dynamics and predict the future state \(s'\);
[0195] The input of the environment model includes the current state \(s\) and the action \(a\), and the output is the predicted future state \(s'\);
[0196] b) Prediction error
[0197] Calculate the error between the predicted state \(s'\) and the actual state \(s''\) for optimizing the environment model;
[0198] The prediction error can use loss functions such as mean - squared error (MSE) or cross - entropy (Cross - Entropy).
[0199] Among them, the algorithm flow of the collaborative control algorithm based on deep reinforcement learning in step 2) is as follows:
[0200] 1) Initialize the state \(s\), select an action \(a\) to execute;
[0201] 2) Observe the reward \(r\) and the next state \(s'\);
[0202] 3) Update the local Q - function \(Q_i(s,a)\);
[0203] 4) Regularly upload the local model parameter \(\theta_i\) and download the latest global model parameter \(\theta_{global}\);
[0204] 7) Use the environment model to predict the future state \(s'\) and calculate the prediction error;
[0205] 6) Repeat steps 1 - 5 until convergence.
[0206] Among them, the specific formula of the collaborative control algorithm based on deep reinforcement learning in step 2) further includes:
[0207] 1) Multi - agent reinforcement learning
[0208] a) Q - learning update formula
[0209] For each robot i, the update formula for its local Q-function is as follows:
[0210] Qi(s,a)←Qi(s,a)+α[r+γmaxa′Qi(s′,a′)-Qi(s,a)]
[0211] Where:
[0212] Qi(s,a) is the action-value function for robot i to execute action a in state s;
[0213] α is the learning rate, which controls the update step size;
[0214] r is the current reward;
[0215] γ is the discount factor, which measures the importance of future rewards;
[0216] maxa′Qi(s′,a′) is the maximum Q-value in the next state s′;
[0217] b) Deep Q-Network (DQN) loss function
[0218] If a deep Q-network is used, the loss function is:
[0219] L(θi)=E(s,a,r,s′)~D[(r+γmaxa′Qi(s′,a′;θi - )-Qi(s,a;θi)) 2
[0220] Where:
[0221] θi are the local Q-network parameters of robot i;
[0222] θi - are the parameters of the target network (copied from θi periodically);
[0223] D is the experience replay buffer, which stores historical data (s,a,r,s′);
[0224] c) Multi-Agent Policy Gradient (MAPG)
[0225] For multi-agent collaborative control, a policy gradient method can be used. The policy gradient update formula for each robot i is:
[0226]
[0227] Where:
[0228] πi(ai∣si) is the policy function of robot i;
[0229] Qπ(s,a) is the joint action value function, representing the expected cumulative reward for executing the joint action a in state s;
[0230] 2) Distributed training
[0231] a) Local model update
[0232] The update formula for the local Q-function of each robot i is:
[0233] Qi(s,a;θi) ← Qi(s,a;θi) + α[r + γmaxa′Qi(s′,a′;θi - ) - Qi(s,a;θi)]
[0234] b) Federated learning parameter aggregation
[0235] In distributed training, federated learning is used to aggregate model parameters. The update formula for the global model parameter θglobal is:
[0236]
[0237] Where:
[0238] N is the number of robots;
[0239] ni is the local data volume of robot i;
[0240] ni is the total data volume;
[0241] c) Local model synchronization
[0242] Each robot i periodically downloads the parameters from the global model:
[0243] θi ← θglobal
[0244] 3) Environment modeling and prediction
[0245] a) Environment model prediction
[0246] Use a deep learning model (such as LSTM or Transformer) to model the environment dynamics. The prediction formula for the environment model is:
[0247]
[0248] Where:
[0249] is the predicted next state;
[0250] fenv is the environment model with parameters φ;
[0251] s is the current state;
[0252] a is the current action;
[0253] b) Environment model loss function
[0254] The training objective of the environment model is to minimize the prediction error, and the loss function is:
[0255]
[0256] Where:
[0257] s′ is the actual next state;
[0258] is the predicted next state;
[0259] c) Q - function update combined with the environment model
[0260] In the Q - function update, the environment model is used to generate simulated data:
[0261]
[0262] Where is the next state predicted by the environment model.
[0263] Among them, the algorithm flow of the collaborative control algorithm based on deep reinforcement learning in step 2) further includes:
[0264] 1) Initialization:
[0265] Initialize the local Q - function Q i (s,a;θ i ) and the environment model f env (s,a;φ) for each robot i;
[0266] Initialize the global model parameter θ global ;
[0267] 2) Local training:
[0268] Each robot i collects data (s,a,r,s′);
[0269] Update the local Q - function:
[0270] Q i (s,a;θ i )←Q i (s,a;θ i )+α[r + γmax a′ Q i (s′,a′;θ i - ) - Q i (s,a;θ i )]
[0271] Update the environmental model:
[0272]
[0273] 3) Parameter aggregation:
[0274] Regularly upload the local parameter θ i to the cloud;
[0275] Update the global model parameters:
[0276]
[0277] 4) Model synchronization:
[0278] Each robot i downloads the latest global model parameters:
[0279] θ i ←θ global
[0280] 5) Repeat:
[0281] Repeat steps 2 - 4 until convergence.
[0282] After adopting the above technical solutions, local decision-making is achieved through edge computing, reducing data transmission latency; distributed training and federated learning are adopted to avoid single-point failures and enhance system reliability; through environmental modeling and prediction, the swarm of robots can better adapt to the dynamic environment; multi-agent reinforcement learning optimizes the collaborative control strategy of the swarm of robots, which can effectively improve the task completion efficiency.
[0283] The present invention can be widely applied to the following scenarios:
[0284] Disaster rescue: The swarm of robots collaborates to search for survivors, transport relief supplies, build temporary shelters, etc.
[0285] Environmental monitoring: The swarm of robots collaborates to monitor environmental pollution, collect environmental data, draw environmental maps, etc.
[0286] Military reconnaissance: The swarm of robots collaborates to perform reconnaissance tasks, collect intelligence, conduct target tracking, etc.
[0287] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.
Claims
1. A real-time control method for a swarm of robots based on edge computing, characterized in that: It includes the following steps: Step 1): Deploy edge computing nodes on each robot using a distributed edge computing architecture, which is responsible for real-time processing of local sensor data and generation of control instructions. Communication and cooperation are carried out between the edge computing nodes through a wireless network to achieve distributed decision-making and control; Step 2): Use a cooperative control algorithm based on deep reinforcement learning to enable each robot to autonomously learn a cooperative control strategy according to its own sensor data and the state information of surrounding robots. This algorithm can adapt to changes in the dynamic environment and optimize the overall performance of the robot swarm, such as task completion efficiency and resource utilization; Step 3): Adopt data compression and transmission optimization technologies to reduce the amount of sensor data transmitted, relieve the network bandwidth pressure, and dynamically adjust the data transmission priority according to the importance and real-time requirements of the data to ensure the timely transmission of key data.
2. The real-time control method for a swarm of robots based on edge computing according to claim 1, characterized in that: The specific formula of the cooperative control algorithm based on deep reinforcement learning in the said Step 2) is as follows: 1) Multi-agent reinforcement learning a) State Space The state s_i of each robot i includes its own sensor data and the state information of surrounding robots; The overall state space S = {s_1, s_2,..., s_N}, where N is the number of robots; b) Action Space The action a_i of each robot i includes control instructions; The overall action space A = {a_1, a_2,..., a_N}; c) Reward Function The reward function R(s, a) is used to evaluate the overall performance of the robot swarm, such as task completion progress, resource utilization, and collision avoidance; The design of the reward function needs to be adjusted according to the specific application scenario; d) Value Function The state value function V(s) represents the long-term cumulative reward expectation of the robot swarm in state s; The action value function Q(s, a) represents the long-term cumulative reward expectation of the robot swarm after executing action a in state s; e) Policy The policy π(a|s) represents the action selection probability distribution of the robot swarm in state s; The goal is to find the optimal policy π* to maximize the long-term cumulative reward of the robot swarm; 2) Distributed training a) Local model update Each edge computing node i updates the local Q function Q_i(s, a) using local data (s_i, a_i, r_i, s_i'); The update formula can adopt the Q-learning or Deep Q-Network (DQN) algorithm; b) Model parameter sharing Periodically upload the local model parameters θ_i to the cloud or other edge computing nodes; Adopt the Federated Learning algorithm to achieve distributed aggregation and update of model parameters; c) Global model update Use the aggregated model parameters θ_global to update the global Q function Q_global(s, a); Each edge computing node regularly downloads the latest global model parameters θ_global from the cloud or other edge computing nodes and updates the local model; 3) Environment Modeling and Prediction a) Environment Model Use a deep learning model to model the environmental dynamics and predict the future state s'; The input of the environment model includes the current state s and the action a, and the output is the predicted future state s'; b) Prediction Error Calculate the error between the predicted state s' and the actual state s'', which is used to optimize the environment model; The prediction error can adopt the mean squared error (MSE) or cross-entropy loss function.
3. The real-time control method for a swarm of robots based on edge computing according to claim 1, wherein: The algorithm flow of the collaborative control algorithm based on deep reinforcement learning in step 2) is as follows: 1) Initialize the state s and select an action a to execute; 2) Observe the reward r and the next state s'; 3) Update the local Q function Q_i(s,a); 4) Regularly upload the local model parameters θ_i and download the latest global model parameters θ_global; 5) Use the environment model to predict the future state s' and calculate the prediction error; 6) Repeat steps 1-5 until convergence.
4. The real-time control method for a swarm of robots based on edge computing according to claim 1, wherein: The specific formula of the collaborative control algorithm based on deep reinforcement learning in step 2) further includes: 1) Multi-Agent Reinforcement Learning a) Q-learning Update Formula For each robot i, the update formula of its local Q function is: Qi(s,a)←Qi(s,a)+α[r+γmaxa′Qi(s′,a′)-Qi(s,a)] Where: Qi(s,a) is the action value function of robot i executing action a in state s; α is the learning rate, which controls the update step size; r is the current reward; γ is the discount factor, which measures the importance of future rewards; maxa′Qi(s′,a′) is the maximum Q value in the next state s′; b) Deep Q Network (DQN) Loss Function If a deep Q network is used, the loss function is: L(θi) = E(s,a,r,s′) ∼ D[(r + γ maxa′Qi(s′,a′;θi−) - Qi(s,a;θi)) 2 Where: θi is the local Q network parameter of robot i; θi - are the parameters of the target network (copied from θi periodically); D is the experience replay buffer, which stores historical data (s,a,r,s′); c) Multi-Agent Policy Gradient (MAPG) For multi-agent collaborative control, a policy gradient method can be used. The policy gradient update formula for each robot i is: Where: πi(ai∣si) is the policy function of robot i; Qπ(s,a) is the joint action value function, which represents the expected cumulative reward of executing the joint action a in state s; 2) Distributed Training a) Local Model Update The update formula of the local Q function for each robot i is: Qi(s,a;θi)←Qi(s,a;θi)+α[r+γmaxa′Qi(s′,a′;θi - )-Qi(s,a;θi)] b) Federated Learning Parameter Aggregation In distributed training, federated learning is used to aggregate the model parameters. The update formula of the global model parameters θglobal is: Where: N is the number of robots; ni is the local data volume of robot i; ni is the total amount of data; c) Local Model Synchronization Each robot i regularly downloads the parameters from the global model: θi←θglobal 3) Environment Modeling and Prediction a) Environment Model Prediction Use a deep learning model (such as LSTM or Transformer) to model the environmental dynamics. The prediction formula of the environment model is: Where: is the predicted next state; The fenv is an environment model with parameter φ; s is the current state; a is the current action; b) Environment model loss function The training objective of the environment model is to minimize the prediction error, and the loss function is: Where: s′ is the actual next state; is the predicted next state; c) Q-function update combined with the environment model In the Q-function update, simulated data is generated using the environment model: wherein is the next state predicted by the environmental model.
5. The real-time control method for a swarm of robots based on edge computing according to claim 1, characterized in that: The algorithm flow of the collaborative control algorithm based on deep reinforcement learning in step 2) further includes: 1) Initialization: Initialize the local Q - function Qi(s,a;θi) and the environment model f env (s,a;φ); Initialize the global model parameter θ global ; 2) Local training: Each robot i collects data (s, a, r, s′); Update the local Q-function: Q i (s, a; θ i ) ← Q i (s, a; θ i ) + α[r + γ max a′ Q i (s′, a′; θ i - ) - Q i (s, a; θ i )] Update the environment model: 3) Parameter aggregation: Regularly upload the local parameter θ i to the cloud; Update the global model parameters: 4) Model synchronization: Each robot i downloads the latest global model parameters: θ i ←θ global 5) Repeat: Repeat steps 2 - 4 until convergence.
Citation Information
Cited By
Interaction control method and system for intelligent device with body
CN122363939A