Multi-robot control and communication collaborative optimization method based on nonlinear information age
By constructing a nonlinear information age model and an adaptive handshake frequency hopping mechanism, combined with multi-agent reinforcement learning, the communication and control of multi-robot systems are optimized, solving the problems of channel conflict, information timeliness and energy consumption in multi-robot collaborative systems, and improving the collaborative performance and stability of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-03-31
AI Technical Summary
Multi-robot collaborative systems face challenges such as communication delays, information lags, and inefficient collaboration. Traditional methods have failed to effectively address issues like channel conflicts, insufficient information timeliness, and excessive energy consumption. Furthermore, traditional linear information aging models cannot accurately characterize the nonlinear exacerbation effect of delays on system performance.
A collaborative optimization method for multi-robot control and communication based on nonlinear information age is constructed. By building a control-communication coupling model, designing an adaptive handshake frequency hopping mechanism and a nonlinear AoI penalty function, and combining it with a multi-agent reinforcement learning algorithm, communication efficiency, control stability and energy consumption are optimized.
It achieves global collaborative optimization of multi-robot systems, improving reliability, real-time performance and economy in complex scenarios, and is suitable for multi-robot collaborative operation scenarios such as industrial manufacturing and intelligent warehousing.
Smart Images

Figure CN121764052A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-robot cooperative control technology, and more specifically, to a multi-robot control and communication cooperative optimization method based on nonlinear information age. Background Technology
[0002] With the rapid development of industrial intelligence and logistics automation technologies, intelligent robots have become core equipment for material handling, path tracking, and collaborative operations in industrial manufacturing, intelligent warehousing, and other scenarios. In this process, the collaborative control and wireless communication requirements of multiple robots have become more complex. They not only need to adapt to dynamically changing operating environments (such as dense equipment and signal interference), but also need to handle real-time information exchange, motion state synchronization, and task collaborative execution between robots. In modern industrial and logistics scenarios, the communication reliability and control stability of multiple robots are of great significance for improving operational efficiency, ensuring system safety, and reducing operating costs.
[0003] Currently, many multi-robot collaborative scenarios suffer from frequent channel conflicts, insufficient information timeliness, and excessive energy consumption in communication and control coordination. For example, different robots have significantly different requirements for communication bandwidth and transmission power, while channel resources are subject to competition and sharing; dynamic changes in robot motion states (speed, position) can lead to communication link instability, resulting in signal attenuation and data loss; information lag between multiple robots directly affects control accuracy, leading to path deviations or task execution errors. Furthermore, the coupling effect of communication delay, information staleness, and control errors has not been fully quantified and optimized. Under conditions of high interference and multi-device collaboration, designing a collaborative optimization method that balances communication efficiency, control stability, and energy economy is a significant technical challenge.
[0004] Traditional multi-robot control methods typically optimize the motion accuracy of a single robot, focusing only on trajectory tracking errors or local motion control. This fails to comprehensively reflect the interaction between inter-robot communication states and global collaborative performance. Traditional communication optimization methods often focus on channel rate or delay optimization, neglecting the impact of communication parameter adjustments on robot control performance. While information age is used to measure information timeliness, traditional linear information age models cannot accurately characterize the "nonlinear aggravation effect" of delay on system performance. When delay exceeds a threshold, information staleness leads to an exponential increase in control error, and it is difficult to comprehensively consider the coupling relationship between robot motion states, channel dynamics, and control errors. Therefore, multi-robot system performance optimization faces bottlenecks. Summary of the Invention
[0005] To address the challenges of communication delay, information lag, and efficient collaboration in existing multi-robot collaborative systems, this invention proposes a multi-robot control and communication collaborative optimization method based on nonlinear information age. By constructing a control-communication coupling model, designing an adaptive handshake frequency hopping mechanism, and a nonlinear AoI penalty function, global collaborative optimization of communication efficiency, control stability, and energy consumption is achieved.
[0006] This invention first constructs a multi-robot dynamics model, communication model, nonlinear information aging model, and control model, forming a deeply coupled control-communication framework to accurately characterize the interaction between robot motion state, communication link characteristics, and information timeliness. Second, it designs an adaptive handshake frequency hopping mechanism, dynamically adjusting the handshake cycle based on robot motion and communication states, and employing a complementary strategy of multi-channel frequency hopping and time-division multiplexing to reduce channel conflicts. Then, it applies an exponential penalty to high-latency information using a nonlinear AoI penalty function, prioritizing the updating of critical timeliness data. Subsequently, it constructs a constrained collaborative optimization problem with the objectives of minimizing system energy consumption, nonlinear information aging penalty, and maximizing control stability. Finally, the problem is transformed into a partially observable Markov decision process (POMDP), solved using a multi-agent reinforcement learning algorithm based on nearest neighbor policy optimization (PPO) (PPO-CCBNA), outputting the optimal communication and control strategy. This invention achieves global collaborative optimization of multi-robot communication efficiency, control stability, and energy consumption, significantly improving the reliability, real-time performance, and economy of multi-robot collaborative operations in complex scenarios, and is applicable to multi-robot collaborative operation scenarios such as industrial manufacturing and intelligent warehousing.
[0007] The technical solution adopted by this invention to achieve the above objectives is: a multi-robot control and communication cooperative optimization method based on nonlinear information age, comprising the following steps:
[0008] System modeling: Construct a multi-robot dynamics model, communication model, nonlinear information aging model, and control model to form a multi-robot system framework that couples control and communication;
[0009] Construction of an adaptive handshake frequency hopping mechanism: Based on the robot's motion state and communication state, an adaptive handshake cycle, a multi-channel frequency hopping strategy, and a time-division multiplexing complementary mechanism are constructed to reduce channel conflicts;
[0010] Nonlinear information aging optimization: Construct a nonlinear AoI penalty function to guide a multi-robot system to prioritize updating high-latency information;
[0011] Cooperative optimization problem construction: To minimize the energy consumption of the multi-machine system, the nonlinear AoI penalty, and maximize the control stability, a constrained optimization problem is constructed.
[0012] Multi-agent reinforcement learning solution: The optimization problem is transformed into a partially observable Markov decision process, which is solved by a cooperative optimization algorithm based on nearest neighbor policy optimization, and the optimal communication and control policy is output.
[0013] The system modeling includes:
[0014] (1.1) Dynamic model:
[0015]
[0016] in, For the first A robot is positioned in a two-dimensional space in time slot t. For the robot's facing angle, and These are the robot's linear velocity and angular velocity, respectively. For time step;
[0017] (1.2) The specific communication model is as follows:
[0018] Communication rate calculation:
[0019]
[0021] in, For the first The robot and the first A robot in the time slot Channel The handshake success indicator For channel bandwidth, For transmission power, For channel gain, For noise power, For data block length, Block error rate, It is the inverse Gaussian function. For channel dispersion, These represent the robot's currently selected communication channel and the entire set of available channels, respectively. Both represent robot serial numbers, and: ;
[0022] Communication delay calculation:
[0023]
[0024] in, For data group size, Minimal positive numbers to ensure Not zero;
[0025] Communication energy consumption calculation:
[0026]
[0027] (1.3) Construction of nonlinear information age model:
[0028] AoI definition: , For the first The robot moved towards the first The generation time of the information sent by each robot; The current time;
[0029] Non-linear AoI penalty function: ,right An exponential penalty is imposed for exceeding the threshold, i.e. , This is the penalty coefficient;
[0030] (1.4) Control model construction:
[0031] State estimation delay: ;in, The current age of the information;
[0032] State estimation:
[0033] in, For the first A robot in the time slot The estimated location;
[0034] Trajectory error:
[0035] in, For the first The reference trajectory of the robot;
[0036] Control Law:
[0037] in, For proportional gain; , These are the coordinates for the trajectory error.
[0038] Control stability function:
[0039] in, As a weighting factor, Use extremely small positive numbers to ensure Bounded.
[0040] The adaptive handshake frequency hopping mechanism design includes the following steps:
[0041] (2.1) Adaptive handshake cycle adjustment: The handshake cycle is dynamically adjusted by combining robot speed, position changes, nonlinear AoI, and control stability.
[0042]
[0043] in, As a weighting factor, For the first Adjustment coefficient for the handshake cycle of each robot;
[0044] (2.2) Multi-channel frequency hopping strategy: Construct a handshake success determination function by combining channel state and robot orientation angle:
[0045]
[0046] in, The function value indicates whether the handshake was successful. This represents the channel selected by the i-th robot. For the first The robot pointed to the first The robot's orientation angle, The threshold for determining the orientation angle. This is a handshake frequency counter;
[0047] (2.3) Time Division Multiplexing Complementary: When the channel collision rate is detected to exceed the set threshold, it automatically switches to time division multiplexing mode and allocates dedicated time slots for different robot communication pairs on the time axis.
[0048] The nonlinear information age optimization, through a nonlinear information age penalty function, specifically includes:
[0049] Set AoI threshold ;
[0050] when When applying a linear penalty: ; The information's age is represented by AoI;
[0051] when At that time, apply an exponential penalty: ,in, , Let be the coefficient, and , This is the penalty coefficient.
[0052] The construction of the collaborative optimization problem includes the following steps:
[0053] (1) Construct the objective function:
[0054]
[0055] in, The total number of robots, For the weighting coefficients, satisfying , This represents the nonlinear AoI penalty function. Represents the control stability function;
[0056] (2) Constructing constraints:
[0057] Transmit power constraints: , This represents the robot's maximum transmission power.
[0058] Handshake cycle constraints: , This is the maximum handshake period;
[0059] Location boundary constraints: ; , For the minimum boundary coordinates, , The coordinates of the maximum boundary;
[0060] Speed constraints: , For the maximum linear velocity, This is the maximum angular velocity;
[0061] Time constraints: , This is the maximum duration of the task.
[0062] The multi-agent deep reinforcement learning solution includes the following steps:
[0063] (5.1) Definition of state space:
[0064]
[0065] in, For the first The currently selected channel for the robot For the first The average information age of each robot;
[0066] (5.2) Action Space Construction:
[0067]
[0068] in, , Set of available channels This represents the power allocated to the i-th robot; ;
[0069] (5.3) Construction of reward function:
[0070]
[0071] in, Penalty for failed handshake matching Penalty for load imbalance For channel collision penalty, These are the weighting coefficients; This represents the control stability function. Indicates communication energy consumption. This represents the nonlinear AoI penalty function;
[0072] (5.4) Perform multi-agent deep reinforcement learning training and execution; the optimal action obtained is used as the optimal communication and control strategy.
[0073] The objective function for multi-agent deep reinforcement learning training and execution is as follows:
[0074] The objective function of the Actor network is:
[0075]
[0076] The objective function of the Critic network is:
[0077]
[0078] in, For batch size, The total number of robots, For the duration of trajectory collection, This is the cutting factor. The regularization coefficient is the policy entropy. Let entropy be the policy function. The value of the normalized state; Indicates the update strategy ratio. This represents the generalized advantage estimation. This represents the clipping function. The coefficients of the policy entropy regularization term are represented. Represents the value function. This represents the value function after normalization.
[0079] A multi-robot control and communication cooperative optimization system based on nonlinear information age includes:
[0080] The system modeling module is used to construct multi-robot dynamics models, communication models, nonlinear information aging models, and control models, forming a system framework that couples control and communication.
[0081] The handshake frequency hopping mechanism module is used to construct an adaptive handshake cycle, a multi-channel frequency hopping strategy, and a time-division multiplexing complementary mechanism based on the robot's motion state and communication state, so as to reduce channel conflicts.
[0082] The nonlinear information age optimization module is used to guide the system to prioritize updating high-latency information through a nonlinear AoI penalty function;
[0083] The optimization problem construction module is used to construct constrained optimization problems with the objectives of minimizing system energy consumption, nonlinear information age penalty, and maximizing control stability.
[0084] The multi-agent deep reinforcement learning optimization module is used to transform the optimization problem into a partially observable Markov decision process. It solves the problem through a cooperative optimization algorithm based on nearest neighbor policy optimization and outputs the optimal communication and control strategy to guide the robot to execute in real time.
[0085] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the multi-robot control and communication cooperative optimization method based on nonlinear information age.
[0086] The present invention has the following beneficial effects and advantages:
[0087] 1. This invention constructs a deeply coupled control-communication model, integrating robot dynamics, state estimation, communication lag, and control error models to accurately capture the mutual influence between motion control and communication performance, overcoming the limitations of traditional single-dimensional optimization. Compared to methods that only optimize communication rate or control accuracy, it effectively avoids the contradiction of "efficient communication but inaccurate control" or "stable control but delayed communication," significantly improving the global collaborative performance of multi-robot collaborative systems.
[0088] 2. This invention introduces a nonlinear information age penalty mechanism, which, compared to traditional linear information age, more accurately characterizes the "nonlinear aggravation effect" of latency on system performance. It imposes an exponential penalty on high-latency information, guiding the system to prioritize the updating of critical time-sensitive data. Simultaneously, by combining adaptive handshake frequency hopping and time-division multiplexing complementary strategies, it dynamically adjusts channel selection, handshake period, and transmit power, effectively reducing channel collision rate, lowering communication latency, and ensuring information freshness and communication reliability for multiple robots in high-interference environments.
[0089] 3. This invention achieves optimization solutions based on multi-agent deep reinforcement learning, employing a "centralized training, distributed execution" architecture: centralized training ensures global policy optimization, while distributed execution meets the robot's real-time decision-making needs. The algorithm balances training stability and exploration efficiency through generalized advantage estimation, PPO-Clip constraints, and policy entropy terms. Even in scenarios with dynamically changing robot numbers, it maintains control stability, reduces system energy consumption, and exhibits strong robustness and scalability, making it suitable for complex multi-robot collaborative scenarios such as industrial manufacturing and intelligent logistics. Attached Figure Description
[0090] Figure 1 This is a schematic diagram of the method of the present invention;
[0091] Figure 2 This is a diagram illustrating the algorithm architecture of an embodiment of the present invention. Detailed Implementation
[0092] The present invention will now be described in detail with reference to the accompanying drawings.
[0093] This invention provides a multi-robot control and communication collaborative optimization method and system based on nonlinear information aging. It offers a novel optimization scheme to address the technical challenges of communication delay, information lag, frequent channel conflicts, and insufficient control and communication coordination in multi-robot collaborative tasks. By constructing a deeply coupled control-communication model, integrating robot dynamics, state estimation, communication lag, and control error models, the system accurately captures the mutual influence between motion control and communication performance. An adaptive handshake frequency hopping mechanism and a nonlinear information age penalty function are designed to prioritize updating key time-sensitive data and reduce channel conflicts. Combined with a multi-agent deep reinforcement learning framework, the communication and control strategies are optimized, further improving the system's robustness, real-time performance, and energy economy. This invention is applicable to multi-robot collaborative operation scenarios such as industrial manufacturing and intelligent warehousing, meeting the requirements for high real-time performance, anti-interference capabilities, and global collaborative optimization.
[0094] This invention addresses tasks such as material handling and path tracking performed collaboratively by multiple robots in industrial control scenarios. By performing system modeling, parameter calculation, and algorithm training within an edge server, and utilizing a multi-agent deep reinforcement learning framework, it optimizes parameters such as robot channel selection, transmission power, and handshake cycle, thereby achieving coordinated optimization of control and communication.
[0095] like Figure 1 As shown, a multi-robot control and communication cooperative optimization method based on nonlinear information age includes the following steps:
[0096] 1) System modeling: Construct a multi-robot dynamics model, communication model, nonlinear information age model and control model to form a system framework with deep control-communication coupling.
[0097] 2) Adaptive handshake frequency hopping mechanism design: Based on the robot's motion state and communication state, an adaptive handshake cycle, multi-channel frequency hopping strategy, and time-division multiplexing complementary mechanism are designed to reduce channel conflicts.
[0098] 3) Nonlinear information age optimization: The system is guided to prioritize updating high-latency information through a nonlinear AoI penalty function.
[0099] 4) Construction of Cooperative Optimization Problem: To minimize system energy consumption, nonlinear AoI penalty and maximize control stability, a constrained optimization problem is constructed.
[0100] 5) Multi-agent reinforcement learning solution: The optimization problem is transformed into a partially observable Markov decision process (POMDP), which is solved by a cooperative optimization algorithm based on nearest neighbor policy optimization (PPO-CCBNA) to output the optimal communication and control policy.
[0101] In step 1), the system modeling includes the following steps:
[0102] (1.1) Dynamics Model Construction: Describes the evolution of the robot's position, attitude, and motion state, expressed as:
[0103]
[0104] in, For the first A robot is positioned in a two-dimensional space in time slot t. For the robot's facing angle, and These are the robot's linear velocity and angular velocity, respectively. For time step.
[0105] (1.2) Communication Model Construction: Using a multi-channel frequency hopping device-to-device (D2D) communication method, combined with a finite block length transmission model, the actual communication rate, latency, and energy consumption between robots are calculated.
[0106] Communication rate calculation:
[0107]
[0109] in, For the first The robot and the first A robot in the time slot Channel The handshake success indicator ( This indicates success. (indicating failure) For channel bandwidth, For transmission power, For channel gain, For noise power, For data block length, Block error rate, It is the inverse Gaussian function. For channel dispersion, and: .
[0110] Communication delay calculation:
[0111]
[0112] in, For data group size, It is a very small positive number (to ensure that the denominator is non-zero and improve the robustness of numerical calculation).
[0113] Communication energy consumption calculation:
[0114]
[0115] (1.3) Construction of Nonlinear Information Age Model
[0116] Define nonlinear information age, encompassing the entire process of "information generation - communication pairing - network establishment - data transmission," and quantify the nonlinear negative impact of information obsolescence:
[0117] Basic AoI definition: ( For the first The robot moved towards the first (Generation time of information sent by each robot);
[0118] Non-linear AoI penalty function: ,right Apply an exponential penalty (e.g., if the threshold is exceeded) to cases where the threshold is exceeded. , (As a penalty coefficient), it guides the system to prioritize updating high-latency information.
[0119] (1.4) Control Model Construction: Based on the state estimation delay and trajectory error, the control law is designed, and the control stability function is introduced to evaluate the control performance.
[0120] State estimation delay:
[0121] in, The current age of the information;
[0122] State estimation:
[0123] in, For the first A robot in the time slot The estimated location;
[0124] Trajectory error:
[0125] in, For the first The reference trajectory of the robot;
[0126] Control Law:
[0127] in, For proportional gain;
[0128] Control stability function:
[0129] in, As a weighting factor, It is a very small positive number (ensuring that the denominator is bounded).
[0130] Step 2) involves the following steps in designing the adaptive handshake frequency hopping mechanism:
[0131] (2.1) Adaptive handshake cycle adjustment: The handshake cycle is dynamically adjusted by combining robot speed, position changes, nonlinear AoI, and control stability.
[0132]
[0133] in, As a weighting factor, For the first The handshake cycle adjustment coefficient for each robot (the smaller the value, the shorter the handshake cycle).
[0134] (2.2) Multi-channel frequency hopping strategy: Design a handshake success determination criterion by combining channel state and robot orientation angle to avoid channel conflicts.
[0135]
[0136] in, For the first The robot pointed to the first The robot's orientation angle, The threshold for determining the orientation angle. This is a handshake frequency counter.
[0137] (2.3) Time Division Multiplexing Complementary: When the system detects that the channel conflict rate exceeds the set threshold, it automatically switches to time division multiplexing mode, allocates dedicated time slots for different robot communication pairs on the time axis, avoids time domain conflicts, and ensures the real-time transmission of key task information.
[0138] In step 3), the nonlinear information age optimization is specifically as follows:
[0139] The system is guided to prioritize updating high-latency information through a non-linear information age penalty function, specifically including:
[0140] Set AoI threshold (Configure according to the real-time requirements of the task, such as settings for industrial material handling scenarios;)
[0141] when When applying a linear penalty (such as...) );when When, apply an exponential penalty (such as) ),in, The penalty coefficient is... This ensures that high-latency information is processed first.
[0142] Step 4) involves constructing the collaborative optimization problem, which includes the following steps:
[0143] (1) Construction of the objective function:
[0144]
[0145] in, The total number of robots, Weighting coefficients (satisfying) Such as industrial scene design Prioritize ensuring control stability).
[0146] (2) Construction of constraints:
[0147] C1: Transmit power constraint ( (This refers to the robot's maximum transmission power).
[0148] C2: Handshake Cycle Constraints ( (Maximum handshake period);
[0149] C3: Location Boundary Constraints ;
[0150] C4: Speed Constraint ( For the maximum linear velocity, (Maximum angular velocity);
[0151] C5: Time Constraints ( (Maximum duration of the task).
[0152] like Figure 2As shown, step 5), the multi-agent deep reinforcement learning solution, includes the following steps:
[0153] (5.1) State space definition: includes robot motion state, channel state, and information age, comprehensively reflecting the system dynamics:
[0154]
[0155] in, For the first The currently selected channel for the robot For the first The average information age of each robot.
[0156] (5.2) Action Space Design: Focusing on key decision variables affecting communication and control performance, including channel selection and transmit power control:
[0157]
[0158] in, ( (Set of available channels) .
[0159] (5.3) Reward Function Design: Taking into account control stability, energy consumption, nonlinear AoI penalty, and various conflict penalties, the agent is guided to learn the globally optimal policy.
[0160]
[0161] in, Penalty for failed handshake matching Penalty for load imbalance Penalty for channel collisions.
[0162] (5.4) Algorithm Training and Execution: An Actor-Critic architecture of "centralized training and distributed execution" is adopted.
[0163] Actor network: Shares parameters and outputs channel selection and transmit power decisions for each robot;
[0164] Critic Network: Evaluates the value of the current state by calculating the advantage function through generalized advantage estimation (GAE), thereby improving training stability. GAE is defined as follows:
[0165]
[0166] Time error for:
[0167] in As a discount factor, The state value function;
[0168] Strategy optimization: The PPO-Clip technique is used to constrain the policy update magnitude, and the policy update ratio is:
[0169] The objective function of the Actor network is:
[0170]
[0171] The objective function of the Critic network is:
[0172]
[0173] in, For batch size, The total number of robots, For the duration of trajectory collection, This is the cutting factor. The regularization coefficient is the policy entropy. Let entropy be the policy function. The value of the normalized state; Indicates the update strategy ratio. This represents the generalized advantage estimation. This represents the clipping function. The coefficients of the policy entropy regularization term are represented. Represents the value function. This represents the value function after normalization. Execution phase: Each robot makes distributed decisions based on its trained strategy, adjusting communication and control parameters in real time to achieve dynamic collaborative optimization.
[0174] A multi-robot control and communication cooperative optimization system based on nonlinear information age includes:
[0175] The system modeling module is used to construct multi-robot dynamics models, communication models, nonlinear information aging models, and control models, forming a control-communication coupling framework and providing basic model support for subsequent optimization.
[0176] The handshake frequency hopping mechanism module is used to design adaptive handshake cycles, multi-channel frequency hopping strategies, and time-division multiplexing complementary mechanisms to reduce channel conflicts and improve communication reliability and real-time performance.
[0177] The nonlinear information aging optimization module is used to design a nonlinear information aging penalty function, which prioritizes the updating of key time-sensitive information and quantifies and mitigates the negative impact of information obsolescence on system performance.
[0178] The optimization problem construction module is used to construct constrained optimization problems with the objectives of minimizing system energy consumption, nonlinear information age penalty, and maximizing control stability, and to clarify the optimization direction and boundary conditions.
[0179] The multi-agent deep reinforcement learning optimization module is used to transform the optimization problem into a POMDP problem. It trains and outputs the optimal communication and control strategy through a PPO-based collaborative optimization algorithm to guide the robot to execute in real time.
Claims
1. A multi-robot control and communication cooperative optimization method based on nonlinear information age, characterized in that, Includes the following steps: System modeling: Construct a multi-robot dynamics model, communication model, nonlinear information aging model, and control model to form a multi-robot system framework that couples control and communication; Construction of an adaptive handshake frequency hopping mechanism: Based on the robot's motion state and communication state, an adaptive handshake cycle, a multi-channel frequency hopping strategy, and a time-division multiplexing complementary mechanism are constructed to reduce channel conflicts; Nonlinear information aging optimization: Construct a nonlinear AoI penalty function to guide a multi-robot system to prioritize updating high-latency information; Cooperative optimization problem construction: To minimize the energy consumption of the multi-machine system, the nonlinear AoI penalty, and maximize the control stability, a constrained optimization problem is constructed. Multi-agent reinforcement learning solution: The optimization problem is transformed into a partially observable Markov decision process, which is solved by a cooperative optimization algorithm based on nearest neighbor policy optimization, and the optimal communication and control policy is output.
2. The multi-robot control and communication cooperative optimization method based on nonlinear information age as described in claim 1, characterized in that, The system modeling includes: (1.1) Dynamic model: ; in, For the first A robot is positioned in two dimensions in time slot t. For the robot's facing angle, and These are the robot's linear velocity and angular velocity, respectively. For time step; (1.2) The specific communication model is as follows: Communication rate calculation: ; in, For the first The robot and the first A robot in the time slot Channel The handshake success indicator For channel bandwidth, For transmission power, For channel gain, For noise power, For data block length, Block error rate, It is the inverse Gaussian function. For channel dispersion, These represent the robot's currently selected communication channel and the entire set of available channels, respectively. Both represent robot serial numbers, and: ; Communication delay calculation: ; in, For data group size, Minimal positive numbers to ensure Not zero; Communication energy consumption calculation: ; (1.3) Construction of nonlinear information age model: AoI definition: , For the first The robot moved towards the first The generation time of the information sent by each robot; The current time; Non-linear AoI penalty function: ,right An exponential penalty is imposed for exceeding the threshold, i.e. , This is the penalty coefficient; (1.4) Control model construction: State estimation delay: ;in, The current age of the information; State estimation: ; in, For the first A robot in the time slot The estimated location; Trajectory error: ; in, For the first The reference trajectory of the robot; Control Law: ; in, For proportional gain; , These are the coordinates for the trajectory error. Control stability function: ; in, As a weighting factor, Use extremely small positive numbers to ensure Bounded.
3. The multi-robot control and communication cooperative optimization method based on nonlinear information age as described in claim 1, characterized in that, The adaptive handshake frequency hopping mechanism design includes the following steps: (2.1) Adaptive handshake cycle adjustment: The handshake cycle is dynamically adjusted by combining robot speed, position changes, nonlinear AoI, and control stability. ; in, As a weighting factor, For the first Adjustment coefficient for the handshake cycle of each robot; (2.2) Multi-channel frequency hopping strategy: Construct a handshake success determination function by combining channel state and robot orientation angle: ; in, The function value indicates whether the handshake was successful. This represents the channel selected by the i-th robot. For the first The robot pointed to the first The robot's orientation angle, The threshold for determining the orientation angle. This is a handshake frequency counter; (2.3) Time Division Multiplexing Complementary: When the channel collision rate is detected to exceed the set threshold, it automatically switches to time division multiplexing mode and allocates dedicated time slots for different robot communication pairs on the time axis.
4. The multi-robot control and communication cooperative optimization method based on nonlinear information age according to claim 1, characterized in that, The nonlinear information age optimization, through a nonlinear information age penalty function, specifically includes: Set AoI threshold ; when When applying a linear penalty: ; The information's age is represented by AoI; when At that time, apply an exponential penalty: ,in, , Let be the coefficient, and , This is the penalty coefficient.
5. The multi-robot control and communication cooperative optimization method based on nonlinear information age according to claim 1, characterized in that, The construction of the collaborative optimization problem includes the following steps: (1) Construct the objective function: ; in, The total number of robots, Let be the weighting coefficient, satisfying , This represents the nonlinear AoI penalty function. Represents the control stability function; (2) Constructing constraints: Transmit power constraints: , This represents the robot's maximum transmission power. Handshake cycle constraints: , This is the maximum handshake period; Location boundary constraints: ; , For the minimum boundary coordinates, , The coordinates of the maximum boundary; Speed constraints: , For the maximum linear velocity, This is the maximum angular velocity; Time constraints: , This is the maximum duration of the task.
6. The multi-robot control and communication cooperative optimization method based on nonlinear information age according to claim 1, characterized in that, The multi-agent deep reinforcement learning solution includes the following steps: (5.1) Definition of state space: ; in, For the first The currently selected channel for the robot For the first The average information age of each robot; (5.2) Action space construction: ; in, , Set of available channels This represents the power allocated to the i-th robot; ; (5.3) Construction of reward function: ; in, Penalty for failed handshake matching Penalty for load imbalance For channel collision penalties, These are the weighting coefficients; This represents the control stability function. Indicates communication energy consumption. This represents a nonlinear AoI penalty function; (5.4) Perform multi-agent deep reinforcement learning training and execution; the optimal action obtained is used as the optimal communication and control strategy.
7. The multi-robot control and communication cooperative optimization method based on nonlinear information age according to claim 6, characterized in that, The objective function for multi-agent deep reinforcement learning training and execution is as follows: The objective function of the Actor network is: ; The objective function of the Critic network is: ; in, For batch size, The total number of robots, For the duration of trajectory collection, This is the cutting factor. The regularization coefficient is the policy entropy. Let entropy be the policy function. The value of the normalized state; Indicates the update strategy ratio. This represents the generalized advantage estimation. This represents the clipping function. The coefficients of the policy entropy regularization term are represented. Represents the value function. This represents the value function after normalization.
8. A multi-robot control and communication cooperative optimization system based on nonlinear information age, including: The system modeling module is used to construct multi-robot dynamics models, communication models, nonlinear information aging models, and control models, forming a system framework that couples control and communication. The handshake frequency hopping mechanism module is used to construct an adaptive handshake cycle, a multi-channel frequency hopping strategy, and a time-division multiplexing complementary mechanism based on the robot's motion state and communication state, so as to reduce channel conflicts. The nonlinear information age optimization module is used to guide the system to prioritize updating high-latency information through a nonlinear AoI penalty function; The optimization problem construction module is used to construct constrained optimization problems with the objectives of minimizing system energy consumption, nonlinear information age penalty, and maximizing control stability. The multi-agent deep reinforcement learning optimization module is used to transform the optimization problem into a partially observable Markov decision process. It solves the problem through a cooperative optimization algorithm based on nearest neighbor policy optimization and outputs the optimal communication and control strategy to guide the robot to execute in real time.
9. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the multi-robot control and communication cooperative optimization method based on nonlinear information age as described in claims 1-7.