Unmanned aerial vehicle assisted large-scale mobile robot intelligent scheduling method and device

By optimizing the UAV-assisted collaborative scheduling architecture and Markov decision process, the problem of inefficient communication resource utilization of reconfigurable intelligent metasurfaces in dynamic environments is solved. This achieves deterministic communication with low latency and low jitter, reduces trajectory tracking errors, and improves the stability and reliability of large-scale mobile robot scheduling.

CN121501014BActive Publication Date: 2026-04-17SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
Filing Date
2025-09-28
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In large-scale mobile robot scheduling, existing reconfigurable intelligent metasurface systems are suitable for static deployment, but their application in dynamic environments is limited. Furthermore, they do not consider the coupling between wireless communication and mobile robot motion control, resulting in inefficient use of communication resources and an inability to meet the requirements for low-error control.

Method used

A collaborative scheduling architecture assisted by unmanned aerial vehicles (UAVs) is constructed. A reconfigurable intelligent metasurface is used to establish a dynamically reconfigurable deterministic communication link between the mobile robot and the edge server. The joint action of control and communication is optimized through Markov decision process, and the scheduling and phase configuration of the mobile robot control command, UAV flight path and reconfigurable intelligent metasurface are adjusted in real time.

Benefits of technology

It achieves deterministic communication with low latency and low jitter, significantly reducing the trajectory tracking error of mobile robots and improving the stability and reliability of the scheduling system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121501014B_ABST
    Figure CN121501014B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of large-scale mobile robot scheduling, and provides a method and device for intelligent scheduling of large-scale mobile robots assisted by unmanned aerial vehicles, comprising: constructing a collaborative scheduling architecture comprising unmanned aerial vehicles, multiple mobile robots and edge servers; in the collaborative scheduling architecture, based on the communication delay, jitter and control error of each mobile robot, a joint optimization problem is constructed with the objective of minimizing real-time tracking error; the joint optimization problem is modeled as a Markov decision process; the control and communication joint action of the Markov decision process is determined by using an intelligent scheduling algorithm; and the mobile robot control instruction, the unmanned aerial vehicle flight path, the scheduling and phase configuration of the reconfigurable intelligent metasurface of the unmanned aerial vehicle, and the transmission power of the edge server in the collaborative scheduling architecture are adjusted in real time according to the control and communication joint action. The embodiment reduces the communication delay, jitter and trajectory tracking error of the mobile robot, and improves the stability and reliability of the system scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of large-scale mobile robot scheduling technology, and more specifically, to a method and apparatus for intelligent scheduling of large-scale mobile robots assisted by unmanned aerial vehicles (UAVs). Background Technology

[0002] In the development of smart factories, the scheduling of large-scale mobile robots has become a key link in ensuring efficient and orderly production. Through high-precision positioning and perception, and decision-making by edge servers, mobile robots enhance their adaptability to dynamic environments within a hierarchical scheduling architecture at the terminal and edge.

[0003] However, mobile robots need to interact with the edge in real time during task execution, which consumes a lot of communication resources. Due to the influence of limited communication bandwidth, environmental noise interference, obstacle occlusion and multipath fading, the channel exhibits strong randomness, which not only increases communication latency and jitter, but may also lead to the accumulation of trajectory tracking errors, thereby affecting the motion accuracy and stability of the mobile robot.

[0004] To address this issue, reconfigurable smart metasurface technology achieves low latency, low jitter, and high reliability communication through intelligent configuration of reflective elements. By adjusting the phase offset of the reflective elements, a more stable wireless channel can be constructed. However, despite some progress, this technology still faces two major limitations in large-scale mobile robot scheduling: first, most reconfigurable smart metasurface systems are only suitable for static deployment, limiting their application in dynamic environments; second, existing reconfigurable smart metasurface-assisted architectures do not consider the coupling between wireless communication and mobile robot motion control, which may lead to inefficient utilization of communication resources and fail to meet the stringent requirements for low-error control. Summary of the Invention

[0005] This disclosure provides a method and apparatus for intelligent scheduling of large-scale mobile robots assisted by unmanned aerial vehicles (UAVs), which achieves deterministic communication with low latency and low jitter, significantly reduces the trajectory tracking error of mobile robots, and improves the stability and reliability of the scheduling system.

[0006] This disclosure provides a method for intelligent scheduling of large-scale mobile robots assisted by unmanned aerial vehicles (UAVs), including:

[0007] Construct a collaborative scheduling architecture that includes drones, multiple mobile robots, and edge servers; wherein, the drones are equipped with reconfigurable smart metasurfaces, and by adjusting the reflective elements and beam phase of the reconfigurable smart metasurfaces, a dynamically reconfigurable deterministic communication link is established between each of the mobile robots and the edge server;

[0008] In the cooperative scheduling architecture, a joint optimization problem is constructed with the goal of minimizing real-time tracking error, based on the communication latency, jitter, and control error of each mobile robot; and the joint optimization problem is modeled as a Markov decision process.

[0009] The established Markov decision process is solved using a trained intelligent scheduling algorithm to determine the joint control and communication actions for the cooperative scheduling architecture.

[0010] Based on the combined control and communication actions, the mobile robot control commands, UAV flight paths, scheduling and phase configuration of the UAV reconfigurable intelligent metasurface, and edge server transmission power in the collaborative scheduling architecture are adjusted in real time.

[0011] This disclosure provides an unmanned aerial vehicle (UAV)-assisted intelligent scheduling device for large-scale mobile robots, comprising:

[0012] An architecture building module is used to build a collaborative scheduling architecture that includes drones, multiple mobile robots, and edge servers. The drones are equipped with reconfigurable smart metasurfaces, and by adjusting the reflective elements and beam phase of the reconfigurable smart metasurfaces, dynamic and reconfigurable deterministic communication links are established between each mobile robot and the edge server.

[0013] The model building module is used to construct a joint optimization problem with the goal of minimizing real-time tracking error based on the communication latency, jitter, and control error of each mobile robot in the cooperative scheduling architecture; and to model the joint optimization problem as a Markov decision process.

[0014] The action determination module is used to solve the established Markov decision process using a trained intelligent scheduling algorithm to determine the joint control and communication actions for the cooperative scheduling architecture.

[0015] The action execution module is used to adjust the mobile robot control commands, UAV flight paths, UAV reconfigurable smart metasurface scheduling and phase configuration, and edge server transmission power in the collaborative scheduling architecture in real time according to the control and communication joint actions.

[0016] This disclosure provides a computer device including a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the UAV-assisted large-scale mobile robot intelligent scheduling method as described in any of the above possible embodiments.

[0017] This disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the unmanned aerial vehicle-assisted intelligent scheduling method for large-scale mobile robots as described in any of the possible embodiments above.

[0018] The UAV-assisted intelligent scheduling method and apparatus for large-scale mobile robots provided in this disclosure overcomes the communication latency and jitter problems caused by limited communication resources and channel randomness in traditional mobile robot scheduling by constructing a collaborative scheduling architecture including UAVs, mobile robots, and edge servers, and establishing a dynamically reconfigurable deterministic communication link using a reconfigurable intelligent metasurface mounted on the UAV. This provides stable and reliable communication between the mobile robot and the edge server. Based on the communication latency, jitter, and control errors of each mobile robot, a joint optimization problem with the goal of minimizing real-time tracking error is constructed and modeled as a Markov decision process. This enables collaborative design of the motion control and communication resource allocation of the mobile robot, achieving joint optimization of control and communication. The established Markov decision process is solved using a trained intelligent scheduling algorithm to determine the joint action of control and communication, and the scheduling and phase configuration of the mobile robot control commands, UAV flight paths, edge server transmission power, and reconfigurable intelligent metasurface are adjusted in real time. This allows the entire scheduling system to dynamically adjust its strategy according to the real-time environment, improving the system's adaptability to complex dynamic environments.

[0019] Thus, this embodiment achieves low-latency, low-jitter deterministic communication, significantly reduces trajectory tracking errors of mobile robots, and improves the stability and reliability of the scheduling system. It also collaboratively optimizes control and communication resources, thereby improving the overall performance and efficiency of large-scale mobile robot scheduling.

[0020] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings referenced in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.

[0022] Figure 1A flowchart of a drone-assisted intelligent scheduling method for large-scale mobile robots provided in an embodiment of this disclosure is shown;

[0023] Figure 2 A schematic diagram of a collaborative scheduling architecture provided by an embodiment of this disclosure is shown;

[0024] Figure 3 A flowchart of a control and communication joint action determination method provided by an embodiment of this disclosure is shown;

[0025] Figure 4 A flowchart of a training method for an intelligent scheduling algorithm provided in an embodiment of this disclosure is shown;

[0026] Figure 5 A flowchart of a control and communication joint action execution method provided in an embodiment of this disclosure is shown;

[0027] Figure 6 A schematic diagram of an intelligent scheduling algorithm provided in an embodiment of this disclosure is shown;

[0028] Figure 7 This diagram illustrates the structure of a drone-assisted intelligent scheduling device for large-scale mobile robots provided in an embodiment of the present disclosure.

[0029] Figure 8 A schematic diagram of the structure of a computer device provided in an embodiment of this disclosure is shown. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0031] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0032] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0033] To facilitate understanding of this embodiment, the executing entity of the UAV-assisted large-scale mobile robot intelligent scheduling method provided in this disclosure will first be described in detail. The executing entity of the UAV-assisted large-scale mobile robot intelligent scheduling method provided in this disclosure is a computer device. This computer device can be a terminal device or a server. The terminal device can also be a mobile device, user terminal, terminal, handheld device, computing device, vehicle-mounted device, wearable device, etc. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms. Optionally, this method can also be applied to an implementation environment composed of computer devices and servers.

[0034] The intelligent scheduling method for large-scale mobile robots assisted by unmanned aerial vehicles (UAVs) provided in this application will be described in detail below with reference to the accompanying drawings. See also: Figure 1 The diagram shows a flowchart of a drone-assisted intelligent scheduling method for large-scale mobile robots provided in this disclosure. The method includes the following steps S101 to S104:

[0035] S101 constructs a collaborative scheduling architecture that includes drones, multiple mobile robots, and edge servers.

[0036] Understandably, the foundation of the collaborative scheduling architecture lies in the fact that drones can carry reconfigurable intelligent metasurfaces and establish dynamic, reconfigurable communication links between various mobile robots and edge servers by adjusting the reflective elements and beam phase. A reconfigurable intelligent surface (RIS) is an intelligent device composed of multiple reflective elements. By controlling the reflective properties of these elements, the propagation characteristics of wireless signals can be adjusted to adapt to environmental changes and optimize the transmission quality of wireless signals.

[0037] Specifically, a drone carrying a reconfigurable smart metasurface with multiple reflective elements can autonomously move in three-dimensional space, expanding communication coverage and enabling mobile robots in more areas to access the communication network. Simultaneously, it can effectively avoid obstacle obstruction, preventing communication interruptions or signal quality degradation caused by obstructions. Here, the reconfigurable smart metasurface can reconstruct the wireless channel environment by dynamically adjusting the scheduling matrix and phase configuration matrix. By regulating the reflective element scheduling matrix and beam phase configuration matrix of the reconfigurable smart metasurface, enhanced channel services can be provided for the uplink and downlink between each mobile robot and the edge server. The scheduling matrix is ​​used to rationally arrange the working state and time of each reflective element, while the phase configuration matrix is ​​used to adjust the phase of the reflected signal. The two work together to provide enhanced channel services for the uplink and downlink between each mobile robot and the edge server, improving communication efficiency and data transmission quality. In the collaborative scheduling architecture, the mobile robot undertakes tasks such as material handling, receiving and executing control commands, and transmitting information to the drone carrying the reconfigurable smart metasurface. As the core data processing and control unit of the entire collaborative scheduling architecture, the edge server undertakes the crucial task of collecting status information from each mobile robot. It can acquire real-time data such as the robot's position, speed, and operational status, and execute scheduling algorithms based on this data. Through complex calculations and analysis, it formulates the optimal scheduling strategy for the mobile robots. Finally, the edge server sends the generated control commands to the corresponding devices, forming a closed-loop control system with the mobile robots, drones, and other components, enabling real-time monitoring and dynamic adjustment of the entire scheduling process.

[0038] For example, refer to Figure 2 The diagram shown illustrates a collaborative scheduling architecture provided in this disclosure. Taking the scenario of the collaborative scheduling architecture in the diagram as an example, it may include a mobile robot control system and an airborne reconfigurable intelligent metasurface-assisted communication system. The mobile robot control system includes M mobile robots B, represented as a set. The three-dimensional coordinates of the m-th mobile robot B can be represented as: Where τ m Let be the communication delay between the m-th mobile robot B and the edge server A, k represent the k-th time slot, and T represent transpose. In an airborne reconfigurable intelligent metasurface-assisted communication system, there is an edge server A deployed at the origin, possessing abundant computing resources, capable of deploying mobile robot scheduling algorithms, analyzing the mobile robot's state perception information, and issuing control information; and a drone D, which can improve the coverage of the mobile robots. Drone D moves according to a drone trajectory F, and its three-dimensional coordinates can be represented as P. u (k)=[x u (k),y u (k),zu (k)] T The drone was equipped with A reconfigurable smart metasurface E (where I and J represent the row and column numbers, respectively) with reflective elements serves as a bridge for the interaction between the mobile robot's state perception information and the cloud server's control information. The reflective element R of the reconfigurable smart metasurface in the i-th row and j-th column... i,j The three-dimensional coordinates are d x and d y These are the row spacing and column spacing between the reflective elements, respectively. In addition, multiple obstacles C are also present in this scene.

[0039] The mobile robot control system is composed of the kinematic model of the mobile robot. Specifically, the m-th mobile robot... The kinematic model is modeled as a nonlinear discrete-time system with time delay characteristics. Here, the kinematic model of the m-th mobile robot can be represented as:

[0040]

[0041] in, This represents the motion state vector of the m-th mobile robot in the k-th time slot; It is the orientation angle of the m-th mobile robot in the k-th time slot. It is the floor function, q m (·) is the control input, T s Indicates the sampling period. It is the Jacobian matrix affected by time delay, and its expression is: The trajectory tracking error of the m-th mobile robot It can be represented by a rotation matrix as follows:

[0042]

[0043] in, This represents the desired reference state for the mobile robot.

[0044] Here, the airborne reconfigurable smart metasurface-assisted communication system consists of a channel model between the edge server and the mobile robot, as well as a time delay derivation. The channel model between the edge server and the mobile robot includes the wireless access system from the edge server to the UAV and the wireless access system from the mobile robot to the UAV. The two interact through the reconfigurable smart metasurface mounted on the UAV.

[0045] Specifically, the channel model G of the wireless access system from the edge server to the drone. c It can be represented as: in, R represents the reflective element R of the reconfigurable smart metasurface in the i-th row and j-th column. i,j Channel gain to the edge server. This represents the open-ground loss between the edge server and the reflective element of the reconfigurable smart metasurface in the i-th row and j-th column. Follow the small-scale Ruili weakness.

[0046] Specifically, the channel model H of the wireless access system from mobile robot to drone. m It can be represented as: in, It is the scheduling matrix of a reconfigurable intelligent metasurface; R represents the reflective element R of the reconfigurable smart metasurface in the i-th row and j-th column. i,j Whether it is applied to the signal reflection of the m-th mobile robot; diag(·) denotes a diagonal matrix; R represents the reflective element R of the reconfigurable smart metasurface in the i-th row and j-th column. i,j Channel gain up to the m-th mobile robot; R represents the reflective element R of the m-th mobile robot and the reconfigurable smart metasurface in the i-th row and j-th column. i,j Path loss between; Follows small-scale Ricean decay.

[0047] Here, the reconfigurable smart metasurface utilizes the phase configuration matrix Φ to facilitate the interaction between the state perception information of the m-th mobile robot and the control information of the edge server. The phase configuration matrix Φ can be expressed as:

[0048]

[0049] in, This represents the beam phase of the reflective element of the l-th reconfigurable smart metasurface.

[0050] Specifically, the delay derivation consists of uplink signal-to-noise ratio (SNR) derivation, downlink SNR derivation, and round-trip delay derivation. In the uplink SNR derivation, the state perception information of the m-th mobile robot is obtained through the channel matrix H from the mobile robot to the UAV. m The phase configuration matrix Φ of the reconfigurable smart metasurface and the channel matrix G from the edge server to the UAV. c The data is transmitted to the edge server, forming an uplink. Accordingly, the edge server receives the status information of the m-th mobile robot as follows:

[0051]

[0052] Where, p m It is the transmission power of the m-th mobile robot, and n m The noise power is represented by σ.m Additive white Gaussian noise.

[0053] Thus, the uplink signal-to-noise ratio for large-scale mobile robot scheduling scenarios assisted by airborne reconfigurable intelligent metasurfaces can be expressed as:

[0054] Accordingly, in the derivation of the downlink signal-to-noise ratio, the control information q′ generated by the edge server... m Channel matrix G from edge server to drone c The reflection matrix Φ of the reconfigurable smart metasurface and the channel matrix H from mobile robot to drone. m The data is transmitted to the m-th mobile robot, forming a downlink. Here, the control information received by the m-th mobile robot is: Where, p′ m It is the transmission power of the edge server, and n′ m The noise power is represented by σ' m Additive white Gaussian noise.

[0055] Thus, the downlink signal-to-noise ratio for large-scale mobile robot scheduling scenarios assisted by airborne reconfigurable intelligent metasurfaces can be expressed as:

[0056] Specifically, in the derivation of the round-trip delay, the bidirectional interaction between the state perception information published by the mobile robot and the control information published by the edge server establishes a closed-loop control system for the m-th mobile robot, with a round-trip delay of τ. m , can be represented as:

[0057]

[0058] in, and These represent the uplink bandwidth and the amount of state perception information of the m-th mobile robot, respectively. and This represents the downlink bandwidth and the amount of data related to the state perception information of the m-th mobile robot.

[0059] S102, in the cooperative scheduling architecture, based on the communication latency, jitter and control error of each mobile robot, a joint optimization problem is constructed with the goal of minimizing the real-time tracking error; and the joint optimization problem is modeled as a Markov decision process.

[0060] Furthermore, to effectively improve the spatial coverage of mobile robots, reduce information path loss, and mitigate noise interference in complex wireless spectrum, this disclosure addresses a joint optimization problem aimed at minimizing real-time tracking error, comprehensively considering communication latency, jitter, and control errors of mobile robots. A joint optimization objective function and its corresponding constraints are then constructed. In a multi-mobile robot collaborative scheduling environment, minimizing real-time tracking error refers to optimizing the efficiency of the entire scheduling system by further reducing the difference between the actual trajectory and the predetermined trajectory of the mobile robot during operation, under deterministic communication conditions with low latency and low jitter.

[0061] Here, the constraints corresponding to the joint optimization objective function include control input constraints, control stability constraints, UAV operating speed constraints, transmit power constraints, reconfigurable smart metasurface reflector scheduling constraints, and reconfigurable smart metasurface beam phase configuration constraints. Specifically, the control input constraint indicates that the control input of the mobile robot should be within the upper and lower limits of the control input caused by the physical structure of the mobile robot, so as to avoid equipment damage, motion instability and other problems caused by exceeding the physical hardware bearing capacity; the control stability constraint indicates that the round-trip time delay of wireless transmission should be within the scheduling deadline of the mobile robot, so as to avoid the transmission delay being too large, which would damage the phase characteristics of the control system and cause system oscillation or instability; the UAV operating speed constraint indicates that the operating speed of the UAV should be within the maximum speed range, so as to ensure that the UAV will not cause power overload or control failure due to speeding when operating within this range; the transmit power constraint indicates the maximum transmit power of the edge server and the mobile robot, so as to reduce unnecessary energy consumption; the reconfigurable smart metasurface reflective element scheduling constraint indicates that each reconfigurable smart metasurface element serves at most one mobile robot in any time slot, so as to avoid beam interference between multiple users and improve the spectral efficiency and signal quality of reconfigurable smart metasurface assisted communication; the reconfigurable smart metasurface constant mode constraint indicates that the phase configuration matrix of the reconfigurable smart metasurface only changes the phase and does not change the beam amplitude.

[0062] Specifically, the joint optimization problem includes a joint optimization objective function and constraints. Here, the joint optimization objective function can be expressed as:

[0063]

[0064] Constraints may include:

[0065] C1:q min ≤q m ≤q max ;

[0066]

[0067] C3:‖Pu (k)-P u (k-1)‖2≤v max ;

[0068]

[0069]

[0070] In the formula, U, Z, P, W, and Θ are the set of variables to be optimized in the problem; U represents the control input of the mobile robot. Z represents the path planning of the drone; P represents the transmission power of the edge server. W represents the reflective element scheduling matrix of the reconfigurable smart metasurface. Θ represents the reflection angle of the reconfigurable smart metasurface reflective element; Represented as the control objective function The normalization function; Represented as the communication objective function The normalization function; δ1 and δ2 represent the first weighting coefficient and the second weighting coefficient, respectively; M represents the total number of mobile robots; Let q represent a set of mobile robots; C1 and C2 are the control constraints of the mobile robots; min and q max q represents the minimum and maximum values ​​of the control input for the mobile robot, respectively; m τ represents the control input for the m-th mobile robot. m This represents the communication latency between the m-th mobile robot and the edge server. Let C1 represent the scheduling deadline for the m-th mobile robot; C2 represent the trajectory planning constraints for the UAV; P represents the time limit for scheduling the m-th mobile robot. u (k) represents the three-dimensional coordinates of the UAV at time k; P u (k-1) represents the three-dimensional coordinates of the UAV at time k-1; v max C4 represents the maximum flight speed of the drone; C5 and C6 represent the communication constraints of the edge server; p m p represents the transmission power from the edge server to the m-th mobile robot. max This represents the maximum transmission power of the edge server; The reflective element R of the reconfigurable smart metasurface is represented as the element in the i-th row and j-th column. i,j Whether the signal reflection is applied to the m-th mobile robot, 0 indicates no application, 1 indicates application; θ l This represents the beam phase configuration of the l-th reflecting element of the reconfigurable smart metasurface. This represents the total number of reflective elements on the reconfigurable smart metasurface.

[0071] Here, the control objective function It can be represented as:

[0072]

[0073] Q1, Q2, and Q3 are all constant positive definite matrices. This control objective function jointly considers the control input and the control error, ensuring both the smoothness of the control input and the goal of minimizing the trajectory tracking error of the mobile robot.

[0074] Here, the communication objective function It can be represented as:

[0075]

[0076] Wherein, the communication objective function By calculating the reciprocal of the time delay, 1 / τ m coefficient of variation (1 / τ) m The ratio of the standard deviation to the mean is used to quantify communication determinism, where 1 / τ m The standard deviation reflects jitter, while 1 / τ m The average value reflects the latency. By minimizing... This function not only considers the actual round-trip time delay, but also uses the delay standard deviation to ensure the determinism of communication delay, thereby achieving a highly reliable, low-latency, and low-jitter communication mechanism.

[0077] Specifically, the joint optimization problem comprehensively considers multiple factors affecting the trajectory tracking of mobile robots, such as communication quality and the accuracy of control commands. These factors are integrated into a unified optimization objective in the form of mathematical functions to effectively improve the path planning accuracy of the mobile robot, thereby optimizing the efficiency of the entire scheduling system. To achieve this objective, the joint optimization problem can be modeled as a Markov Decision Process (MDP). A Markov Decision Process is a mathematical model used to describe decision-making in uncertain environments. It is based on the Markov property, which states that the next state of the system depends only on the current state and the action taken, and is independent of past states. By modeling the joint optimization problem as a Markov Decision Process, the theory and methods of this model can be used to analyze and solve for the optimal decision strategy to minimize communication latency and jitter, further reduce trajectory errors, and optimize the entire scheduling process.

[0078] Specifically, because the optimization problem involves high-dimensional system states and strong coupling constraints, its complexity increases exponentially. Therefore, the joint optimization problem can be modeled as a Markov decision process, which includes a state space, an action space, and a reward function. The state space includes the motion states of each mobile robot in the k-th time slot. Reference status Uplink / Downlink packet size Broadband transmission and signal-to-noise ratio Specifically, it can be expressed as:

[0079]

[0080] in, This represents the motion state of the m-th mobile robot in the k-th time slot. This represents the reference state of the m-th mobile robot in the k-th time slot. This represents the packet size of the uplink (a=↑) or downlink (a=↓) transmission when the m-th mobile robot transmits information with the edge server in the k-th time slot. This represents the uplink (a=↑) or downlink (a=↓) transmission bandwidth when the m-th mobile robot transmits information with the edge server in the k-th time slot. This represents the signal-to-noise ratio of the uplink (a = ↑) or downlink (a = ↓) when the m-th mobile robot transmits information with the edge server in the k-th time slot.

[0081] Accordingly, the motion space includes the control input q of each mobile robot in the k-th time slot. m (k) Trajectory planning vector P of the UAV u (k) The transmit power p of the edge server m (k) Reconfigurable intelligent metasurface scheduling matrix W m (k) and the phase configuration matrix θ of the reconfigurable smart metasurface. l (k), which can be specifically represented as:

[0082]

[0083] Where, q m (k) represents the control input of the m-th mobile robot in the k-th time slot, P u (k) represents the trajectory planning vector of the UAV, p m (k) represents the transmission power of the edge server when the m-th mobile robot transmits information with the edge server in the k-th time slot, in W. m (k) represents the scheduling matrix of the reflective elements of the reconfigurable smart metasurface when the m-th mobile robot transmits information with the edge server in the k-th time slot, θ l (k) represents the phase configuration matrix of the reconfigurable smart metasurface.

[0084] Specifically, the reward function of the Markov decision process model is used to evaluate the actions taken by the agent. In each time slot, the agent receives a corresponding reward after executing the action policy based on the state. The reward function can be expressed as the average of a normalized communication objective function considering communication delay and jitter, and a normalized control objective function considering control error, after being adjusted by weighting factors and then taking their exponents.

[0085]

[0086] Here, Exp(·) represents the exponential function.

[0087] Understandably, after the above Markov decision process modeling, the optimization objective of the joint optimization objective function is ultimately transformed into the agent's reward maximization problem, which can be expressed as:

[0088]

[0089] S103, using the trained intelligent scheduling algorithm to solve the established Markov decision process, and determine the joint control and communication actions for the cooperative scheduling architecture.

[0090] It is understandable that control and communication joint actions refer to a series of coordinated actions taken, taking into account the control needs of mobile robots and the requirements of communication systems, aiming to achieve optimal operation of the entire scheduling architecture. Control and communication joint actions include: control commands for each mobile robot (specific operational commands issued to the mobile robot, such as forward, backward, and turning commands); the next flight position of the UAV (UAV), which, as a communication node or auxiliary device, affects communication coverage and coordination with other devices; the transmission power of the edge server to the corresponding mobile robot (the edge server provides computing and communication support for mobile robots and other devices, and its transmission power directly affects signal strength and communication quality); and the scheduling matrix and phase configuration matrix of the reconfigurable smart metasurface's reflective elements. The reconfigurable smart metasurface is a novel artificial electromagnetic material surface that can change the propagation characteristics of electromagnetic waves by dynamically adjusting the scheduling matrix and beam phase configuration matrix of the reflective elements. The reflective element scheduling matrix determines which reflective elements participate in the operation, while the phase configuration matrix controls the phase changes of the reflective elements. Through these real-time adjustments, efficient operation and optimized control of the entire collaborative scheduling architecture can be achieved.

[0091] Specifically, refer to Figure 3 As shown, in order to achieve efficient and stable operation of the cooperative scheduling architecture in complex dynamic environments, when using a trained intelligent scheduling algorithm to solve the established Markov decision process based on the objective function and constraints of the joint optimization problem, the following steps S301 to S302 may be included:

[0092] S301, using the communication link established by the UAV equipped with the reconfigurable smart metasurface, the motion state, reference state, transmission packet size, transmission bandwidth and signal-to-noise ratio of the mobile robot are obtained as the state input of the Markov decision process.

[0093] Here, the communication link is the channel for information transmission between the drone and mobile robots, edge servers, and other devices. Drones equipped with reconfigurable smart metasurfaces can enhance communication signals and expand communication range. The motion state of the mobile robot can include current speed, acceleration, and direction of travel. The reference state is a pre-set target state, such as target position and target speed. The transmission packet size is the size of the data packets that the mobile robot needs to transmit. The transmission bandwidth is the amount of data that the communication link can transmit per unit time. The signal-to-noise ratio is the ratio of signal power to noise power, reflecting the signal quality.

[0094] S302, the trained intelligent scheduling algorithm outputs the control and communication joint action regarding the cooperative scheduling architecture based on the state input and the constraints.

[0095] Here, the trained intelligent scheduling algorithm outputs the joint control and communication actions of the collaborative scheduling architecture based on the state input and constraints. The intelligent scheduling algorithm is a machine learning model that can analyze the input state information, comprehensively consider various factors such as the current state of the mobile robot and communication quality, and then output the optimal joint control and communication actions to achieve the efficient operation of the entire scheduling architecture.

[0096] Specifically, to further improve the decision-making ability and generalization performance of intelligent scheduling algorithms in complex scenarios, this disclosure proposes an experience-enhanced group relative policy optimization scheduling algorithm. The experience enhancement involves constructing virtual experience data with a distribution consistent with real experience data using a Wasserstein generative adversarial network, which is then used to optimize the training process of the intelligent scheduling algorithm. (Refer to...) Figure 4 The intelligent scheduling algorithm training method provided in this disclosure may include the following steps S401 to S406:

[0097] S401, Initialize the network parameters of the policy network and the simulation environment of the cooperative scheduling architecture.

[0098] Here, initializing the network parameters of the policy network sets a starting point for the subsequent training process. Different initial parameters may affect the convergence speed and final performance of the training. The simulation environment is an abstraction and simulation of the actual collaborative scheduling architecture. It includes the operating rules and interaction relationships of devices such as mobile robots, drones, and edge servers. Through the simulation environment, intelligent scheduling algorithms can be trained and tested without relying on actual hardware.

[0099] S402, in the simulated environment, the current policy network outputs an action based on the current environment state of the simulated environment and applies the action to the simulated environment; the new state of the simulated environment after executing the action is determined, and the corresponding reward value is calculated according to the reward function; real experience data consisting of the current state, action, reward value and new state is collected and stored in the real experience pool.

[0100] Here, the policy network can make decisions and output actions based on the current state of the simulation environment. After the simulation environment executes the action, it will enter a new state and calculate a reward value based on the reward function. The reward function is used to measure the quality of the action. The higher the reward value, the more conducive the action is to achieving the goal.

[0101] S403, based on the experience enhancement mechanism of Wasserstein generative adversarial networks, generates virtual experience data consistent with the distribution of real experience data through adversarial training between the generator network and the discriminator network, and stores it in the virtual experience pool.

[0102] Specifically, the experience enhancement mechanism of Wasserstein Generative Adversarial Networks (GANs) is based on a dynamic game between two competing networks, which can better handle the distributional differences between generated and real data. It mainly consists of a generator network and a discriminator network. The goal of Wasserstein GANs is to generate virtual experience that approximates the real experience distribution for specific noise sources. This not only reduces dependence on the environment but also facilitates the agent's exploration of sparse reward regions. The discriminator network is responsible for determining whether the input data is real or virtual experience data, and its goal is to maximize the Wasserstein distance between the virtual and real experience distributions. Through adversarial training, the generator network can gradually generate virtual experience data consistent with the distribution of real experience data. This virtual experience data can expand the training dataset and improve the generalization ability of the policy network.

[0103] For example, generating network Γ ω1 The network parameters are ω1, and the discriminant network Γ ω2 The network parameters are ω2, and its training process mainly includes initializing the generator network Γ. ω1 The network parameters are ω1 and the discriminant network is Γ.ω2 The network parameters are ω2 and the virtual experience set D'; then, sampling is performed through the real experience pool D and the virtual experience pool D', and the objective function of the discriminant network is calculated using the following formula. The parameters of the discriminant network are updated by maximizing its objective function; the formula is as follows:

[0104]

[0105] Where o∈D represents the distribution b from real experience. e (o) represents real experience sampled from the virtual experience distribution b′; o′∈D′ represents real experience sampled from the virtual experience distribution b′. e Virtual experience sampled in (o′).

[0106] Furthermore, the objective function of the generator network is calculated using the following formula. The parameters of the generator network are updated by minimizing its objective function, as shown in the following formula:

[0107]

[0108] Where o′∈D′ represents the virtual experience distribution b′ e Virtual experience sampled in (o′).

[0109] Thus, the above steps are repeated for training until the loss of the generator network and the loss of the discriminator network are both stable at a relatively constant value. In this way, the network parameters of the Wasserstein generative adversarial network that can solve the problem proposed in this disclosure can be obtained.

[0110] S404, Sample experience data from the hybrid experience pool, which is composed of the real experience pool and the virtual experience pool.

[0111] Here, sampling from a hybrid experience pool can make full use of both real and virtual experience data, avoiding over-reliance on any one type of data, thereby improving the stability and effectiveness of training.

[0112] S405, based on the sampled empirical data, the group relative policy optimization algorithm is used to calculate the group relative advantage function value and the target loss function value including the KL divergence penalty term, and the network parameters of the policy network are updated according to the calculation results.

[0113] Specifically, the group relative policy optimization algorithm is a reinforcement learning algorithm that evaluates the relative merits of different actions by calculating the group relative advantage function value. It also introduces a KL divergence penalty term to limit the magnitude of policy updates, which acts as a regularization mechanism to prevent the policy from deviating too far from a stable reference point and avoids excessively large policy update magnitudes during training, thus ensuring training stability. The network parameters of the policy network are updated based on the calculated target loss function value, enabling the policy network to gradually learn better decision-making policies.

[0114] Here, the formula for calculating the group's relative advantage function value can be expressed as:

[0115]

[0116] in, Indicates in T g The set of rewards collected along the trajectory.

[0117] Here, the target loss function value can be expressed as:

[0118]

[0119] in, It is the current policy network and previous policy networks The ratio; and These are the parameters of the current policy network and the parameters of the previous policy network, respectively. This is represented by the cutoff function; β is the regularization factor, used to balance the stability and diversity of policy updates; KL(.) represents the current policy network. and previous policy networks The Kullback-Leibler divergence.

[0120] The formula for calculating the truncation function can be expressed as:

[0121]

[0122] In the formula, ε1 represents the cutting coefficient.

[0123] Here, we can maximize the target loss function in the policy network. This is to update the network parameters of the policy network.

[0124] S406. Repeat S402 to S405 to iteratively optimize the policy network parameters until the performance index of the policy network converges, thus obtaining the trained intelligent scheduling algorithm.

[0125] Understandably, through continuous iterative training, the performance of the policy network will gradually improve. When the performance indicators converge (such as the reward value stabilizing), it means that the policy network has learned a relatively stable decision-making policy. At this point, the trained intelligent scheduling algorithm can be applied to the actual collaborative scheduling architecture to achieve efficient joint action decision-making of control and communication.

[0126] S104, adjust the scheduling and phase configuration of the mobile robot control commands, UAV flight paths, edge server transmission power, and reconfigurable smart metasurface in the collaborative scheduling architecture in real time according to the control and communication joint action.

[0127] Specifically, refer to Figure 5 As shown, after obtaining the joint control and communication actions of the cooperative scheduling architecture using the trained intelligent scheduling algorithm, the following steps S501 to S504 can be executed to achieve real-time control of the cooperative scheduling architecture, specifically including:

[0128] S501, based on the control instructions of each mobile robot in the control and communication joint action, generate corresponding motion control signals and send them to each mobile robot through the edge server.

[0129] Understandably, mobile robots, as key execution units in the collaborative scheduling architecture, undertake important responsibilities such as material handling and task execution. Control commands encompass crucial operational information such as the robot's direction of travel, speed, and start / stop. For example, if a control command instructs a mobile robot to move goods to a designated area, motion control signals need to be generated to control its direction of travel, its appropriate speed, and its precise stopping upon arrival at the destination. Edge servers act as the information hub and computing core in the collaborative scheduling architecture. Through edge servers, the generated motion control signals can be distributed to each mobile robot, ensuring stable transmission and timely reception of the signals. Upon receiving the motion control signals, the mobile robot's internal control system precisely controls the motors, steering mechanisms, and other components according to the signal commands, thereby achieving movement according to predetermined requirements and ensuring the smooth execution of various tasks.

[0130] S502, based on the drone's next flight position, plan the drone's real-time flight path and control the drone to fly along the real-time flight path.

[0131] Here, determining the next flight position of the UAV is based on the global optimization of the entire collaborative scheduling architecture and real-time requirements. For example, this could be to better cover the mobile robot's activity area to ensure communication quality, or to travel to a specific location for data collection. Planning the UAV's real-time flight path requires comprehensive consideration of various factors, including the UAV's flight performance (such as maximum speed, acceleration, and endurance), obstacle information in the surrounding environment, and aerodynamic conditions. Path planning algorithms, such as the A* algorithm, improved versions of Dijkstra's algorithm, or reinforcement learning-based path planning methods, can generate a safe, efficient, and mission-compliant real-time flight path. Controlling the UAV to fly along this real-time flight path can be achieved through the UAV's internal flight controller, which simultaneously controls the UAV's flight attitude and power output, ensuring that the UAV can fly accurately and stably along the planned trajectory to adapt to the dynamic changes of the collaborative scheduling architecture.

[0132] S503, based on the reflective element scheduling matrix and beam phase configuration matrix of the reconfigurable smart metasurface, dynamically adjust the service object and beam phase of the reflective element to reconstruct the communication link between the mobile robot and the edge server.

[0133] Specifically, the reflector scheduling matrix determines which reflectors participate in the operation and their operating modes, while the beam phase configuration matrix controls the phase change of each reflector. By dynamically adjusting the service targets of the reflectors, the reconfigurable smart metasurface can concentrate more reflected energy in the direction of the mobile robot requiring communication, enhancing the signal strength in that direction. Simultaneously, precise adjustment of the beam phase can change the propagation direction and beam shape of electromagnetic waves, enabling flexible reconfiguration of the communication link. For example, when the mobile robot's position changes or interference occurs in the communication environment, timely adjustment of the reflector scheduling matrix and phase configuration matrix can re-optimize the communication link, allowing the signal to bypass obstacles or avoid interference areas, thereby improving communication reliability and quality. This ability to dynamically reconfigure the communication link allows the collaborative scheduling architecture to better adapt to complex and changing environments, ensuring a stable and efficient communication connection between the mobile robot and the edge server, and guaranteeing the stable operation of the entire collaborative scheduling architecture.

[0134] S504, adjust the transmission power allocation of the edge server to each corresponding mobile robot according to the transmission power of the edge server to the corresponding mobile robot.

[0135] Specifically, the communication quality between the edge server and the mobile robot directly affects the information transmission efficiency and reliability of the collaborative scheduling architecture. Transmission power is one of the key factors influencing communication quality. Appropriate transmission power ensures sufficient signal strength to overcome noise interference during transmission, while avoiding excessive power that leads to energy waste and unnecessary electromagnetic interference. Adjusting the transmission power allocation in real-time based on the edge server's requirements for the corresponding mobile robot requires consideration of several aspects. Firstly, the power must be dynamically adjusted based on the distance between the mobile robot and the edge server, and the communication environment (such as the presence of obstacles or multipath effects). Transmission power should be increased appropriately when the distance is greater or the communication environment is poor, and decreased conversely. Secondly, the communication needs of other devices in the entire collaborative scheduling architecture must be considered to avoid adverse effects on other communication links due to power adjustments of a single mobile robot. Through precise power allocation adjustments, the communication link between the edge server and the mobile robot can be optimized, improving the stability and efficiency of information transmission and ensuring the normal operation of the collaborative scheduling architecture.

[0136] To achieve a better understanding of the system in the UAV-assisted intelligent scheduling method for large-scale mobile robots proposed in this application, the following section combines... Figure 6 This application provides a detailed description of the intelligent scheduling algorithm for solving the joint control and communication actions of a cooperative scheduling architecture. The intelligent scheduling algorithm shown in the figure consists of two main parts: a Wasserstein generative adversarial network and group relative policy optimization, and is applied to the scheduling of mobile robots in a cooperative scheduling architecture.

[0137] The Wasserstein generative adversarial network (GAN) component mainly consists of a generator network and a discriminator network. The generator network receives the input and generates the output o′∈D′, and calculates the loss function using the RMSProp optimization algorithm. The discriminant network distinguishes between the generator network output o′∈D′ and the real sample o∈D, and also uses the RMSProp optimization algorithm to calculate the loss function. The generator network and the discriminator network are trained adversarially. The generator network aims to generate more realistic samples to deceive the discriminator network, while the discriminator network tries to distinguish between real samples and generated samples. The two work together to form a mixed set of experiences.

[0138] The group-relative policy optimization section includes a policy network and a reference network. Batch sampling is performed from a mixed experience set, and the policy network updates its parameters using the Adam optimization algorithm, with the loss function being... Output strategy and group relative advantage function Reference networks provide reference strategies This is used to optimize the auxiliary policy network. Ultimately, the intelligent scheduling algorithm combines state... action and rewards Information such as these enable the scheduling of mobile robots within a collaborative scheduling architecture.

[0139] The UAV-assisted intelligent scheduling method and apparatus for large-scale mobile robots provided in this disclosure constructs a collaborative scheduling architecture including UAVs, mobile robots, and edge servers. It utilizes a reconfigurable intelligent metasurface mounted on the UAV to establish a dynamically reconfigurable deterministic communication link, effectively overcoming the communication latency and jitter problems caused by limited communication resources and channel randomness in traditional mobile robot scheduling. This provides stable and reliable communication guarantees between the mobile robot and the edge server. Based on the communication latency, jitter, and control errors of each mobile robot, a joint optimization problem with the goal of minimizing real-time tracking errors is constructed and modeled as a Markov decision process. This enables collaborative design of the motion control and communication resource allocation for the mobile robot, achieving joint optimization of control and communication. A trained intelligent scheduling algorithm is used to solve the established Markov decision process, determining the joint control and communication actions. Real-time adjustments are made to the mobile robot control commands, UAV flight paths, edge server transmission power, and the scheduling and phase configuration of the reconfigurable intelligent metasurface. This allows the entire scheduling system to dynamically adjust its strategy according to the real-time environment, improving the system's adaptability to complex dynamic environments.

[0140] Thus, this embodiment achieves low-latency, low-jitter deterministic communication, significantly reduces trajectory tracking errors of mobile robots, and improves the stability and reliability of the scheduling system. It also collaboratively optimizes control and communication resources, thereby improving the overall performance and efficiency of large-scale mobile robot scheduling.

[0141] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0142] Based on the same inventive concept, this disclosure also provides a drone-assisted large-scale mobile robot intelligent scheduling device corresponding to the drone-assisted large-scale mobile robot intelligent scheduling method. Since the principle of the device in this disclosure for solving the problem is similar to the drone-assisted large-scale mobile robot intelligent scheduling method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0143] Reference Figure 7 The diagram shown is a schematic of a drone-assisted intelligent scheduling device 700 for large-scale mobile robots provided in this embodiment of the present disclosure. The device includes:

[0144] Architecture building module 701 is used to build a collaborative scheduling architecture that includes drones, multiple mobile robots and edge servers; wherein, the drones are equipped with reconfigurable smart metasurfaces, and by adjusting the reflective elements and beam phase of the reconfigurable smart metasurfaces, dynamic reconfigurable deterministic communication links are established between each of the mobile robots and the edge servers.

[0145] The model building module 702 is used to construct a joint optimization problem with the goal of minimizing real-time tracking error based on the communication latency, jitter, and control error of each mobile robot in the cooperative scheduling architecture; and to model the joint optimization problem as a Markov decision process.

[0146] Action determination module 703 is used to solve the established Markov decision process using a trained intelligent scheduling algorithm to determine the joint control and communication actions of the cooperative scheduling architecture.

[0147] The action execution module 704 is used to adjust the mobile robot control commands, UAV flight paths, UAV reconfigurable smart metasurface scheduling and phase configuration, and edge server transmission power in the collaborative scheduling architecture in real time according to the control and communication joint action.

[0148] In some possible embodiments, the cooperative scheduling architecture specifically includes:

[0149] The UAV is equipped with a reconfigurable smart metasurface with multiple reflective elements, and moves autonomously in three-dimensional space to expand communication coverage and avoid obstacles. By adjusting the reflective element scheduling matrix and beam phase configuration matrix of the reconfigurable smart metasurface, enhanced channel services are provided for the uplink and downlink between each mobile robot and the edge server.

[0150] The mobile robot is used to receive and execute control commands and transmit information to a drone equipped with a reconfigurable smart metasurface.

[0151] The edge server is used to collect the status of each mobile robot, train and execute scheduling algorithms, and issue control commands to form a closed-loop control system.

[0152] In some possible embodiments, the joint optimization problem includes a joint optimization objective function and constraints; the action determination module 703 is specifically used for:

[0153] Using a trained intelligent scheduling algorithm, the Markov decision process established based on the joint optimization problem is solved according to the joint optimization objective function and constraints, thereby determining the joint control and communication actions for the cooperative scheduling architecture.

[0154] In some possible embodiments, the model building module 702 is specifically used for:

[0155] A Markov decision process model is established based on the joint optimization problem; wherein, the Markov decision process model includes a state space, an action space, and a reward function;

[0156] The state space includes the motion state of each mobile robot, reference state, uplink / downlink transmission packet size, transmission bandwidth, and signal-to-noise ratio;

[0157] The action space includes the control inputs of each mobile robot, the trajectory planning vector of the UAV, the transmission power of the edge server, the scheduling matrix of the reflective elements of the reconfigurable smart metasurface, and the beam phase configuration matrix.

[0158] The reward function is expressed as a normalized communication objective function that considers communication delay and jitter, and a normalized control objective function that considers control error. After being adjusted by weighting factors, the exponents are taken, and then the average value is calculated.

[0159] In some possible embodiments, the action determination module 703 is specifically used for:

[0160] The communication link established by the UAV equipped with a reconfigurable smart metasurface is used to obtain the motion state, reference state, transmission packet size, transmission bandwidth and signal-to-noise ratio of the mobile robot, which are used as the state input of the Markov decision process.

[0161] Based on the state input and the constraints, the trained intelligent scheduling algorithm outputs the control and communication joint actions of the collaborative scheduling architecture. The control and communication joint actions include the control commands of each mobile robot, the next flight position of the UAV, the transmission power of the edge server to the corresponding mobile robot, and the scheduling matrix of the reflective elements and the beam phase configuration matrix of the UAV's reconfigurable intelligent metasurface.

[0162] In some possible embodiments, the action determination module 703 is further configured to perform:

[0163] Step 1: Initialize the network parameters of the policy network and the simulation environment of the cooperative scheduling architecture;

[0164] Step 2: In the simulated environment, the current policy network outputs an action based on the current state of the simulated environment and applies the action to the simulated environment; the new state of the simulated environment after executing the action is determined, and the corresponding reward value is calculated according to the reward function; real experience data consisting of the current state, action, reward value, and new state is collected and stored in the real experience pool;

[0165] Step 3: Based on the experience enhancement mechanism of Wasserstein generative adversarial network, virtual experience data consistent with the distribution of real experience data is generated through adversarial training between the generator network and the discriminator network, and stored in the virtual experience pool;

[0166] Step 4: Sample experience data from the hybrid experience pool, which is composed of the real experience pool and the virtual experience pool;

[0167] Step 5: Based on the sampled empirical data, the group relative strategy optimization algorithm is used to calculate the group relative advantage function value and the target loss function value including the KL divergence penalty term, and the network parameters of the policy network are updated according to the calculation results;

[0168] Step 6: Repeat steps 2 to 5 to iteratively optimize the policy network parameters until the performance index of the policy network converges, thus obtaining the trained intelligent scheduling algorithm.

[0169] In some possible embodiments, the action determination module 703 is specifically used for:

[0170] Based on the control commands of each mobile robot in the joint control and communication action, corresponding motion control signals are generated and sent to each mobile robot through the edge server;

[0171] Based on the drone's next flight position, plan the drone's real-time flight path and control the drone to fly along the real-time flight path;

[0172] Based on the reflective element scheduling matrix and beam phase configuration matrix of the reconfigurable smart metasurface, the service object and beam phase of the reflective element are dynamically adjusted to reconstruct the communication link between the mobile robot and the edge server.

[0173] Based on the transmission power of the edge server to the corresponding mobile robot, adjust the transmission power allocation of the edge server to each corresponding mobile robot.

[0174] Based on the same technical concept, this disclosure also provides a computer device. (See also...) Figure 8 The diagram shows the structure of a computer device 800 provided in this embodiment of the present disclosure, including a processor 801, a memory 802, and a bus 803. The memory 802 stores execution instructions and includes a main memory 8021 and an external memory 8022. The main memory 8021, also called internal memory, is used to temporarily store computational data in the processor 801 and data exchanged with external memory 8022 such as a hard disk. The processor 801 exchanges data with the external memory 8022 through the main memory 8021.

[0175] In this embodiment, the memory 802 is specifically used to store application code that executes the solution of this application, and its execution is controlled by the processor 801. That is, when the computer device 800 is running, the processor 801 communicates with the memory 802 through the bus 803, causing the processor 801 to execute the application code stored in the memory 802, thereby executing the method described in any of the foregoing embodiments. The memory 802 may be, but is not limited to, random access memory, read-only memory, programmable read-only memory, erasable read-only memory, electrically erasable read-only memory, etc.

[0176] Processor 801 may be an integrated circuit chip with signal processing capabilities. The aforementioned processor can be a general-purpose processor, including a central processing unit (CPU), network processor, etc.; it can also be a digital signal processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor.

[0177] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the computer device 800. In other embodiments of this application, the computer device 800 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0178] This disclosure also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program performs the steps of the unmanned aerial vehicle-assisted large-scale mobile robot intelligent scheduling method described in the above-described method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.

[0179] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the UAV-assisted large-scale mobile robot intelligent scheduling method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here. The computer program product can be implemented specifically through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied as a computer storage medium; in another optional embodiment, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0180] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0181] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, the functional units in the various embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0182] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0183] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. A method for intelligent scheduling of large-scale mobile robots assisted by unmanned aerial vehicles (UAVs), characterized in that, include: Construct a collaborative scheduling architecture that includes drones, multiple mobile robots, and edge servers; wherein, the drones are equipped with reconfigurable smart metasurfaces, and by adjusting the reflective elements and beam phase of the reconfigurable smart metasurfaces, a dynamically reconfigurable deterministic communication link is established between each of the mobile robots and the edge server; In the cooperative scheduling architecture, a joint optimization problem is constructed with the goal of minimizing real-time tracking error, based on the communication latency, jitter, and control error of each mobile robot; and the joint optimization problem is modeled as a Markov decision process. The established Markov decision process is solved using a trained intelligent scheduling algorithm to determine the joint control and communication actions for the cooperative scheduling architecture. Based on the joint control and communication actions, the mobile robot control commands, UAV flight paths, scheduling and phase configuration of the UAV reconfigurable intelligent metasurface, and edge server transmission power in the collaborative scheduling architecture are adjusted in real time. The joint optimization problem includes a joint optimization objective function and constraints, wherein the joint optimization objective function is expressed as: ; ; ; The constraints include: ; ; ; ; ; ; In the formula, , , , , The set of variables to be optimized in the problem; This represents the control input for the mobile robot. ; This is represented as the path planning for the drone; This is expressed as the transmit power of the edge server. ; Represented as the reflective element scheduling matrix of a reconfigurable smart metasurface. ; Represented as the reflection angle of a reconfigurable smart metasurface reflective element; Represented as the control objective function The normalization function; Represented as the communication objective function The normalization function; , These are represented as the first weighting coefficient and the second weighting coefficient, respectively; M represents the total number of mobile robots; Let be the trajectory tracking error of the m-th mobile robot at time k; This represents the control input of the m-th mobile robot at time k; This represents a set of mobile robots; C1 and C2 are the control constraints of the mobile robots. and These are the minimum and maximum values ​​of the control input for the mobile robot, respectively. This represents the control input for the m-th mobile robot; This represents the communication latency between the m-th mobile robot and the edge server. C1 represents the scheduling deadline for the m-th mobile robot; C2 represents the trajectory planning constraints for the UAV. This is represented by the three-dimensional coordinates of the UAV at time k; This represents the three-dimensional coordinates of the UAV at time k-1; K represents the total number of time steps. This represents the maximum flight speed of the drone; C4, C5, and C6 are the communication constraints for the edge server. This represents the transmission power from the edge server to the m-th mobile robot. This represents the maximum transmission power of the edge server; Represented as a reflective element of a reconfigurable smart metasurface in the i-th row and j-th column. Is it applied to the first The signal reflection of a mobile robot: 0 indicates no application, 1 indicates application; The first reconfigurable smart metasurface is denoted as the... a beam phase configuration of the reflective elements, a total number of reflective elements representing a reconfigurable intelligent metasurface, , and are constant positive definite matrices; The step of modeling the joint optimization problem as a Markov decision process includes: A Markov decision process model is established based on the joint optimization problem; wherein, the Markov decision process model includes a state space, an action space, and a reward function; The state space includes the motion state of each mobile robot, reference state, uplink / downlink transmission packet size, transmission bandwidth, and signal-to-noise ratio; The action space includes the control inputs of each mobile robot, the trajectory planning vector of the UAV, the transmission power of the edge server, the scheduling matrix of the reflective elements of the reconfigurable smart metasurface, and the beam phase configuration matrix. The reward function is expressed as a normalized communication objective function that considers communication delay and jitter, and a normalized control objective function that considers control error. After being adjusted by weighting factors, the exponents are taken, and then the average value is calculated.

2. The method of claim 1, wherein, The collaborative scheduling architecture specifically includes: The UAV is equipped with a reconfigurable smart metasurface with multiple reflective elements, and moves autonomously in three-dimensional space to expand communication coverage and avoid obstacles. By adjusting the reflective element scheduling matrix and beam phase configuration matrix of the reconfigurable smart metasurface, enhanced channel services are provided for the uplink and downlink between each mobile robot and the edge server. The mobile robot is used to receive and execute control commands and transmit information to a drone equipped with a reconfigurable smart metasurface. The edge server is used to collect the status of each mobile robot, train and execute scheduling algorithms, and issue control commands to form a closed-loop control system.

3. The method of claim 1, wherein, The process of solving the established Markov decision process using a trained intelligent scheduling algorithm to determine the joint control and communication actions for the cooperative scheduling architecture includes: Using a trained intelligent scheduling algorithm, the Markov decision process established based on the joint optimization problem is solved according to the joint optimization objective function and constraints, thereby determining the joint control and communication actions for the cooperative scheduling architecture.

4. The method of claim 3, wherein, The step of using a trained intelligent scheduling algorithm to solve the Markov decision process established based on the joint optimization problem according to the joint optimization objective function and constraints includes: The communication link established by the UAV equipped with a reconfigurable smart metasurface is used to obtain the motion state, reference state, transmission packet size, transmission bandwidth and signal-to-noise ratio of the mobile robot, which are used as the state input of the Markov decision process. Based on the state input and the constraints, the trained intelligent scheduling algorithm outputs the control and communication joint actions of the collaborative scheduling architecture. The control and communication joint actions include the control commands of each mobile robot, the next flight position of the UAV, the transmission power of the edge server to the corresponding mobile robot, and the scheduling matrix of the reflective elements and the beam phase configuration matrix of the UAV's reconfigurable intelligent metasurface.

5. The method of claim 4, wherein, The intelligent scheduling algorithm is trained through the following steps: Step 1: Initialize the network parameters of the policy network and the simulation environment of the cooperative scheduling architecture; Step 2: In the simulated environment, the current policy network outputs an action based on the current state of the simulated environment and applies the action to the simulated environment; the new state of the simulated environment after executing the action is determined, and the corresponding reward value is calculated based on the reward function; Collect real-world experience data consisting of the current state, actions, reward values, and new states, and store it in the real-world experience pool; Step 3: Based on the experience enhancement mechanism of Wasserstein generative adversarial network, virtual experience data consistent with the distribution of real experience data is generated through adversarial training between the generator network and the discriminator network, and stored in the virtual experience pool; Step 4: Sample experience data from the hybrid experience pool, which is composed of the real experience pool and the virtual experience pool; Step 5: Based on the sampled empirical data, the group relative strategy optimization algorithm is used to calculate the group relative advantage function value and the target loss function value including the KL divergence penalty term, and the network parameters of the policy network are updated according to the calculation results; Step 6: Repeat steps 2 to 5 to iteratively optimize the policy network parameters until the performance index of the policy network converges, thus obtaining the trained intelligent scheduling algorithm.

6. The method of claim 4, wherein, The real-time adjustment of mobile robot control commands, UAV flight paths, scheduling and phase configuration of UAV reconfigurable intelligent metasurfaces, and edge server transmit power in the collaborative scheduling architecture based on the joint control and communication actions includes: Based on the control commands of each mobile robot in the joint control and communication action, corresponding motion control signals are generated and sent to each mobile robot through the edge server; Based on the drone's next flight position, plan the drone's real-time flight path and control the drone to fly along the real-time flight path; Based on the reflective element scheduling matrix and beam phase configuration matrix of the reconfigurable smart metasurface, the service object and beam phase of the reflective element are dynamically adjusted to reconstruct the communication link between the mobile robot and the edge server. Based on the transmission power of the edge server to the corresponding mobile robot, adjust the transmission power allocation of the edge server to each corresponding mobile robot.

7. An unmanned aerial vehicle-assisted large-scale mobile robot intelligent scheduling apparatus, characterized in that, include: An architecture building module is used to build a collaborative scheduling architecture that includes drones, multiple mobile robots, and edge servers. The drones are equipped with reconfigurable smart metasurfaces, and by adjusting the reflective elements and beam phase of the reconfigurable smart metasurfaces, dynamic and reconfigurable deterministic communication links are established between each mobile robot and the edge server. The model building module is used to construct a joint optimization problem with the goal of minimizing real-time tracking error based on the communication latency, jitter, and control error of each mobile robot in the cooperative scheduling architecture; and to model the joint optimization problem as a Markov decision process. The action determination module is used to solve the established Markov decision process using a trained intelligent scheduling algorithm to determine the joint control and communication actions for the cooperative scheduling architecture. The action execution module is used to adjust the mobile robot control commands, UAV flight paths, UAV reconfigurable smart metasurface scheduling and phase configuration, and edge server transmission power in the collaborative scheduling architecture in real time according to the control and communication joint actions. The joint optimization problem includes a joint optimization objective function and constraints, wherein the joint optimization objective function is expressed as: ; ; ; The constraints include: ; ; ; ; ; ; In the formula, , , , , The set of variables to be optimized in the problem; This represents the control input for the mobile robot. ; This is represented as the path planning for the drone; This is expressed as the transmit power of the edge server. ; Represented as the reflective element scheduling matrix of a reconfigurable smart metasurface. ; Represented as the reflection angle of a reconfigurable smart metasurface reflective element; Represented as the control objective function The normalization function; Represented as the communication objective function The normalization function; , These are represented as the first weighting coefficient and the second weighting coefficient, respectively; M represents the total number of mobile robots; Let be the trajectory tracking error of the m-th mobile robot at time k; This represents the control input of the m-th mobile robot at time k; This represents a set of mobile robots; C1 and C2 are the control constraints of the mobile robots. and These are the minimum and maximum values ​​of the control input for the mobile robot, respectively. This represents the control input for the m-th mobile robot; This represents the communication latency between the m-th mobile robot and the edge server. C1 represents the scheduling deadline for the m-th mobile robot; C2 represents the trajectory planning constraints for the UAV. This is represented by the three-dimensional coordinates of the UAV at time k; The coordinates of the UAV at time k-1 are represented by 3D coordinates; K represents the total number of time steps. This represents the maximum flight speed of the drone; C4, C5, and C6 are the communication constraints for the edge server. This represents the transmission power from the edge server to the m-th mobile robot. This represents the maximum transmission power of the edge server; Represented as a reflective element of a reconfigurable smart metasurface in the i-th row and j-th column. Is it applied to the first The signal reflection of a mobile robot: 0 indicates no application, 1 indicates application; The first reconfigurable smart metasurface is denoted as the... Beam phase configuration of each reflective element This represents the total number of reflective elements on the reconfigurable smart metasurface. , and All are constant positive definite matrices; Specifically, the model building module is used for: A Markov decision process model is established based on the joint optimization problem; wherein, the Markov decision process model includes a state space, an action space, and a reward function; The state space includes the motion state of each mobile robot, reference state, uplink / downlink transmission packet size, transmission bandwidth, and signal-to-noise ratio; The action space includes the control inputs of each mobile robot, the trajectory planning vector of the UAV, the transmission power of the edge server, the scheduling matrix of the reflective elements of the reconfigurable smart metasurface, and the beam phase configuration matrix. The reward function is expressed as a normalized communication objective function that considers communication delay and jitter, and a normalized control objective function that considers control error. After being adjusted by weighting factors, the exponents are taken, and then the average value is calculated.

8. A storage medium having stored thereon a computer program, characterized in that When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.

9. A computer device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Optimization method for communication system of auxiliary advancing vehicle of unmanned aerial vehicle

    CN113747397A

  • Unmanned aerial vehicle-mounted RIS auxiliary vehicle network communication method and system

    CN115915069A