Edge cloud cooperative control method for multi-axis servo system

By adopting the edge-cloud collaborative control framework based on deep reinforcement learning in a networked multi-axis servo system, dynamically selecting the control solution, the delay and loss problems of the system in the edge and cloud coordinated control are solved, and more efficient multi-axis coordinated control performance is achieved.

CN120215419APending Publication Date: 2025-06-27NANJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510306784.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-15
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing networked multi-axis servo systems have shortcomings in the coordinated control of edge and cloud, especially in dynamic process control, and the problems of network delay and packet loss have not been fully solved.

Method used

Adopting an edge-cloud collaborative control framework based on deep reinforcement learning, by building an edge-cloud collaborative model of multi-axis servo system, dynamically selecting appropriate control solutions, combining the low latency of edge devices and high computing power in the cloud, scheduling strategies are optimized to improve coordinated control performance.

Benefits of technology

By optimizing scheduling strategies, effectively leveraging the advantages of the edge and cloud, the target tracking performance of multi-axis systems is improved, control costs are reduced, and the overall performance of the system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215419A_ABST
    Figure CN120215419A_ABST
Patent Text Reader

Abstract

The invention discloses an edge cloud cooperative control method for a multi-axis servo system, and provides an optimized edge cloud cooperative control framework combining the advantages of edge computing and cloud computing aiming at the condition that a plurality of motion axes in the multi-axis cooperative system track different target trajectories during cooperative work. Considering the condition that the state and the control sequence are transmitted through a limited shared channel, an estimated state, an estimated covariance, cache content and a future instruction are further constructed as a system state based on an enhanced Q backpack method, a steady-state Markov decision problem is constructed, and the decision of each step is converted into a packet backpack problem of an estimated Q value, so that the system performance is improved. Through uplink of a scheduling state and downlink of a control sequence, the overall performance of the system is maximized under the constraint that the number of distributable time slots is limited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer control systems and artificial intelligence, and particularly relates to a coordinated control framework for a networked multi-axis servo system, and more particularly to an edge-cloud collaborative control framework based on a deep reinforcement learning method constructed for the differences in target trajectories between axes. Background Art

[0002] As a key component of the industrial Internet of Things, a multi-axis servo system is a servo system with multiple independent control axes and has been widely used in fields such as numerically controlled machine tools (CNCs), robotic arms, and automated assembly lines. A typical multi-axis servo system drives each axis to accurately move along a predetermined target trajectory through a trajectory tracking control scheme and completes complex tasks through coordinated work. With the introduction of networked control, the controller receives the status information of each axis through the cloud and transmits control signals to the corresponding axis through the network, enabling the multi-axis servo system to be deployed and implemented more flexibly and cost-effectively in an industrial environment.

[0003] Although networked multi-axis servo systems have the potential to offer many advantages, there are still some key issues that have not been fully studied. Specifically, most current research on networked multi-axis servo systems focuses on optimizing limited network resources to reduce network latency and packet loss, thereby ensuring cloud-based control performance. However, the introduction of edge computing is often overlooked. In addition, the collaborative control between the edge and the cloud has potential for multi-axis servo systems and is worthy of in-depth study. Specifically, edge devices perform well in steady-state control, but due to limited computing power, they can only execute relatively simple control schemes, while the cloud can run more complex control schemes and perform excellently in dynamic process control, but it will bring inevitable network latency and packet loss. This indicates that it is necessary to dynamically select appropriate control schemes, combine the respective advantages of the edge and the cloud, to adapt to different target trajectories and control subtasks, thereby improving the coordinated control performance of the multi-axis servo system. Summary of the Invention

[0004] Object of the Invention: The present invention aims to provide an edge-cloud collaborative control method for a multi-axis servo system, which gives full play to the high computing power of the cloud and the low latency and anti-interference ability of the edge for industrial Internet of Things applications, and dynamically selects a suitable control scheme according to the target trajectory of each axis, thereby improving the coordinated control performance.

[0005] Technical Solution: An edge-cloud collaborative control method for a multi-axis servo system, which includes the following steps:

[0006] (1)Construct a multi-axis servo system model with edge-cloud collaboration, where each independent axis tracks a time-varying target trajectory. The edge device equipped with each axis calculates the edge control signal through a suboptimal LQT controller, and the cloud calculates the optimal control sequence through MPC. The system decides at each moment whether to upload the state signal of each axis to the cloud or whether to download the control sequence to the edge device;

[0007] (2)By constructing an observable state space, the scheduling problem is reformulated as a steady-state Markov decision process in the cloud. Subsequently, the reinforcement Q-knapsack method based on reinforcement learning is used to optimize the scheduling strategy. The optimal action selection problem at each time step is converted into a grouped Q-value knapsack problem, and dynamic programming is used to determine the optimal action combination. Finally, DDQN is used to estimate the Q-value of the extended state-action pair, thereby optimizing the long-term performance of the system.

[0008] In step (1) of the above method, the cloud first calculates the control signal sequence of each axis in a future period of time, and transmits the control signal sequence to the corresponding edge device through a shared network channel and stores it in the buffer of the edge device; for each axis, when the buffer is available, the control signal is first taken out from the buffer and applied to the axis; otherwise, the edge device uses a suboptimal tracker to calculate the corresponding control signal;

[0009] At each time step, for each axis, it is necessary to judge whether to upload the state of the axis to the cloud, whether to download the control signal sequence calculated by the cloud to the buffer of the edge device, and the length of the control signal sequence, in order to maximize the coordinated control performance.

[0010] Furthermore, in order to describe the behavior of the multi-axis system, network communication, and edge-cloud collaboration mechanism in the edge-cloud collaborative control method for a multi-axis coordinated servo system, the method is modeled through the following process and decision-making in step (1):

[0011] (11)Controlled object It is described by a discrete linear time-invariant stochastic system equation, and its expression is:

[0012]

[0013] where and respectively represent the state vector, input vector, and output vector at time step k. Matrices and respectively represent the state transition matrix, input matrix, and output matrix; for the initial state x l (0) and the disturbance w l(k)(k ∈ N) are independent and identically distributed (i.i.d.) random variables, following a normal distribution with mean 0 and covariance Ξ l In the actual system (A l , B l ) is controllable;

[0014] (12) Define as the multi-axis coordinated task instruction, which can be decomposed into independent target tracking tasks for each axis where, represents the reference output of the axis at time step k. The total task duration K is large but finite. Although the tracking targets of each axis vary with time and are different, the tracking targets of a single axis remain unchanged during certain time periods. Therefore, for each axis satisfy the following conditions:

[0015]

[0016] where, P l is a very small subset, i.e., |P l | << K, and its element k satisfies In many cases, the tracking target of a single axis remains a constant vector for some time.

[0017] (13) At time step k, the method assumes that for all transmitting the state information x l (k) occupies one time slot in IEEE 802.15.4, while transmitting a control sequence containing h l (k) control steps occupies time slots, where α ∈ (0, 1]. In particular, when h l (k) = 0, it also occupies one time slot.

[0018] Define the decision variable to indicate whether to allocate a time slot for uploading x l (k). Similarly, define the decision variable to indicate whether to allocate a time slot for transmitting . h l (k) is also regarded as a decision variable. Crucially, at any time step, the total number of time slots occupied by all state information and control information used for decision-making shall not exceed the maximum value M ∈ Z + allowed by the system.

[0019] (14) The edge ε l mainly focuses on the performance of the controlled object , and its control objective is to minimize the cost function as follows:

[0020]

[0021] Obtain a function of x through LQT l (k) feedback gain and regarding Sub-optimal tracker of feedforward gain:

[0022]

[0023] (15) Cloud Use MPC to calculate Optimal control signal in the future time domain, define the optimal control sequence as The optimal control problem (OCP) is formulated as:

[0024]

[0025] (16) At time step k, if the decision variable then The first h l (k) control signals (denoted as ) are transmitted to the edge ε through GTS l , otherwise no transmission is performed. Among them, h l (k) ∈ {0, 1, …, H} c , H c represents the maximum transmission length.

[0026] In the edge ε l A buffer is provided for storing the subsequent control sequence sent by the cloud Let the matrix represent the content of the buffer at time step k. The update mechanism of and the actual input action of the actuator can be expressed as:

[0027]

[0028] Among them, the symbol [] represents an empty matrix, the matrix b is used to extract the first (i.e., the current time step) control signal from the buffer , and the matrix H is used to remove the first control signal from .

[0029] Cloud Utilize and ACK frame information to simulate the buffer update process, so as to maintain the mirror image of in real time at the cloud end Due to the existence of transmission delay T s When the cloud Only the true state x at time step k-1 can be obtained l Let and respectively represent the state estimate and input estimate of at time step k, which can be calculated as:

[0030]

[0031] The system state prediction at time step k+1 is

[0032]

[0033] (17) The goal of this method is to maximize the expected overall performance of the multi-axis system at each time step by making decisions on the upload of the system state x l (k), the downlink transmission of the control sequence and its length h l (k), while ensuring that the number of superframe time slots used within each time step does not exceed the system constraint M

[0034] Mathematically, this constrained scheduling problem can be formulated as:

[0035]

[0036] Furthermore, to optimize the above constrained scheduling problem, the method converts the original problem into a Markov problem in step (2) and uses the reinforcement Q-knapsack method based on reinforcement learning to optimize the scheduling strategy, where

[0037] (21) Define as the estimation error of the cloud for at time step k, and let represent the closed-loop system matrix used to control ε l .

[0038] The present invention proves Lemma 1: E l (k) follows a Gaussian distribution with a mean of 0 and a covariance of Σ l (k), where Σ l (k) satisfies:

[0039]

[0040] (22) Define states, actions, and rewards for analysis

[0041] Considering the information available in the cloud, the present invention defines the observable state at time step k as s(k), and the set of all feasible states is denoted as Here, is constructed as:

[0042]

[0043] s(k) = {s1(k), …, s L (k)}

[0044] where represents the control command sequence from time step k to K, expressed as:

[0045]

[0046] Define as the action taken by the system at time step k, where is the set of all possible actions:

[0047] a(k) = {a1(k), …, a L (k)}

[0048]

[0049] Define r l (k) as the uniaxial reward obtained by the controlled object at time step k, and r(k) as the overall system reward at time step k:

[0050]

[0051] Prove Theorem 1: The decision-making process represented by s(k), a(k), and r(k) constitutes a stationary Markov decision process (SMDP).

[0052] (23) Convert it into a policy optimization problem in the reinforcement learning (RL) framework.

[0053] Define as a policy function that maps the state s(k) to the action a(k), and V π is the state value function under the policy π:

[0054]

[0055] where, when k = 0, obviously V π (s(0)) is the negative value of the objective function under the policy π. Therefore, the original problem is equivalent to:

[0056]

[0057] (24) Considering that the sequence a l only affects the sequence r l , the present invention will combine the action combinations of each axis while maintaining the independence of each axis.

[0058] Define decision variables It corresponds to a l (k) as follows:

[0059] Calculate If Then If Then

[0060] Calculate If Then Otherwise, And

[0061] Among them, the set of valid actions a + (k) is defined as

[0062] (25) Inspired by the knapsack problem, a grouped Q-knapsack algorithm based on dynamic programming is proposed in the present invention. First, define As when , the number of time slots consumed by the action, and this value can be pre-calculated and remain unchanged during the execution process. Next, define As the cumulative expected reward of axis l obtained by forcing the action under the system state s(k) according to the policy π, and its formula is:

[0063]

[0064] Regard all actions of an axis as a group of items, where the value and volume of each item at time k are represented by and respectively.

[0065] The objective of the present invention is to select one item from each of the L groups so that the total value of these L items is maximized while their total volume does not exceed M. The grouped Q-knapsack algorithm iteratively calculates the optimal solution of the above problem by dynamically deriving two optimal substructures dp and parent.

[0066] (26) Represent the mapping from s(k) to a + (k) as π. Denote:

[0067]

[0068] Abbreviate the grouped Q-knapsack as

[0069] Theorem 2: The policy is better than π(·), that is, V π'(s(0)) ≥ V π (s(0)).

[0070] (27) Due to the high-dimensionality of the state s(k) and its complex relationship with the Q-value, a deep neural network is adopted as the Q-value estimator.

[0071] Furthermore, the extended state is defined as:

[0072]

[0073] where represents the potential impact of the estimation error on the expected reward. n l (k) ∈ 0, 1, …, H c represents the length of the buffer at time step k . represents the i-th element of the buffer , padded with H c -n l (k) zeros at the end to keep the length consistent.

[0074] (28) and its weights are represented by the parameter θ. Since s + (k) contains variable-length sequence data (y r , V * and V ◇ ), the present invention first uses an LSTM encoder to extract features from these sequence data. Subsequently, the extracted features are concatenated with the fixed-length state features and further encoded through a multi-layer perceptron (MLP). Since the output consists of L sets of Q-values, the present invention uses L independent linear heads, each head outputting 2(H c +2) Q-values.

[0075] Furthermore, the DDQN algorithm is adopted to estimate the Q-value and iteratively optimize the policy, denoted as π θ . Different from the typical optimization objective, the multi-axis environment allows the calculation of independent single-axis rewards, expressed as:

[0076] r + (k) = {r1(k), …, r L (k)}

[0077] To reduce the variance of the Q-value estimation, the loss is calculated for the Q-values output by each head; the total loss function is defined as follows:

[0078]

[0079] where θ -represents the parameters of the target network, which remain unchanged for a certain period. γ is the discount factor for stabilizing training, and ζ is the learning rate. represents the replay buffer, which is used to store the experiences collected during the interaction with the environment.

[0080] Beneficial effects: The present invention addresses the edge-cloud collaborative control problem of a multi-axis coordination system and solves the control problem where each axis in the system has different targets. Specifically, to make full use of the high computing power of the cloud and the real-time advantages of edge computing, the present invention is based on the IEEE 802.15.4-based edge-cloud collaborative network framework. This framework uploads the state of the controlled object to the cloud or downloads the optimal control sequence calculated by the cloud to the buffer at the edge, aiming to minimize the control cost of the multi-axis target tracking task. The present invention provides a novel reinforcement learning Q-knapsack method, which transforms the scheduling problem at each time step into a combination of selecting a set of action Q-values and uses DDQN to optimize the long-term decision-making performance. The simulation results show that the proposed RQK method is superior to similar algorithms in edge-cloud collaborative scheduling. This method not only ensures that the decision-making satisfies the communication constraints but also improves the target tracking performance of the multi-axis system by making full use of the advantages of edge and cloud computing. Description of the Drawings

[0081] Figure 1 is a schematic diagram of a multi-axis servo system taking a four-axis CNC machine tool as an example in the present invention;

[0082] Figure 2 is a schematic diagram of the structure of a superframe in the IEEE 802.15.4 protocol used for network communication in the present invention;

[0083] Figure 3 is a schematic diagram of the network architecture of the Q-value estimation neural network in the present invention;

[0084] Figure 4 is a convergence graph of the rewards obtained during the training process of the reinforcement Q-knapsack method in the present invention;

[0085] Figure 5 is a comparison graph of the four-dimensional target trajectory and the response trajectory in the demonstration experiment of the present invention;

[0086] Figure 6 is for the present invention using the reinforcement Q-knapsack method target / response output trajectory and decision-making scheduling graph;

[0087] Figure 7 is a comparison graph of the control costs of different scheduling methods under different disturbance coefficients in the present invention;

[0088] Figure 8 is a comparison graph of the control costs of different scheduling methods under different numbers of axes in the present invention. Detailed implementation manners

[0089] To elaborate on the technical solution disclosed by the present invention in detail, the present invention will be further described below in conjunction with the accompanying drawings and specific implementation manners.

[0090] First of all, the key problem to be solved by the method of the present invention is the problem that edge computing is often ignored in the control of networked multi-axis servo systems, and it focuses on the characteristics that the edge and the cloud have their own advantages to adapt to different target trajectories, and uses a data-driven edge collaboration method to improve the coordinated control performance of the multi-axis servo system.

[0091] The present invention first mathematically models the multi-axis servo system, its trajectory tracking task, communication network, and edge-cloud collaboration mechanism, and then constructs an optimization problem for the coordinated control performance of the multi-axis servo system. To optimize this problem, the present invention constructs state, action, and reward functions to transform the above optimization problem into a policy optimization problem under the framework of a Markov decision problem. Then, based on the enhanced Q-knapsack algorithm, which regards the combined action optimization problem at each time step as a grouped knapsack problem about the number of consumed slots and the action Q value, and uses the dynamic programming method to solve the single-step optimal combined action, and then uses DDQN to iteratively optimize the estimation of the action Q value. The present invention mathematically proves that the enhanced Q-knapsack method can effectively optimize and improve the coordinated control performance of the multi-axis servo system. Finally, demonstration and comparative experiments are used to show that the edge-cloud collaborative control method for multi-axis coordinated servo systems invented is superior to other methods.

[0092] To describe the steps of the method of the present invention in detail, the present invention defines the following symbols: N represents the set of all natural numbers. G T represents the transpose of matrix G. When matrix G is positive definite (semi-positive definite), it is denoted as G > 0 (G ≥ 0). For vector x, it is defined as Tr(A) is defined as the trace of square matrix A. Given sets P and Q, where the difference set of the two sets is defined as |P| represents the number of elements in set P. Given random variable X, its expectation is denoted as E[X], and its covariance matrix is denoted as Conv(X). The list concatenation operator is defined as

[0093] An edge-cloud collaborative control method for a multi-axis servo system, the specific implementation includes:

[0094] Step 1: Construct a multi-axis servo system model for edge-cloud collaboration

[0095] The multi-axis servo system consists of L axes, where represents the set of axis labels. The multi-axis coordination task It can be decomposed into trajectory tracking tasks for each axis, where the target trajectories of each axis are often different. As an example, Figure 1 shows a multi-axis CNC milling machine, where the servo motors of different axes track their respective reference trajectories, enabling the cutting head to machine complex curves in three-dimensional space.

[0096] In a multi-axis servo system, although each axis is dynamically independent and physically far apart from each other, for each axis controlled object sensor actuator and the edge device ε l are usually highly integrated. At time step k, ε l can obtain the axis state x detected by l (k) without delay and error, and use a suboptimal discrete linear quadratic tracker (LQT) to calculate the edge control signal cloud communicates with each ε l through the shared channel of the IEEE 802.15.4 wireless personal area network protocol, while ε l do not communicate directly with each other. Due to communication delays and limited channel resources, the cloud runs an estimator to estimate the system state of all axes, denoted as and runs a high-computing-power-demand model predictive control (MPC). The cloud solves the optimal control problem (OCP) according to (k) and the future target sequence to obtain the optimal control sequence

[0097] controlled object is described by a discrete linear time-invariant stochastic system equation, and its expression is:

[0098]

[0099] where, and represent the state vector, input vector, and output vector at time step k, respectively. Matrices and represent the state transition matrix, input matrix, and output matrix, respectively. For the initial state x l (0) and the disturbance w l (k) (k ∈ N) are independent and identically distributed (i.i.d.) random variables, following a normal distribution with a mean of 0 and a covariance of Ξl. In an actual system, (A l , B l ) is controllable.

[0100] The present invention defines as a multi-axis coordinated task instruction, which can be decomposed into independent target tracking tasks for each axis wherein, represents the reference output of the axis at time step k, and the target sequence has been sent to the corresponding ε via the downlink before the start of control l . Considering the actual requirements of the multi-axis coordinated task, the total task duration K is large but finite. Although the tracking targets of each axis vary with time and are different from each other, the tracking targets of a single axis remain unchanged during certain time periods. Therefore, for each axis satisfies the following conditions:

[0101]

[0102] wherein, P l is a very small subset, i.e., |P l | << K, and its element k satisfies In many cases, the tracking target of a single axis remains a constant vector for a period of time.

[0103] The multi-axis coordinated target tracking task must comprehensively consider the output error cost and control input cost of all axes. Therefore, for a given tracking target the total cost is expressed as:

[0104]

[0105] wherein, is the error vector. G l ≥ 0 is the state weight matrix, and R l ≥ 0 is the control input weight matrix. In a practical system, is observable.

[0106] IEEE 802.15.4 is a communication standard specifically designed for low-rate wireless personal area networks (PANs), mainly used to support short-distance, low-power, and low-data-rate wireless communications. In the invention, each ε l is equipped with a terminal device for communicating with the having a PAN coordinator. The network topology is star-shaped with the cloud as the center.

[0107] The MAC layer of IEEE 802.15.4 adopts the beacon-enabled mode to achieve time synchronization between the PAN coordinator and the terminal device through repeated superframes. As Figure 2As shown, each superframe starts with a beacon frame sent by the PAN coordinator. The beacon frame provides the duration of the superframe and the channel allocation information during this period. The present invention assumes that the time interval between consecutive beacon frames, i.e., the beacon interval (BI), is the same as the sampling period T of the discrete system. s Same.

[0108] The calculation formula for the beacon interval is BI = aBaseSuperframeDuration × 2^BO, where BO is the beacon order, and its value range is from 0 to 14. Therefore, the shortest BI is aBaseSuperframeDuration. For the 2.4GHz frequency band, this duration is usually equal to 960 symbol times, i.e., 15.36 milliseconds. This time interval is sufficient to meet the requirements of applications with high real-time requirements and can be used as the sampling period when discretizing the linear system. BO

[0109] During the inactive period of the superframe, communication between PAN devices is not carried out to save energy, while the active period of the superframe is divided into 16 time slots and further divided into a contention access period (CAP) and a contention-free period (CFP). The PAN coordinator allocates several guaranteed time slots (GTS) within the CFP and broadcasts the allocation information to all terminal devices through the beacon frame. Each GTS occupies one or more consecutive time slots to ensure communication with a specific terminal device. Many literatures discuss and modify the limitations of the IEEE802.15.4 MAC protocol, allowing the number of time slots allocated to the CAP and CFP to be adjustable. In the edge-cloud collaboration framework, the system performance is improved by scheduling the upload of servo state information and the downlink of control sequences. Given the limited number of time slots, the scheduling of the CFP is mainly concerned, while the CAP can be used for other data streams. For the 2.4GHz frequency band, the data rate of IEEE 802.15.4 is 250 kbps. The duration of one time slot is 0.24 milliseconds and can carry 60 bits of data.

[0110] At time step k, the present invention assumes that for all transmitting the state information x(k) of a single servo system occupies one time slot, while transmitting a control sequence containing h(k) control steps l occupies l αh(k) time slots, where α ∈ (0,1]. In particular, when h(k) = 0, it also occupies one time slot. Define the decision variable to represent whether a time slot is allocated for uploading x(k). Similarly, define the decision variable to represent whether a time slot is allocated for transmitting l l ​​​​​​l (k) is also regarded as a decision variable. Importantly, at any time step, the total number of time slots occupied by all state information and control information for decision-making shall not exceed the maximum value M ∈ Z allowed by the system. + .

[0111] Considering the significant impact of packet loss in wireless transmission on system performance, the present invention uses a Bernoulli process to model the packet loss problem. Specifically, indicator variables and are defined to represent the packet loss event from time k - 1 to k, where represents the packet loss from ε l to , represents the packet loss from to ε l . Its expression is:

[0112]

[0113] where, and are the packet loss probabilities in the two directions respectively.

[0114] The present invention defines indicator variables and to represent whether the state information and control sequence of are successfully transmitted at time step k respectively. When or , it indicates successful transmission; when the value is 0, it indicates transmission failure.

[0115] The edge ε l mainly focuses on the performance of the controlled object , and its control objective is to minimize the cost function According to the certainty equivalence property, for a linear stochastic system, the optimal control solution is the same as that of the system without additive perturbation. Therefore, for any axis in, ε l the following discrete linear quadratic tracking (LQT) problem needs to be solved:

[0116]

[0117] According to the optimal control law, the optimal affine control input of axis at time step k can be rewritten as:

[0118]

[0119] where, the feedback gain and the feedforward gain The calculation formula is:

[0120]

[0121] Auxiliary sequence S l (k) and v l (k) is obtained by backward recursion from the boundary conditions at time step K, and the boundary conditions are and Their generation formulas are:

[0122]

[0123] Although the sequences and v l (k) can be calculated offline, this computational requirement is unacceptable for edge devices that need to start tasks immediately after receiving commands. Considering that for most time k, the command sequence for a relatively long future period (e.g., from time step k to N) is constant (denoted as ), the gains and of the tracker will converge rapidly to the steady-state values and as N - k increases. At the same time, v l (k) will converge to a constant linear relationship with y r , expressed as Therefore, at time step k, a sub-optimal tracker κ l with fixed coefficients and simple calculations can be calculated, and its expression is:

[0124]

[0125] where represents the affine gain.

[0126] Cloud uses MPC to calculate the optimal control signal over the future time domain. However, since the control signal is transmitted in superframes at intervals of T s , it is necessary to calculate the optimal control sequence starting from time step k + 1 at time step k. Let and w l (i|k) represent the system state prediction, output prediction, target, input prediction, and external disturbance for time step k + i based on the information available at time step k, respectively. The total predicted cost between time steps k + 1 and K is defined as:

[0127]

[0128] Due to the existence of random perturbations, MPC minimizes the objective function to the expected value of the total cost. The present invention defines the optimal control sequence as The optimal control problem (OCP) is formulated as:

[0129]

[0130] The above problem can be transformed into an unconstrained quadratic programming problem through matrix augmentation, and thus solved using an analytical formula

[0131] Next, the specific edge-cloud collaboration mechanism is described. At time step k, if the decision variable then the first h l (k) control signals (denoted as ) are transmitted to the edge ε through GTS l , otherwise no transmission is performed. Among them, h l (k) ∈ {0, 1, …, H} c , and H c represents the maximum transmission length. This process can be expressed as

[0132]

[0133] At the edge , a buffer is provided to store the subsequent control sequence sent by the cloud . Let the matrix represent the content of the buffer at time step k. The update mechanism of

[0134]

[0135] and the actual input action of the actuator can be expressed as: where the symbol [] represents an empty matrix, the matrix b is used to extract the first (i.e., the current time step) control signal in the buffer , and the matrix H is used to remove the first control signal from

[0136] The edge-cloud collaboration mechanism operates as follows: when is non-empty, the first control action in is preferentially selected as the actual control action u l (k). Otherwise, u l (k) is set to this value is calculated by κ l . When , the old content of is replaced by the new control sequence .

[0137] It should be noted that when k = 0, or when the control actions in but when is still an empty matrix until a non-empty is transmitted to

[0138] In the IEEE 802.15.4 protocol, after a data frame is transmitted, the receiving device will send an acknowledgment (ACK) frame to confirm the successful reception of the data frame. Therefore, it is possible to utilize and the ACK frame information to simulate the buffer update process, thereby maintaining the mirror image of in real time at the end. Due to the existence of the transmission delay T s when when only the true state x l at time step k - 1 can be obtained. Let and respectively represent the state estimate and input estimate of at time step k, and they can be calculated as:

[0139]

[0140] The system state prediction at time step k + 1 is:

[0141]

[0142] The objective of the present invention is to maximize the expected overall performance of the multi-axis system at each time step by making decisions on the upload of the system state x l (k), the downlink transmission of the control sequence and its length h l (k), on the premise of ensuring that the superframe time slots used within each time step do not exceed the system constraint M. Mathematically, this constrained scheduling problem can be expressed as:

[0143]

[0144] Step 2: Optimize the system performance using the enhanced Q-knapsack method

[0145] The present invention defines as the estimation error of the cloud for at time step k, and let represent the closed-loop system matrix used to control .

[0146] The present invention proves Lemma 1: El (k) follows a Gaussian distribution with mean 0 and covariance Σ l where Σ l (k) satisfies:

[0147]

[0148] The state of the edge-cloud collaborative control mechanism at time step k-1 can be discussed in four different cases.

[0149] Case 1: Assume k = 0,

[0150] Case 2: Assume k > 0 and Since it is known that If then Otherwise, Therefore,

[0151] Case 3: Assume k > 0 and and Since Therefore, E[E l (k)] = A l [El(k - 1)],

[0152] Case 4: Assume k > 0 and and Since Therefore,

[0153] Assume k = 0 or then E[E l (k)] = 0. According to the recurrence formula, it is known that Since in Case 3 and Case 4, E l (k) is a linear combination of E l (k - 1), so E l (k) is still a Gaussian distribution. Therefore, Lemma 1 is proven.

[0154] Next, the present invention defines states, actions, and rewards for analysis. Considering the information available in the cloud, the present invention defines the observable state at time step k as s(k), and the set of all feasible states is denoted as Here, is constructed as:

[0155]

[0156] s(k) = {s1(k), …, s L (k)}

[0157] Wherein, represents the control command sequence from time step k to K, denoted as:

[0158]

[0159] Define as the action taken by the system at time step k, where is the set of all feasible actions.

[0160] a(k) = {a1(k), …, a L (k)}

[0161]

[0162] Define r l (k) as the uniaxial reward obtained at time step k, and r(k) is the overall system reward at time step k.

[0163]

[0164] The present invention proves Theorem 1: The decision-making process represented by s(k), a(k), and r(k) constitutes a stationary Markov decision process (SMDP). The proof process is as follows:

[0165] First, the present invention proves that the above process has a definite state transition probability function Wherein represents the probability of transitioning to state s(k + 1) when the system is in state s(k) and takes action a(k).

[0166] For when determined by a(k) and the fixed packet loss probability is equal to 0, is a fixed value Wherein is determined by and ; when and when obeys a Gaussian distribution with mean and covariance Ψ T Σ l (k); when and when obeys a Gaussian distribution with mean and covariance A T Σ l (k)A.

[0167] For Σ l (k + 1): According to Lemma 1, Σ l (k + 1) is completely determined by Σ l (k), a(k), and the fixed packet loss rate.

[0168] For when when is determined by ; otherwise, the length of is determined by h l (k), and its value depends on and is determined by and .

[0169] For is a subset of, and thus is determined.

[0170] Therefore, the probability distribution of s(k + 1) depends only on s(k) and a(k), exists and is determined. Next, the present invention proves that r(k) is a deterministic function of s(k).

[0171] Case 1: then

[0172]

[0173] where

[0174] Case 2: then

[0175]

[0176] Therefore, in any case, r l (k) is determined by s l(k) Decision. According to the above formula, it can be seen that \(r(k)\) is a deterministic function of \(s(k)\), denoted as

[0177] Generally speaking, the next state \(s(k + 1)\) and the reward \(r(k)\) are only determined by the current state \(s(k)\) and the action \(a(k)\). The original process is a steady-state Markov decision process (SMDP), denoted as Thus, Theorem 1 is proved.

[0178] After proving the good properties of the original problem, the present invention attempts to transform it into a policy optimization problem in the reinforcement learning (RL) framework. Define as a policy function that maps the state \(s(k)\) to the action \(a(k)\), and \(V\) π is the state value function under the policy \(\pi\).

[0179]

[0180] Among them, when \(k = 0\), obviously \(V\) π (s(0)) is the negative value of the objective function under the policy \(\pi\). Therefore, the original problem is equivalent to:

[0181]

[0182] As the number of axes \(L\) and the number of allocable time slots \(M\) increase, the number of effective action combinations will grow explosively. Therefore, appropriate simplification is crucial for supporting policy optimization.

[0183] Considering that the sequence \(a\) l only affects the sequence \(r\) l , the present invention will merge the action combinations of each axis while maintaining the independence of each axis. Obviously, for \(a\) l (k), there are a total of \(2(H\) c +2) action combinations. Therefore, the present invention defines the decision variable whose corresponding relationship with \(a\) l (k) is as follows:

[0184] Calculate If then If then

[0185] Calculate If then Otherwise, and

[0186] For the single-axis action All possible values, each action consumes a certain number of time slots and affects the rewards obtained for this axis in the future. Only when the total number of time slots consumed by the action combination is less than M does it satisfy the constraint. The set of valid actions a + (k) is defined as

[0187] Inspired by the knapsack problem, the present invention is based on the grouped Q-knapsack algorithm of dynamic programming. First, define

[0188] as the number of time slots consumed by the action when . This value can be pre-computed and remain unchanged during execution. Next, define as the cumulative expected reward of axis l obtained after executing the forced action in the system state s(k) according to the policy π. Its formula is:

[0189]

[0190] Regarding all actions of an axis as a group of items, where the value and volume of each item at time k are represented by and respectively. The objective of the present invention is to select one item from each of the L groups such that the total value of these L items is maximized while their total volume does not exceed M. Algorithm 1 iteratively calculates the optimal solution to the above problem by dynamically deriving two optimal substructures dp and parent. Specifically, dp[l][m] represents the maximum total value that can be achieved when using the first l groups of items to fill a knapsack with a capacity of m, while parent[l][m] stores the list of indices of the selected items in each group. Therefore, parent[L][M] contains the list of optimal action indices for this problem. The optimality of this algorithm can be easily proven by induction.

[0191]

[0192]

[0193] Next, the present invention will prove that the invented Algorithm 1 can find a superior policy π' based on the current policy π. Since and there is a one-to-one correspondence between them, the present invention still represents the mapping from s(k) to a + (k) as π. Denote Abbreviate Algorithm 1 as

[0194] Theorem 2: The policy is superior to π(·), i.e., V π' (s(0)) ≥ Vπ (s(0)). The proof process is as follows:

[0195] Let represent the expected total reward after taking action a + (k), which can be expressed as:

[0196]

[0197] First, the present invention proves that the reward obtained by using policy π' at k = 0 and then always following policy π is greater than the reward of always following policy π. V π (s(0)) can be expressed as:

[0198]

[0199] Due to the optimality of Algorithm 1, there is:

[0200]

[0201] Therefore,

[0202] Next, the present invention proves that the reward obtained by always using policy π' is greater than the reward obtained by using policy π' only at k = 0.

[0203]

[0204] Therefore, V π' (s(0)) ≥ V π (s(0))

[0205] For any policy π, as long as the set For Algorithm 1 can be used to find an improved policy π'. This process can be iteratively repeated, and π' is further optimized to obtain π'', and so on, thus achieving iterative optimization. In the following part, the present invention uses deep reinforcement learning (DRL) to implement this iterative optimization process.

[0206] Given the high-dimensionality of the state s(k) and its complex relationship with the Q-value, a deep neural network is adopted as the Q-value estimator. Although s(k) contains all the observable information of the system, it is not appropriate to directly use it as the input of the neural network:

[0207] The buffer has a variable size, ranging from 0 to H c , which poses a challenge to neural network processing.

[0208] The relationship between the reward r(k) and the state s(k) Mathematically very complex and highly dependent on system constants such as A l , B l and G l . Since this prior knowledge cannot be provided to the network, it increases the difficulty of reinforcement learning training.

[0209] Therefore, the present invention defines the extended state as:

[0210]

[0211] where represents the potential impact of the estimation error on the expected reward. n l (k) ∈ 0, 1, …, H c represents the length of the buffer at time step k. represents the i-th element of c -n l (k) zeros are padded at the end to keep the length consistent. represents the predicted single-step cost of the system at time step k + i when controlling using buffer data in the noise-free case. Similarly, H c -n l (k) zeros are padded at the end. They can be calculated as:

[0212]

[0213] represents the predicted single-step cost of the cloud at time step k + 1 + i, which is obtained by solving the optimal control problem (OCP) in using only the control actions obtained from time step k + 1 to K - 1 calculated. They can be expressed as:

[0214]

[0215] represents the predicted cost of the cloud for the system at time step k + 1 + i, which is calculated by the calculated control actions and calculated using only the control actions from time k + 1 to K - 1. They can be expressed as:

[0216]

[0217] Since s + (k) contains ∑ l (k) and and The length and value can be determined by and together, so s + (k) completely contains all the state information in s(k). Compared with s(k), the auxiliary states added to s + (k) can all be determined by s(k). Therefore, the state transition probability of s + (k) is still well-defined and determined, ensuring that the problem remains a steady-state Markov decision process (SMDP). In addition, since the newly added states incorporate the prior knowledge of the system, the training of the policy network will be easier.

[0218] After state reconstruction, the present invention defines a neural network framework and represents its weights with parameter θ. Since s + (k) contains variable-length sequence data (y r , V * and V ◇ ), the present invention first uses an LSTM encoder to extract features from these sequence data. Subsequently, the extracted features are concatenated with the state features of a fixed length and further encoded through a multi-layer perceptron (MLP). Since the output consists of L sets of Q values, the present invention uses L independent linear heads, and each head outputs 2(H c +2) Q values. The neural network framework is as Figure 3 shown.

[0219] Next, the present invention uses the DDQN algorithm to estimate the Q values and iteratively optimize the policy, denoted as π θ . Different from the typical optimization objective, the multi-axis environment allows the calculation of independent single-axis rewards, expressed as r + (k) = r1(k), …, r L (k)}. To reduce the variance of Q value estimation, the present invention calculates the loss for the Q values output by each head. The total loss function is defined as follows:

[0220]

[0221] where θ - represents the parameters of the target network, which remain unchanged for a certain period. γ is the discount factor for stable training, ζ is the learning rate, represents the replay buffer, which is used to store the experiences collected during the interaction with the environment.

[0222] When this algorithm is deployed on the cloud, only the above neural network needs to be used to calculate the corresponding + for the current observable state s Then the grouped Q-knapsack algorithm is applied to determine the scheduling actions

[0223] Step 3: Verify the effectiveness of the edge-cloud collaborative control and the enhanced Q-knapsack method of the multi-axis coordinated servo system through demonstration and comparative experiments.

[0224] In this section, the present invention simulates a specific physical system to evaluate the invented edge-cloud collaborative framework and the RQK method. In addition, the present invention compares its performance with other constrained scheduling methods under different system parameters to verify the superiority of the RQK method.

[0225] Multi-axis system parameter settings: The present invention simulates a four-axis gantry crane control system, where respectively represent the DC motors that control the horizontal movements of the X-axis and Y-axis and the vertical movement of the Z-axis. is the DC motor on the lifting mechanism that controls the rotating shaft Φ of the lifting equipment. The relevant system parameters are defined as follows:

[0226] C1 = C2 = C3 = [1.0 0.0],

[0227] G1 = G2 = G3 = [1], R1 = R2 = R3 = [0.2

[0228] A4 = [0.9], C4 = [1.0], B4 = [1.0], Ξ4 = [0.01], G4 = [0.5], R4 = [0.1]

[0229] where the motor input is the DC voltage (V), to outputs the equivalent linear displacement (m), outputs the equivalent rotational speed (rpm).

[0230] Tracking target parameter settings: For construction and assembly tasks, the crane needs to track different target trajectories on each axis to achieve complex four-dimensional movements under load. Three commonly used dynamic target trajectories are predefined in the experiment: step, ramp, and sine wave. The total task length is set to K = 50, and the size of the control change point set P l is set to approximately 3% to 10% of K.

[0231] Edge-cloud collaborative network parameter settings: Configure a time slot to transmit two control signals, and set α = 0.5. Set M = 2 and H c = 4. The packet loss rates of both the uplink and downlink are set to 5%.

[0232] Neural network training parameter settings: The hidden layer size of the bidirectional LSTM encoder is set to 64. The MLP encoder consists of four layers with layer sizes [256, 256, 128, 64], and the activation function uses tanh. During training, the Adam optimizer is used, and the initial learning rate is set to ζ = 1×10 -4 , and the discount factor γ = 0.9 is set. Data collection uses 8 parallel environments, and the replay buffer has a memory size of 2000, and the mini-batch size is 64. The random exploration probability starts from the initial value of 2% and decays exponentially.

[0233] To emphasize the coordination between axes, the present invention excludes instruction change point configurations with less change in the demonstration experiment. The X-axis is set to track a ramp target from 0 to 5 during k = 0 to k = 50. The Y-axis is set to mutate at k = 7, 20, 30, 45 and track a step target. The Z-axis and Φ-axis are respectively set to track a ramp target from 0 to -5 and a sine instruction with an amplitude of 2.0 and a frequency of 0.06 during k = 10 to k = 40.

[0234] Figure 4 Shows the average cumulative reward obtained by the RQK method in each episode in eight parallel environments during training. Although random interference causes slight fluctuations in the reward for each episode, the method converges after 500 training rounds, demonstrating the effectiveness of Theorem 2.

[0235] In Figure 5 , the four-dimensional target curve and response curve under edge-cloud collaborative control under the RQK method are plotted, where the value of the Φ-axis is represented by color intensity. Figure 6 Shows Figure 5 the time-line projection of each axis in and are decision variables and indicators. Circles indicate that the corresponding value is 1; otherwise, it is 0. If is an empty circle, indicating h l (k - 1) = 0; if has a connection line, indicating that the control sequence corresponding to contains multiple control signals. Therefore, the time steps k covered by the solid circles and connection lines of

[0236] are controlled by the cloud, while other steps are controlled by the edge controller. Figure 6 It can be seen from that for axes with constant instructions or small changes within a certain time, the RQK method does not preferentially schedule control signals. This is because, in this case, l (k), may reduce performance. On the other hand, for the case where instructions change rapidly, the RQK method will actively schedule to take over the dynamic control process.

[0237] Regarding the scheduling of the state information x l due to to the interference is relatively large compared to the state upload of is arranged before the control signal is downloaded. This ensures that the cloud controller updates its state estimate to improve the performance of the control signal. In contrast, does not require state update because its state estimate remains highly accurate.

[0238] From Figure 5 it can be seen that the edge-cloud collaboration framework based on RQK scheduling can predict future target positions, reduce control costs, and at the same time utilize the real-time responsiveness of the edge controller for steady-state control, thus improving the overall performance of the gantry crane system.

[0239] To prove the superiority of the invented RQK method, the following methods are selected as benchmarks for comparison:

[0240] Edge Controller (EC): The EC method only relies on the sub-optimal controller κ at the edge l and always satisfies the constraints because edge-cloud communication is not required.

[0241] Greedy: Randomly select one of the following two operations to execute at each time step:

[0242] (1) Select the axis l with the largest value at time step k to update the state to reduce the estimation error and set all other decision variables to 0.

[0243] (2) Select the axis l that satisfies and is the largest, and transmit the control sequence downstream, that is and and set all other decision variables to 0.

[0244] Since each operation only consumes one time slot, this method always satisfies the constraints.

[0245] PPO-Lagrange (PPO-lag): A constrained proximal policy optimization (PPO) method based on the Lagrange multiplier method. This method uses the neural network encoder shown in Figure 3 and uses 2L heads to directly predict the upstream and downstream actions of all axes. Define a constrained reward During the training process, the network parameters and Lagrange multipliers are iteratively optimized to ensure that the action combination output by the network satisfies the constraints while maximizing the expected system reward.

[0246] It can be observed from the demonstration experiment that the disturbance level of the system has a significant impact on the scheduling strategy. In order to quantify this impact, the present invention defines a disturbance coefficient ν (i.e. The disturbance covariance of l ). Figure 7 The total control cost of each scheduling method under different ν values ​​is shown, where ν∈

[0247] 0.01,0.1,1.0,10.0}.

[0248] It can be observed that the EC method performs the worst among all the methods due to its passive response to dynamic instructions. However, since edge computing has strong real-time performance and can alleviate the impact of noise to a certain extent, the performance of the EC method does not decrease significantly with the increase of the disturbance coefficient. The PPO-lag method performs well under low disturbance conditions, but since this method contains Lagrange multiplier terms to reduce the probability of constraint violations, the trained network tends to converge to a mode that always chooses the same action that satisfies the constraints. Therefore, when the system disturbance coefficient increases, the PPO-lag method tends to avoid state updates to minimize the probability of constraint violations (although the probability of constraint violations cannot be guaranteed to be zero). In contrast, the invented RQK method outperforms these three methods in overall performance.

[0249] Figure 8 The control cost of each scheduling method under different number of axes is shown. The parameters of all axes are the same as The same, and the instructions of all axes are random step signals. It can be seen that due to the same instruction pattern, the total cost of the EC method increases almost linearly with the number of axes. At the same time, the Greedy and PPO-lag methods often only schedule a single axis or some fixed axes, so the overall performance is only slightly improved. The invented RQK method optimizes the scheduling strategy on a longer time scale based on the status of all axes, thus performing best in overall performance.

[0250] In summary, in the edge-cloud collaborative control method for a multi-axis servo system of the present invention, each axis is connected to an edge device, and the edge device is equipped with a sub-optimal tracker, while the cloud performs computationally intensive model predictive control (MPC). The cloud first calculates the control signal sequence for each axis over a future period of time and transmits this control signal sequence through a shared network channel to the corresponding edge device, where it is stored in the buffer of the edge device. For each axis, when the buffer is available, the control signal is first retrieved from the buffer and applied to the axis; otherwise, the edge device uses the sub-optimal tracker to calculate the corresponding control signal. This method enables the control signal source for each axis to be flexibly adjusted while optimizing the utilization of communication resources. At each time step, for each axis, it is necessary to determine whether to upload the state of the axis to the cloud, whether to download the control signal sequence calculated by the cloud to the buffer of the edge device, and the length of the control signal sequence, in order to maximize the coordinated control performance. To make these decisions reasonably, the present method first formulates the state, action, and reward function of the multi-axis servo system at each time step. Based on these formulas, the present invention theoretically proves that the original problem can be transformed into a static Markov decision process (SMDP). Then, in the process of optimizing the policy of this SMDP, the present invention proposes a new method that combines deep reinforcement learning (DRL) with the grouped Q-knapsack algorithm, called reinforcement Q-knapsack (RQK), whose goal is to maximize the coordinated control performance of the system while satisfying the communication constraints. Finally, through demonstration and comparative experiments, it is shown that the edge-cloud collaborative control method for a multi-axis coordinated servo system of the present invention is superior to other methods.

Claims

1. An edge-cloud collaborative control method for a multi-axis coordinated servo system, characterized in that: The method comprises the following steps: (1) Construct an edge-cloud collaborative multi-axis servo system model in which each independent axis tracks the time-varying target trajectory, and the edge device equipped with each axis calculates the edge control signal through a suboptimal LQT controller. The cloud calculates the optimal control sequence through MPC. The multi-axis coordinated servo system decides at each moment whether to upload the status signal of each axis to the cloud or whether to pass the control sequence down to the edge device. Specifically, the cloud first calculates the control signal sequence of each axis in the future, and transmits the control signal sequence to the corresponding edge device through a shared network channel and stores it in the buffer of the edge device; for each axis, when the buffer is available, the control signal is first taken out of the buffer and applied to the axis; otherwise, the edge device uses a suboptimal tracker to calculate the corresponding control signal; At each time step, for each axis, it is necessary to determine whether to upload the state of the axis to the cloud, whether to download the control signal sequence calculated in the cloud to the buffer of the edge device, and the length of the control signal sequence, in order to maximize the coordinated control performance; (2) By constructing an observable state space, the scheduling problem is formulated as a steady-state Markov decision process in the cloud. Subsequently, the scheduling strategy is optimized by the reinforced Q-knapsack method based on reinforcement learning. The optimal action selection problem at each time step is converted into a grouped Q-value knapsack problem, and dynamic programming is used to determine the optimal action combination. Finally, DDQN is used to estimate the Q-value of the extended state-action pair, thereby optimizing the long-term performance of the system.

2. The edge-cloud collaborative control method for a multi-axis coordinated servo system according to claim 1, characterized in that: The multi-axis servo system model is modeled by the following steps: (11) The accused object p l It is described by discrete linear time-invariant stochastic system equations, and its expression is: Among them, x l (k),u l (k) and y l (k) respectively represent the controlled object P l The state vector, input vector, and output vector at time step k, matrix A l , B l and C l Represent the state transfer matrix, input matrix and output matrix respectively; for Initial state x l (0) and the disturbance w l (k)(k∈N) is an independent and identically distributed (iid) random variable with mean 0 and covariance Ξ l Normal distribution, in the actual system, (A l ,B l ) is controllable; (12) Definition Coordinate multi-axis task instructions and decompose them into independent target tracking tasks for each axis in, represents the reference output of the axis at time step k. For each axis l, satisfy: Among them, P l is a relatively small subset, namely |P l |<<K, whose element k satisfies (13) Consider transmitting the status information x of a single servo system l (k) occupies a time slot in IEEE 802.15.4, and the transmission contains h l (k) control sequence of control steps Occupancy time slots, where α∈(0,1]; Defining decision variables To indicate whether to upload x l (k) Allocate time slots and define decision variables To indicate whether it is a transmission Assign time slots, h l (k) is also regarded as a decision variable. At any time step, the total number of time slots occupied by all state information and control information used for decision-making shall not exceed the maximum value M∈Z allowed by the system. + ; (14) Edge Mainly focus on the controlled object P l The control objective is to minimize the cost function Through the LQT controller, we can get a l (k) Feedback gain and The suboptimal tracker with feedforward gain is expressed as: (15) Cloud Using MPC to calculate the controlled object P l The optimal control signal in the future time domain defines the optimal control sequence as (16) On the edge Provide a buffer For storage in the cloud The subsequent control sequence sent, let the matrix Represents the contents of the buffer at time step k, the buffer The update mechanism of and the actual input action of the actuator are expressed as: The symbol [] represents an empty matrix, and matrix b is used to extract the buffer. The first control signal in the buffer, and the matrix H is used to Remove the first control signal; Cloud use and ACK frame information to simulate the buffer update process, so that Real-time maintenance buffer Mirror image of; Due to the transmission delay T s existence, when When, Cloud Only the real state x at time step k-1 can be obtained l ; make and Respectively represent the cloud At time step k, the controlled object The state estimate and input estimate are calculated as: The system state prediction at time step k+1 is: (17) The goal of this method is to determine the system state x at each time step l (k) Upload and control sequence Downlink transmission and its length h l (k) maximize the expected overall performance of the multi-axis system while ensuring that the superframe time slots used in each time step do not exceed the system constraint M; Mathematically, the constrained scheduling problem is expressed as:

3. The edge-cloud collaborative control method for a multi-axis coordinated servo system according to claim 1, characterized in that: In order to optimize the above constrained scheduling problem, the original problem is first converted into a Markov problem, and the scheduling strategy is optimized using the reinforced Q-knapsack method based on reinforcement learning, which includes: (21) Definition For the cloud to control the controlled object The estimated error at time step k, and let Indicates the control edge The closed-loop system matrix; (22) The observable state at time step k is defined as s(k), and the set of all feasible states is denoted as The expression is as follows: in, Represents the controlled object from time step k to K The control command sequence is expressed as: Define a(k) as the action taken by the system at time step k, Represents the set of all possible actions: Definition l (k) is the accused The single-axis reward obtained at time step k, r(k) is the overall system reward at time step k, and the calculation is expressed as: (23) Convert the solution of step (22) into a policy optimization problem in the reinforcement learning framework and define Strategy function, mapping state s(k) to action a(k), V π is the state value function under strategy π: in When k = 0, V π (s(0)) is the negative value of the objective function under strategy π, so the original optimization problem is equivalent to: (24) Considering the sequence a l Only affects sequence r l , which will combine the actions of each axis while maintaining the independence of each axis, defining the decision variables Its l The corresponding relationship of (k) is as follows: calculate like but like but calculate like but otherwise, and The valid action set a + (k) is defined as (25) The solution is based on the grouped Q-knapsack algorithm of dynamic programming, specifically: definition For The number of time slots consumed by the action, which is used for pre-calculation and remains unchanged during execution; definition To be the mandatory action after executing according to the strategy π in the system state s(k) The cumulative expected reward of axis l is obtained, and its formula is: (26) will go from s(k) to a + The mapping of (k) is denoted by π, The grouped Q backpack is abbreviated as (27) An LSTM encoder is used to extract features from these sequence data, the extracted features are concatenated with the fixed-length state features, and further encoded through a multi-layer perceptron; Output It consists of L groups of Q values, using L independent linear heads, each head outputs 2 (H c +2) Q values; (28) The DDQN algorithm is used to estimate the Q value and iteratively optimize the strategy, denoted by π θ , and the multi-axis environment allows the calculation of independent single-axis rewards, denoted as r + (k)={r1(k),…,r L (k)}; In order to reduce the variance of Q value estimation, the DDQN algorithm is used to estimate the Q value and iteratively optimize the strategy, denoted as π θ , the total loss function is defined as follows: Among them, θ - represents the parameters of the target network, which remain unchanged for a certain period of time, γ is the discount factor used to stabilize training, ζ is the learning rate, Represents the replay buffer, which is used to store experience collected during the interaction with the environment.

4. The edge-cloud collaborative control method for a multi-axis coordinated servo system according to claim 1, characterized in that: The method aims at the situation where a multi-axis system tracks different target trajectories in coordinated work, takes into account the different characteristics of edge devices and cloud computing devices in the motion axis, constructs an edge-cloud collaborative control framework, and implements the optimal scheduling strategy under limited channel conditions.

5. The edge-cloud collaborative control method for a multi-axis coordinated servo system according to claim 3, characterized in that: In step (27), in view of the high dimensionality of the state s(k) and its complex relationship with the Q value, a deep neural network is used as the Q value estimator, and the extended state is defined as: in, represents the potential impact of estimation error on expected reward, n l (k)∈0,1,…,H c represents the buffer at time step k Length, express The i-th element of c -n l (k) zeros to keep the length consistent.

6. The edge-cloud collaborative control method for a multi-axis coordinated servo system according to claim 3, characterized in that: Step (28) includes extracting features from the sequence data using an LSTM encoder, then concatenating the extracted features with fixed-length state features and further encoding them through a multi-layer perceptron, and allowing independent single-axis rewards to be calculated for a multi-axis environment.

7. A computing device, characterized in that It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the edge-cloud collaborative control method for a multi-axis coordinated servo system as described in any one of claims 1 to 6.

8. A computer storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a computer, the edge-cloud collaborative control method for a multi-axis coordinated servo system as described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Intelligent actuator adaptive control method and system based on Internet of Things

    CN120704151A

  • Multi-node single-chip microcomputer automatic synchronization control method based on edge collaboration

    CN121477736A

  • Remote driving method

    CN121777177A

  • Remote drive method

    CN121777177B

  • Robot multi-axis motion task time consistency certainty scheduling method and device and storage medium

    CN122463174A