Robust fault-tolerant control method based on deep reinforcement learning
Through the robust fault-tolerant control method of deep reinforcement learning, the generation of states and actions, combined with adversarial training and priority experience replay, the robustness and control accuracy of the bridge crane system in complex environments is improved, and the problem of insufficient robustness in the existing technology is solved.
Patent Information
- Application Number
- CN202510507779.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-15
AI Technical Summary
The existing intelligent control methods of bridge crane systems lack the robustness and stability to external disturbances, parameter perturbation and actuator failure in complex environments, resulting in insufficient control strategies.
A robust fault-tolerant control method based on deep reinforcement learning is adopted, and the original sample generation state is input, and the policy network and value network are used to generate action and reward updates. Combined with adversarial training and priority experience playback, the robustness of the model in dynamic and complex environments is improved.
It improves the robustness and control accuracy of the bridge crane system in complex environments, and enhances the fault tolerance of external disturbances and actuator failures.
Smart Images

Figure CN120491449A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of data processing, and in particular to a robust fault-tolerant control method based on deep reinforcement learning. Background Art
[0002] The current intelligent control method of bridge crane system does not fully consider the external disturbances, parameter perturbations and actuator failures in complex environments. It lacks the analysis of potential disturbance information and the design of robust training network models, resulting in insufficient robustness and stability of the control strategy in dynamic and complex environments. Summary of the Invention
[0003] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.
[0004] The purpose of this application is to solve one of the technical problems existing in the related technology to at least a certain extent. The embodiment of this application provides a robust fault-tolerant control method based on deep reinforcement learning, which improves the robustness of the model in complex environments.
[0005] An embodiment of the first aspect of the present application is a robust fault-tolerant control method based on deep reinforcement learning, comprising:
[0006] Inputting the original sample and the adversarial sample into the bridge crane environment to generate a state, wherein the state includes the displacement, speed, load swing angle and angular velocity of the bridge crane;
[0007] The policy network generates an action according to the state quantity, wherein the action includes a driving force of the bridge crane;
[0008] Applying the action amount to the bridge crane environment to generate a next state amount and reward;
[0009] Store states, actions, rewards, and state transition probabilities into a recycling pool;
[0010] Selecting a first state and a first action from the recycling pool according to a priority;
[0011] The value network outputs a state value reward distribution and an action value reward distribution according to the first state and the first action;
[0012] Update the policy network according to the state value reward distribution and the action value reward distribution;
[0013] Updating the value network according to the policy entropy of the policy network;
[0014] When the preset termination condition is reached, the trained bridge crane control model is obtained.
[0015] According to certain embodiments of the first aspect of the present application, updating the policy network according to the state value reward distribution and the action value reward distribution includes:
[0016] Determine the soft-state action reward distribution under the policy;
[0017] Use KL divergence as the objective function to update the soft state action reward distribution;
[0018] A new soft-state action reward distribution is obtained according to the objective function of the updated soft-state action reward distribution.
[0019] According to certain embodiments of the first aspect of the present application, updating the policy network according to the state value reward distribution and the action value reward distribution includes:
[0020] The policy network is updated by maximizing the objective function based on the soft reward value.
[0021] According to certain embodiments of the first aspect of the present application, updating the policy network according to the state value reward distribution and the action value reward distribution includes: updating the policy network through expected value replacement, dual value distribution learning and variance-based critical gradient adjustment.
[0022] According to certain embodiments of the first aspect of the present application, selecting the first state and the first action from the recycling pool according to the priority includes:
[0023] Determine the priority based on the difference between the actual reward and the expected reward after the bridge crane system performs an action in a state;
[0024] Determine the importance sampling weight according to the non-uniform probability and the importance sampling update coefficient;
[0025] Selecting a first state and a first action from the recycling pool according to the priority and the importance sampling weight;
[0026] The importance sampling update coefficient is modified during the learning process.
[0027] According to certain embodiments of the first aspect of the present application, the adversarial sample is generated according to the following formula: in, is an adversarial sample, s t is the original sample, ε is the amplitude of the adversarial disturbance, is the gradient of the loss function L with respect to the original input sample, π(s t ; θ) is the strategy determined by the original sample and the value return distribution θ, a t For action.
[0028] According to certain embodiments of the first aspect of the present application, the objective function of the bridge crane control model is: Among them, L total is the objective function of the bridge crane control model, α is the weight of the original sample loss, β is the weight of the adversarial sample loss, L(π(s t ;θ),a t ) is the original sample loss, is the adversarial sample loss, is an adversarial sample, s t is the original sample, π(s t ; θ) is the strategy determined by the original sample and the value return distribution θ, is the strategy determined by the adversarial sample and the value return distribution θ, a t For action.
[0029] An embodiment of the second aspect of the present application is a method for controlling a bridge crane, comprising:
[0030] The displacement, speed, load swing angle and angular velocity of the bridge crane are input into the trained bridge crane control model to obtain the driving force of the bridge crane;
[0031] Among them, the trained bridge crane control model is trained according to a robust fault-tolerant control method based on deep reinforcement learning described in an embodiment of the first aspect of the present application.
[0032] An embodiment of the third aspect of the present application is an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements a robust fault-tolerant control method based on deep reinforcement learning as described in the embodiment of the first aspect of the present application.
[0033] An embodiment of the fourth aspect of the present application is a computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute a robust fault-tolerant control method based on deep reinforcement learning as described in the embodiment of the first aspect of the present application.
[0034] The above scheme has at least the following beneficial effects: the original sample and the adversarial sample are input into the bridge crane environment to generate the state; the policy network generates the action according to the state quantity; the action quantity is applied to the bridge crane environment to generate the next state quantity and reward; the state, action, reward and state transition probability are stored in the recycling pool; the first state and the first action are selected from the recycling pool according to the priority; the value network outputs the state value reward distribution and the action value reward distribution according to the first state and the first action; the policy network is updated according to the state value reward distribution and the action value reward distribution; the value network is updated according to the policy entropy of the policy network; adversarial samples are generated for external disturbances, and the model is made more robust in dynamic and complex environments by constructing an adversarial training strategy; the sampling probability is allocated by priority experience recycling to improve the learning efficiency and real-time control performance under complex working conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings are used to provide a further understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.
[0036] Figure 1 This is a step diagram of the robust fault-tolerant control method based on deep reinforcement learning;
[0037] Figure 2 It is a step diagram for selecting the first state and the first action from the recycling pool according to the priority;
[0038] Figure 3 It is a step diagram for updating the policy network based on the state value reward distribution and the action value reward distribution;
[0039] Figure 4 It is a schematic diagram of the one-dimensional bridge crane system model;
[0040] Figure 5 It is a schematic diagram of the bridge crane control model. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0042] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and the like in the specification, claims, or accompanying drawings are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0043] The embodiments of the present application are further described below with reference to the accompanying drawings.
[0044] The embodiments of the present application provide a robust fault-tolerant control method based on deep reinforcement learning.
[0045] Reference Figure 1 , a robust fault-tolerant control method based on deep reinforcement learning, comprising the following steps:
[0046] Step S100, inputting the original sample and the adversarial sample into the bridge crane environment to generate a state;
[0047] Step S200, the policy network generates an action based on the state quantity;
[0048] Step S300, applying the action amount to the bridge crane environment to generate the next state amount and reward;
[0049] Step S400, storing the state, action, reward and state transition probability into a recycling pool;
[0050] Step S500, selecting a first state and a first action from a recycling pool according to a priority;
[0051] Step S600: The value network outputs a state value reward distribution and an action value reward distribution according to the first state and the first action;
[0052] Step S700, updating the policy network according to the state value reward distribution and the action value reward distribution;
[0053] Step S800, updating the value network according to the policy entropy of the policy network;
[0054] Step S900: When the preset termination condition is reached, the trained bridge crane control model is obtained.
[0055] The state includes the displacement, speed, load swing angle and angular velocity of the bridge crane; the action includes the driving force of the bridge crane.
[0056] It should be noted that, referring to Figure 4 , in the one-dimensional bridge crane system model, its dynamic equation is as follows: Where M is the mass of the trolley, m is the mass of the load, l is the length of the suspension rope, u is the driving force in the horizontal direction of the trolley, θ is the swing angle of the load, and g is the acceleration due to gravity.
[0057] The kinetic equation can be rewritten as the following state equation form:
[0058] The discrete state equation form of the dynamic equation is:
[0059]
[0060] Considering the actuator failure fault and bias fault, it can be expressed as follows: t =ρu t +r. Where 0<ρ1≤ρ≤ρ2≤1 represents the loss fault of the actuator, ρ1 and ρ2 are two constants, and r represents the bias fault of the actuator.
[0061] The control problem of the bridge crane is modeled as a Markov decision process (MDP). It is a mathematical description of the interactive learning between the agent and the bridge crane environment. Formally, an MDP can be defined as a five-tuple (S, A, p, R, γ), where S and A represent the state space and action space respectively, p(s′|s, a), s∈S, is the state transition probability, R is the reward function, and γ∈[0,1] is the discount factor. For the bridge crane, at discrete time t, the agent is in state Next Choose an action and transition to the next state s with probability p t+1 , and receive reward R t The Markov decision process terminates when the agent reaches the set termination condition.
[0062] The behavior of the agent is defined by a random policy: π(a t |s t ):S→P(a t ), which is a probability distribution that maps a specific state to an action. The goal of the agent is to learn an optimal policy that maximizes the sum of expected discounted rewards. For example, π * (a t |s t )=argmaxJ(π),
[0063] In addition, two value functions are defined. The first is the state value function, which represents the expected discounted reward after the state s with a strategy of π, which can be expressed as The second is the action value function, which represents the expected discounted reward after taking action a in state s under strategy π, which can be expressed as
[0064] For bridge crane control, the state refers to the state quantity of the bridge crane Specifically: displacement x, speed Load swing angle θ, angular velocity The action refers to the output of the motor, that is, the driving force of the system. The action is defined as at ={u1,u2,···,u n} continuous action space, ranging from {-2,2}.
[0065] The reward function of the system is: Among them, ||·||2 represents the L2 norm of the difference between the four states and the target state, s st is the safety boundary limit of the state, and at is the current action.
[0066] Reference Figure 3 , updating the policy network according to the state value return distribution and action value return distribution, including the following steps:
[0067] Step S710, determining the soft state action reward distribution under the strategy;
[0068] Step S720, using KL divergence as the objective function for updating the soft state action reward distribution;
[0069] Step S730 , obtaining a new soft state action reward distribution according to the objective function of updating the soft state action reward distribution.
[0070] Standard reinforcement learning typically aims to maximize expected reward, but this makes it difficult to achieve global exploration performance. This embodiment aims to maximize the cumulative expected reward while also attempting to maximize policy entropy.
[0071] The objective function is defined as: Where α is the temperature coefficient, which determines the weight of entropy regularization. Η(p) = -∫p(x)logp(x)dx is the policy entropy.
[0072] By adopting a two-step policy iteration method of soft policy evaluation and soft policy improvement, the goal is to learn an optimal policy. The optimal policy is: * =argmaxJ(π).
[0073] Distribution RL is used to capture the information of cumulative return discount rate.
[0074] The soft-state action reward distribution under policy π can be expressed as: in, Indicates from s t The cumulative return of entropy increase.
[0075] Define the distribution value function Z π (Z π (s,a)|s,a): S×A→P(Z π (s,a)), which means from (s t ,a t ) to the soft-state action reward distribution.
[0076] Repeatedly apply the distributed Bellman operator T under the optimal strategy π π The distributed Bellman operator is expressed as: Among them, T π is a Bellman operator, A= D B means that the two random variables follow the law of equal probability.
[0077] set up in It's T π The policy is updated by minimizing the distribution distance between the Bellman operator and the current distribution: Here, d is the KL divergence, which measures the distance between two distributions.
[0078] In the policy evaluation phase, the state-action return distribution is trained and the KL divergence metric is used as the update Z π The objective function is to minimize the loss function: Where c is a constant, B represents the sample replay buffer, θ′ and φ′ are the parameters of the target network return distribution and policy function. KL represents the KL divergence between two distributions.
[0079] The objective function is calculated by the following equation: Among them, y z Is the target value returned randomly, expressed as: y z =r+γ(Z(s t+1 ,a t+1 )-αlogπ φ′ (a t+1 |s t+1 )).
[0080] Update the soft return distribution by the following formula: Among them, Z θ is a parameterized distribution value function, and θ is learned from training.
[0081] In the policy improvement phase, the policy network is updated by maximizing an objective function based on the soft reward value and selecting actions with low variance:
[0082] Update the policy network by: By minimizing the following objective function: The temperature coefficient α is updated. is the expected policy entropy. In the context of bridge crane control, s∈s t and a∈a t . Qθ(s t ,at ) is the learned soft Q-value function.
[0083] By updating the policy network through expected value replacement, dual-value distribution learning, and variance-based critical gradient adjustment, the overestimation after convergence is further suppressed, thereby improving the policy performance.
[0084] Reducing the randomness in the variance-dependent gradient by clipping the random target return value does not solve the high randomness of the mean-dependent gradient caused by the random target return. Using a more stable Q function to replace the random target return to adjust the mean-dependent gradient is: q =r+γ(Q θ′ (s t+1 ,a t+1 )-αlogπ φ′ (a t+1 |s t+1 )).
[0085] The gradient of the critic network is:
[0086]
[0087] Among them, clip is the clipping function, b is the boundary, and σ is the standard deviation.
[0088] According to the distribution change of clipping double Q learning to dual value distribution learning, two value distributions are parameterized, characterized by parameters θ1 and θ2, and the value distribution with a smaller mean is selected to construct the gradient of the critic and the actor.
[0089] For evaluation updates, the index defining the distribution of selected values is:
[0090] The target benefit and target q value are evaluated. The expressions of these target evaluations are as follows:
[0091] Substituting the above formulas into each other, we have:
[0092] The actor's goal adopts a modified bivalued distribution:
[0093] Adopting bivalued distribution minimization to shape the target distribution of critic updates further mitigates the overestimation bias and tends to produce slight underestimation, which helps improve the stability of learning.
[0094] In order to reduce the parameter adjustment work, Reduce sensitivity to reward scaling. Where ξ is a constant parameter that controls the clipping range. In this setting, the bound can adapt to different reward sizes across tasks and training stages.
[0095] Reference Figure 2 , selecting the first state and the first action from the recycling pool according to the priority, including the following steps:
[0096] Step S510, determining a priority based on a difference between an actual reward and an expected reward after the bridge crane system performs an action in a state;
[0097] Step S520, determining the importance sampling weight according to the non-uniform probability and the importance sampling update coefficient;
[0098] Step S530: Select the first state and the first action from the recycling pool according to the priority and importance sampling weight.
[0099] Among them, the importance sampling update coefficient is modified during the learning process.
[0100] In Prioritized Experience Replay, we prioritize using temporal error, which represents the difference between the actual reward observed by the agent after performing an action in a certain state and its expected reward. A larger temporal error means that the experience is more informative for the agent, as it indicates that there is a larger deviation between the agent's prediction and the actual situation, and more learning may be needed.
[0101] The probability of each sample is defined as: Where λ is the priority control coefficient, p j Is the weight coefficient of sample j. Define p j =|δ j |+ε, where ε represents the edge case transition when the time error becomes zero.
[0102] Prioritized replay may introduce biases that change the distribution and the convergence of the probability distribution moments in an uncontrolled way. This bias can be accounted for using importance sampling weights, which are expressed as:
[0103] If β = 1, the sample weights compensate for the non-uniform probability p j These weights are obtained by using w i δ i Incorporate Q-learning updates. Utilize 1 / max i w iNormalization is performed so that updates are scaled upward, resulting in a stable learning process. Unbiased updates are non-ideal in reinforcement learning because the process is unstable due to variability in the policy, state distribution, and guidance objectives. Therefore, the importance sampling update coefficient β is modified over time from its initial value of 1. Importance sampling offers significant advantages when combined with prioritized replay to approximate nonlinear functions. A small global step size is recommended because large step sizes can disrupt the learning process, as first-order gradient approximations are only locally reliable. Therefore, prioritization ensures that high-error transitions are observed multiple times. Simultaneously, the importance sampling correction reduces the gradient magnitude, and with the continuous reapproximation of the Taylor expansion, the algorithm follows the curvature of the highly nonlinear optimization landscape.
[0104] To enhance the model's robustness to external perturbations and dynamic environmental changes, adversarial examples are introduced into the pendulum angle update formula to improve the model's generalization and anti-interference performance. Specifically, adversarial examples are generated by adding optimized, small perturbations to the original input state, aiming to maximize the model's output error, thereby forcing the model to learn a more stable control strategy during training.
[0105] The original state is Adversarial Examples The generation formula is: in, is an adversarial sample, s t is the original sample, ε is the amplitude of the adversarial disturbance, is the gradient of the loss function L with respect to the original input sample, π(s t ; θ) is the strategy determined by the original sample and the value return distribution θ, a t For action.
[0106] The generated adversarial examples Substituting the swing angle update formula into the model, the model responds to this input through the training process, making it able to effectively resist deviations caused by external disturbances or model uncertainties, thereby improving robustness.
[0107] During the training process, adversarial samples and original samples participate in the loss calculation together, and the final optimization objective function is: Among them, L total is the objective function of the bridge crane control model, α is the weight of the original sample loss, β is the weight of the adversarial sample loss, L(π(s t ;θ),a t ) is the original sample loss, is the adversarial sample loss, is an adversarial sample, s t is the original sample, π(s t; θ) is the strategy determined by the original sample and the value return distribution θ, is the strategy determined by the adversarial sample and the value return distribution θ, a t For action.
[0108] This adversarial training strategy effectively improves the stability and control accuracy of the model in complex environments, ensuring robust control of the bridge crane under non-ideal working conditions.
[0109] To enhance the policy's memory and enable it to implicitly recognize external perturbations and failures, AL-PDSACT employs LSTM modules as the actor and critic network structures. The policy and value networks employ LSTM networks. LSTM networks are designed to handle long-term sequence dependencies. When trained using a distributed DRL strategy, this architecture implicitly encodes diverse perturbation information in its hidden states, enabling the agent to learn a universal policy that effectively responds to external perturbations and actuator failures.
[0110] The input of LSTM includes the current state s t , the previous hidden state h t-1 and the previous cell state c t-1 . The previous hidden state h t-1 and the previous cell state c t-1 It effectively stores the previous payload quality information, which helps to train the general policy. The forget gate, input gate and output gate are used to control which information should be forgotten, retained or output. The operation of the three gates can be described as: t =ρ(W l [h t-1 ,s t ]+b l );i t =ρ(W i [h t-1 ,s t ]+b i );z t =ρ(W z [h t-1 ,s t ]+b z ).
[0111] Current hidden state h t and cell state c t Expressed as: h t =z t *tanh(c t ); represents a candidate for normalizing the input data, expressed as:
[0112] Reference Figure 5 To address the robustness and fault tolerance of a bridge crane, a bridge crane controller is designed that accounts for unknown disturbances, parameter perturbations, and actuator failures. AL-PDSACT effectively integrates distributed reinforcement learning, prioritized experience replay, and adversarial training. During the training phase, the algorithm executes multiple cycles. In each cycle, an action is first randomly selected. The agent then interacts with the environment, generating new state sequences and adversarial examples through an adversarial training strategy. Experience is collected and stored in an experience replay buffer. Next, experience is sampled from the experience replay buffer based on the time error and used to update the parameters of the value function and policy network. After the parameters are updated, the experience replay buffer is updated based on the priority of the experience. This process is repeated until the maximum number of training steps is reached or convergence conditions are met. After training, the general control strategy is robust and fault-tolerant in a test environment that simulates reality. The agent can be regarded as a controller, making decisions for the bridge crane.
[0113] An embodiment of the present application provides a method for controlling a bridge crane.
[0114] The control method of a bridge crane comprises the following steps: inputting the displacement, speed, load swing angle and angular velocity of the bridge crane into a trained bridge crane control model to obtain the driving force of the bridge crane.
[0115] The trained bridge crane control model is obtained by training according to the robust fault-tolerant control method based on deep reinforcement learning as described above.
[0116] The present application also provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the aforementioned robust fault-tolerant control method based on deep reinforcement learning and the bridge crane control method. The electronic device can be any smart terminal, including a tablet computer and an in-vehicle computer.
[0117] The processor can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application; the memory can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory and is called by the processor to execute a robust fault-tolerant control method based on deep reinforcement learning and a control method for a bridge crane in the embodiments of the present application.
[0118] The input / output interface is used to realize information input and output; the communication interface is used to realize communication interaction between this device and other devices. Communication can be achieved through wired methods (such as USB, network cable, etc.) or wireless methods (such as mobile network, WIFI, Bluetooth, etc.); the bus transmits information between the various components of the device (such as processor, memory, input / output interface and communication interface); among them, the processor, memory, input / output interface and communication interface realize communication connection with each other within the device through the bus.
[0119] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned robust fault-tolerant control method based on deep reinforcement learning and the control method of the bridge crane.
[0120] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0121] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0122] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0123] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0124] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0125] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0126] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0127] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0128] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0129] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0130] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0131] The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present application, and these equivalent modifications or substitutions are all included in the scope defined by the claims of the present application.
Claims
1. A robust fault-tolerant control method based on deep reinforcement learning, characterized in that: include: Inputting the original sample and the adversarial sample into the bridge crane environment to generate a state, wherein the state includes the displacement, speed, load swing angle and angular velocity of the bridge crane; The policy network generates an action according to the state quantity, wherein the action includes a driving force of the bridge crane; Applying the action amount to the bridge crane environment to generate a next state amount and reward; Store states, actions, rewards, and state transition probabilities into a recycling pool; Selecting a first state and a first action from the recycling pool according to a priority; The value network outputs a state value reward distribution and an action value reward distribution according to the first state and the first action; Update the policy network according to the state value reward distribution and the action value reward distribution; Updating the value network according to the policy entropy of the policy network; When the preset termination condition is reached, the trained bridge crane control model is obtained.
2. A robust fault-tolerant control method based on deep reinforcement learning according to claim 1, characterized in that: The updating of the policy network according to the state value reward distribution and the action value reward distribution includes: Determine the soft-state action reward distribution under the policy; Use KL divergence as the objective function to update the soft state action reward distribution; A new soft-state action reward distribution is obtained according to the objective function of the updated soft-state action reward distribution.
3. The robust fault-tolerant control method based on deep reinforcement learning according to claim 1, characterized in that: The updating of the policy network according to the state value reward distribution and the action value reward distribution includes: The policy network is updated by maximizing the objective function based on the soft reward value.
4. The robust fault-tolerant control method based on deep reinforcement learning according to claim 1, characterized in that: The updating of the policy network according to the state value reward distribution and the action value reward distribution includes: updating the policy network through expected value replacement, dual value distribution learning and variance-based critical gradient adjustment.
5. The robust fault-tolerant control method based on deep reinforcement learning according to claim 1, characterized in that: The selecting the first state and the first action from the recycling pool according to the priority includes: Determine the priority based on the difference between the actual reward and the expected reward after the bridge crane system performs an action in a state; Determine the importance sampling weight according to the non-uniform probability and the importance sampling update coefficient; Selecting a first state and a first action from the recycling pool according to the priority and the importance sampling weight; The importance sampling update coefficient is modified during the learning process.
6. The robust fault-tolerant control method based on deep reinforcement learning according to claim 1, characterized in that: The adversarial sample is generated according to the following formula: in, is an adversarial sample, s t is the original sample, ε is the amplitude of the adversarial disturbance, is the gradient of the loss function L with respect to the original input sample, π(s t ; θ) is the strategy determined by the original sample and the value return distribution θ, a t For action.
7. The robust fault-tolerant control method based on deep reinforcement learning according to claim 1, characterized in that: The objective function of the bridge crane control model is: Among them, L total is the objective function of the bridge crane control model, α is the weight of the original sample loss, β is the weight of the adversarial sample loss, L(π(s t ;θ),a t ) is the original sample loss, is the adversarial sample loss, is an adversarial sample, s t is the original sample, π(s t ; θ) is the strategy determined by the original sample and the value return distribution θ, is the strategy determined by the adversarial sample and the value return distribution θ, a t For action.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, a robust fault-tolerant control method based on deep reinforcement learning according to any one of claims 1 to 7 is implemented.
9. A computer storage medium, characterized in that Computer-executable instructions are stored, and the computer-executable instructions are used to execute a robust fault-tolerant control method based on deep reinforcement learning as described in any one of claims 1 to 7.