Bellman infinite error-based robust Q learning resisting method and system
By constructing an eigenstate neighborhood set and using consistent adversarial robust operators, the problem of decision-making failure in the face of adversarial perturbation is solved, and efficient training and excellent performance in complex environments are achieved.
Patent Information
- Application Number
- CN202510104333.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-03
AI Technical Summary
Existing deep Q networks fail in decision making in the face of small adversarial perturbations and sacrifice performance in clean environments when improving adversarial robustness, resulting in limited trusted deployment in complex environments.
By constructing an eigenstate neighborhood set, using consistency adversarial robust operators for iterative updates, generating the optimal adversarial Q function, and using Bellmann’s infinite norm for error calculation, combining projection gradient descent algorithm and interval boundary propagation method, efficient training of adversarial robust Q network is achieved.
On the premise of ensuring the existence of the strategy, the consistent and excellent performance of the deep Q network in a clean and confrontational environment is achieved, and the adversarial robustness and training stability are improved.
Smart Images

Figure CN120087444A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular, to an adversarial robust Q - learning method and system based on Bellman infinite error. Background Art
[0002] Deep reinforcement learning has achieved remarkable success in solving complex decision - making problems. Among them, the Q - learning method iteratively updates the Q - function based on the Bellman optimal equation and makes decisions. The Deep Q - Network uses a neural network to approximately represent the Q - function and is trained based on the Bellman error. In an ideal situation, when the Bellman error approaches zero, the Q - network can well approximate the optimal Q - function. However, due to limitations such as the expressive power of neural networks and optimization algorithms, the Bellman error can actually only reach a very small non - zero value. Although in a clean environment without any adversarial perturbations, a Q - network with a small Bellman error can exhibit excellent performance.
[0003] However, existing Deep Q - Networks do not have stability guarantees under small Bellman errors. When facing carefully constructed and imperceptible small adversarial perturbations, the decisions based on the Q - network will completely fail. In addition, since adversarial robust policies may not exist under general conditions, existing adversarial robust training methods essentially need to trade - off between adversarial robustness and the existence of policies, sacrificing performance in the clean environment while enhancing adversarial robustness. This vulnerability and lack of adversarial robustness greatly limit the reliable deployment of deep reinforcement learning methods in real - world complex environments. Summary of the Invention
[0004] The present application provides an adversarial robust Q - learning method and system based on Bellman infinite error, which is used to design a deep Q - learning method that is both adversarially robust and has training stability while ensuring the existence of policies, so that the trained agent can exhibit consistent excellent performance in both clean and adversarial environments.
[0005] In a first aspect, the present application provides an adversarial robust Q - learning method based on Bellman infinite error. The adversarial robust Q - learning method based on Bellman infinite error includes:
[0006] According to the state space, action space, and dynamic transition probability function of the Markov decision process, use the Bellman optimal Q - function to construct an eigen - state neighborhood set;
[0007] Based on the eigen - state neighborhood set, perform iterative update operations on the state - action value function through a consistency adversarial robust operator to generate an optimal adversarial Q - function;
[0008] Use the Bellman infinite norm to calculate the error of the optimal adversarial Q - function to form an optimization training objective;
[0009] For the inner maximization problem in the optimized training objective, the projected gradient descent algorithm is used for solving operations to obtain adversarial state samples;
[0010] Based on the adversarial state samples, the upper and lower bounds of the output of the neural network in the local neighborhood are estimated via the interval bound propagation method to obtain the Q-value boundary range, and an alternative upper bound of the inner maximization problem is obtained;
[0011] According to the Q-value boundary range, the network parameters of the adversarial robust Q-network are optimized and updated by the adaptive momentum estimation algorithm to realize the training process of the adversarial robust Q-network.
[0012] In a second aspect, the present application provides an adversarial robust Q-learning system based on the Bellman infinite error. The adversarial robust Q-learning system based on the Bellman infinite error includes:
[0013] A construction module for constructing a set of eigenstate neighborhoods by using the Bellman optimal Q-function according to the state space, action space, and dynamic transition probability function of the Markov decision process;
[0014] An update module for iteratively updating and operating on the state-action value function through a consistent adversarial robust operator based on the set of eigenstate neighborhoods to generate an optimal adversarial Q-function;
[0015] A calculation module for calculating the error of the optimal adversarial Q-function by using the Bellman infinite norm to form an optimized training objective;
[0016] An operation module for solving and operating on the inner maximization problem in the optimized training objective by using the projected gradient descent algorithm to obtain a set of adversarial state samples;
[0017] An estimation module for estimating the upper and lower bounds of the output of the neural network in the local neighborhood via the interval bound propagation method based on the adversarial state samples to obtain the Q-value boundary range, and an alternative upper bound of the inner maximization problem is obtained;
[0018] An optimization module for optimizing and updating the network parameters of the adversarial robust Q-network by the adaptive momentum estimation algorithm according to the Q-value boundary range to realize the training process of the adversarial robust Q-network.
[0019] In the technical solution provided by this application, by modeling the state space and action space in the Markov decision process and using the state transition probability function to perform random perturbation processing on the observed state, the set of eigenstate neighborhoods is successfully constructed, effectively capturing the local structural features in the state space. On this basis, the consistent adversarial robust operator is used to perform iterative update operations on the state value function, significantly improving the stability of the Q-network in the face of adversarial perturbations. By calculating the error of the Q-function using the Bellman infinity norm, it provides a theoretical guarantee for the formation of the optimization objective, making the training process more reliable. For the inner maximization problem in the optimized training objective, the projected gradient descent algorithm is used for solution operations, which not only improves the solution efficiency but also ensures the quality of the solution. Through the interval bound propagation method, the upper and lower bounds of the output of the neural network in the local neighborhood are estimated, accurately obtaining the boundary range of the Q-value, providing a reliable metric standard for the robustness of the network. Finally, the adaptive momentum estimation algorithm is used to optimize and update the network parameters, achieving the efficient training of the adversarial robust Q-network. The entire method ensures the existence of the policy while successfully realizing the unity of adversarial robustness and training stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0021] Figure 1 FIG. is a schematic diagram of an embodiment of the adversarial robust Q-learning method based on the Bellman infinite error in the embodiments of this application;
[0022] Figure 2 FIG. is a schematic diagram of an embodiment of the adversarial robust Q-learning system based on the Bellman infinite error in the embodiments of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The embodiments of the present application provide an adversarial robust Q - learning method and system based on Bellman infinite error. Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims, and the above - mentioned drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that illustrated or described here. In addition, the term "comprising" or "having" and any of its variants are intended to cover non - exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0024] For ease of understanding, the specific process of the embodiments of the present application will be described below. Please refer to Figure 1 One embodiment of the adversarial robust Q - learning method based on Bellman infinite error in the embodiments of the present application includes:
[0025] Step S101: According to the state space, action space, and dynamic transition probability function of the Markov decision process, use the Bellman optimal Q - function to construct the set of eigen - state neighborhoods;
[0026] Step S102: Based on the set of eigen - state neighborhoods, perform iterative update operations on the state - action value function through a consistency adversarial robust operator to generate the optimal adversarial Q - function;
[0027] Step S103: Calculate the error of the optimal adversarial Q - function using the Bellman infinity norm to form an optimization training objective;
[0028] Step S104: For the inner - layer maximization problem in the optimization training objective, use the projected gradient descent algorithm to perform a solution operation to obtain an adversarial state sample set;
[0029] Step S105: Based on the adversarial state samples, estimate the upper and lower bounds of the output of the neural network in the local neighborhood through the interval bound propagation method, obtain the Q - value boundary range, and get an alternative upper bound for the inner - layer maximization problem;
[0030] Step S106: According to the Q - value boundary range, use the adaptive momentum estimation algorithm to perform gradient optimization updates on the network parameters of the adversarial robust Q - network to implement the training process of the adversarial robust Q - network.
[0031] It can be understood that the execution entity of the present application can be an adversarial robust Q - learning system based on Bellman infinite error, or a terminal or a server. Specifically, it is not limited here. The embodiments of the present application will be described by taking the server as the execution entity as an example.
[0032] Specifically, based on the Markov decision process modeling framework, the state space and action space are modeled and analyzed. The state space in the Markov decision process represents the set of all possible states that an agent can be in the environment, and the action space represents the set of all possible actions that the agent can take in each state. The state transition probability function describes the probability distribution of the agent transitioning to the next state given the current state and action. To achieve adversarial robustness training, the observed state is randomly perturbed to generate a set of neighborhoods of the eigenstate. The neighborhood of the eigenstate refers to the subset of states for which the optimal action remains unchanged in the perturbation set of any state. The initial state space is sampled through the state visitation distribution, and the immediate reward for each state-action pair is calculated using the state-action value function. The state-action value function measures the long-term value of taking a specific action in a certain state. Then, temporal difference calculation is performed on the state transition sequence based on the discount factor to obtain the state transition trajectory. Temporal difference calculation takes into account the decay effect of future rewards, and the discount factor usually takes values between 0 and 1. Next, the Bellman optimal operator is used to evaluate the state-action pair, and the Bellman optimal operator associates the value function of the current state with the maximum expected value function of the successor state.
[0033] After obtaining the set of neighborhoods of the eigenstate, the state value function is iteratively updated through the consistency adversarial robustness operator. The consistency adversarial robustness operator first calculates the conditional transition probability of the state-action pair, then performs Bellman iteration on the value function, and on this basis, searches for the optimal action space. The policy gradient method is used to optimize the parameters of the state value function. The policy gradient directly optimizes the policy parameters, avoiding the instability of value function approximation. Subsequently, a neural network is used to approximate the Q function, and the robustness of the network is improved through adversarial value function training. For the trained adversarial robust Q-network model, the Bellman infinity norm is used to calculate the error of the Q function. The Bellman infinity norm measures the maximum error of the Q function over the entire state-action space, and this error measurement method helps to ensure the global stability of the learned policy. Specifically, training data is sampled through the state-action visitation distribution, the Bellman error between the temporal difference target and the current Q value is calculated, and the maximum error value is extracted as the optimization target. For batch data, the errors are accumulated in a weighted average manner, and a surrogate loss function is used for convex approximation to finally obtain an optimizable training target.
[0034] For the inner maximization problem in the optimization training objective, the projected gradient descent algorithm is used to solve it. The algorithm first randomly initializes the perturbation state, and then iteratively updates the perturbation state through the gradient ascent method to continuously increase the value of the objective function. During the iteration process, the step size is adaptively adjusted to balance the convergence speed and stability. The projection operator ensures that the updated perturbation state still satisfies the constraint conditions, and the neighborhood determination ensures that the perturbation range is within the preset boundary. Finally, the obtained perturbation samples are clustered and sorted to form the final adversarial state sample set. When estimating the upper and lower bounds of the output of the neural network on the adversarial state sample set, the interval bound propagation method is used. This method first performs linear relaxation calculation on the weights of the neural network to obtain the value range of the output of each layer of the network. Then, through the interval mapping characteristics of the activation function, the output boundary is propagated and updated layer by layer. For the non-linear layer, the convex relaxation method is used for boundary propagation, and finally the exact boundary range of the Q value is obtained.
[0035] Finally, based on the obtained Q value boundary range, the adaptive momentum estimation algorithm is used to optimize and update the network parameters. The algorithm adaptively adjusts the learning rate of each parameter by calculating the first-order moment and second-order moment information of the gradient. Specifically, the exponential weighted average of the gradient information is used to obtain the momentum value, and the gradient variance is estimated to adaptively adjust the step size. The stability in the early training stage is ensured through bias correction, and finally the optimized network parameters are obtained.
[0036] For example, in the autonomous driving scenario, the state space includes information such as vehicle position, speed, and surrounding obstacles, and the action space includes operations such as acceleration, deceleration, and steering. For a certain state s=(x = 10m, v = 20km / h), the probability that the next state s'=(x' = 12m, v' = 25km / h) after taking the acceleration action a is calculated by the state transition probability function is 0.8. Based on this native cost state neighborhood, a state set with a speed perturbation range of ±2km / h and a position perturbation range of ±0.5m is included. The consistency adversarial robust operator updates the value function within this neighborhood to obtain an estimated value of Q(s, a)=85. When calculating the error using the Bellman infinity norm, the maximum error on 1000 state-action samples is 0.15. The projected gradient descent algorithm obtains the adversarial sample s_adv=(x = 10.3m, v = 21.5km / h) after 20 rounds of iteration. The interval bound propagation determines the upper and lower bounds of the Q value on this adversarial sample to be [80, 90]. Finally, through the adaptive momentum estimation algorithm, the network parameters are updated with a learning rate of 0.001, so that the output value of the model on the adversarial sample converges to the range of 85±2.
[0037] In the embodiments of the present application, by modeling the state space and action space in the Markov decision process and using the state transition probability function to perform random perturbation processing on the observed state, the set of eigenstate neighborhoods is successfully constructed, effectively capturing the local structural features in the state space. On this basis, the consistency adversarial robust operator is used to perform iterative update operations on the state value function, significantly improving the stability of the Q-network in the face of adversarial perturbations. The error of the Q-function is calculated through the Bellman infinity norm, providing a theoretical guarantee for the formation of the optimization objective and making the training process more reliable. For the inner maximization problem in the optimization training objective, the projected gradient descent algorithm is used for solution operations, which not only improves the solution efficiency but also ensures the quality of the solution. The upper and lower bounds of the output of the neural network in the local neighborhood are estimated through the interval bound propagation method, accurately obtaining the boundary range of the Q-value and providing a reliable metric for the robustness of the network. Finally, the network parameters are optimized and updated by the adaptive momentum estimation algorithm, realizing the efficient training of the adversarial robust Q-network. The entire method ensures the existence of the policy while successfully achieving the unity of adversarial robustness and training stability.
[0038] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0039] (1) Based on the state space, action space, and dynamic transition probability function of the Markov decision process, the state-action value function is iteratively calculated using the Bellman optimal equation to obtain the Bellman optimal Q-function;
[0040] (2) According to the Bellman optimal Q-function, the eigenstate is constructed, and the neighborhood division operation is performed on the perturbed state through the eigenstate consistency determination to generate the set of eigenstate neighborhoods.
[0041] Specifically, the adversarial robust Q-learning method based on the Bellman infinite error starts from the Markov decision process. The Markov decision process is defined as a six-tuple (S, A, r, P, γ, μ 0 ), where: S represents the state space, representing the set of all possible states of the agent; A represents the action space, containing all possible action choices; r: is the reward function, assigning an immediate reward value to each state-action pair; P: S × A → Δ(S) represents the transition dynamic function; γ is the discount factor; μ 0 ∈ Δ(S) is the initial state distribution.
[0042] For the probability distribution operation of the initial state space, the state-action visit distribution is adopted:
[0043]
[0044] where: Denote the state-action visitation probability under the initial distribution μ 0 and the policy π; s 0 Denote the initial state; s t , a t respectively denote the state and action at time t; Pr π Denote the transition probability under the policy π.
[0045] Based on the state visitation distribution, calculate the reward through the state-action value function:
[0046]
[0047] where: Q π (s, a) denotes the long-term value of taking action a in state s under the policy π; E π,P Denotes the expectation under the policy π and the transition dynamics P; r(s t , a t ) denotes the immediate reward obtained at time t.
[0048] The temporal difference target is defined as:
[0049] y t = r t + γQ(s t+1 , argmax a′ Q(s t+1 , a′; θ′); θ′)
[0050] where: y t Denotes the temporal difference target value; r t Denotes the immediate reward; θ′ denotes the target network parameter; θ denotes the current network parameter.
[0051] The Bellman optimal operator is defined as:
[0052]
[0053] Then the adversary perturbation function searches within the preset perturbation range B ε (s):
[0054]
[0055] The eigenstate consistency determination is defined as:
[0056]
[0057] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0058] (1) Based on the set of intrinsic state neighborhoods, use the consistency adversarial robust operator to perform Bellman iteration on the value function to obtain the state-action value update sequence;
[0059] (2) According to the state-action value update sequence, perform an optimal policy search on the action space via the maximization operator to obtain the optimal action set;
[0060] (3) Use the optimal action set to perform parameter optimization calculation on the state value function through policy gradient to form the policy iteration result;
[0061] (4) Based on the policy iteration result, use a neural network to perform an approximate representation operation on the Q function to obtain the Q-network parameter set;
[0062] (5) According to the Q-network parameter set, perform robustness training on the Q network through the adversarial value function to generate the optimal adversarial Q function.
[0063] Specifically, calculate the state transition probability. For each state-action pair (s, a), the conditional probability function P(s'|s, a) defines the probability of transitioning to the next state s' after performing action a in the current state s. This transition probability is represented by the conditional distribution P(·|s, a), where the conditional probability calculation takes into account the dimensionality characteristics of the state space and the constraints of the action space. After obtaining the state transition parameters, the consistency adversarial robust (CAR) operator is used for the Bellman iteration of the value function. The CAR operator defines how to update the state-action value function, taking into account the perturbation effects of the opponent. For each state s and action a, the CAR operator combines the immediate reward and the expected value of future discounted rewards, where the future rewards consider the worst-case opponent perturbations. By repeatedly applying the CAR operator, a state value update sequence is generated, which reflects the change trend of the value function during the iteration process.
[0064] Based on the state value update sequence, the maximization operator is used to search for the optimal policy in the action space. For each state s, the optimal action is determined by comparing the Q-values of different actions. This process generates an optimal action set, which contains the optimal actions to be taken in each state. This optimal action set is continuously optimized as the value function is updated. Using the obtained optimal action set, the policy gradient method is used to optimize the parameters of the state value function. The policy gradient calculation is based on the derivative of the performance objective function with respect to the policy parameters, and the parameters are updated along the gradient direction to improve the policy performance. This process produces the policy iteration result, which includes the updated policy parameters and the corresponding performance evaluation.
[0065] The results of policy iteration are then used to train a neural network to approximate the Q-function. The structure design of the neural network needs to consider the dimension of the state space and the characteristics of the action space, usually including multiple hidden layers, and each layer uses an appropriate activation function. The network parameters are optimized through the backpropagation algorithm, and finally a set of Q-network parameters is obtained, which can better approximate the true Q-function. Finally, the adversarial value function is used for the robust training of the Q-network. During the adversarial training process, for each training sample, the loss value under the worst-case perturbation is calculated, and this loss is used to update the network parameters. This training method can improve the robustness of the Q-network when facing opponent perturbations.
[0066] Taking the navigation control of an intelligent warehousing robot as an example for illustration. Consider a scenario with a 4-dimensional state space, including the position coordinates (x, y) of the robot, the orientation angle θ, and the current speed v. The action space includes 5 discrete actions: forward, backward, left turn, right turn, and stop.
[0067] First, for the state s = (2.5m, 1.8m, 30°, 0.5m / s), after executing the "forward" action, the conditional probability calculation shows that there is an 80% probability of transferring to the expected next state s' = (3.0m, 2.1m, 30°, 0.5m / s), a 15% probability of a small deviation, and a 5% probability of a large deviation. These transition probabilities form the state transition parameter matrix. Then, the iterative calculation of the CAR operator on this state-action pair considers an opponent perturbation range of ε = 0.1m. The Q-value obtained in the first round of iteration is 8.5, it drops to 8.2 in the second round, reaches 7.9 in the third round, and finally converges to 7.8, forming a state value update sequence {8.5, 8.2, 7.9, 7.8}.
[0068] The Q-values of all actions are compared through the maximization operator: [7.8, 6.5, 7.2, 7.0, 5.5], and "forward" is determined as the optimal action. This process is carried out on all states to construct the optimal action set. Based on this, the policy gradient calculates the gradient vector [0.3, -0.2, 0.1, 0.4] for updating the policy network parameters. The neural network adopts a three-layer structure (4 nodes in the input layer, 64 nodes in the hidden layer, and 5 nodes in the output layer), and uses the ReLU activation function. After 1000 rounds of training, the network parameters converge to a stable state, and the average prediction error on the validation set drops below 0.15.
[0069] Finally, adversarial training is carried out to search for the most adversarial perturbations within the ε-neighborhood of each state. For example, for the above state, the found adversarial sample is (2.58m, 1.75m, 32°, 0.48m / s), which maximizes the Q-value difference. After adversarial training, the performance of the model on adversarial samples is significantly improved, and the Q-value prediction error is reduced from the original 0.8 to 0.3, showing good robustness.
[0070] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0071] (1) Sample and calculate the training data through the state-action visitation distribution to obtain a state-action sample set;
[0072] (2) According to the state-action sample set, use the temporal difference target to estimate the Bellman error of the optimal adversarial Q-function to obtain a Bellman error sequence;
[0073] (3) Based on the Bellman error sequence, perform a maximum extraction operation on the error terms via the infinity norm to form an error boundary value;
[0074] (4) Use the error boundary value to calculate the error accumulation of the batch samples through weighted averaging to obtain a batch error statistic;
[0075] (5) According to the batch error statistic, use the surrogate loss function to perform a convex approximation operation on the training target to obtain a softened training target;
[0076] (6) According to the softened training target, perform an optimization iteration calculation on the network parameters through gradient update to form an optimized training target.
[0077] Specifically, sample and calculate the training data through the state-action visitation distribution. The state-action visitation distribution reflects the probability of accessing each state-action pair when executing a certain policy, and this probability value is jointly determined by the initial state distribution and the policy parameters. Sampling through this distribution ensures the representativeness and diversity of the samples. The sampling process records the actions taken, the rewards obtained, and the next state transferred to under each state, forming a complete state-action sample set. Based on the sampled state-action sample set, use the temporal difference target to estimate the Bellman error of the Q-function. The temporal difference target combines the immediate reward obtained currently and the estimation of the future value, where the future value is calculated through the target Q-network. For each sample, calculate the difference between the actually observed temporal difference target and the predicted value of the current Q-network, and this difference is the Bellman error. Arrange the Bellman errors of all samples in time sequence to form a Bellman error sequence.
[0078] The Bellman error sequence reflects the prediction errors of the Q-network on different state-action pairs. By performing a maximum extraction operation on these error terms using the infinity norm, the maximum error value is found. This maximum error value serves as the error boundary value, representing the performance of the Q-network in the worst-case scenario. The use of the infinity norm ensures the effective capture of the maximum error, providing a clear direction for improvement in subsequent optimization. For each sample in a training batch, a weighted average calculation is performed based on the error boundary value. The weights assigned to different samples are related to the magnitude of their errors, with samples having larger errors receiving higher weights during training. This weighting method makes the training process pay more attention to those samples with poor prediction effects, and the error statistic for the entire batch is obtained through cumulative calculation.
[0079] To make the training process more stable and easier to optimize, a surrogate loss function is used to perform a convex approximation operation on the training objective. The surrogate loss function maintains the main characteristics of the original objective while having better mathematical properties. Through this approximation process, a smoother softened training objective is obtained, facilitating gradient calculation and parameter update. Finally, based on the softened training objective, the network parameters are optimized iteratively through gradient updates. This process continuously adjusts the parameters of the Q-network to make its predicted values as close as possible to the temporal difference objective. Through multiple rounds of iterative optimization, an optimized training objective with good generalization performance is finally formed.
[0080] For example, in the intelligent medical diagnosis scenario, the state space includes various vital sign indicators of the patient, such as body temperature of 37.5°C, blood pressure of 125 mmHg, heart rate of 75 beats per minute, etc. The action space includes different treatment plan selections. First, 100 training samples are obtained through sampling from the state-action visitation distribution, and each sample includes the current state, the treatment plan adopted, the observed immediate effect, and the transferred state. For a specific sample, treatment plan A is selected in the given state, and it is observed that the patient's indicators improve, with an immediate reward value of 0.8 points, and the state transfers to a body temperature of 37.2°C, blood pressure of 120 mmHg, and heart rate of 72 beats per minute. The Bellman error for this sample is calculated to be 0.315, indicating the gap between the current prediction of the Q-network and the actual observed value. Among all the samples in the batch, the maximum error value is 0.42, and the average error is 0.28. After softening, this sample obtains a weight coefficient of 0.012 during training, and the corresponding loss contribution is 0.00378. After parameter update, the Q-value of this state-action pair increases from 8.9 to 8.923, and the error decreases to 0.292, showing an obvious optimization effect. This process continues to iterate until the Q-network can make accurate predictions in various states.
[0081] In a specific embodiment, the process of performing step S104 may specifically include the following steps:
[0082] (1) Initialize the perturbation state through stochastic gradient calculation according to the optimized training objective to obtain the initial perturbation state;
[0083] (2) According to the initial perturbation state, perform a maximization iterative operation on the objective function using gradient ascent to obtain a gradient update sequence;
[0084] (3) Based on the gradient update sequence, perform an adaptive adjustment calculation on the iteration direction through step size decay to form a projected gradient value;
[0085] (4) Use the projected gradient value to perform a constraint mapping operation on the perturbation state through a projection operator to obtain valid perturbation samples;
[0086] (5) According to the valid perturbation samples, perform a boundary constraint calculation on the perturbation range using neighborhood determination to obtain a legal perturbation set;
[0087] (6) According to the legal perturbation set, perform a clustering and sorting operation on the perturbation state through sample screening to form an adversarial state sample set.
[0088] Specifically, the perturbation state is initialized by the stochastic gradient method. The stochastic gradient method is an optimization algorithm based on random sampling, which estimates the gradient direction by randomly sampling the objective function. During the initialization process, random perturbations following a uniform distribution or a Gaussian distribution are added to each state dimension, and the perturbation magnitude is controlled by a preset perturbation range ε. This random initialization method helps to explore a larger state space and increase the diversity of adversarial samples. After obtaining the initial perturbation state, the gradient ascent method is used to perform a maximization iteration on the objective function. The gradient ascent is opposite to the gradient descent direction, and its purpose is to maximize the objective function value. In each iteration, the gradient of the objective function with respect to the perturbation is calculated at the current perturbation state, and the perturbation state is updated along the gradient direction. This process generates a series of gradient update values through multiple iterations to form a gradient update sequence. Each sequence element records the gradient direction and magnitude of the corresponding iteration step.
[0089] Based on the obtained gradient update sequence, it is necessary to adaptively adjust the iteration direction. Step size decay is a commonly used adaptive adjustment method. As the number of iterations increases, the update step size is gradually reduced. The step size decay coefficient is usually set as the ratio of the initial step size to the current number of iterations, or in an exponential decay form. Through the adjustment of step size decay, both the rapid convergence in the early iterations and the stability in the later iterations are ensured. The adjusted gradient value forms the projected gradient value. When using the projected gradient value to perform a constrained mapping on the perturbation state, a projection operator is required. The projection operator projects the updated state into a preset constraint set to ensure that the perturbation state satisfies the constraint conditions. The constraint conditions include perturbation range limits, state space boundaries, etc. The projection process first calculates whether the updated state satisfies the constraints. If it exceeds the constraint range, it is projected to the nearest legal position. In this way, the obtained perturbation samples are all valid, that is, the perturbation states that satisfy all the constraint conditions.
[0090] For the obtained valid perturbation samples, neighborhood determination is required to ensure the legality of the perturbation range. The neighborhood determination process compares the distance between the perturbation state and the original state to ensure that it does not exceed the preset perturbation radius ε. At the same time, it is also necessary to verify whether the perturbed state satisfies the problem-specific constraint conditions. Through this boundary constraint calculation, a series of legal perturbation states are screened out to form a legal perturbation set.
[0091] Finally, the states in the legal perturbation set are clustered and sorted. The clustering process groups similar perturbation states into one group, and the most representative sample is selected from each group. This sorting method not only reduces redundant samples but also retains the diversity of the perturbation distribution. The sample set after clustering and sorting is the final adversarial state sample set.
[0092] In a specific embodiment, the process of executing step S105 may specifically include the following steps:
[0093] (1) Based on the adversarial state sample set, perform a linear relaxation calculation on the neural network weights through interval propagation to obtain the relaxation boundary parameters;
[0094] (2) According to the relaxation boundary parameters, perform an interval mapping operation on the neuron output using an activation function to obtain the inter-layer propagation sequence;
[0095] (3) Based on the inter-layer propagation sequence, perform a boundary estimation calculation on the network output via the upper and lower bound functions to form an output boundary set;
[0096] (4) Using the output boundary set, perform a boundary propagation operation on the non-linear layer through convex relaxation to obtain the propagation constraint value;
[0097] (5) According to the propagation constraint value, perform an accurate estimation calculation on the Q value range through boundary contraction to obtain the Q value constraint set;
[0098] (6) According to the Q - value constraint set, perform upper and lower bound statistical operations on the local neighborhood through boundary merging to form the Q - value boundary range.
[0099] (7) Substitute the Q - value boundary range into the solution of the inner - layer maximization problem to obtain the alternative upper bound of the inner - layer maximization problem.
[0100] Specifically, perform linear relaxation calculations on the neural network weights through the interval propagation method. Interval propagation is a method for estimating the output range of a neural network. By linearizing the weight matrix, the non - linear transformation of the network layer is converted into linear constraints. The specific linear relaxation process requires performing interval arithmetic operations on the weight matrix W and bias vector b of each layer of the network to obtain the upper and lower bound intervals of the weights [Wl, Wu] and the upper and lower bound intervals of the biases [bl, bu]. These intervals constitute the relaxation boundary parameters for subsequent boundary propagation calculations. According to the obtained relaxation boundary parameters, the next step is to perform interval mapping operations on the output of the neurons. Each neuron in the neural network contains an activation function. Common activation functions include ReLU, tanh, and sigmoid, etc. For each activation function, it is necessary to calculate its output range on the input interval. Taking the ReLU function as an example, given the input interval [xl, xu], the output interval is [max(0, xl), max(0, xu)]. Arranging the output intervals of each layer of neurons in the forward propagation order of the network forms the inter - layer propagation sequence.
[0101] Based on the inter - layer propagation sequence, estimate the boundaries of the network's final output through specialized upper and lower bound functions. The upper - bound function uses a layer - by - layer maximization method to calculate the maximum value that each layer's output may reach; the lower - bound function uses a layer - by - layer minimization method to calculate the minimum value that each layer's output may reach. This process needs to consider the positive and negative signs of the weights. The upper bound of the interval is used for positive weights, and the lower bound of the interval is used for negative weights. The finally obtained output boundary set contains the boundaries of all possible output values of the network within the given input perturbation range. For the non - linear layers in the network, boundary propagation needs to be performed through the convex relaxation method. Convex relaxation approximates the non - linear function with a piece - wise linear function, ensuring the computability of the boundary propagation process. Specifically, for each non - linear activation function, construct a convex - up function and a convex - down function on its input interval, and these two functions respectively give the upper and lower bounds of the original function. The propagation constraint values calculated in this way not only ensure the correctness of the boundaries but also avoid the boundaries being too conservative.
[0102] For the obtained propagation constraint values, the boundary contraction technique is used to accurately estimate the Q-value range. The boundary contraction process is iteratively optimized to gradually tighten the upper and lower bounds of the Q-value. Each iteration uses the known constraint conditions to solve for the optimal boundary values through methods such as linear programming or quadratic programming. This process generates a Q-value constraint set that contains the exact value ranges of the Q-value in different states.
[0103] Finally, based on the Q-value constraint set, boundary merging operations are performed on the local neighborhood. Adjacent or overlapping Q-value intervals are merged to obtain a more compact boundary representation. This process needs to consider the intersection of intervals. For intersecting intervals, the union is taken, and for inclusive intervals, the larger interval is taken. The finally formed Q-value boundary range gives the output change range of the neural network under adversarial perturbations. Taking the intelligent logistics distribution scenario as an example, consider a three-layer fully connected neural network. The input is the state information such as the position coordinates, speed, and load of the distribution vehicle, and the output is the Q-value of different distribution routes. For the original state s = (x = 5.0km, y = 3.0km, v = 40km / h, load = 80%), its corresponding adversarial state sample set contains three samples: {s1, s2, s3}. After the weight matrix W1 of the first layer network is linearly relaxed, the weight range of the interval [-0.5, 0.5] is obtained, and the value range of the bias vector b1 is [-0.1, 0.1]. Applying the ReLU activation function to the 64 neurons of the first layer, the input interval [-1, 1] is mapped to the output interval [0, 1]. Calculating layer by layer in this way forms an inter-layer propagation sequence {h1, h2, h3}.
[0104] The upper and lower bound functions estimate that the output range of the first hidden layer is a [-0.4, 0.4] × 64-dimensional vector. The tanh activation function of the second layer is convexly relaxed, and its upper and lower bounds are approximated by two straight lines y = 0.8x + 0.2 and y = 0.8x - 0.2. After boundary propagation, the obtained propagation constraint value is in the range of [-0.6, 0.6]. Through boundary contraction calculation, the final Q-value constraint range is [7.5, 8.5]. Merging the Q-value intervals of the three adversarial samples, it is finally determined that in this local neighborhood, the Q-value boundary range is [7.2, 8.8], and this range reflects the sensitivity of the network output to input perturbations.
[0105] In a specific embodiment, the process of performing step S106 may specifically include the following steps:
[0106] (1) According to the Q-value boundary range, the gradient information is exponentially weighted calculated by the first moment to obtain the gradient momentum value;
[0107] (2) According to the gradient momentum value, the second moment is used to perform an adaptive estimation operation on the gradient variance to obtain a gradient variance sequence;
[0108] (3) Based on the gradient variance sequence, the momentum factor is corrected through bias correction to form a corrected momentum parameter;
[0109] (4) Using the corrected momentum parameter, the learning rate is adaptively updated through step size adjustment to obtain an optimized step size value;
[0110] (5) According to the optimized step size value, gradient descent calculation is performed on the network weights through parameter update to obtain an updated parameter set;
[0111] (6) Based on the updated parameter set, the adversarial robust Q-network is periodically synchronized through the target network to form a trained adversarial robust Q-network model.
[0112] Specifically, the gradient information is calculated by exponential weighted averaging of the first moment. The first moment is the mean estimate of the gradient, and the exponential weighted average smooths the gradient change by assigning a larger weight to the recent gradients. During the calculation, the hyperparameter β1 (usually set to 0.9) is used to control the decay rate of the historical gradient information, and the gradients at each time step are weighted and accumulated to obtain the gradient momentum value. This momentum value reflects the main direction of parameter update and helps to overcome local oscillations in the optimization process. Based on the calculated gradient momentum value, the second moment is then used to adaptively estimate the gradient variance. The second moment calculation considers the squared value of the gradient and is used to estimate the learning rate of each parameter. The hyperparameter β2 (usually set to 0.999) is used to control the decay rate of the variance estimate, and the squared gradient values are exponentially weighted averaged. This process generates a gradient variance sequence, and each element in the sequence corresponds to the degree of gradient change of a parameter.
[0113] Based on the obtained gradient variance sequence, the momentum factor needs to be corrected through a bias correction mechanism. Since the exponential weighted average will produce a bias in the initial stage of training, especially when the initial estimate value is zero, the first moment and second moment estimate values need to be corrected. The bias correction is achieved by dividing by the corresponding bias coefficients, and these coefficients gradually approach 1 as the number of training steps increases. The corrected momentum parameter obtained more accurately reflects the true distribution of the gradient. Using the corrected momentum parameter, the learning rate is adaptively updated through a step size adjustment method. The core idea of step size adjustment is to dynamically adjust the learning rate according to the gradient variance of the parameters. Parameters with large gradient variances use smaller learning rates, and parameters with small gradient variances use larger learning rates. This adaptive mechanism ensures that each parameter can be updated with an appropriate step size, thus obtaining an optimized step size value.
[0114] Based on the calculated optimized step size value, gradient descent calculation is performed on the network weights using the parameter update rule. The parameter update combines gradient momentum and adaptive learning rate, and updates each weight parameter in the network. The update process takes into account the corrected first-order moment estimate, second-order moment estimate, and global learning rate, and generates an updated parameter set. This parameter set contains the updated weight values of all layers of the network.
[0115] Finally, according to the updated parameter set, the Q-network is periodically updated through the target network synchronization mechanism. The target network is a periodically replicated version of the Q-network, and its parameters are synchronized with the current Q-network every fixed number of steps. This periodic synchronization mechanism helps to improve the stability of training and avoid oscillations during the parameter update process. After multiple rounds of training and synchronization, a Q-network model with adversarial robustness is finally formed.
[0116] The above describes the adversarial robust Q-learning method based on Bellman infinite error in the embodiments of the present application. Next, the adversarial robust Q-learning system based on Bellman infinite error in the embodiments of the present application will be described. Please refer to Figure 2 , an embodiment of the adversarial robust Q-learning system based on Bellman infinite error in the embodiments of the present application includes:
[0117] A construction module, configured to construct a set of eigenstate neighborhoods using the Bellman optimal Q-function according to the state space, action space, and dynamic transition probability function of the Markov decision process;
[0118] An update module, configured to iteratively update and calculate the state-action value function through a consistency adversarial robust operator based on the set of eigenstate neighborhoods to generate an optimal adversarial Q-function;
[0119] A calculation module, configured to calculate the error of the optimal adversarial Q-function using the Bellman infinite norm to form an optimized training objective;
[0120] An operation module, configured to solve the inner maximization problem in the optimized training objective using the projected gradient descent algorithm to obtain an adversarial state sample set;
[0121] An estimation module, configured to perform upper and lower bound estimation on the output of the neural network in the local neighborhood via the interval bound propagation method based on the adversarial state samples, obtain the Q-value boundary range, and obtain an alternative upper bound for the inner maximization problem;
[0122] An optimization module, configured to perform gradient optimization and update on the network parameters of the adversarial robust Q-network through the adaptive momentum estimation algorithm according to the Q-value boundary range to implement the training process of the adversarial robust Q-network.
[0123] Through the collaborative cooperation of the above-mentioned various components, by modeling the state space and action space in the Markov decision process and using the state transition probability function to perform stochastic perturbation processing on the observed states, the set of eigenstate neighborhoods is successfully constructed, effectively capturing the local structural features in the state space. On this basis, the consistent adversarial robust operator is used to perform iterative update operations on the state value function, significantly improving the stability of the Q-network in the face of adversarial perturbations. By calculating the error of the Q-function using the Bellman infinity norm, a theoretical guarantee is provided for the formation of the optimization objective, making the training process more reliable. For the inner maximization problem in the optimization training objective, the projected gradient descent algorithm is used for solution operations, which not only improves the solution efficiency but also ensures the quality of the solution. Through the interval bound propagation method, the upper and lower bounds of the output of the neural network in the local neighborhood are estimated, accurately obtaining the boundary range of the Q-value, providing a reliable metric for the robustness of the network. Finally, the network parameters are optimized and updated by the adaptive momentum estimation algorithm, realizing the efficient training of the adversarial robust Q-network. The entire method ensures the existence of the policy while successfully achieving the unity of adversarial robustness and training stability.
[0124] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A robust Q-learning method based on Bellman infinite error, characterized in that: The Bellman infinite error-based adversarial robust Q-learning method includes: According to the state space, action space and dynamic transition probability function of the Markov decision process, the eigenstate neighborhood set is constructed using the Bellman optimal Q function. Based on the eigenstate neighborhood set, the state-action value function is iteratively updated by a consistent adversarial robust operator to generate an optimal adversarial Q function; The Bellman infinite norm is used to calculate the error of the optimal adversarial Q function to form an optimized training target; For the inner layer maximization problem in the optimization training objective, a projected gradient descent algorithm is used to perform a solution operation to obtain an adversarial state sample; Based on the adversarial state samples, the output of the neural network in the local neighborhood is estimated by the interval bound propagation method, the Q value boundary range is obtained, and the alternative upper bound of the inner layer maximization problem is obtained; According to the Q value boundary range, the network parameters of the adversarial robust Q network are gradient optimized and updated through an adaptive momentum estimation algorithm to achieve the training process of the adversarial robust Q network.
2. The Bellman infinite error-based adversarial robust Q-learning method according to claim 1, characterized in that: The method of constructing an eigenstate neighborhood set using the Bellman optimal Q function according to the state space, action space and dynamic transition probability function of the Markov decision process includes: Based on the state space, action space and dynamic transition probability function of the Markov decision process, the state-action value function is iteratively calculated using the Bellman optimal equation to obtain the Bellman optimal Q function. The eigenstate is constructed according to the Bellman optimal Q function, and the neighborhood partition operation is performed on the disturbance state through the eigenstate consistency judgment to generate the eigenstate neighborhood set.
3. The Bellman infinite error-based adversarial robust Q-learning method according to claim 1, characterized in that: The method of performing iterative updating operation on the state-action value function based on the eigenstate neighborhood set by using a consistent adversarial robust operator to generate an optimal adversarial Q function includes: Based on the eigenstate neighborhood set, a consistent adversarial robust operator is used to perform Bellman iterative operation on the value function to obtain a state-action value update sequence; According to the state-action value update sequence, an optimal strategy search is performed on the action space via a maximization operator to obtain an optimal action set; Using the optimal action set, the state value function is optimized and calculated through the policy gradient to form a policy iteration result; Based on the strategy iteration result, a neural network is used to perform an approximate representation operation on the Q function to obtain a Q network parameter set; According to the Q network parameter set, the Q network is robustly trained through an adversarial value function to generate an optimal adversarial Q function.
4. The Bellman infinite error-based adversarial robust Q-learning method according to claim 1, characterized in that: The Bellman infinite norm is used to calculate the error of the optimal adversarial Q function to form an optimization training target, including: The training data is sampled and calculated through the state-action access distribution to obtain a state-action sample set; According to the state-action sample set, a Bellman error estimation is performed on the optimal adversarial Q function using a time difference target to obtain a Bellman error sequence; Based on the Bellman error sequence, performing a maximum value extraction operation on the error term via an infinite norm to form an error boundary value; Using the error boundary value, performing error accumulation calculation on the batch samples by weighted average to obtain batch error statistics; According to the batch error statistic, a convex approximation operation is performed on the training target using a substitute loss function to obtain a softened training target; According to the softened training target, the network parameters are optimized and iteratively calculated through gradient updating to form an optimized training target.
5. The Bellman infinite error-based adversarial robust Q-learning method according to claim 4, characterized in that: The inner layer maximization problem in the optimization training objective is solved by using a projected gradient descent algorithm to obtain adversarial state samples, including: According to the optimization training objective, the disturbance state is initialized and calculated by stochastic gradient to obtain an initial disturbance state; According to the initial disturbance state, the objective function is iterated by maximizing the gradient to obtain a gradient update sequence; Based on the gradient update sequence, adaptively adjusting and calculating the iteration direction via step size attenuation to form a projected gradient value; Using the projected gradient value, a constrained mapping operation is performed on the disturbance state through a projection operator to obtain a valid disturbance sample; According to the effective disturbance samples, the boundary constraint calculation of the disturbance range is performed by using the neighborhood judgment to obtain a legal disturbance set; According to the legal disturbance set, the disturbance states are clustered and sorted through sample screening to form an adversarial state sample set.
6. The Bellman infinite error-based adversarial robust Q-learning method according to claim 5, characterized in that: Based on the adversarial state samples, the upper and lower bounds of the output of the neural network in the local neighborhood are estimated by the interval bound propagation method, the Q value boundary range is obtained, and the alternative upper bound of the inner layer maximization problem is obtained, including: According to the adversarial state sample set, a linear relaxation calculation is performed on the neural network weights through interval propagation to obtain a relaxation boundary parameter; According to the relaxation boundary parameters, an activation function is used to perform interval mapping operation on the neuron output to obtain an inter-layer propagation sequence; Based on the inter-layer propagation sequence, the boundary estimation calculation of the network output is performed via upper and lower bound functions to form an output boundary set; Using the output boundary set, performing boundary propagation operation on the nonlinear layer through convex relaxation to obtain a propagation constraint value; According to the propagation constraint value, the Q value range is accurately estimated and calculated by using boundary contraction to obtain a Q value constraint set; According to the Q value constraint set, upper and lower bound statistical operations are performed on the local neighborhood by boundary merging to form a Q value boundary range; Substitute the Q value boundary range into the inner maximization problem to solve it and obtain the alternative upper bound of the inner maximization problem.
7. The Bellman infinite error-based adversarial robust Q-learning method according to claim 6, characterized in that: The method of performing gradient optimization and updating on the network parameters of the adversarial robust Q network by using an adaptive momentum estimation algorithm according to the Q value boundary range to implement the training process of the adversarial robust Q network includes: According to the Q value boundary range, the gradient information is exponentially weighted calculated by the first-order moment to obtain the gradient momentum value; According to the gradient momentum value, the gradient variance is adaptively estimated using the second-order moment to obtain a gradient variance sequence; Based on the gradient variance sequence, the momentum factor is corrected and calculated through deviation correction to form a corrected momentum parameter; Using the modified momentum parameter, adaptively updating the learning rate by adjusting the step size to obtain an optimized step size value; According to the optimization step value, the network weights are gradient-descent calculated using parameter update to obtain an updated parameter set; According to the update parameter set, the adversarial robust Q network is periodically synchronized through the target network to form a trained adversarial robust Q network model.
8. A Bellman infinite error-based adversarial robust Q-learning system, used to implement the Bellman infinite error-based adversarial robust Q-learning method as claimed in any one of claims 1 to 7, characterized in that: The Bellman infinite error-based adversarial robust Q-learning system includes: A construction module is used to construct an eigenstate neighborhood set using the Bellman optimal Q function according to the state space, action space and dynamic transition probability function of the Markov decision process; An updating module, configured to perform an iterative updating operation on the state-action value function through a consistent adversarial robust operator based on the eigenstate neighborhood set to generate an optimal adversarial Q function; A calculation module, used to calculate the error of the optimal adversarial Q function using the Bellman infinite norm to form an optimization training target; A calculation module, used for solving the inner layer maximization problem in the optimization training objective by using a projected gradient descent algorithm to obtain a set of adversarial state samples; An estimation module is used to estimate the upper and lower bounds of the output of the neural network in the local neighborhood based on the adversarial state samples through the interval bound propagation method, obtain the boundary range of the Q value, and obtain an alternative upper bound of the inner layer maximization problem; The optimization module is used to perform gradient optimization and update on the network parameters of the adversarial robust Q network through an adaptive momentum estimation algorithm according to the Q value boundary range, so as to realize the training process of the adversarial robust Q network.
Citation Information
Cited By
Multi-objective collaborative optimization pipe network automatic transmission and distribution control method
CN121364636A