Off-line reinforcement learning-based deception defense strategy optimization method

By establishing a cloud deception system and a Markov decision process model in network deception defense, and using multilayer perceptrons and Q-networks for policy optimization, the problems of high cost of high-complexity designs and insufficient capture capability of low-complexity designs are solved, and effective deception defense strategy optimization in real-world scenarios is realized.

CN120956480APending Publication Date: 2025-11-14GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511144907.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In existing technologies, high-complexity deception designs are costly to design in network attack and defense scenarios and are difficult to adapt to changes in scenarios. Low-complexity designs are difficult to capture advanced attackers. Traditional game theory-based and online reinforcement learning schemes have limitations in practical applications, and offline reinforcement learning has failed to effectively solve the problem of insufficient samples for network deception designs.

Method used

By establishing a cloud deception system, collecting attack logs, defining a Markov decision process model, training state transition functions using multilayer perceptrons and Q-networks, setting adversarial MDPs for policy optimization, and employing maximum entropy policy modeling and deep reinforcement learning, the deception defense strategy is updated to adapt to real-world scenarios.

Benefits of technology

It achieves effective deception defense strategy optimization in real-world scenarios, improves the real-world usability of deception design and the ability to capture advanced attackers, solves the problem of insufficient samples, and provides a robust deception design scheme.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956480A_ABST
    Figure CN120956480A_ABST
Patent Text Reader

Abstract

The invention provides a deception defense strategy optimization method based on offline reinforcement learning. The method comprises the following steps: establishing a cloud deception system according to a target scene to collect attack logs; defining an MDP according to the attack log; performing training iteration on the multi-layer perceptron according to the attack log to obtain a target MDP; setting an antagonistic MDP and determining agent strategy selection, performing maximum entropy strategy modeling to obtain a target strategy function, defining a state action value function, performing state action value function training to update an action selection strategy, defining a state value function, and updating the state selection strategy; strategy learning training is carried out, and when a training result meets a convergence condition, strategy learning training is stopped, and an optimal learning strategy is obtained; and sampling in the action space according to the optimal learning strategy to obtain an optimal action so as to modify the complexity design of the cloud spoofing system to realize spoofing defense strategy optimization. By applying the method, deception design training with more scene pertinence can be carried out based on real logs, and a robust deception defense strategy with availability is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network security technology, and in particular to an optimization method for deception defense strategies based on offline reinforcement learning. Background Technology

[0002] In network attack and defense scenarios, defensive deception is a key means for defenders to shift from passive to active defense. Network deception is a method of capturing attacker information via the web. Existing deception designs are generally divided into low-complexity and high-complexity designs. Low-complexity designs, such as low-interaction honeypots, record all accessing IPs through a single-level interactive fake interface. However, in real-world scenarios, this type of deception design can usually only capture probing attackers or very low-level attackers. High-complexity deception designs, such as high-interaction honeypots, are typically highly realistic deception designs with complete interactive processes. Theoretically, higher complexity designs mean higher capture levels and capabilities, enabling the capture of advanced attackers. However, high-complexity designs face challenges such as excessively high design costs and difficulty in developing adaptive deception mechanisms for different scenarios.

[0003] To address the problems in deception mechanism design, existing technologies have proposed several solutions, such as honeypot adaptive optimization based on game theory and policy optimization based on online reinforcement learning. However, game theory-based solutions inevitably idealize attack and defense scenarios during deception design policy optimization training, making it difficult to apply the trained deception mechanism effectively in real-world scenarios. Online reinforcement learning-based solutions inevitably face key issues such as the limited number of network attack and defense events and the high actual harm of network attacks, making it impossible to "trial and error" in critical protected environments to train the model under risky conditions. Offline reinforcement learning, as an emerging reinforcement learning framework, seems to offer a feasible solution for reinforcement learning on high-risk events; however, existing offline reinforcement learning optimizations are not adequately suited for network deception. Furthermore, traditional CQL-based offline reinforcement learning methods have also failed to address the critical issue of limited network attack and defense event samples.

[0004] Therefore, it is necessary to provide a new method for optimizing deception design strategies to improve their practical usability. Summary of the Invention

[0005] The purpose of this invention is to provide a deception defense strategy optimization method based on offline reinforcement learning, which can be used to design deception defense strategies that can be effectively applied in real-world scenarios.

[0006] In a first aspect, the deception defense strategy optimization method based on offline reinforcement learning provided by this invention includes: establishing a cloud deception system according to the target scenario, and collecting attack logs using the cloud deception system; defining an MDP based on the attack logs, including defining a state space, action space, reward function, policy space, and state transition function; training and iterating the hidden layer parameters of a multilayer perceptron according to the attack logs to obtain a target MDP, wherein the target MDP is used to output the mean and variance of the Gaussian distribution in the state transition function based on any state and action; setting an adversarial MDP and determining the agent's policy selection; performing maximum entropy policy modeling based on the agent's policy selection to obtain the target policy function; and then... The target policy function defines the state-action value function. A Q-network and a target Q-network are used to train the state-action value function to update the action selection policy. The state value function is defined, and the state selection policy is updated based on the updated action selection policy. Interactive data during the state-action value function training process is stored in a memory buffer. Data is sampled from the memory buffer, and policy learning is performed using the target MDP. Training stops when the results meet the convergence condition, yielding the optimal learned policy. Based on the optimal learned policy, the optimal action is sampled in the action space to modify the complexity design of the cloud deception system and optimize the deception defense strategy.

[0007] The beneficial effects of the deception defense strategy optimization method based on offline reinforcement learning provided by this invention are as follows: By collecting attack logs as the main input, a Markov decision process model is established, and the deception design scenario is used as the model. The state transition function that conforms to the scenario is trained by analyzing the log content. A zero-sum game between the agent and the model selection is established for robust model selection. Then, deep reinforcement learning is performed on the complete Markov decision process to obtain a deception design suitable for effective application in real-world scenarios.

[0008] In one possible embodiment, training the hidden layer parameters of the multilayer perceptron according to the attack log to obtain the target MDP includes: preprocessing the attack log to form a vector input to the input layer of the multilayer perceptron; defining the loss function of the multilayer perceptron as negative log-likelihood; taking the derivative of the loss function with respect to the Gaussian distribution mean and variance of the state transition function; updating the hidden layer parameters of the multilayer perceptron through gradient descent; and training iteratively to obtain the optimal parameters of the hidden layer to obtain the trained multilayer perceptron as the target MDP.

[0009] In another possible embodiment, preprocessing the attack log to form a vector includes: defining field transformation rules for state, action, new state = number of records, complexity, and number of new records; transforming the information recorded in the attack log into a new field representation according to the field transformation rules; and setting the new field representation as a vector.

[0010] In other possible embodiments, setting up an adversarial MDP and determining agent policy selection includes: setting the participants in the adversarial MDP as MDP model state selection and agent policy selection; when the MDP model state selection calls the target MDP to calculate the state selection function, sampling states that bring different benefits to the agent from the Gaussian distribution of the state selection function; according to the zero-sum game principle, the agent policy selection follows the following formula: Where, π * (·|s) represents the agent's policy choice, R(s,π) represents the reward function, and τ represents the coefficients. This represents the entropy term.

[0011] A state-action value function is trained using a Q-network and a target Q-network to update the action selection policy. The state value function is defined, and the state selection policy is updated using the updated action selection policy. This includes: using the agent's policy selection determined by an adversarial MDP as the current state to interact with the environment and obtain interaction data, storing the interaction data in a memory buffer; sampling and calculating the Q-value and target Q-value from the memory buffer using the Q-network and the target Q-network; when the loss function of the Q-network converges, selecting the smaller of the Q-value and the target Q-value to update the action selection policy; and finally, defining the state value function and updating the state selection policy using the updated action selection policy.

[0012] The process of sampling data from the memory buffer and training the policy network using the target MDP, stopping when the training results meet the convergence condition, includes: repeatedly executing the interaction between the current policy and the environment through the target MDP to obtain state transition samples, storing the state transition samples in the memory buffer, sampling data from the memory buffer to train the policy network, and stopping training when the convergence condition is met; meeting the convergence condition means that the objective function of the policy network converges, or the policy output is stable, or the cumulative reward exceeds a preset convergence threshold; wherein, the objective function of the policy network is obtained by establishing KL divergence based on the state-action value function and the state-value function.

[0013] In the memory buffer, data is stored in the form of triples, which are represented as (s, a, s'), where s represents the current state, a represents the action performed in the current state, and s' represents the updated state.

[0014] Secondly, this invention also provides a deception defense strategy optimization device based on offline reinforcement learning, comprising: a log collection unit, used to establish a cloud deception system according to a target scenario and collect attack logs using the cloud deception system; an MDP definition unit, used to define an MDP according to the attack logs, including defining a state space, action space, reward function, policy space, and state transition function; an MDP training unit, used to iteratively train the hidden layer parameters of a multilayer perceptron according to the attack logs to obtain a target MDP, wherein the target MDP is used to output the mean and variance of the Gaussian distribution in the state transition function according to any state and action; and a policy training unit, used to set an adversarial MDP and determine the agent's policy selection, and to perform maximum entropy policy construction based on the agent's policy selection. The system obtains the target policy function, defines the state-action value function based on the target policy function, and trains the state-action value function using a Q-network and a target Q-network to update the action selection policy. It also defines the state value function and updates the state selection policy based on the updated action selection policy. Interactive data during the state-action value function training process is stored in a memory buffer. A reinforcement learning unit samples data from the memory buffer and trains the policy using the target MDP. Training stops when the convergence condition is met, yielding the optimal learning policy. A deception design unit samples the optimal action in the action space based on the optimal learning policy to modify the complexity of the cloud deception system and optimize the deception defense strategy.

[0015] Thirdly, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described method for optimizing deception defense strategies based on offline reinforcement learning.

[0016] Fourthly, the present invention also provides an electronic device, comprising: a processor and a memory; the memory being used to store a computer program; the processor being used to execute the computer program stored in the memory, so that the electronic device executes the above-described offline reinforcement learning-based deception defense strategy optimization method.

[0017] For the beneficial effects of the second to fourth aspects mentioned above, please refer to the description of the first aspect mentioned above. Attached Figure Description

[0018] Figure 1 A flowchart illustrating a deception defense strategy optimization method based on offline reinforcement learning provided in an embodiment of the present invention;

[0019] Figure 2 A schematic diagram of the architecture of an optimization method for deception defense strategy based on offline reinforcement learning provided in an embodiment of the present invention;

[0020] Figure 3 A schematic diagram of a deception defense strategy optimization device based on offline reinforcement learning provided in an embodiment of the present invention;

[0021] Figure 4 This is a schematic diagram of an electronic device structure provided in an embodiment of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention. Unless otherwise defined, the technical or scientific terms used herein should have the ordinary meaning understood by those skilled in the art. The terms "comprising" and similar expressions used herein mean that the element or object preceding the word covers the element or object listed following the word and its equivalents, but do not exclude other elements or objects.

[0023] This embodiment provides a method for optimizing deception defense strategies based on offline reinforcement learning. See also... Figure 1 and Figure 2 The method includes:

[0024] S101: Establish a cloud deception system based on the target scenario and use the cloud deception system to collect attack logs.

[0025] In a specific embodiment, the target scenario refers to the scenario that is protected by the deception defense strategy, that is, the scenario that the final deception defense strategy will be applied to and needs to be protected.

[0026] A cloud-based deception system is built based on the target scenario. This system is deployed in the cloud, isolated from the real business systems. It attracts attackers by simulating different interfaces within the target scenario, thus preventing attacks on the business systems in advance. While attracting attacker access, the cloud-based deception system also collects and logs attack information, enabling the collection of attack logs within the target scenario. The logs record fields including the recording time, IP address, and target attack interface.

[0027] S102: Define the MDP based on the attack log, including defining the state space, action space, reward function, policy space, and state transition function.

[0028] Designing an MDP (Markov Decision Process) as the environment for reinforcement learning training of deception defense strategies involves defining the state space, action space, reward function, policy space, and state transition function. Defining the MDP based on attack logs collected in the target scenario makes the optimized deception defense strategy obtained through reinforcement learning more scenario-specific.

[0029] In a specific embodiment, defining an MDP based on attack logs specifically includes:

[0030] Define state space The definition of the state space is related to the number of attackers captured by the cloud deception system. In the historical attack log dataset, if the maximum number of simultaneous records across all time periods is defined as n, then the state S∈{s1,…,s} n It's important to note that the number of states and attackers are not strictly the same. To better measure the states under different detailed parameter settings, the state space is defined as a continuous space, meaning that states may exist in various forms, such as s. i =1.23.

[0031] Define action space The action space defines all possible actions an agent can take during training in reinforcement learning. The actions of an agent are defined in relation to complexity, with the complexity range defined as [1, m]. Therefore, the action space of the agent is defined as A = {a1, ..., a...}. m In the action space, different actions correspond to different levels of complexity. Complexity is defined as the breadth of the deception system design; higher complexity means a wider range of interactive possibilities, and vice versa. For example, in the deception design of a web website, complexity represents the number of interactive behaviors on the page, such as sub-interfaces or buttons.

[0032] Define the policy space π. When an agent chooses an action, it does so non-deterministically, meaning the agent selects an action based on a probability distribution of that action, i.e., a ~ π(·|s). Where the policy π(·|s) = [π(a1|s), ..., π(a2|s)]. m [s)] covers the probabilities of different actions across all action spaces.

[0033] Define state transition function When an agent is in state s, and chooses action 'a' through policy π, it will enter a new state with a probability distribution, i.e., s ~ P(s′|s,a). It's important to note that the state resulting from a particular policy, i.e., the number of attackers at a given complexity, follows a Gaussian distribution. Therefore, the probability of entering a new state is highest in a particular state, and decreases sequentially in adjacent states. Thus, the state transition function is defined as a Gaussian distribution, i.e. in, Let μ and σ represent a Gaussian distribution. 2 This represents the mean and variance of a Gaussian distribution.

[0034] Define the payoff function The reward function represents the penalty and reward of an agent's actions. When the agent is in state s, it will choose an action according to the policy, enter a new state s', and receive a reward, which is expressed as: in, Let N(s) represent the joint probability distribution of state a and state s′, N(s) represent the number of attackers captured in state s, ΔN(s,s′) represent the increase in the number of attackers from state s to s′, and Cost is the probability distribution of state a. a Let represent the cost of the agent choosing action 'a', including complexity costs, etc., and φ, λ, and ψ represent weight coefficients. The reward function calculates the joint probability expectation using two probability distribution policies π(·|s) and state transition P(·|s,a) to obtain the feedback reward of reinforcement learning.

[0035] S103: Train the hidden layer parameters of the multilayer perceptron (MLP) iteratively based on the attack log to obtain the target MDP. The target MDP is used to output the mean and variance of the Gaussian distribution in the state transition function according to any state and action.

[0036] In one possible implementation, in order to closely approximate the state transition situation in real-world scenarios and establish a more robust relationship between complexity and the captured situation, deep learning will be used to set the mean and variance, i.e., {μ(s,a),δ(s,a)}=Gaussian MLP(s,a).

[0037] In one possible embodiment, training the hidden layer parameters of the multilayer perceptron according to the attack log to obtain the target MDP includes: preprocessing the attack log to form a vector input to the input layer of the multilayer perceptron; defining the loss function of the multilayer perceptron as negative log-likelihood; taking the derivative of the loss function with respect to the Gaussian distribution mean and variance of the state transition function; updating the hidden layer parameters of the multilayer perceptron through gradient descent; and training iteratively to obtain the optimal parameters of the hidden layer to obtain the trained multilayer perceptron as the target MDP.

[0038] In a specific embodiment, preprocessing the attack log to form a vector includes: defining field transformation rules for state, action, new state = number of records, complexity, and number of new records; transforming the information recorded in the attack log into a new field representation according to the field transformation rules; and setting the new field representation as a vector.

[0039] For example, the data in the attack log is processed into MDP form. That is, according to the rule that state, action, new state = number of records, complexity, number of new records, the fields of the attack log are transformed into MDP form. All records after field transformation are set as vectors and input to the input layer of the multilayer perceptron, with the vector format satisfying: x = [x1, x2, x3]. The loss function of the MLP is defined as negative log-likelihood (NLL) to correlate the content of the MLP model with the real scene as much as possible. The specific loss function satisfies the following formula: Here, y represents the observed value. NLL focuses on the entire Gaussian distribution, aiming to make μ close to the true value and predict the uncertainty σ. There are two objectives to optimize: μ and σ. The optimization of μ is achieved by differentiating the loss function, i.e. This serves two purposes: the mean of the Gaussian distribution will approach the true value, and if the value of σ is too large, the mean will not be required to approximate the true value. Taking the derivative with respect to σ, we get... When the error is large, σ is increased to avoid the impact of incorrect predictions. If the error is small, σ is decreased to train the model more accurately. The parameters of the hidden layer are updated using gradient descent. The optimal parameters are obtained through iterative training of the MLP, resulting in the target MLP. The target MLP can output a Gaussian distribution μ in the state transition function for any input state and action. * and σ * Among them, μ * For the optimal representation of the mean of a Gaussian distribution, σ * This is the optimal representation of the variance of the Gaussian distribution.

[0040] S104: Set up an adversarial MDP and determine the agent's policy selection. Based on the agent's policy selection, perform maximum entropy policy modeling to obtain the target policy function. Define the state-action value function based on the target policy function. Use a Q-network and a target Q-network to train the state-action value function to update the action selection policy. Define the state value function. Update the state selection policy through the state value function based on the updated action selection policy. The interactive data during the training of the state-action value function is stored in a memory buffer.

[0041] In one possible embodiment, setting up an adversarial MDP and determining agent policy selection includes: defining the participants in the adversarial MDP as MDP model state selection and agent policy selection; when the MDP model state selection calls the target MDP to calculate the state selection function, sampling states that bring different rewards to the agent from the Gaussian distribution of the state selection function; based on the zero-sum game principle, the agent policy selection follows the following formula: Where, π * (·|s) represents the agent's policy choice, R(s,π) represents the reward function, and τ represents the coefficients. This represents the entropy term.

[0042] In one possible embodiment, a state-action value function is trained using a Q-network and a target Q-network to update the action selection policy. The state value function is defined, and the state selection policy is updated using the updated action selection policy through the state value function. This includes: using the agent's policy selection determined by an adversarial MDP as the current state to interact with the environment and obtain interaction data, storing the interaction data in a memory buffer; using the Q-network and the target Q-network to sample and calculate the Q-value and the target Q-value from the memory buffer; when the loss function of the Q-network converges, selecting the smaller value between the Q-value and the target Q-value to update the action selection policy; defining the state value function, and updating the state selection policy using the updated action selection policy through the state value function.

[0043] In a specific embodiment, due to the limitations of offline reinforcement learning models—namely, the difficulty in predicting state transition functions for states and actions not provided in the dataset—a robust method is needed for state transition prediction. In this embodiment, an adversarial MDP is defined using a two-player zero-sum game. Two adversarial game participants are identified as the agent's policy choice participant and the MDP model's state choice participant. The agent's state choice aims to maximize its own gain, while the MDP model's state choice aims to minimize its own gain. Specifically, when the MDP model predicts the state choice function by calling the target MDP, it samples K times from the Gaussian distribution of the state choice function to obtain states s that bring different gains to the agent. Then, based on the principle that Minimax in a two-player zero-sum game is a Nash equilibrium, the agent's policy choice follows: Where, π * (·|s) represents the agent's policy choice, R(s,π) represents the reward function, and τ represents the coefficients. This represents the entropy term. The Minimax algorithm can be used to obtain robust reinforcement learning policy selection based on model selection. This section considers entropy as a policy factor. When K is sufficiently large, if the lowest-yield state selection algorithm is consistently used pessimistically, MDP will always choose the lowest-yield state because the Gaussian distribution always has a certain probability of sampling the lowest-yield state. By considering the randomness and uncertainty of entropy, the trained model will not always choose the lowest-yield state, ensuring the effectiveness and convergence of the training.

[0044] To model the maximum entropy policy, entropy is incorporated into the policy considerations. The objective policy function is specifically expressed as: In this context, state s is selected based on the agent's policy determined by the set adversarial MDP, and α represents the coefficient of entropy. As an entropy term, it is used to measure the rarity of an action. Since a smaller negative logarithm indicates a higher probability of selection, the entropy increases, leading to a preference for actions with lower probabilities during action selection, thereby increasing the exploratory nature of the process.

[0045] Based on the objective policy function, the state-action value function is defined using the Bellman equation as follows:

[0046]

[0047] Among them, Q π (s,a) represents the state-action value, s′ represents the random variable under the state transition function p(·|s,a), and a′ represents the random variable under the policy π(·|s′). The state s is selected according to the agent policy determined by the adversarial MDP. γ∈[0,1] represents the discount coefficient, which considers the impact of future expected returns. A Q-network and a target Q-network are used to train the state-action value function to update the action selection policy: the agent policy selection determined by the adversarial MDP is used as the current state to interact with the environment to obtain interaction data. After each interaction, the interaction data is saved in a memory buffer. Then, the Q-network samples from the memory buffer for training, with the training objective being to make the Q-value approximate the target Q-value. In the memory buffer, data is stored in the form of triples, represented as (s,a,s'), where s represents the current state, a represents the action performed in the current state, and s' represents the updated state. The loss function of the Q-network is defined as: θ represents the target parameters to be trained, (s,a,s′)~D represents the memory replay mechanism, and Q θ (s,a) represents the state-action value estimated by the current Q-network, V θ′(s′) represents the target state value, and θ′ is obtained from the soft update of the target Q-network, i.e., θ′=τθ+(1-τ)θ′. When the loss function of the Q-network converges, the smaller of the Q-value and the target Q-value is selected for action selection policy update. In defining and quantifying the state value, the Q-values ​​of all possible actions in the current state (the smaller of the Q-values ​​obtained when training the state-action value function using the Q-network and the target Q-network) and the entropy term are considered. The state corresponding to the agent policy determined according to the set adversarial MDP is selected as the current state, and the state value is defined as: V(s)=E a~π(·|S) [Q(s,a)+αH(π(·|s))], where Q(s,a) represents the state action function.

[0048] S105: Sample data from the memory buffer, train the policy through the target MDP, and stop when the training result meets the convergence condition to obtain the optimal learning policy.

[0049] In one possible embodiment, sampling data from a memory buffer and training the policy network through a target MDP, stopping when the training results meet the convergence condition, includes: repeatedly executing the process of using the current policy to interact with the environment through the target MDP to obtain state transition samples, storing the state transition samples in a memory buffer, sampling data from the memory buffer to train the policy network, until the training stops when the convergence condition is met; meeting the convergence condition means that the objective function of the policy network converges, or the policy output is stable, or the cumulative reward exceeds a preset convergence threshold; wherein, the objective function of the policy network is obtained by establishing KL divergence based on the state-action value function and the state-value function.

[0050] For example, policy output stability means that the policy distribution in a given state is not too uniform; when the policy output is stable, a unique policy with the highest return can be found in each specific state. The cumulative return is derived by summing the returns from each training iteration. Policy network training converges when one of the following criteria is met: the policy network's objective function converges, the policy output is stable, or the cumulative return exceeds a preset convergence threshold. Once the policy network converges, the optimal learned policy can be obtained, meaning the optimal policy should be determined for each state.

[0051] In a specific embodiment, to better measure whether an action is more advantageous in a specific state, an advantage function is used instead of the Q value. The advantage function can be expressed as A(s,a) = Q(s,a) - V(s). It is important to note that the influence of the entropy term has already been considered when defining the state value function V(s) and the state-action value function Q(s,a). Therefore, policy selection is related to the advantage function, expressed as π(a|s)∝exp(Q(s,a)-V(s)). Based on this, the KL divergence is established as the objective function of the policy network. The objective function of the policy network satisfies the following formula: Where φ represents the parameters of the policy network, J π (φ) represents the objective function, D KL This indicates the calculation of the KL divergence.

[0052] S106: Based on the optimal learning strategy, sample the optimal action in the action space to modify the complexity design of the cloud deception system and implement the deception defense strategy optimization.

[0053] In one possible implementation, the optimal learning strategy π * (a|s) maximizes expected return and considers entropy during training. The optimal learning strategy with entropy can be directly used in the target scenario, and applying this strategy enables robust deception design within the target scenario. By using the strategy, the optimal action (a ~ π(·|s)) is sampled in the action space based on the learned strategy. The optimal action is then selected according to the action design in the Markov decision process, achieving optimal deception complexity and optimizing the deception defense strategy.

[0054] This invention provides a deception defense strategy optimization method based on offline reinforcement learning. It uses attack logs as the primary input to establish a Markov decision process model under network deception, employing the deception design scenario as the model. The method trains a state transition function that fits the scenario by analyzing the log content. A robust model selection is achieved through a zero-sum game between the agent and model selection, overcoming the key challenges of sample constraints and excessively high payoff estimation in traditional offline reinforcement learning. Then, deep reinforcement learning is applied to the complete Markov decision process to derive a deception design suitable for effective application in real-world scenarios.

[0055] The offline reinforcement learning-based deception defense strategy optimization method, based on a Markov decision model built from real log structures, offers greater practical usability than existing idealized game-theoretic models. Offline reinforcement learning using attack logs from real-world scenarios enables the development of more scenario-specific deception designs in a risk-free manner compared to the high-risk training methods employed in online reinforcement learning. The application of the trained target MDP model addresses the critical challenge of limited training samples in traditional CQL offline reinforcement learning. By modeling a zero-sum game in the scenario selection strategy, robust policy results are maintained even in unknown scenarios, thus overcoming the challenge of overestimating rewards in offline reinforcement learning.

[0056] See the instruction manual appendix Figure 3 This embodiment also provides a deception defense strategy optimization device based on offline reinforcement learning, which is used to implement the above method embodiment. The device includes:

[0057] Log collection unit 201 is used to establish a cloud deception system based on the target scenario and to collect attack logs using the cloud deception system.

[0058] MDP definition unit 202 is used to define an MDP based on attack logs, including defining the state space, action space, reward function, policy space, and state transition function.

[0059] MDP training unit 203 is used to train the hidden layer parameters of the multilayer perceptron iteratively based on the attack log to obtain the target MDP. The target MDP is used to output the mean and variance of the Gaussian distribution in the state transition function based on any state and action.

[0060] The policy training unit 204 is used to set up an adversarial MDP and determine the agent's policy selection. Based on the agent's policy selection, it performs maximum entropy policy modeling to obtain the target policy function. Based on the target policy function, it defines a state-action value function. It uses a Q-network and a target Q-network to train the state-action value function to update the action selection policy. It defines a state value function and updates the state selection policy based on the updated action selection policy. The interactive data during the state-action value function training process is stored in a memory buffer.

[0061] The reinforcement learning unit 205 is used to sample data from the memory buffer, perform policy learning training through the target MDP, and stop when the training results meet the convergence condition to obtain the optimal learning policy.

[0062] The deception design unit 206 is used to sample the optimal action in the action space according to the optimal learning strategy to modify the complexity design of the cloud deception system and realize the optimization of the deception defense strategy.

[0063] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0064] In other embodiments of this application, an electronic device is disclosed, such as... Figure 4 As shown, the electronic device 300 may include: one or more processors 301; a memory 302; a display 303; one or more application programs (not shown); and one or more computer programs 304. These devices can be connected via one or more communication buses 305. The one or more computer programs 304 are stored in the memory and configured to be executed by the one or more processors 301. The one or more computer programs 304 include instructions that can be used to perform actions such as... Figure 1 And the various steps in the corresponding embodiments.

[0065] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0066] In the embodiments of this application, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0067] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as flash memory, portable hard disk, read-only memory, random access memory, magnetic disk, or optical disk.

[0068] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.

Claims

1. A method for optimizing deception defense strategies based on offline reinforcement learning, characterized in that, include: A cloud deception system is established based on the target scenario, and attack logs are collected using the cloud deception system. The attack log is used to define an MDP, including defining the state space, action space, reward function, policy space, and state transition function. The target MDP is obtained by training the hidden layer parameters of the multilayer perceptron according to the attack log. The target MDP is used to output the mean and variance of the Gaussian distribution in the state transition function according to any state and action. An adversarial MDP is set up and the agent policy selection is determined. The target policy function is obtained by performing maximum entropy policy modeling based on the agent policy selection. The state-action value function is defined based on the target policy function. The state-action value function is trained using a Q network and a target Q network to update the action selection policy. The state value function is defined. The state selection policy is updated through the state value function based on the updated action selection policy. The interactive data during the training process of the state-action value function is stored in a memory buffer. Data is sampled from the memory buffer, and policy learning is performed through the target MDP. The process stops when the training results meet the convergence condition, and the optimal learning policy is obtained. The optimal action is sampled in the action space according to the optimal learning strategy to modify the complexity design of the cloud deception system and optimize the deception defense strategy.

2. The method according to claim 1, characterized in that, Based on the attack logs, the target MDP is obtained by iteratively training the hidden layer parameters of the multilayer perceptron, including: The attack logs are preprocessed to form vectors that are input to the input layer of the multilayer perceptron. The loss function of the multilayer perceptron is defined as negative log-likelihood. The loss function is used to differentiate the mean and variance of the Gaussian distribution of the state transition function. The hidden layer parameters of the multilayer perceptron are updated by gradient descent. The optimal parameters of the hidden layer are obtained through training iterations to obtain the trained multilayer perceptron as the target MDP.

3. The method according to claim 2, characterized in that, Preprocessing the attack logs to form vectors includes: Define the field conversion rules for state, action, and new state as the number of records, complexity, and number of new records; The information recorded in the attack log is converted into a new field representation according to the field conversion rules. Set the new field representation as a vector.

4. The method according to claim 1, characterized in that, Setting up an adversarial MDP and determining agent policy selection includes: The participants in the adversarial MDP are defined as the MDP model state selection and agent policy selection. When the MDP model calls the target MDP to calculate the state selection function, it samples the states that bring different benefits to the agent from the Gaussian distribution of the state selection function. Based on the zero-sum game principle, the agent's strategy selection follows the formula below: Where, π * (·|s) represents the agent's policy choice, R(s,π) represents the reward function, and τ represents the coefficients. This represents the entropy term.

5. The method according to claim 1, characterized in that, A Q-network and a target Q-network are used to train the state-action value function to update the action selection policy. The state value function is defined, and the state selection policy is updated using the state value function based on the updated action selection policy. This includes: The agent's policy selection, determined through adversarial MDP, is used as the current state to interact with the environment and obtain interaction data, which is then stored in a memory buffer. The Q-value and target Q-value are calculated by sampling from the memory buffer using a Q-network and a target Q-network. When the loss function of the Q-network converges, the smaller value between the Q-value and the target Q-value is selected for action selection policy update. Define a state value function, and update the state selection policy based on the updated action selection policy using the state value function.

6. The method according to claim 1, characterized in that, Data is sampled from the memory buffer, and policy learning is performed using the target MDP. The process stops when the training results meet the convergence condition. Repeatedly execute the process of using the current policy to interact with the environment through the target MDP to obtain state transition samples, store the state transition samples in the memory buffer, sample data from the memory buffer to train the policy network, and stop training when the convergence condition is met. Meeting the convergence condition means that the objective function of the policy network converges, or the policy output is stable, or the cumulative reward exceeds the preset convergence threshold. The objective function of the policy network is obtained by establishing KL divergence based on the state-action value function and the state-value function.

7. The method according to claim 1, characterized in that, In the memory buffer, data is stored in the form of triples, which are represented as (s, a, s'), where s represents the current state, a represents the action performed in the current state, and s' represents the updated state.

8. A deception defense strategy optimization device based on offline reinforcement learning, characterized in that, The device includes: A log collection unit is used to establish a cloud deception system based on the target scenario and to collect attack logs using the cloud deception system. The MDP definition unit is used to define an MDP based on the attack log, including defining the state space, action space, reward function, policy space, and state transition function. The MDP training unit is used to train and iterate the hidden layer parameters of the multilayer perceptron according to the attack log to obtain the target MDP. The target MDP is used to output the mean and variance of the Gaussian distribution in the state transition function according to any state and action. The policy training unit is used to set up an adversarial MDP and determine the agent's policy selection. Based on the agent's policy selection, it performs maximum entropy policy modeling to obtain the target policy function. Based on the target policy function, it defines a state-action value function. It uses a Q-network and a target Q-network to train the state-action value function to update the action selection policy. It defines a state value function and updates the state selection policy based on the updated action selection policy. The interactive data during the state-action value function training process is stored in a memory buffer. The reinforcement learning unit is used to sample data from the memory buffer, learn and train the policy through the target MDP, and stop when the training result meets the convergence condition to obtain the optimal learning policy. The deception design unit is used to sample the optimal action in the action space according to the optimal learning strategy to modify the complexity design of the cloud deception system and realize the optimization of the deception defense strategy.

9. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the deception defense strategy optimization method based on offline reinforcement learning as described in any one of claims 1 to 7.

10. An electronic device, characterized in that, include: Processor and memory; The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory to cause the electronic device to perform the deception defense strategy optimization method based on offline reinforcement learning as described in any one of claims 1 to 7.