Dynamic strategy real-time optimization method based on DQN and mixed Nash equilibrium

By introducing a dynamic strategy optimization method based on DQN and hybrid Nash equilibrium in the decision system, the problem that traditional decision systems are difficult to adapt to and optimize decisions in dynamic changes and uncertain environments is solved, and more flexible and accurate intelligent decision support is achieved.

CN120163068APending Publication Date: 2025-06-17XIAN UNIV OF TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510380694.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Traditional decision-making systems are difficult to achieve adaptive and optimized decision-making in the face of dynamic changes and uncertainties, especially in complex game environments.

Method used

A dynamic strategy real-time optimization method based on DQN and hybrid Nash equilibrium is adopted to realize intelligent decision support through deep neural network initialization, building a benefit matrix, iteratively solve mixed Nash equilibrium, extract mixed Nash equilibrium vectors and generate final decision solutions.

Benefits of technology

It improves the flexibility and accuracy of the decision-making system, and can realize adaptive and real-time optimization of decision-making in complex game environments, and is suitable for business negotiations, resource allocation and network security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163068A_ABST
    Figure CN120163068A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic strategy real-time optimization method based on DQN and hybrid Nash equilibrium, and belongs to the technical field of deep learning. The method comprises the following steps: estimating a participant Q value by initializing a deep neural network, and inputting game information to construct a benefit matrix; and utilizing DQN iterative optimization network parameters to approach a mixed Nash equilibrium solution, extracting an equilibrium vector to generate a dynamic strategy, and realizing a real-time adaptive decision. The method solves the technical problems of strategy dynamic adjustment and global optimization in a complex game scene, and is suitable for the fields of resource allocation, intelligent confrontation deduction and the like. The innovation point lies in that efficient decision support is provided for the high-dimensional dynamic game by combining the autonomous learning ability of the DQN and a multi-benefit balance mechanism of mixed Nash equilibrium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning, and provides a dynamic strategy real-time optimization method based on DQN (Deep Q-Learning) and mixed Nash equilibrium. Background Technique

[0002] With the rapid development of artificial intelligence technology, the applications of deep learning and reinforcement learning in complex decision-making problems are becoming increasingly widespread. Traditional decision-making systems usually rely on predefined rules or static models, and it is difficult to cope with dynamic environments and uncertainties. Especially in the field of game theory, how to achieve optimal decisions in complex games with multiple participants has always been the focus of attention in the academic and industrial communities.

[0003] Game theory provides a theoretical basis for strategic choices among multiple decision-makers, while deep learning provides new technical means for game decision-making through its powerful data learning and pattern recognition capabilities. As a representative algorithm of deep reinforcement learning, the Deep Q-Network (DQN) can continuously optimize strategies by interacting with the environment and has achieved remarkable results in fields such as games and robot control. However, how to combine DQN with game theory to solve the adaptive decision-making problem in complex games remains a challenging research direction.

[0004] Deep learning is a machine learning technology that mimics the neural network structure of the human brain and learns and understands data through multi-layer neural networks. These neural networks include an input layer, hidden layers, and an output layer, and continuously optimize the weights through the backpropagation algorithm to minimize the gap between the predicted result and the actual result. Deep learning has achieved remarkable achievements in fields such as image recognition, speech recognition, and natural language processing, and is widely used in the field of artificial intelligence.

[0005] DQN (Deep Q-Network) is a reinforcement learning algorithm aimed at learning and implementing the value function (Q function) through a deep neural network, enabling the agent to make optimal decisions. It combines the ideas of deep learning and reinforcement learning. In DQN, the neural network is used to estimate the Q value of each possible action, representing the long-term reward expectation of this action for the current state. Through techniques such as experience replay and target network, DQN improves stability and convergence and overcomes some challenges in reinforcement learning. This algorithm was initially applied by DeepMind to learn to play video games and has achieved success in multiple fields.

[0006] Game theory is a mathematical theory that studies the strategic choices made by decision-makers in an interactive environment. Game methods include various models and strategies, and some of the main concepts are as follows: (1) Game models: Describe the game participants, available strategies, and possible outcomes. Common models include zero-sum games, cooperative games, and non-cooperative games. (2) Strategies: The action plans of participants, with the goal of optimizing their own interests. Strategies can be pure strategies (deterministic) or mixed strategies (random). (3) Nash equilibrium: In non-cooperative games, it refers to the combination of strategies chosen by each participant such that no one has an incentive to change their strategy given the strategies of others. (4) Coordination and cooperation: Coordination games focus on how to align the interests of participants, while cooperative games study how participants negotiate to allocate resources. Game methods have wide applications in fields such as economics, computer science, and biology, and help to understand the interactions between decision-makers and optimal decision-making strategies.

[0007] The mixed-strategy Nash equilibrium is a theoretical solution in game theory for game participants in situations of uncertainty or incomplete information. In this equilibrium state, players no longer simply choose deterministic pure strategies, but instead randomly select different pure strategies through a probability distribution. Such randomness can provide a mixed possibility for players' decisions in different situations. Specifically, assume that the mixed strategies of two players are probability distributions p and q respectively, and each player will choose different pure strategies with these probabilities. The key to the mixed-strategy Nash equilibrium is that when all players adopt mixed strategies, no single player has an incentive to change their probability distribution because such a change cannot increase their expected utility. The concept of the mixed-strategy Nash equilibrium makes game theory more closely approximate real-world situations, especially in situations of information asymmetry or incomplete information. By introducing randomness, the mixed-strategy Nash equilibrium provides a more flexible game model that can better describe the decision-making behavior of game participants in uncertain environments. Summary of the Invention

[0008] The objective of the present invention is to provide a dynamic strategy adaptive generation and real-time optimization method based on DQN and hybrid Nash equilibrium. This method can be applied to multiple fields, including business negotiation, resource allocation, network security, etc., providing intelligent decision-making support for complex decision-making scenarios. Since traditional decision-making systems show limitations in the face of increasing uncertainty and dynamic changes, there is an urgent need for more intelligent methods to achieve immediate decision-making. The uniqueness of this method lies in the introduction of deep reinforcement learning (DQN), enabling the system to continuously learn and optimize decision-making strategies based on real-time environmental feedback, bringing new possibilities to the field of intelligent decision-making. By integrating DQN, the system can actively adapt to evolving scenarios, thereby endowing intelligent decision-making with more flexible and accurate characteristics. This innovation provides unprecedented opportunities for the intelligentization of decision-making systems, enabling them to better cope with the challenges of dynamic and uncertain modern society.

[0009] The technical solution adopted by the present invention is as follows:

[0010] A dynamic strategy real-time optimization method based on DQN and hybrid Nash equilibrium specifically includes the following steps:

[0011] Step 1, initialization of the DQN network model: Initialize the deep neural network for estimating the Q value in the game process;

[0012] Step 2, input of game participants and strategy sets: Provide information on game participants and their respective strategy sets, providing basic data for the input of the DQN network model algorithm;

[0013] Step 3, construction of the payoff matrix: Establish the payoff matrix of the game process, reflecting the payoffs of each participant under different strategy combinations;

[0014] Step 4, iterative solution of the DQN network model: Use the DQN network model for iteration, and continuously adjust the parameters of the DQN network model through the backpropagation algorithm to approximate the optimal solution of the hybrid Nash equilibrium in the game process;

[0015] Step 5, obtaining the hybrid Nash equilibrium vector: Extract the hybrid Nash equilibrium vector from the trained DQN network model, representing the optimal strategy combination of the participants in the game;

[0016] Step 6, generation of game strategy decisions: Use the obtained hybrid Nash equilibrium vector to generate the final game strategy decision plan.

[0017] Furthermore, the initialization of the DQN network model in Step 1 is to assign values to the following parameters:

[0018] Configure the state space dimension parameter observation_size to cover the state, action, and context information required for game environment interaction;

[0019] Set the number of nodes in the policy output layer, strategy_count, which strictly corresponds to the total number of optional strategies of the game participants;

[0020] Select a combination of loss function criterion and optimizer optimizer, supporting gradient descent variants such as Adam or SGD;

[0021] Define the learning rate lr, batch size batch_size, and reward discount factor reward_decay_ratio, which control the parameter update step size, training efficiency, and long-term reward weight respectively;

[0022] Set the capacity of the experience replay buffer max_experience_size and the update interval update_experience_interval to ensure the diversity and timeliness of training data;

[0023] Introduce action selection randomness through the pos_action_prob parameter to balance policy exploration and exploitation.

[0024] Furthermore, the information set of the game participants in step 2 includes state observations and the opponent's historical action sequence; the policy set supports discrete policy labels or vectorized representations of continuous policy spaces, adapting to the requirements of different game scenarios; the policy set is extended through a dynamic update mechanism, supporting online learning of new policies or elimination of inefficient policies.

[0025] Furthermore, the construction method of the payoff matrix in step 3 is as follows:

[0026] Define the payoff matrix

[0027] where N is the number of participants, and Ai is the size of the strategy space of the i-th participant;

[0028] Obtain the immediate reward ri of the strategy combination (a1,…,aN) through Monte Carlo sampling;

[0029] Use the Q-network to predict the long-term reward Qi(s,a1,…,aN), and calculate the comprehensive reward M[i,a1,…,aN] = αri + (1 - α)Qi, where α ∈ [0,1] is the reward fusion coefficient.

[0030] Furthermore, the formula of the DQN network model in step 1 is:

[0031] Q(s,a;θ) = f θ (s) a

[0032] where f θFor a deep neural network, the input is the state s, and the output layer is the Q-value vector for all policies a, f θ (s) a represents the Q-value corresponding to policy a.

[0033] Furthermore, the iteration of the DQN network model in step 4 is specifically as follows:

[0034] 1. Experience data sampling: Extract a batch of training data from the experience replay buffer according to the priority, where the priority is determined by the absolute value of the prediction error of each state transition sample, and the samples with larger errors have a higher probability of being selected;

[0035] 2. Target Q-value calculation: For each state transition data sampled, if the next state is a game termination state, the target Q-value directly takes the current immediate reward; if it is a non-termination state, calculate the maximum Q-value of all possible policies for the next state through the target network, and add it to the current reward to obtain the target Q-value;

[0036] 3. Policy network parameter update: Calculate the mean square error between the predicted Q-value of the policy network and the target Q-value, and at the same time introduce the regret value between the current policy and the Nash equilibrium policy as a constraint term, and sum the two weighted as the loss function, and update the policy network parameters through the gradient descent algorithm;

[0037] 4. Target network parameter synchronization: After each completion of the set number of updates of the policy network, copy the parameters of the policy network completely to the target network to ensure the stability of the target Q-value calculation;

[0038] 5. Exploration rate dynamic adjustment: Set a higher exploration rate at the beginning of training to fully explore the policy space randomly, and decay the exploration rate exponentially as the number of iterations increases, gradually transitioning to the exploitation stage mainly based on the maximum Q-value policy.

[0039] Furthermore, the specific backpropagation algorithm for step 4 is as follows:

[0040] 1. Calculate the target Q-value:

[0041]

[0042] where θ- is the target network parameter, γ is the reward discount factor, r is the immediate reward obtained after executing action a in the current state s, and y is the Q-value to be obtained.

[0043] 4. Calculate the loss function

[0044]

[0045] where B is the batch data sampled from the experience replay buffer;

[0046] 5. Perform gradient descent:

[0047]

[0048] where β is the learning rate, is the gradient of the loss function with respect to the network parameter θ, which is calculated by the backpropagation algorithm.

[0049] The features of the present invention also lie in that,

[0050] Step 1, initialize through a deep neural network, and this model aims to accurately estimate the Q value during the game process. This enables the system to make more accurate decisions based on the learned experience.

[0051] Step 2, provide the information of the game participants and the individual strategy sets, which provides diverse input data for the algorithm and enables the system to more comprehensively understand the decision-making environment.

[0052] Step 3, by establishing a payoff matrix, the system can comprehensively understand the payoff situation of each participant under different strategy combinations, providing an effective feedback signal for the DQN algorithm.

[0053] Step 4, through the iterative learning of the DQN network model, the system can continuously optimize the parameters and approach the optimal solution of the mixed Nash equilibrium during the game process. This ensures that the learning process of the system is more refined and efficient.

[0054] Step 5, extract the mixed Nash equilibrium vector from the trained DQN network, and this vector represents the optimal strategy combination of the participants in the game. This enables the system to understand and cope with complex game scenarios involving multiple parties.

[0055] Step 6, use the obtained mixed Nash equilibrium vector to generate the final game strategy decision plan. This ensures that the system can make more intelligent and rational decisions in the game to maximize the overall payoff.

[0056] The beneficial effects of the present invention are as follows: The beneficial effects of this invention are reflected at multiple levels. First, by adopting an intelligent game decision-making method based on the DQN algorithm, the decision-making system can continuously adapt to changing situations, thereby improving the flexibility of decision-making. This means that in business negotiations, the system can better handle changes in transaction conditions; in resource allocation, the system can more effectively adjust resource allocation strategies to meet different needs; in network security, the system can respond in a timely manner to prevent new threats. Second, the introduction of deep reinforcement learning enables the decision-making system to learn from real-time environmental feedback and continuously optimize decision-making strategies. This self-learning feature makes the system more intelligent and can continuously improve its performance level in specific fields. In a business environment, this means more accurate market forecasts and more optimized business strategies; in resource management, the system can more effectively adjust strategies to improve resource utilization; in network security, the system can learn and adapt to new threats and enhance the defense level. Finally, this innovation provides unprecedented opportunities for the decision-making system to better cope with the challenges of modern social dynamics and uncertainties. Whether facing market fluctuations, resource limitations, or cyberattacks, the system can make decisions in a more intelligent and predictive manner, bringing greater confidence and the possibility of success to practitioners and decision-makers in various fields.

[0057] The differences between the present invention and traditional methods are as follows:

[0058] 1. Differences in the overall process

[0059] In the model initialization stage, traditional methods rely on predefined rules or simple Q-table structures and lack the ability to model complex state spaces. In contrast, the method of the present invention initializes through a deep Q-network (DQN) to construct a neural network architecture including an input layer, a hidden layer, and an output layer, which can effectively process high-dimensional game states. In the input data processing link, traditional methods adopt a static policy set with fixed and non-expandable policy selection. The method of the present invention supports the real-time update of the dynamic policy set. For example, when new game participants or strategies are added, the system can adaptively adjust the policy library, significantly enhancing flexibility.

[0060] In the construction of the payoff matrix, traditional methods require manual definition, relying on expert experience and being difficult to adjust dynamically. In contrast, the method of the present invention automatically constructs and optimizes the payoff matrix through real-time game data (such as environmental feedback), making the revenue calculation more in line with the actual scenario. During the strategy optimization process, traditional methods are based on greedy algorithms or random searches, with slow convergence speeds and a tendency to fall into local optima. The method of the present invention introduces the experience replay mechanism and target network separation technology of DQN, accelerating convergence and approaching the mixed Nash equilibrium through the efficient reuse of historical experience and stable Q-value estimation. In addition, traditional methods only support the extraction of pure strategy equilibrium solutions and cannot handle mixed strategy scenarios, while the method of the present invention directly extracts the mixed strategy probability distribution from the DQN output to generate a mixed Nash equilibrium vector, providing theoretical support for complex games. In the final decision-making generation stage, traditional methods output fixed strategies and are difficult to adapt to environmental changes, while the method of the present invention dynamically adjusts strategies based on the mixed Nash equilibrium vector, responding in real time to changes in the game environment and achieving intelligent decision-making.

[0061] 2. Differences in key technical details

[0062] In terms of the model architecture, traditional methods lack neural network support and only rely on rules or Q-tables, making it difficult to handle high-dimensional state spaces. The method of the present invention uses a deep Q-network (DQN), capturing complex game relationships through multi-layer non-linear mappings and combining an experience replay buffer to break data correlations, improving training stability. The introduction of the target network further reduces the fluctuations in Q-value estimation, solving the problem of non-convergence in training caused by direct overwrite updates in traditional methods.

[0063] At the level of strategy optimization, traditional methods use a fixed ε-greedy strategy, with low exploration efficiency and only focusing on immediate rewards while ignoring long-term benefits. In contrast, the method of the present invention balances exploration and exploitation through an adaptive exploration rate (such as decaying ε) and introduces a discount factor (γ) to optimize long-term cumulative rewards, making the strategy more focused on the global optimum. In addition, traditional methods rely on mathematical derivations (such as linear programming) to solve equilibria, with high computational complexity, while the method of the present invention uses DQN to directly approximate the mixed Nash equilibrium, significantly reducing the computational cost.

[0064] In terms of dynamic adaptability, traditional methods require manual intervention to cope with environmental changes, with obvious response delays and difficulty in handling multi-party games and unknown states. The method of the present invention automatically adjusts strategies by learning environmental feedback in real time, using the generalization ability of deep networks to model multi-party strategy interactions and effectively coping with complex game scenarios and uncertainties. For example, in a resource allocation scenario, the fixed priority allocation of traditional methods leads to low resource utilization, while the method of the present invention dynamically optimizes the allocation strategy, increasing the utilization rate to over 90% and shortening the response time to the minute level.

[0065] The intelligent game decision - making analysis method based on the DQN algorithm combines the ideas of deep learning and game theory. By integrating the strategic choices of game players and the analysis of decision - maker payoff values, it seeks the optimal Nash equilibrium solution throughout the iterative process of the DQN algorithm. This method not only considers the maximization of individual decision - making benefits but also fully utilizes the concepts of game theory to ensure reasonable decision - making coordination in an interactive environment. Through in - depth analysis of input information, this decision - making framework enables decision - makers to more effectively formulate optimal decision - making strategies in dynamic games, providing an innovative solution that combines deep learning and game theory for intelligent game problems in complex environments.

[0066] In summary, the limitations of traditional methods are reflected in static rules, inefficient calculations, and manual dependence. In contrast, the method of the present invention realizes dynamic policy optimization, efficient calculation, and end - to - end automation through the integration of deep reinforcement learning and game theory. Its core innovations are as follows:

[0067] 1. Deep network architecture: Supports complex state modeling and extraction of mixed - strategy equilibria;

[0068] 2. Adaptive learning mechanism: Improves training efficiency through experience replay, target networks, and dynamic exploration rates;

[0069] 3. Real - time response ability: Generates dynamic decisions based on mixed Nash equilibria to adapt to changing environments.

[0070] This technological breakthrough provides an efficient, flexible, and scalable solution for intelligent decision - making in fields such as business negotiation, resource allocation, and network security. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 is a flowchart of a dynamic strategy real - time optimization method based on DQN and mixed Nash equilibrium of the present invention;

[0072] Figure 2 is a system architecture diagram of a dynamic strategy real - time optimization method based on DQN and mixed Nash equilibrium of the present invention;

[0073] Figure 3 is an interface diagram of a dynamic strategy real - time optimization method based on DQN and mixed Nash equilibrium of the present invention.

[0074] Figure 4 is a diagram comparing the efficiency of a dynamic strategy real - time optimization method based on DQN and mixed Nash equilibrium of the present invention with traditional methods. DETAILED DESCRIPTION OF THE INVENTION

[0075] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0076] A specific embodiment of the application of the present invention can be the optimal solution of the strategy that each department of an enterprise company can take when facing certain abnormal index problems.

[0077] Figure 1 It is a flowchart of a dynamic strategy real-time optimization method based on DQN and mixed Nash equilibrium of the present invention; the following steps are as shown in the figure.

[0078] Figure 2 It is a system architecture diagram of a dynamic strategy real-time optimization method based on DQN and mixed Nash equilibrium of the present invention; the required parameters and network initialization are as shown in the figure.

[0079] Figure 3 It is an interface diagram of a dynamic strategy real-time optimization method based on DQN and mixed Nash equilibrium of the present invention. The interface shown is the operation interface for network initialization and the start of the game.

[0080] Figure 4 It is a comparison chart of the efficiency of a dynamic strategy real-time optimization method based on DQN and mixed Nash equilibrium of the present invention and traditional methods. It shows the advantages of this method and traditional methods in game decision-making.

[0081] The formula for solving the loss function in the present invention is: L = (r + γmaxQ(s', a'; θ') - Q(s, a; θ))^2, where the meanings and functions of each variable are as follows:

[0082] L represents the loss function. In machine learning and reinforcement learning, the loss function is used to measure the difference between the predicted value and the actual value of the model. In DQN (Deep Q-Network), the loss function guides the training of the neural network to enable it to more accurately predict the Q value.

[0083] r represents the immediate reward. In reinforcement learning, it is the reward obtained by the agent from the environment after executing a certain action. This reward is immediate and reflects the quality of the current action.

[0084] γ is the discount factor. It is a constant between 0 and 1, used to weigh the importance of immediate rewards and future rewards. A larger γ value means that future rewards are more important when calculating the total return, while a smaller γ value places more emphasis on the current immediate rewards.

[0085] The part maxQ(s', a'; θ') represents the maximum predicted Q-value among all possible actions a' in the next state s'. Q(s′, a′; θ′) is the Q-value calculated using the Q-network with parameters θ′. Here, θ′ usually represents the parameters of the target network, which are updated more slowly during training to maintain stability. maxa′ means taking the maximum of the Q-values for all possible actions a'.

[0086] The part Q(s, a; θ) represents the predicted Q-value when taking action a in the current state s. Q(s, a; θ) is the Q-value calculated using the Q-network with parameters θ. Here, θ represents the parameters of the current value network, which are continuously updated during training to minimize the loss function.

[0087] θ and θ′ represent the parameters of the current value network and the target network respectively. In DQN, two networks are usually used to represent the current value function and the target value function respectively. The current value network is used to calculate the Q-value under the current state and action, while the target value network is used to calculate the Q-value under the next state and the optimal action (i.e., the target Q-value). The target network is usually updated more slowly than the current network, which helps to stabilize the training process. Step 1: When initializing the deep Q-network to solve the reinforcement learning problem, the following parameters are assigned: (1) ˋobservation_sizeˋ: This is the dimension of the observable data of the game participant, usually including all relevant information about interacting with the game environment, such as states, actions, etc. (2) ˋstrategy_countˋ: Represents the total number of strategies available to the game participant, which determines the number of nodes in the output layer of the model. (3) ˋcriterionˋ: The loss function, which defines the performance metric of the model and is the objective that the model needs to minimize during training. (4) ˋoptimizerˋ: The optimizer is responsible for adjusting the weights of the model to minimize the loss. It may be a variant of gradient descent, such as Adam, SGD, etc. (5) ˋdeviceˋ: Determines whether the model runs on devices such as CPU or GPU. (6) ˋlrˋ (learning rate): The learning rate controls the size of the adjustment of the model parameters at each update and is a key parameter in the training process. (7) ˋbatch_sizeˋ: The number of samples used for each update of the model parameters, which is the core concept of batch gradient descent. (8) ˋpos_action_probˋ: This parameter affects whether to follow the gradient to perform parameter gradient descent, introducing a certain degree of randomness. (9) ˋreward_decay_ratioˋ: The discount factor in reinforcement learning, which is used to determine the weight of future rewards and affects the model's attention to long-term rewards. (10) ˋmax_experience_sizeˋ: The upper limit of the capacity of the experience replay buffer, which is used to store previous observations and actions for more effective training of the model. (11) ˋupdate_experience_intervalˋ: The time interval for periodically updating the experience buffer to ensure that the model uses the latest experience for training. (12) ˋupdate_model_intervalˋ: Controls the frequency of model weight updates to balance training efficiency and stability. The purpose of this initialization function is to provide a deep Q-network with well-configured parameters so that it can effectively learn and perform well in a given reinforcement learning task;

[0088] Step 2: To provide the basic data for the algorithm, we need to define the information sets of the game players and their respective strategy sets, that is, the departments affected by the abnormal indicators and the response strategies of each department to this abnormal indicator information. Specifically, the information set of the game players can be expressed as {Player 1, Player 2,...}, and each player has an independent strategy set, expressed as {Strategy 1, Strategy 2,...}, and these strategies can form the decision-making library of the enterprise. In this context, the information set of the player may include information about its state observations, the historical action sequences of opponents, and / or other environmental variables. For example, for the players in the game, the information set may include the state of the current situation, the actions of the opponents, etc. This provides the environmental context required by the algorithm in learning. The strategy set represents the alternative action plans available to each player. These strategies may be discrete actions or continuous action spaces. The algorithm will select actions from these strategies to maximize the cumulative reward or achieve a specific goal. Generally speaking, to improve the input of the algorithm, it is necessary to clearly define the information set and strategy set of each player so that the deep Q-network can effectively process the observation data and action selection in the game;

[0089] Step 3: In the dynamic strategy real-time optimization method based on DQN and mixed Nash equilibrium, constructing the payoff matrix of the game process is a crucial step. Specifically, we define a payoff matrix M with the dimension of R^(N×A1×…×AN), where N represents the number of departments participating in the game, and Ai represents the size of the strategy space that the i-th department can choose. To fill this matrix, we take the following steps:

[0090] First, using the Monte Carlo sampling technique, we obtain the immediate reward ri obtained by performing actions under a specific strategy combination (a1,…,aN). These immediate rewards reflect the direct benefits that each department can obtain by adopting a specific strategy combination in the current state.

[0091] Second, we use the trained Q-network to predict the long-term reward Qi(s,a1,…,aN). The Q-network can learn the future rewards that can be obtained by adopting different strategy combinations in different states through deep learning.

[0092] Finally, we combine the immediate reward and the long-term reward and fill the payoff matrix by calculating the comprehensive reward M[i,a1,…,aN]=αri+(1-α)Qi. Among them, α, as the reward fusion coefficient, whose value is in the range of [0,1], is used to balance the proportion of the immediate reward and the long-term reward in the comprehensive reward.

[0093] Through this method, we successfully established the payoff matrix of the game process. This matrix covers all possible strategy combinations, such as (Player 1 - Strategy 1, Player 2 - Strategy 1), (Player 1 - Strategy 2, Player 2 - Strategy 1), (Player 1 - Strategy 1, Player 2 - Strategy 2), and (Player 1 - Strategy 2, Player 2 - Strategy 2), etc. The comprehensive establishment of the payoff matrix enables the system to clearly grasp the revenue situation of each participating department under different strategy combinations, thus providing valuable feedback signals for the DQN algorithm. These feedback signals will further guide the learning process of the DQN algorithm, helping it continuously approach the optimal solution of the mixed Nash equilibrium of the game.

[0094] Step 4: Use the DQN network model for iteration, and continuously adjust the network parameters through the backpropagation algorithm to approach the optimal solution of the mixed Nash equilibrium of the game process. In this process, the DQN network structure is initialized, and then a random exploration strategy is executed in the complex game environment. This step is crucial for accumulating a diverse dataset covering states (s), actions (a), immediate rewards (r), and subsequent states (s'). These valuable data are carefully stored in the experience replay buffer, laying a solid foundation for subsequent efficient training. During the training process, the core of DQN is to continuously optimize the network parameters by calculating the error between the predicted Q-value and the target Q-value. The target Q-value is defined as:

[0095]

[0096] where θ - represents the parameters of the target network, and γ is the reward discount factor, which ensures a balanced consideration between long-term and immediate rewards. The predicted Q-value is directly determined by the current network parameters θ, reflecting the model's expectation of the optimal action selection for a given state.

[0097] To minimize the prediction error, we introduce the loss function:

[0098]

[0099] where B is a batch of data randomly sampled from the experience replay buffer. Through the backpropagation algorithm, we calculate the gradient of the loss function with respect to the network parameters θ and perform a gradient descent step accordingly to optimize the network parameters:

[0100]

[0101] where β is the learning rate, which controls the step size of parameter update.

[0102] As the iterative training progresses, the DQN model gradually learns to select the optimal action in a given state. This process not only significantly improves the model's prediction accuracy but also enables it to gradually approach the mixed Nash equilibrium of the game. The mixed Nash equilibrium, a core concept in game theory, describes the stable state when all players adopt the probability distribution of optimal strategies. At this time, no player can increase their own payoff by unilaterally changing their strategy.

[0103] Taking the two - department game as an example, the mixed Nash equilibrium can be reflected by calculating the expected payoffs of each department under different strategy combinations, as shown in the formula (assuming a simplified case):

[0104] E = q·E(B1(p))+(1 - q)·E(B2(p))

[0105] Where q is the probability that department 1 adopts a certain strategy, and E(B1(p)) and E(B2(p)) represent the expected payoffs of department 1 and department 2 respectively under the given strategy probability distribution p.

[0106] Throughout the training process, the DQN model continuously samples experiences from the game environment and efficiently trains through experience replay. This strategy not only accelerates the model's convergence speed but also ensures that the model can comprehensively explore and utilize the state space. Finally, through continuous iteration and optimization, the DQN model can accurately learn the complex strategy interactions in the game environment and successfully reach the optimal solution of the mixed Nash equilibrium, providing strong support for game theory research and practice.

[0107] Step 5: Through the trained DQN network, the mixed Nash equilibrium vector can be extracted. The DQN learns through repeated trial and error, continuously optimizing its action selection strategy in various states. When the training reaches a stable state, the strategy distribution shown by the DQN essentially reflects the mixed Nash equilibrium of the game, that is, each participant adopts an optimal probability distribution to select actions to maximize its own benefits while considering the possible strategies of other participants. Therefore, by analyzing the output probabilities of the DQN in different states, we can construct a strategy vector representing the mixed Nash equilibrium, providing strong support for game theory research and practice. This vector reflects the optimal strategy combination of each department in the game. The extraction of this vector allows the system to better understand and handle complex game scenarios involving multiple parties. After training, the parameters of the DQN network are adjusted to enable it to estimate the values of various strategies selected by game participants in different states. The mixed Nash equilibrium vector is derived from these learned strategy values, representing the probability distribution of the most favorable strategy choices for participants in the game. This vector provides strategic guidance for the system when facing multi-party interactions. By understanding the mixed Nash equilibrium, the system can predict the strategies of other players and accordingly adjust its own strategies to maximize its expected benefits. This enables the system to make more intelligent and adaptable decisions in the game environment, enhancing its ability to handle complex game scenarios;

[0108] Step 6: Through the obtained mixed Nash equilibrium vector, the system can analyze the optimal strategy probability distribution that each participant should adopt in the game. This vector not only reveals the importance of each strategy but also ensures the optimal response of participants in the face of the uncertainty of their opponents' strategies. The system then intelligently generates a final game strategy decision plan based on these probability distributions, thus ensuring that more intelligent and rational decisions can be made in the game to maximize the overall benefits of the enterprise. The mixed Nash equilibrium vector reflects the probability distribution of the optimal strategy combination of each department in the game. Transforming this vector into a final decision plan involves mapping the probability distribution to specific actions or strategies. This can be achieved by sampling the probability distribution or selecting the strategy with the highest probability. The process of generating the final game strategy decision plan ensures that the system makes wise decisions in complex game scenarios. The system can predict the strategies of other players based on the mixed Nash equilibrium vector and then adopt the optimal response strategy to maximize the overall benefits.

Claims

1. A dynamic strategy real-time optimization method based on DQN and hybrid Nash equilibrium, characterized in that: The specific steps include: Step 1, DQN network model initialization: Initialize the deep neural network to estimate the Q value in the game process; Step 2: Input game participants and strategy sets: Provide game participant information and their respective strategy sets to provide basic data for DQN network model algorithm input; Step 3: Construct a benefit matrix: Construct a benefit matrix for the game process to reflect the benefits of each participant under different strategy combinations; Step 4, iterative solution of DQN network model: using DQN network model iteration, the DQN network model parameters are continuously adjusted through the back propagation algorithm to approach the optimal solution of the mixed Nash equilibrium of the game process; Step 5, hybrid Nash equilibrium vector acquisition: extract the hybrid Nash equilibrium vector from the trained DQN network model to represent the optimal strategy combination of the participants in the game; Step 6, game strategy decision generation: Use the obtained mixed Nash equilibrium vector to generate the final game strategy decision plan.

2. The dynamic strategy real-time optimization method based on DQN and hybrid Nash equilibrium according to claim 1 is characterized in that: The initialization of the DQN network model in step 1 is to assign values ​​to the following parameters: Configure the state space dimension parameter observation_size to cover the state, action, and context information required for game environment interaction; Set the number of strategy output layer nodes strategy_count, which strictly corresponds to the total number of optional strategies available to game participants; Select a combination of loss function criterion and optimizer optimizer, supporting gradient descent variants such as Adam or SGD; Define the learning rate lr, batch size batch_size and reward discount factor reward_decay_ratio to control the parameter update step size, training efficiency and long-term benefit weight respectively; Set the experience replay buffer capacity max_experience_size and update interval update_experience_interval to ensure the diversity and timeliness of training data; The pos_action_prob parameter is used to introduce randomness in action selection, balancing strategy exploration and exploitation.

3. The dynamic strategy real-time optimization method based on DQN and hybrid Nash equilibrium according to claim 1 is characterized in that: The game participant information set in step 2 includes state observation values ​​and opponent historical action sequences; the strategy set supports discrete strategy labels or continuous strategy space vector representation to adapt to the requirements of different game scenarios; the strategy set is expanded through a dynamic update mechanism to support online learning of new strategies or elimination of inefficient strategies.

4. The dynamic strategy real-time optimization method based on DQN and hybrid Nash equilibrium according to claim 1 is characterized in that: The method for constructing the benefit matrix in step 3 is: Defining the Benefit Matrix Where N is the number of participants, Ai is the size of the strategy space of the i-th participant; Obtain the immediate reward ri of the strategy combination (a1,…,aN) through Monte Carlo sampling; Use the Q network to predict the long-term return Qi(s,a1,…,aN) and calculate the comprehensive return M[i,a1,…,aN]=αri+(1-α)Qi, where α∈[0,1] is the return fusion coefficient.

5. The dynamic strategy real-time optimization method based on DQN and hybrid Nash equilibrium according to claim 1 is characterized in that: The DQN network model formula in step 1 is: Q(s,a;θ)=f θ (s) a Among them, f θ is a deep neural network, the input is state s, and the output layer is the Q value vector of all strategies a, f θ (s) a Indicates the Q value corresponding to strategy a.

6. The dynamic strategy real-time optimization method based on DQN and hybrid Nash equilibrium according to claim 1 is characterized in that: The DQN network model iteration in step 4 is specifically as follows: (1) Experience data sampling: training data batches are extracted from the experience replay buffer according to priority, where the priority is determined by the absolute value of the prediction error of each state transition sample. The sample with a larger error has a higher probability of being selected. (2) Target Q value calculation: For each state transition data sampled, if the next state is the game termination state, the target Q value is directly taken as the current instant reward; If it is not a terminal state, the maximum Q value of all possible strategies for the next state is calculated through the target network, and added to the current reward to obtain the target Q value; (3) Policy network parameter update: Calculate the mean square error between the policy network's predicted Q value and the target Q value. At the same time, introduce the regret value of the current policy and the Nash equilibrium policy as constraints. Take the weighted sum of the two as the loss function and update the policy network parameters using the gradient descent algorithm. (4) Target network parameter synchronization: After completing the set number of policy network updates, the policy network parameters are completely copied to the target network to ensure the stability of the target Q value calculation; (5) Dynamic adjustment of exploration rate: In the early stage of training, a higher exploration rate is set to fully and randomly explore the strategy space. As the number of iterations increases, the exploration rate decays exponentially, gradually transitioning to a stage where the maximum Q-value strategy is used as the main strategy.

7. The dynamic strategy real-time optimization method based on DQN and hybrid Nash equilibrium according to claim 1 is characterized in that: The back propagation algorithm in step 4 is specifically as follows: (1) Calculate the target Q value: Among them, θ- is the target network parameter, γ is the reward discount factor, r is the immediate reward obtained after performing action a in the current state s, and y is the desired Q value; (2) Calculate the loss function Where B is the batch data sampled from the experience replay buffer; (3) Perform gradient descent: Among them, β is the learning rate, is the gradient of the loss function with respect to the network parameters θ, calculated by the back-propagation algorithm.

Citation Information

Cited By

  • Front-end large file uploading method and device

    CN120416239A

  • Load balancing method based on reinforcement learning and game theory

    CN121743064A

  • Man-machine dynamic game control method driven by double reinforcement learning

    CN122035054A

  • Satellite orbit pursuit game control method and device, electronic equipment and storage medium

    CN122411227A