Online generation method and system of maneuvering game countermeasure
By establishing a game-adversarial situational equation system for discrete differential equations and using differential game countermeasure theory, combining game subprocess generation and open-loop control strategies, the problem that game countermeasure generation in the existing technology is difficult to adapt to complex time-varying conditions, and a game strategy of rapid online generation and real-time update is realized.
Patent Information
- Application Number
- CN202411914852.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-13
AI Technical Summary
The existing game confrontation countermeasure generation method is difficult to achieve rapid online generation under complex time-varying conditions, consumes a lot of computing power and is difficult to adapt to uncertain time-varying environments, affecting the effectiveness of the strategy.
By establishing a game-confrontation situational equation system in the form of discrete differential equations, using the differential game countermeasure theory to obtain feedback control forms, and through the game subprocess generation and open-loop control strategy establishment, the game strategy is updated online.
It realizes the rapid online generation of game strategies in an uncertain time-varying environment, improves the real-time and speed of decision-making, reduces calculation parameters, and enhances application value.
Smart Images

Figure CN119989626A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of confrontation technology, and in particular to an online generation method and system for maneuverable game strategies. Background Art
[0002] Maneuver game strategy refers to the motion control strategy decided by the participating parties within the limits of their own mobility in a certain confrontational environment in order to gain situational advantage in the confrontational environment.
[0003] Existing game countermeasure generation can be divided into traditional methods and intelligent methods. Traditional methods use expert knowledge deduction, performance simulation, and simulation deduction to achieve the process of flexible decision-making, focusing on prior knowledge such as industry knowledge deduction and historical game information accumulation. At the strategy generation level, traditional methods use deduction methods such as logic gates and flow charts to give strategies based on fixed feedback conditions. However, traditional methods often do not have the ability to optimize autonomously, the final generation results are relatively fixed, and the real-time adjustment ability is poor. It is easy for the other party to predict and reverse the game results, resulting in game failure.
[0004] Intelligent methods are represented by machine learning, fuzzy algorithms, reinforcement learning, and heuristic algorithms. They are characterized by having certain self-learning and self-optimization capabilities, and can realize the autonomous generation of mobile countermeasures according to the actual confrontation situation. For example, reinforcement learning sets the situation space, experience pool, and reward function of the intelligent agent in the environment, inputs information into the reinforcement learning network, and continuously updates the network parameters in multiple rounds of training until the number of iterations is reached to output the generated countermeasures. Although intelligent methods have their specific advantages, the strategies generated in a single training environment are difficult to adapt to the uncertainty and time-varying nature of the actual use environment, have poor migration capabilities to new environments, and are prone to problems with the difficulty of converging reward functions during the training process. How to achieve rapid online generation of robust, autonomous, and real-time game strategies has become a major technical difficulty.
[0005] Patent document CN117828286A discloses a multi-agent adversarial decision-making method and device based on deep reinforcement learning. The decision-making method includes: obtaining a preset task goal; inputting the preset task goal into a pre-trained multi-agent adversarial decision-making reinforcement learning model to obtain the corresponding optimal strategy; the multi-agent adversarial decision-making reinforcement learning model is trained through the following steps: based on the micro-manipulation environment of the real-time strategy game, using the multi-objective Markov decision process and the observation characteristics of the agent corresponding to each role in the real-time strategy game, constructing a sample space, and dividing the sample space into multiple sub-experience pools; calculating the importance of each sub-experience pool respectively; based on the importance of each sub-experience pool, using the experience replay mechanism to train the multi-agent adversarial decision-making reinforcement learning model.
[0006] However, when existing methods are applied under complex time-varying conditions, they will consume a lot of computing power in the online learning and strategy updating stages, and take a long time to train. They are difficult to apply to uncertain time-varying environments with timeliness requirements, which affects the effectiveness of the strategy. Summary of the invention
[0007] In view of the defects in the prior art, the purpose of the present invention is to provide an online generation method and system for mobile game strategies.
[0008] According to the present invention, an online generation method of a maneuverable game strategy includes:
[0009] Step S1: Based on the perceived game information, a game confrontation situation equation group represented by a mathematical model in the form of discrete differential equations is established;
[0010] Step S2: Based on the game confrontation situation equations, according to the differential game strategy theory, a feedback control form of the maneuvering game is obtained;
[0011] Step S3: construct a game sub-process generation method, divide the game sub-process, discretize the total process into N short game sub-processes, and obtain situation information before each sub-process starts;
[0012] Step S4: Establishment of open-loop control strategy. According to the game sub-process determined in step S3, in the period when information is not updated, the feedback control constraints are relaxed and the open-loop strategy solution is generated. Then, in each sub-process, the input is adjusted in real time according to the situation change to realize the online update of the game strategy.
[0013] Preferably, the motion parameters, control parameters and control coefficients at discrete time points are used to construct a game confrontation situation equation group;
[0014] The terminal fixed conditions of the game process state include: the start time of the game is t0, the end time is t f =NΔt, where N is used to control the total duration of the game; x0 is the initial condition of the game situation, is the initial value of the control parameter of one party in the mobile game, and N represents the number of short game sub-processes.
[0015] Preferably, step S1 includes: establishing a discrete differential equation form mathematical model of the state transfer equation to obtain the state transfer equation of the motion parameter under the discrete state, as shown in the following formula:
[0016]
[0017] Among them, x t represents the motion state of the maneuvering party in the game, represents the control parameter of a party in the mobile game, n represents the total number of participants in the mobile game, t represents the current time, B tThe coefficient representing the control quantity;
[0018] Initial conditions x0 and u0 of the motion state i is a fixed value, which determines the decision step length Δt of the discrete game.
[0019] Preferably, step S2 comprises:
[0020] Step S2.1: Define the payment function of each party in the game, the payment function J i , as follows:
[0021]
[0022] in, represents the weight coefficient of the state quantity, Represents the weighted coefficient of the control quantity, x t represents the motion state of the maneuvering party in the game, Represents the control parameter of one party in the maneuver game.
[0023] Step S2.2: Calculate the feedback control strategy of the action game through the quadratic payment function form.
[0024] Preferably, the step S2.2 comprises:
[0025] Step S2.2.1: Based on the Pontryagin minimization principle, for each participant at each time step, construct the Hamiltonian function
[0026]
[0027] in, represents the coefficient function;
[0028] Step S2.2.2: In the feedback control form, the motion constraint equations are strictly followed. Substituting the game confrontation situation equations into the Hamiltonian function, we get:
[0029]
[0030] The coefficient function in the payment function is and is a positive definite matrix;
[0031] Step S2.2.3: Based on the derivative value of the payoff function, let the derivative value be 0, and assume that the state quantity and the co-state quantity in the game equation remain fixed, and calculate:
[0032]
[0033] The optimal control result of each step in the feedback control form is obtained through the above formula.
[0034] Preferably, step S3 comprises:
[0035] Step S3.1: According to the feedback control formula obtained in step S2, set the rolling time window length t h =sΔt, where s represents the number of discrete steps corresponding to the subprocess length. Let t h The length is much smaller than t f And s is in a time suitable for the execution of game control, so the time length of the generated game sub-process is t h ;
[0036] Step S3.2: At the beginning of each sub-process, update the real state parameter information and reset the initial condition x of the maneuvering game. k0 and
[0037] Step S3.3: Based on the process of step S2, the maneuvering game strategy under the feedback control form within the sub-process is obtained:
[0038]
[0039] in, Represents a discrete time point within a subprocess.
[0040] Preferably, at the kth decision time step, the total duration of game planning is
[0041] Preferably, the step S4 comprises:
[0042] Step S4.1: Convert the feedback control form into open-loop control, relax the discrete state transition form, and rewrite the Hamiltonian function:
[0043]
[0044] Taking the derivative of the open-loop Hamiltonian function, the new open-loop control result is expressed as:
[0045]
[0046] in, Satisfies the relationship:
[0047]
[0048] Step S4.2: At t=t k At the moment, the final state terminal condition of the sub-process is used to calculate the rolling time window period. The control strategy for each step in is generated online according to the control strategy Execute current motion control decisions;
[0049] Step S4.3: Advance one time step backward and update the motion state information x0 and the weighted coefficient of the game and The new countermeasure time period is set as According to the new initial quantity and parameter value, the game strategy generation plan is repeatedly executed, and the step S4.1 is returned to generate the game open-loop strategy plan for the next sub-process to realize the real-time update of the maneuver strategy.
[0050] Preferably, in each sub-process, the input is adjusted in real time according to the situation changes to realize online updating of the game strategy, including: taking the fixed initial information of the sub-process as input, obtaining the open-loop prediction form output within each cycle, thereby converting the infinite-length game into a rolling forward form, and advancing in real time based on the updated sub-process information, and quickly obtaining the results of the mobile game strategy in real time.
[0051] According to the present invention, an online generation system of a maneuverable game strategy includes:
[0052] Module M1: Based on the perceived game information, establish a game confrontation situation equation group represented by a mathematical model in the form of discrete differential equations;
[0053] Module M2: Based on the game confrontation situation equations, according to the differential game strategy theory, the feedback control form of the maneuvering game is obtained;
[0054] Module M3: Construct a method for generating game sub-processes, divide the game sub-processes, discretize the total process into N short game sub-processes, and obtain situation information before each sub-process starts;
[0055] Module M4: Establishment of open-loop control strategy. According to the game sub-process determined by module M3, in the period when information is not updated, the feedback control constraints are relaxed and the open-loop strategy solution is generated. Then, in each sub-process, the input is adjusted in real time according to the situation change to realize the online update of the game strategy.
[0056] Compared with the prior art, the present invention has the following beneficial effects:
[0057] 1. The present invention realizes the generation and update of differential game maneuver strategies through technical methods such as discrete game sub-processes, relaxed feedback control constraints, and open-loop strategy establishment, solving the timeliness problem of online generation of strategies in uncertain time-varying environments.
[0058] 2. The present invention adopts the differential game method to obtain the feedback control form of the game, and utilizes the sub-process generation division strategy to transform the traditional fixed-period differential game with a fixed game cycle into a rolling forward multi-process decision-making process, thereby ensuring the real-time online generation of decisions; the sub-process generation reduces the processing data of each game window and improves the decision-making speed of each step.
[0059] 3. The present invention transforms the game feedback strategy into an open-loop strategy, obtains the suboptimal solution of the strategy by relaxing the constraints, reduces the calculation parameters and feedback items, and achieves the rapid generation of mobile strategies with only a partial loss of accuracy, thereby improving the practical application value of the project. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0061] Figure 1 It is a schematic diagram of the workflow of the present invention;
[0062] Figure 2 A result graph of the preliminary strategy generation for the game strategy simulation in the present invention;
[0063] Figure 3 Generate a result graph for the game strategy simulation update strategy in the present invention;
[0064] Figure 4 Generate a result graph for the game strategy simulation ending strategy in the present invention;
[0065] Figure 5 It is a curve diagram of the relative distance between pursuit and escape in the present invention. DETAILED DESCRIPTION
[0066] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0067] The present invention senses the other party's information through sensors and generates a group of game confrontation state equations; discretizes the total process into N short game sub-processes, and obtains situation information before the start of each sub-process game; obtains the optimal payment control strategy under the game situation based on differential game theory; reduces the calculation parameters required in the strategy and improves the strategy generation efficiency by relaxing feedback control constraints and generating open-loop strategy solutions; adjusts the input in real time according to the situation changes in each sub-process to achieve online rapid update of the game strategy.
[0068] According to an online generation method of a mobile game strategy adopted by the present invention, Figure 1 As shown, the following steps are included:
[0069] Step S1: Based on the perceived game information, establish a game confrontation situation equation group represented by a mathematical model in the form of discrete differential equations, and use motion parameters, control parameters and control coefficients at discrete time points to construct. Determine the time step of the decision, clarify the feasible strategy space of the maneuver game, and confirm the start and end time of the game. The step S1 includes: establishing a mathematical model in the form of discrete differential equations of the state transfer equation, and obtaining the state transfer equation of the motion parameters under discrete states, as follows:
[0070]
[0071] Among them, x t represents the motion state of the maneuvering party in the game, represents the control parameter of a party in the mobile game, n represents the total number of participants in the mobile game, t represents the current time, B t Represents the coefficient of the control quantity.
[0072] Initial conditions x0 and u0 of the motion state i is a fixed value, which determines the decision step length Δt of the discrete game.
[0073] The terminal fixed conditions of the game process state include: the start time of the game is t0, the end time is t f =NΔt, where N is used to control the total duration of the game; x0 is the initial condition of the game situation, is the initial value of the control parameter of one party in the mobile game, and N represents the number of short game sub-processes.
[0074] Step S2: Based on the game confrontation situation equations, according to the differential game strategy theory, the feedback control form of the maneuvering game is obtained. The step S2 includes:
[0075] Step S2.1: Define the payoff function of each player. For each player in the maneuvering game, the quadratic payoff function is J i :
[0076]
[0077] in, represents the weight coefficient of the state quantity, Represents the weighted coefficient of the control quantity, x t represents the motion state of the maneuvering party in the game, Represents the control parameter of one party in the maneuvering game.
[0078] In order to ensure the continuity of the payment function, the payment function J is rewritten asi :
[0079]
[0080] Since the initial conditions of the game are fixed, the game payment value corresponding to time t=1 is also constant, which makes the two payment functions equivalent.
[0081] Step S2.2: Calculate the feedback control strategy of the action game in the form of the quadratic payment function. Step S2.2 includes:
[0082] Step S2.2.1: Based on the Pontryagin minimization principle, for each participant at each time step, construct the Hamiltonian function
[0083]
[0084] in, represents the coefficient function.
[0085] Step S2.2.2: In the feedback control form, the motion constraint equations are strictly followed. Substituting the game confrontation situation equations into the Hamiltonian function, we get:
[0086]
[0087] The coefficient function in the payment function is and is a positive definite matrix.
[0088] Step S2.2.3: Based on the derivative value of the payoff function, let the derivative value be 0, and assume that the state quantity and the co-state quantity in the game equation remain fixed, and calculate:
[0089]
[0090] The optimal control result of each step in the feedback control form is obtained through the above formula.
[0091] Step S3: Construct a method for generating game sub-processes. Divide the game sub-processes. To reduce the number of decision steps, the total process is discretized into N short game sub-processes according to the discrete decision step length. Before the sub-process starts, obtain the enemy's real-time information and select the appropriate segmented rolling time window length. In each segment time, fix the initial game position, speed and other information.
[0092] Consider the coefficient function in the Hamiltonian function In the form of:
[0093]
[0094] Consider terminal conditions In the form of:
[0095]
[0096] The above formula shows that the optimal form of the control strategy can be inferred forward from the final state.
[0097] The step S3 comprises:
[0098] Step S3.1: Determine the time length of the game sub-process. According to the feedback control formula obtained in step S2, set the rolling time window length t h =sΔt, where s represents the number of discrete steps corresponding to the subprocess length. Let t h The length is much smaller than t f And s is a time length suitable for the execution of game control. Therefore, the time length of the generated game sub-process is t h .
[0099] Step S3.2: Rewrite the initial conditions of the sub-process game. At the beginning of each sub-process, update the real state parameter information and reset the initial conditions x of the maneuvering game. k0 and At the kth decision time step, the total duration of game planning is
[0100] Step S3.3: Based on the process of step S2, the maneuvering game strategy under the feedback control form within the sub-process is obtained:
[0101]
[0102] in, Represents a discrete time point within a subprocess.
[0103] Step S4: Establishment of open-loop control strategy. According to the game sub-process determined in step S3, in the period when the information is not updated, the feedback control constraints are relaxed and an open-loop strategy is established; then, the fixed initial information of the sub-process is used as input to obtain the open-loop prediction output in each period, thereby converting the infinite-length game into a rolling forward form, and advancing in real time based on the updated sub-process information, and quickly and in real time obtaining the results of the maneuverable game strategy. The open-loop control strategy relaxes the feedback control constraints, improves the vector control parameter items in the feedback strategy, reduces the calculation parameters, improves the calculation speed, and continuously advances the strategy generation with a rolling strategy. The step S4 includes:
[0104] Step S4.1: Convert the feedback control form into open-loop control, relax the discrete state transition form, and rewrite the Hamiltonian function:
[0105]
[0106] Taking the derivative of the open-loop Hamiltonian function, the new open-loop control result is expressed as:
[0107]
[0108] in, Satisfies the relationship:
[0109]
[0110] Step S4.2: Execute the rolling strategy advancement process. According to the control strategy of step S4.1, at t=t k At the moment, the final state terminal condition of the sub-process is used to calculate the rolling time window period. The control strategy for each step in is generated online according to the control strategy Execute the current motion control decision.
[0111] Step S4.3: Advance one time step backward and update the motion state information x0 and the weighted coefficient of the game and The new countermeasure time period is set as According to the new initial quantity and parameter value, the game strategy generation plan is repeatedly executed, and the step S4.1 is returned to generate the game open-loop strategy plan for the next sub-process to realize the real-time update of the maneuver strategy.
[0112] The present invention aims to solve the problems of insufficient autonomous optimization capability of existing game methods, long time consumption of strategy generation, and weak real-time update and migration capabilities. The present invention utilizes discrete state representation, game subprocess generation, differential strategy and open-loop strategy design to perform online real-time generation of mobile game strategies, which is practical in uncertain time-varying environments.
[0113] Furthermore, combined with Figure 2 To Attachment Figure 5 The online generation method of the mobile game strategy of the present invention is specifically described as follows:
[0114] Step 1: Consider a maneuver strategy of pursuit between two parties. The discrete form of the state transfer equation can be written as:
[0115]
[0116] where x t is the overall situation of the game space maneuver, which indicates the relative distance between the two parties in the game. represents the speed of the pursuer, Indicates the speed of the fleeing party.
[0117] Initial conditions x0 and u0 of the motion state iis a fixed value, which determines the decision step length Δt of the discrete game. The game starts at t0 and ends at t f =NΔt.
[0118] Step 2: For both parties in the maneuverable confrontation, the quadratic payment function is:
[0119]
[0120] in Indicates the relative distance between the two parties in pursuit and escape. Represents the energy consumed by each, all expressed in square form.
[0121] The decision step length of the discrete game is set to Δt, and the Hamiltonian function of each time step is expressed as:
[0122]
[0123] Step 3: Taking the escaped party as an example, the length of the rolling time window is set to t h =sΔt.
[0124] At the start t0, the initial state is x0, The optimal control solution for the next moment can be obtained:
[0125]
[0126] Step 4: Determine the value of M2 at discrete time points based on the length of the rolling time window, and obtain
[0127] use The mobile game strategy is formulated, and the escape strategy is advanced one time step forward to obtain the actual x1 result. The above calculation process is repeated to update the overall game strategy in real time. Figure 2 As shown in the figure, the preliminary strategy generation result diagram of the game strategy simulation; Figure 3 As shown in , the game strategy simulation update strategy generates the result graph; Figure 4 As shown in the figure, the game strategy simulation ends and the strategy generates the result graph; Figure 5 As shown, the pursuit and escape relative distance curve.
[0128] The present invention also provides an online generation system for mobile game strategies, which can be implemented by executing the process steps of the online generation method for mobile game strategies, that is, those skilled in the art can understand the online generation method for mobile game strategies as a preferred implementation of the online generation system for mobile game strategies.
[0129] According to the present invention, an online generation system of a maneuverable game strategy includes:
[0130] Module M1: Based on the perceived game information, establish a game confrontation situation equation group represented by a mathematical model in the form of discrete differential equations. Use the motion parameters, control parameters and control coefficients at discrete time points to construct the game confrontation situation equation group. The terminal fixed conditions of the game process state include: the start time of the game is t0, the end time is t f =NΔt, N is used to control the total duration of the game. x0 is the initial condition of the game situation, is the initial value of the control parameter of one party in the maneuvering game, and N represents the number of short game sub-processes. Module M1 includes: establishing a discrete differential equation form mathematical model of the state transfer equation, and obtaining the state transfer equation of the motion parameter under the discrete state, as follows:
[0131]
[0132] Among them, x t represents the motion state of the maneuvering party in the game, represents the control parameter of a party in the mobile game, n represents the total number of participants in the mobile game, t represents the current time, B t The coefficient representing the control quantity. The initial conditions of the motion state are x0 and u0 i is a fixed value, which determines the decision step length Δt of the discrete game.
[0133] Module M2: Based on the game confrontation situation equations, according to the differential game strategy theory, the feedback control form of the maneuvering game is obtained. The module M2 includes: Module M2.1: Define the payment function of each party in the game, the payment function J i , as follows:
[0134]
[0135] in, represents the weight coefficient of the state quantity, Represents the weighted coefficient of the control quantity, x t represents the motion state of the maneuvering party in the game, Represents the control parameter of a party in the maneuvering game. Module M2.2: Calculate the feedback control strategy of the maneuvering game through the quadratic payment function form. The module M2.2 includes: Module M2.2.1: Based on the Pontryagin minimization principle, for each time step of each participant, a Hamiltonian function is constructed
[0136]
[0137] in, Represents the coefficient function. Module M2.2.2: In the feedback control form, the motion constraint equations are strictly followed. Substituting the game confrontation situation equations into the Hamiltonian function, we get:
[0138]
[0139] The coefficient function in the payment function is and is a positive definite matrix. Module M2.2.3: Based on the derivative value of the payoff function, let the derivative value be 0, and assume that the state quantity and the co-state quantity in the game equation remain fixed, and calculate:
[0140]
[0141] The optimal control result of each step in the feedback control form is obtained through the above formula.
[0142] Module M3: Construct a method for generating game sub-processes, divide the game sub-processes, discretize the total process into N short game sub-processes, and obtain situation information before each sub-process starts. The module M3 includes: Module M3.1: According to the feedback control formula obtained in module M2, set the rolling time window length t h =sΔt, where s represents the number of discrete steps corresponding to the subprocess length. Let t h The length is much smaller than t f And s is in a time suitable for the execution of game control, so the time length of the generated game sub-process is t h Module M3.2: At the beginning of each sub-process, update the real state parameter information and reset the initial condition x of the maneuvering game. k0 and Module M3.3: Based on the process of module M2, the maneuvering game strategy under the form of feedback control within the sub-process is obtained:
[0143]
[0144] in, represents a discrete time point within the subprocess. At the kth decision time step, the total duration of the game planning is
[0145] Module M4: Establishment of open-loop control strategy. According to the game subprocess determined by module M3, in the period when information is not updated, the feedback control constraints are relaxed and the open-loop strategy solution is generated. Then, in each subprocess, the input is adjusted in real time according to the situation change to achieve online update of the game strategy. The module M4 includes: Module M4.1: Convert the feedback control form into open-loop control, relax the discrete state transition form, and rewrite the Hamiltonian function:
[0146]
[0147] Taking the derivative of the open-loop Hamiltonian function, the new open-loop control result is expressed as:
[0148]
[0149] in, Satisfies the relationship:
[0150]
[0151] Module M4.2: At t = t k At the moment, the final state terminal condition of the sub-process is used to calculate the rolling time window period. The control strategy for each step in is generated online according to the control strategy Execute the current motion control decision. Module M4.3: Advance one time step backward and update the motion state information x0 and the weighted coefficient of the game and The new countermeasure time period is set as According to the new initial quantity and parameter value, the game strategy generation plan is repeatedly executed, and the module M4.1 is returned to generate the game open-loop strategy plan for the next sub-process to achieve real-time update of the maneuver strategy. In each sub-process, the input is adjusted in real time according to the situation change, and the online update of the game strategy includes: taking the fixed initial information of the sub-process as input, obtaining the open-loop prediction form output within each cycle, thereby converting the infinite-length game into a rolling forward form, and advancing in real time based on the updated sub-process information, and quickly obtaining the maneuver game strategy results in real time.
[0152] Those skilled in the art know that, in addition to realizing the system and its various devices, modules, and units provided by the present invention in a purely computer-readable program code, it is entirely possible to realize the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a hardware component, and the devices, modules, and units included therein for realizing various functions can also be regarded as structures within the hardware component; the devices, modules, and units for realizing various functions can also be regarded as both software modules for realizing the method and structures within the hardware component.
[0153] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A method for online generation of maneuverable game strategies, characterized in that: include: Step S1: Based on the perceived game information, a game confrontation situation equation group represented by a mathematical model in the form of discrete differential equations is established; Step S2: Based on the game confrontation situation equations, according to the differential game strategy theory, a feedback control form of the maneuvering game is obtained; Step S3: construct a game sub-process generation method, divide the game sub-process, discretize the total process into N short game sub-processes, and obtain situation information before each sub-process starts; Step S4: Establishment of open-loop control strategy. According to the game sub-process determined in step S3, in the period when information is not updated, the feedback control constraints are relaxed and the open-loop strategy solution is generated. Then, in each sub-process, the input is adjusted in real time according to the situation change to realize the online update of the game strategy.
2. The online generation method of mobile game strategy according to claim 1, characterized in that: The motion parameters, control parameters and control coefficients at discrete time points are used to construct the game confrontation situation equations. The terminal fixed conditions of the game process state include: the start time of the game is t0, the end time is t f =NΔt, where N is used to control the total duration of the game; x0 is the initial condition of the game situation, is the initial value of the control parameter of one party in the mobile game, and N represents the number of short game sub-processes.
3. The online generation method of mobile game strategy according to claim 2, characterized in that: Step S1 includes: establishing a discrete differential equation form mathematical model of the state transfer equation to obtain the state transfer equation of the motion parameter under the discrete state, as shown in the following formula: Among them, x t represents the motion state of the maneuvering party in the game, represents the control parameter of a party in the mobile game, n represents the total number of participants in the mobile game, t represents the current time, B t The coefficient representing the control quantity; Initial conditions x0 and u0 of the motion state i is a fixed value, which determines the decision step length Δt of the discrete game.
4. The online generation method of mobile game strategy according to claim 1, characterized in that: The step S2 comprises: Step S2.1: Define the payment function of each party in the game, the payment function J i , as follows: in, represents the weight coefficient of the state quantity, Represents the weighted coefficient of the control quantity, x t represents the motion state of the maneuvering party in the game, Represents the control parameter of one party in the maneuver game. Step S2.2: Calculate the feedback control strategy of the action game through the quadratic payment function form.
5. The online generation method of mobile game strategy according to claim 4, characterized in that: The step S2.2 comprises: Step S2.2.1: Based on the Pontryagin minimization principle, for each participant at each time step, construct the Hamiltonian function in, represents the coefficient function; Step S2.2.2: In the feedback control form, the motion constraint equations are strictly followed. Substituting the game confrontation situation equations into the Hamiltonian function, we get: The coefficient function in the payment function is and is a positive definite matrix; Step S2.2.3: Based on the derivative value of the payoff function, let the derivative value be 0, and assume that the state quantity and the co-state quantity in the game equation remain fixed, and calculate: The optimal control result of each step in the feedback control form is obtained through the above formula.
6. The online generation method of mobile game strategy according to claim 1, characterized in that: The step S3 comprises: Step S3.1: According to the feedback control formula obtained in step S2, set the rolling time window length t h =sΔt, where s represents the number of discrete steps corresponding to the subprocess length. Let t h The length is much smaller than t f And s is in a time suitable for the execution of game control, so the time length of the generated game sub-process is t h ; Step S3.2: At the beginning of each sub-process, update the real state parameter information and reset the initial condition x of the maneuvering game. k0 and Step S3.3: Based on the process of step S2, the maneuvering game strategy under the feedback control form within the sub-process is obtained: in, Represents a discrete time point within a subprocess.
7. The online generation method of mobile game strategy according to claim 6, characterized in that: At the kth decision time step, the total duration of game planning is 8. The online generation method of mobile game strategy according to claim 1, characterized in that: The step S4 comprises: Step S4.1: Convert the feedback control form into open-loop control, relax the discrete state transition form, and rewrite the Hamiltonian function: Taking the derivative of the open-loop Hamiltonian function, the new open-loop control result is expressed as: in, Satisfies the relationship: Step S4.2: At t=t k At the moment, the final state terminal condition of the sub-process is used to calculate the rolling time window period. The control strategy for each step in is generated online according to the control strategy Execute current motion control decisions; Step S4.3: Advance one time step backward and update the motion state information x0 and the weighted coefficient of the game and The new countermeasure time period is set as According to the new initial quantity and parameter value, the game strategy generation plan is repeatedly executed, and the step S4.1 is returned to generate the game open-loop strategy plan for the next sub-process to realize the real-time update of the maneuver strategy.
9. The online generation method of mobile game strategy according to claim 1, characterized in that: In each sub-process, the input is adjusted in real time according to the changes in the situation to achieve online updating of the game strategy, including: taking the fixed initial information of the sub-process as input, obtaining the open-loop prediction output within each cycle, thereby converting the infinite-length game into a rolling forward form, and advancing in real time based on the updated sub-process information, and quickly obtaining the results of the mobile game strategy in real time.
10. An online generation system for maneuverable game strategies, characterized in that: include: Module M1: Based on the perceived game information, establish a game confrontation situation equation group represented by a mathematical model in the form of discrete differential equations; Module M2: Based on the game confrontation situation equations, according to the differential game strategy theory, the feedback control form of the maneuvering game is obtained; Module M3: Construct a method for generating game sub-processes, divide the game sub-processes, discretize the total process into N short game sub-processes, and obtain situation information before each sub-process starts; Module M4: Establishment of open-loop control strategy. According to the game sub-process determined by module M3, in the period when information is not updated, the feedback control constraints are relaxed and the open-loop strategy solution is generated. Then, in each sub-process, the input is adjusted in real time according to the situation change to realize the online update of the game strategy.
Citation Information
Patent Citations
Multi-agent confrontation decision-making method and device based on deep reinforcement learning
CN117828286A