A Method for Locating the Input-Output Boundary of a Game Agent for Inverting Time-Series Data
By building a game environment in deep reinforcement learning game agents, obtaining time series data, calculating cumulative reward values and performing unsupervised clustering, successfully positioning and inverting the input and output boundaries of the agents, the problems of opacity and low fidelity of boundary positioning in the existing technology are solved, and a more efficient and credible decision-making process is achieved.
Patent Information
- Application Number
- CN202510000217.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-01-02
AI Technical Summary
The prior art is difficult to effectively position the input and output boundaries of deep reinforcement learning game agents, especially in air combat environments with high dynamic and real-time requirements, resulting in low opacity, fidelity and credibility of decision-making.
By constructing a game environment for aircraft confrontation between the enemy and us, obtaining time series data, calculating cumulative reward values, determining the performance boundary moment set, and obtaining the network boundary moment set through unsupervised clustering, the positioning and inversion of the input and output boundary boundaries of the agent is realized.
It realizes the real-time boundary interpretability of game agents at the maneuvering strategy level, improves the fidelity and credibility of decisions, enhances the generalization and adaptability of strategies, and provides more efficient real-time response capabilities.
Smart Images

Figure CN119378695B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep reinforcement learning game agents, and particularly relates to a method for locating the input-output boundary of a game agent by inverting time-series data. Background Art
[0002] In scenarios such as combat simulations, the game agent method trained based on deep reinforcement learning is a commonly used method. Its main purpose is to learn an optimal maneuver decision-making strategy so that the agent (such as our UAV) can maximize the cumulative reward during the interaction with the real or simulated environment. However, the neural network decision model obtained by its training is a black-box model with unknown internal execution logic. The mapping relationship between the network input and output is complex, and the maneuver strategy has poor interpretability. When the autonomous decision-making network of the agent is installed on actual equipment, it is necessary to obtain the maneuver input-output boundary of the game agent and construct a failure boundary scenario to enhance the reliability and security of its intelligent decision-making.
[0003] Currently, the definition of the autonomous input-output boundary of game agents usually performs hyperparameter sampling based on the task-level environment. For example, in "Lu H, Wang S, Cheng S, et al. Surrogate-Based Decision Boundary Evolutionary Sampling Method for Autonomous Systems[J]. IEEE Transactions on Emerging", sampling is carried out among the environment variables such as the battlefield range, the number of air defense equipment and fire coverage, the number of enemy aircraft and equipment, etc. that are known in advance, and task simulations are carried out in the simulation system. The autonomous decision-making performance boundary of the agent is judged according to the success or failure of the task execution. However, even under the same environmental conditions, in each task of the agent, it is very likely that the simulated time-series autonomous decision-making is different from the task result, and this boundary definition does not consider the real-time time-series maneuver autonomous decision-making. And the existing interpretable reinforcement learning game methods only target the action decision-making of the agent. For example, in "Air Combat Maneuver Decision-Making Method Based on Interpretable Reinforcement Learning, Yang Shuheng, Zhang Dong, Xiong Wei, Ren Zhi, Tang Shuo", only the reward function is designed through linear regression and the strategy of the game agent is repaired, and the network decision-making of the agent is classified into interpretable combined actions such as diving, hovering, half-roll inverted, loop, etc., and the output strategy of the model is summarized at the maneuver intention level. Patent CN202410815001.5 converts the incomplete information dynamic game model into a complete information dynamic game model through the Harsanyi transformation to obtain the optimal maneuver strategy of the game agent, and also specifically classifies and explains the maneuver decisions output by the model.
[0004] Due to the complexity and opacity of the model of the deep reinforcement learning game agent, it is difficult to interpret the maneuver decision-making strategy generated by the network, it is difficult to model the complex mapping of input and output features, and the availability of the generalization scenario of the strategy is poor. Existing methods for the interpretability of deep reinforcement learning mainly learn interpretable strategies by fitting a transition probability matrix, or translate the learned strategy into an interpretable form by interpreting the reward function. In the autonomous decision-making of the game agent, existing methods simply classify the decision output of the network into pre-set combined actions, and do not have the judgment on the rationality of the agent network's decision-making in what battlefield situation, and cannot meet the boundary certainty requirements of the high dynamics and strong real-time nature of air combat. For the boundary positioning of the deep reinforcement learning agent, existing technologies perform simulation according to the pre-set environmental variable parameters of sampling, and only perform boundary positioning on the adjustable hyperparameters according to the final task results. The disadvantages are as follows: (1) There is a lack of interpretation of the maneuver decision variables with strong temporal correlation in the game, and this definition does not consider the temporal loop relationship of the autonomous decision-making network and does not combine the real-time battlefield situation for boundary judgment; (2) There are defects in the boundary definition method. Under the same environmental settings, the maneuver decisions and final results of the agent for the game task in each simulation will largely be inconsistent, and only the boundary positioning based on the success or failure of the simulation task does not meet the requirements of the high real-time response of the game agent in terms of basic concepts. Summary of the Invention
[0005] In order to solve the above problems existing in the prior art, the present invention provides a method for boundary positioning of the input and output of a game agent by inverting time series data. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0006] A method for boundary positioning of the input and output of a game agent by inverting time series data, comprising:
[0007] S1, constructing a game environment for the confrontation between the friendly and enemy aircraft, obtaining the input situation data, output maneuver decision data of both the friendly and enemy sides after the game task ends, as well as the input features and output features of the decision-making network based on deep reinforcement learning in our intelligent agent; wherein, our aircraft is the intelligent agent; the input features of the decision-making network are obtained according to the input situation data of our intelligent agent, and the output maneuver decision data of our intelligent agent is obtained according to the output features of the decision-making network;
[0008] S2, calculating the cumulative reward value of the maneuver decision of our intelligent agent at each moment during the execution of the entire game task according to the input situation data;
[0009] S3. Determine the set of performance boundary moments for the change of our side's situation based on the cumulative reward values of the maneuver decisions of our intelligent agent at each moment; perform inversion positioning at each moment according to the set of performance boundary moments to obtain the performance boundary input situation data and performance boundary output maneuver decision data of our side;
[0010] S4. Obtain the set of network boundary moments corresponding to the boundary features of the decision network by performing unsupervised clustering on the output features; perform inversion positioning at each moment according to the set of network boundary moments to obtain the network boundary input situation data and network boundary output maneuver decision data of our side;
[0011] S5. Iteratively execute S1 to S4, and calculate the boundary feature coverage rate based on the obtained performance boundary input situation data, performance boundary output maneuver decision data, network boundary input situation data, and network boundary output maneuver decision data, and stop the iteration until the boundary feature coverage rate is greater than or equal to the preset feature coverage threshold, to obtain the final input-output boundary of the intelligent agent of our side.
[0012] The method for positioning the input-output boundary of a game intelligent agent for time-series data inversion provided by the embodiments of the present invention constructs, for the first time, the positioning methods for the performance boundary and network boundary of the game intelligent agent. Through the feature-level unsupervised clustering interpretation method and the behavior-level real-time environment situation perception and discrimination, driven by the time-series data of the actual or simulation environment, the boundary scene input of the intelligent agent is inverted from the boundary features, realizing the real-time boundary interpretability of the game intelligent agent at the maneuver strategy level. Specifically, it has the following beneficial effects:
[0013] 1. Deeper decision interpretation ability: Through the feature-level unsupervised clustering interpretation method, the present invention can penetrate into the internal feature level of the intelligent agent's decision-making. By solving the first derivative and second derivative of the cumulative reward function, it can accurately locate the interval of the intelligent agent's situation reversal and reveal the internal logic of the intelligent agent's decision-making in different battlefield situations. This in-depth interpretation ability goes beyond the action decision-level interpretation usually provided by the prior art, making the decision-making process more transparent and facilitating the understanding and trust of the operators.
[0014] 2. Higher decision fidelity and credibility: The prior art usually defines the boundary based on the task level. In the same environment setting, the simulation results of the intelligent agent for each task may be inconsistent, resulting in low fidelity and credibility of the boundary positioning. The present invention provides a new positioning method for the performance boundary and network boundary of the game intelligent agent through time-series maneuver input-output boundary positioning, which can more accurately reflect the performance of the intelligent agent in actual games, provide a more comprehensive battlefield situation perception, and thus improve the fidelity and credibility of the decision-making.
[0015] 3. Stronger strategy generalization and adaptability: By inversely inferring the boundary scenario input of the agent through boundary features, the present invention can identify the situations where the agent performs poorly in battlefield situations, and thus optimize the strategy accordingly. This ability enables the strategy of the agent to be effective not only in the training environment, but also in a wider range of actual game environments, improving the generalization ability and adaptability of the strategy.
[0016] 4. More efficient real-time response ability: The present invention can real-time judge the input-output boundary of the agent according to the battlefield situation perception, and this real-time property is crucial for the highly dynamic game environment. In contrast, the existing technologies usually require post-analysis and cannot provide real-time decision support. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic flow chart of a method for locating the input-output boundary of a game agent by inverting time-series data provided by an embodiment of the present invention;
[0018] Figure 2 It is a schematic principle diagram of a method for locating the input-output boundary of a game agent by inverting time-series data provided by an embodiment of the present invention;
[0019] Figure 3 It is a schematic diagram of the battlefield situation function provided by an embodiment of the present invention;
[0020] Figure 4 It is a feature map for locating the performance boundary of the agent for single-air combat data provided by an embodiment of the present invention;
[0021] Figure 5 It is a feature map for locating the network boundary of the agent for single-air combat data provided by an embodiment of the present invention;
[0022] Figure 6 It is a feature map for locating the performance boundary of the agent provided by an embodiment of the present invention;
[0023] Figure 7 It is a feature map for locating the network boundary of the agent provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The following further describes the present invention in detail with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0025] The current method for positioning the boundary of a game agent is to search for settable environmental variables and determine the success or failure of a task. The adjacent variables of the environmental variables corresponding to success and failure are defined as the boundary. However, even under the same environmental settings, the simulation process and results of the agent cannot be exactly the same each time. Therefore, the boundary definition based on the task level has poor fidelity and low credibility. The present invention proposes a brand-new method for positioning the performance boundary and network boundary of a game agent according to the sequential maneuver decision of the agent.
[0026] At the same time, most of the current game agents trained based on deep reinforcement learning have models with a large number of parameters and complex network structures, and the maneuver decision mechanism of their non-linear characteristics is difficult to explain. The present invention is based on time-series data-driven, locates the feature boundary of the agent's situation reversal at the level of maneuver strategy features, and constructs the boundary scenario of the agent through inversion.
[0027] Specifically, an embodiment of the present invention provides a method for positioning the input and output boundaries of a game agent by inverting time-series data. Please refer to Figure 1 the method steps shown in Figure 2 and the method principle shown in
[0028] S1. Construct a game environment for the confrontation between the friendly and enemy aircraft, obtain the input situation data and output maneuver decision data of the friendly and enemy sides after the game task ends, as well as the input features and output features of the decision network based on deep reinforcement learning in our own agent.
[0029] In the embodiment of the present invention, the game environment of the friendly and enemy sides can be constructed by means of real environment interaction or simulation environment interaction. For example, it can be constructed by using a game simulation platform, but the game simulation platform used is not limited here. In this game environment, our side is the red side and the enemy side is the blue side, and the two sides use their respective aircraft for game confrontation. Our aircraft is the agent, and it is not limited whether the enemy aircraft is an agent or not. A trained decision network is deployed in our agent, and this decision network is a deep reinforcement learning network, which can be implemented by using any existing game decision network of an agent based on deep reinforcement learning, and is not limited here.
[0030] The input features of the decision network are obtained according to the input situation data of our agent, and the output maneuver decision data of our agent is obtained according to the output features of the decision network.
[0031] The decision-making network can convert the input situation data perceived by our agent through environmental interaction into the input features of the decision-making network, and through the deep learning decision-making of the decision-making network, obtain the output features, and then convert the output features into the output maneuver decision-making data of our agent to guide the action execution of our agent, so as to conduct adversarial games. For the concepts and interrelationships of the input situation data of our agent, the input features of the decision-making network, the output features of the decision-making network, and the output maneuver decision-making data of our agent, please refer to the relevant content in the field of deep reinforcement learning game agents for understanding, and no detailed description will be given here.
[0032] For S1, first, use the game simulation platform to construct the game environment for both sides of the enemy and us. During the construction process, mainly determine the input and output variables of our agent and the intelligent game simulation environment, etc. Among them, the input variables of our agent usually include the position information of the agent (such as longitude, latitude, altitude), state information (such as pitch angle, heading angle, roll angle), observation data (such as the position, speed, state of the enemy, etc.), and the output variables include the decision-making actions of the agent, such as parameters that can control the aileron, elevator, rudder, throttle, etc. of the agent. The corresponding relationship between the input variables and the output variables is the non-linear mapping of the agent's decision-making network. The game simulation environment usually includes an aerodynamic model (such as the JSBSim aerodynamic platform, etc.), a combat map, etc. The main role of the intelligent game simulation environment is to ensure that the decisions made during the agent's simulation conform to the constraints of the real environment.
[0033] Secondly, start the game tasks of both sides. Specifically, start timing. Through the interaction between our agent and the environment, obtain the input situation data, update the state space of our agent, and output decision-making data, and repeat this until the game task terminates and the timing ends. The termination of the game task can be, for example, reaching the maximum overload of the aircraft, the aircraft being shot down, etc.;
[0034] At the same time, with the help of the game simulation platform, obtain and save the input situation data of both sides of the enemy and us and the output maneuver decision-making data , as well as the input features and the output features of the decision-making network of our agent. This part corresponds to Figure 2 the time-series input perception data and the time-series output decision-making data. The time-series input perception data is the input situation data at each moment during the execution of the game task, and the time-series output decision-making data is the output maneuver decision-making data at each moment during the execution of the game task.
[0035] Among them, the superscript represents our data, the superscript represents the enemy's data, and the subscript and respectively represent the input and output. The input situation data and the output situation data can be specifically set according to the number of aircraft of both sides in the game, the game range, the type of combat, etc. Among them, regarding the number of aircraft of both sides in the game, it can be a 1-on-1 game, or a 2-on-2 game, multi-agent collaborative game, etc. Regarding the game range, it can be a short-range game, a medium-range game, a beyond-visual-range game, etc. Regarding the type of combat, it can be air combat, electronic warfare, air patrol and reconnaissance, etc.
[0036] Taking the 1-on-1 game as an example, in an optional implementation manner, the input situation data may include the altitude, longitude, latitude, heading angle, pitch angle, roll angle and speed of the aircraft; the output maneuver decision data may include the aileron state, elevator state, rudder state, throttle state of the aircraft, and whether to fire a missile. Of course, the input situation data and the output maneuver decision data include but are not limited to the above examples.
[0037] Further, after obtaining the input situation data, output maneuver decision data of both sides of the enemy and us after the game task ends, and the input features and output features of the decision network based on deep reinforcement learning in our intelligent agent, the method may further include:
[0038] Determine the mapping relationship between the input features and the output features, and the corresponding relationship between the input features and the input situation data of our intelligent agent, and the corresponding relationship between the output maneuver decision data of our intelligent agent and the output features.
[0039] Among them, the input features of the decision network of our intelligent agent and the output features The mapping relationship can be expressed as: , represents the decision network of our intelligent agent. This mapping relationship reflects the processing process of the decision network. Through this mapping relationship, given one of the input features and the output features , the other feature can be solved.
[0040] Determining the corresponding relationship between the input features and the input situation data of our intelligent agent, and the corresponding relationship between the output maneuver decision data of our intelligent agent and the output features is to be able to use the corresponding relationship between the input features and the input situation data of our intelligent agent in the subsequent process to obtain the input features corresponding to the decision network for any input situation data of our intelligent agent, and use the corresponding relationship between the output maneuver decision data of our intelligent agent and the output features to obtain the output features corresponding to the decision network for any output maneuver decision data of our intelligent agent, for intelligent agent input-output boundary positioning. Regarding this part of the content, it will be specifically described later.
[0041] S2. According to the input situation data, calculate the cumulative reward value of the maneuver decisions of our agent at each moment during the execution of the entire game task;
[0042] This step is to calculate the real-time battlefield situation of our agent during the execution of the entire game task.
[0043] Specifically, the decision network of our agent is set with a reward function, and the cumulative reward value of the maneuver decisions of our agent at each moment during the execution of the entire game task can be calculated according to the input situation data at each moment and the reward function preset by the decision network.
[0044] In the actual calculation process, the cumulative reward value of the maneuver decisions at each moment is continuously accumulated by calculating the real-time reward value of the maneuver decisions at each moment using the preset reward function. That is to say, assuming that the execution process of the entire game task corresponds to moments 1 to moment, then for each moment in it, the real-time reward value of the maneuver decision of our agent at this moment can be calculated , and this real-time reward value is summed with the real-time reward values of all previous accumulated moments to obtain the cumulative reward value of the maneuver decision of our agent at moment , which represents the sum of the real-time reward values from 1 to moment. Regarding the process of calculating the real-time reward value of the maneuver decision of the agent at each moment given a reward function, reference can be made to related technologies for understanding. Taking a 1-on-1 game as an example, the preset reward function can be combined and set according to the aircraft positions, attitudes, whether in the attack range, whether locked, whether entering the inescapable area, whether shot down, etc. of both sides, and there is no limitation here.
[0045] In a case without considering radar detection, in an optional implementation manner, for a 1-on-1 game, the preset reward function is obtained by adding up the distance reward, attitude advantage reward, and shot-down reward; for example, the sum of the three reward values can be directly calculated. Among them,
[0046] The distance reward is expressed by the formula:
[0047] The distance reward is expressed by the formula:
[0048] ;
[0049] Among them, represents the distance reward; represents calculating the relative distance; represents the position of our aircraft; Indicates the position of the enemy aircraft; the position of the aircraft can be represented by the altitude, accuracy, and latitude and longitude of the aircraft in the geographic coordinate system. The closer our aircraft is to the enemy aircraft, the greater the distance reward is set, with the aim of making our agent actively contact the enemy agent.
[0050] The attitude advantage reward is expressed by the formula:
[0051] ;
[0052] Where, represents the attitude advantage reward; represents the azimuth angle of our aircraft relative to the enemy aircraft; represents the azimuth angle of the enemy aircraft relative to our aircraft; when the sum of the azimuth angles of the enemy and our aircraft is 360°, it means that our agent is in a tail-chase attack posture, and at this time our angle attitude advantage is the greatest; on the contrary, when the sum of the azimuth angles of the enemy and our aircraft is 0°, it means that our agent is in a posture of being tail-chased and attacked, and at this time our angle attitude advantage is the smallest;
[0053] The shoot-down reward is added with represents that the shoot-down reward reflects the confrontation result. A positive reward is given if our aircraft wins, a negative reward is given if the enemy aircraft wins, and a zero reward is given in case of a draw or the game is not over.
[0054] Of course, the above reward function is only an example. In the case of adopting different methods for battlefield situation awareness acquisition, the reward function can be different and can be set accordingly according to the situation. For example, according to different game scenarios, the corresponding battlefield situation judgment criteria can be replaced, such as increasing the radar detection range, radar reflection interface, etc., and then the reward function that meets the requirements can be designed accordingly.
[0055] S3. According to the cumulative reward values of the maneuver decisions of our agent at each moment, determine the set of performance boundary moments of the change of our situation; perform inversion positioning at each moment according to the set of performance boundary moments to obtain the input situation data of the performance boundary and the output maneuver decision data of the performance boundary of our side;
[0056] This step corresponds to Figure 2 in the time-series battlefield situation - situation reversal positioning - performance boundary positioning and the corresponding inversion part.
[0057] In an optional implementation manner, according to the input situation data, calculating the cumulative reward values of the maneuver decisions of our agent at each moment during the execution of the entire game task may include:
[0058] 1), According to the cumulative reward values of the maneuver decisions of our agent at each moment, determine the cumulative reward function;
[0059] According to the time from 1 to during the execution of the entire game task, and the cumulative reward value of the maneuver decisions of our agent at each moment , by means of data fitting and the like, the functional relationship between the time and the cumulative reward value of the maneuver decision can be determined, which is defined as the cumulative reward function and expressed as: ; ;
[0060] Please refer to Figure 3 , Figure 3 , which is a schematic diagram of the battlefield situation function of the embodiment of the present invention. The battlefield situation function is the cumulative reward function. The red curve represents the cumulative reward function of our agent . Figure 3 . The horizontal axis of
[0061] represents time, and the vertical axis represents the cumulative reward value of the maneuver decisions at each moment.
[0062] Specifically, take the first derivative and the second derivative of the cumulative reward function , and take the second derivative ; ;
[0063] If for any point on the curve of the cumulative reward function , it satisfies , and , it can be understood that this point is an inflection point at the peak on the curve of the cumulative reward function . At this time, it is determined that this point is the reversal point where our situation changes from disadvantage to advantage, and the corresponding time of this point is the reversal time point where our situation changes from disadvantage to advantage. Thus, all such points are combined into a set to obtain the performance boundary set , and the set of time points corresponding to the performance boundary set is combined into a set to obtain the first reversal time set;
[0064] If for any point on the curve of the cumulative reward function , it satisfies , and , it can be understood that this point is an inflection point at the trough on the curve of the cumulative reward function . At this time, it is determined that this point is the reversal point where our situation changes from advantage to disadvantage, and the corresponding time of this point is the reversal time point where our situation changes from advantage to disadvantage. Thus, all such points are combined into a set to obtain the performance boundary set , and the set of time points corresponding to the performance boundary set The corresponding time points are grouped to form a set to obtain the second reverse time set;
[0065] 3), combining the first reverse time set and the second reverse time set to obtain a performance boundary time set.
[0066] Combining the performance boundary sets and to merge and obtain a performance boundary interval ; all the time points in the first reverse time set and the second reverse time set are combined in order to obtain a performance boundary time set, that is, the corresponding time series coding interval of the boundary feature is obtained, which can be expressed as . Among them,
[0067] It can be seen that the present invention proposes a method for discriminating the input and output boundary situation of a game agent, judging the battlefield situation information of the game agent according to the cumulative reward value, and combining the first derivative and the second derivative of the situation function (i.e., the cumulative reward function) to obtain the situation reversal interval of the agent.
[0068] Please refer to Figure 4 , Figure 4 which is the intelligent agent performance boundary localization feature map of the single air combat data in the embodiment of the present invention, where the red represents the performance boundary points, and the rest are green non-performance boundary points (normal situation points of the agent).
[0069] Next, inversion localization is performed at each time according to the performance boundary time set to obtain performance boundary input situation data and performance boundary output maneuver decision data, which is based on each time in the performance boundary time set in the input situation data and the output maneuver decision data of the friendly agent saved in S1 for inversion localization, that is, data acquisition at the corresponding time is performed to obtain performance boundary input situation data and performance boundary output maneuver decision data .
[0070] Obtaining performance boundary input situation data and performance boundary output maneuver decision data That is, the performance boundary localization of the friendly side is completed.
[0071] Furthermore, for the convenience of visualizing the localization display, after inversion localization is performed at each time according to the performance boundary time set to obtain the performance boundary input situation data and the performance boundary output maneuver decision data of the friendly side, the method may further include:
[0072] ①For the performance boundary input situation data and the performance boundary output maneuver decision data, respectively perform time extension according to the given visualization range of the performance boundary positioning to obtain the visualization input situation data and the visualization output situation data corresponding to the performance boundary;
[0073] Since the performance boundaries represented by the performance boundary input situation data and the performance boundary output maneuver decision data reflect the short-term changes in our situation, in order to clearly display the specific battle scenarios before and after the situation changes and understand the change process, the embodiments of the present invention can specify a time period to represent the visualization range of the performance boundary positioning for extending the situation change moment. The visualization range of the performance boundary positioning is represented by and the specific size can be set as needed, for example, it can be 20 seconds.
[0074] For the performance boundary input situation data and the performance boundary output maneuver decision data , respectively perform time extension according to the given visualization range of the performance boundary positioning to obtain the visualization input situation data and the visualization output situation data , that is, add time period before and after each moment. The obtained visualization input situation data and the visualization output situation data can be combined into the performance boundary visualization input and output data .
[0075] ②Replicate the visualization input situation data and the visualization output situation data corresponding to the performance boundary through the game simulation platform used when constructing the game environment, and generate the performance boundary battlefield scenario.
[0076] Replicating the performance boundary visualization input and output data through the game simulation platform can visually reproduce the performance boundary battlefield scenario during the corresponding time period, which helps users intuitively understand and feel. For example, for the boundary battlefield scenario where our agent locks the enemy, if there is only the boundary moment set, only the states of both sides of the enemy and us at the current moment can be replicated on the simulation platform, and it is impossible to directly see the sequential dynamic decisions of both sides of the enemy and us before and after the lock. However, this problem can be solved by replicating through the game simulation platform after time extension. This part of the content can be referred to Figure 2 for the understanding of the performance boundary battlefield scenario.
[0077] S4. By performing unsupervised clustering on the output features, obtain the set of network boundary moments corresponding to the boundary features of the decision network; perform inversion positioning at each moment according to the set of network boundary moments to obtain network boundary input situation data and network boundary output maneuver decision data;
[0078] This step corresponds to Figure 2 the unsupervised clustering-network boundary positioning and the corresponding inversion part in
[0079] Unsupervised clustering can adopt existing density clustering, principal component analysis, hierarchical clustering, spectral clustering, etc., which are not limited here. The purpose is to obtain the boundary features in the output features. Through unsupervised clustering, the decision features of the agent can be clustered into two clusters. The decision features corresponding to the cluster with fewer numbers are defined as boundary features, which can be understood as features with poor generalization of the decision network.
[0080] In an optional implementation manner, obtaining the set of network boundary moments corresponding to the boundary features of the decision network by performing unsupervised clustering on the output features may include:
[0081] 1), adopt the density clustering method for the output features to screen out the set of boundary features therein;
[0082] Before performing this step, first perform data preprocessing on the output features of the decision network of our agent saved in S1, specifically remove the corresponding variables with missing values and perform normalization processing to maintain consistent data dimensions.
[0083] The process of performing density clustering on the obtained output features may include the following steps:
[0084] Step a1, input the output feature set composed of all output features of the decision network of our agent, a preset clustering radius and the minimum number of clusters ; ;
[0085] Among them, and are both greater than 0;
[0086] The clustering radius can be set as needed, for example, it can be 3.
[0087] In the boundary positioning of a one-on-one game, the minimum number of clusters is set to 2, which means whether an output feature is a boundary feature.
[0088] Step a2, for each output feature in the output feature set, calculate the The number of points included in the neighborhood is the number of output features, and the obtained number is used as the density of the output feature. ; and judge the density of the output feature whether it satisfies . If so, determine that the output feature is a core feature; otherwise, determine that the output feature is a non-core feature.
[0089] Through step a2, all core features can be screened out.
[0090] Step a3, for each core feature, find all the points in the neighborhood of the core feature to form a cluster.
[0091] Step a4, for each point in the neighborhood of the core feature, if the point is also a core point, then add the points in the neighborhood of the point to the current cluster, and keep recursing until all the points in the neighborhood of the core points have been judged and no new points can be added to the current cluster. Step a5, output the two clusters generated by the clustering. The one with fewer points is defined as the boundary feature cluster, and the other is the regular feature cluster, so as to obtain the boundary feature set according to the boundary feature cluster.
[0092] . .
[0093] 2), according to the moments corresponding to the boundary features in the boundary feature set, obtain the network boundary moment set.
[0094] Combine the moments corresponding to the boundary features in the boundary feature set to obtain the network boundary moment set, which represents the corresponding time series coding interval of the boundary features, so as to represent.
[0095] Please refer to Figure 5 , Figure 5 which is the intelligent agent network boundary positioning feature map of the single air combat data in the embodiment of the present invention, where the red represents the network boundary points and the green represents the points corresponding to the remaining output features.
[0096] Next, perform inversion positioning at each moment according to the network boundary moment set to obtain the network boundary input situation data and the network boundary output maneuver decision data, which are based on each moment in the network boundary moment set, and perform inversion positioning in the input situation data and the output maneuver decision data saved by our side's intelligent agent in S1, that is, obtain the corresponding moment data to obtain the network boundary input situation data and the network boundary output maneuver decision data 。
[0097] Obtain the input situation data of the network boundary and the output maneuver decision data of the network boundary That is, the network boundary positioning of our side is completed.
[0098] It can be seen that the present invention proposes a method for discriminating the generalization of the network boundary of a game agent. By unsupervised clustering, features with poor network generalization are formed into equivalent clusters, and through inversion, the input situation data and output decision data are located to generate a game boundary scenario.
[0099] Furthermore, for the convenience of visual positioning display, after performing inversion positioning at each moment according to the set of network boundary moments and obtaining the input situation data of the network boundary of our side and the output maneuver decision data of the network boundary, the method may further include:[[]]
[0100] ① For the input situation data of the network boundary and the output maneuver decision data of the network boundary, respectively perform moment expansion according to the given visual range of network boundary positioning to obtain the visual input situation data and visual output situation data corresponding to the network boundary;
[0101] Since the network boundary represented by the input situation data of the network boundary and the output maneuver decision data of the network boundary reflects the short - term changes of our side's situation, in order to clearly display the specific battle scenarios before and after the situation changes and understand the change process, the embodiment of the present invention can give a time period to represent the visual range of network boundary positioning for expanding the situation change moments. The visual range of network boundary positioning is represented by and the specific size can be set as needed, for example, it can be 20 seconds.
[0102] For the input situation data of the network boundary and the output maneuver decision data of the network boundary respectively perform moment expansion according to the given visual range of network boundary positioning to obtain the visual input situation data and the visual output situation data , that is, add the time period before and after each moment. The obtained visual input situation data and the visual output situation data Can be combined into the input and output data for network boundary visualization 。
[0103] ② Replicate the visualized input situation data and visualized output situation data corresponding to the network boundary through the game simulation platform used when constructing the game environment, and generate a network boundary battlefield scenario.
[0104] Replicate the network boundary visualization input and output data Through the game simulation platform, it can visually reproduce the network boundary battlefield scenario during the corresponding time period, which helps users understand and feel intuitively. Similar to the performance boundary before, for example, for the boundary battlefield scenario where the decision-making network of our agent makes abnormal maneuver decisions, if there is only the network boundary time set, only the states of both sides of the enemy and us at the current moment can be replicated on the simulation platform, and it is impossible to directly observe the sequential dynamic decision-making of the agent's abnormal maneuver decision. However, this problem can be solved by replicating through the game simulation platform after time expansion. This part of the content can be referred to Figure 2 for the understanding of the network boundary battlefield scenario in
[0105] For S3 and S4, the present invention conducts boundary analysis from the level of decision network features, locates and expands the time series visualization range through inversion, maps it to the corresponding input situation data and output decision data of the agent, replicates and generates the input and output boundary battlefield scenarios through the game simulation platform, thereby visually replicating and displaying the performance boundary and network boundary, which is conducive to intuitive understanding.
[0106] S5. Iteratively execute S1~S4, and calculate the boundary feature coverage rate according to the obtained performance boundary input situation data, performance boundary output maneuver decision data, network boundary input situation data, and network boundary output maneuver decision data, and stop the iteration until the boundary feature coverage rate is greater than or equal to the preset feature coverage threshold, and obtain the final input and output boundary of the agent.
[0107] This step corresponds to Figure 2 the understanding of the iteration in
[0108] Iteratively execute S1~S4 once, and the performance boundary input situation data, performance boundary output maneuver decision data, network boundary input situation data, and network boundary output maneuver decision data of the current iteration can be obtained.
[0109] In an optional implementation manner, the calculating the boundary feature coverage rate according to the obtained performance boundary input situation data, performance boundary output maneuver decision data, network boundary input situation data, and network boundary output maneuver decision data includes:
[0110] 1), Merge the performance boundary input situation data and the network boundary input situation data obtained in the current iteration to obtain the merged input situation data for the current iteration; merge the performance boundary output situation data and the network boundary output situation data obtained in the current iteration to obtain the merged output situation data for the current iteration.
[0111] For each iteration, merge the two types of boundary input situation data and merge the two types of boundary output situation data to obtain the merged input situation data and the merged output situation data.
[0112] 2), According to the established correspondence between the input features and the input situation data of our agent, and the correspondence between the output maneuver decision data of our agent and the output features, calculate the boundary input features corresponding to the merged input situation data for the current iteration, and the boundary output features corresponding to the merged output situation data for the current iteration.
[0113] As mentioned above, in the relevant steps of S1, the correspondence between the input features of the decision network of our agent and the input situation data of our agent, and the correspondence between the output maneuver decision data of our agent and the output features of the decision network of our agent have been determined. Then, the correspondence between the input features of the decision network of our agent and the input situation data of our agent can be used to determine the input features of the decision network corresponding to the merged input situation data for the current iteration, that is, the boundary input features. The correspondence between the output maneuver decision data of our agent and the output features of the decision network of our agent can be used to determine the output features of the decision network corresponding to the merged output situation data for the current iteration, that is, the boundary output features.
[0114] 3), According to the boundary input features and boundary output features obtained in the current iteration, and the boundary input features and boundary output features accumulated before the current iteration, calculate the intersection to obtain the boundary feature coverage rate.
[0115] The formula for calculating the boundary feature coverage rate is as follows:
[0116] ;
[0117] Among them, represents the boundary feature coverage rate; represents the boundary input features and boundary output features obtained in the current iteration; represents the boundary input features and boundary output features accumulated before the current iteration; represents finding the intersection.
[0118] The preset feature coverage threshold can be represented by and Greater than 0, and the specific value can be set according to needs, for example, it can be 0.95.
[0119] If then stop the iteration.
[0120] Specifically, stop the iteration until the boundary feature coverage rate is greater than or equal to the preset feature coverage threshold, and obtain the final input-output boundary of our agent, including:
[0121] According to the boundary input features and boundary output features obtained in the corresponding iteration when the iteration stops, obtain the final input-output boundary of our agent.
[0122] That is to say, the boundary input features and boundary output features obtained in the corresponding iteration when the iteration stops are the final input-output boundary of our agent.
[0123] Please refer to Figure 6 and Figure 7 understand, Figure 6 is the agent performance boundary localization feature map of the embodiment of the present invention, Figure 7 is the agent network boundary localization feature map of the embodiment of the present invention. Specifically, Figure 4 The corresponding single air combat data agent performance boundary localization feature map is the result after executing S1-S3 once, while Figure 6 is the result after executing S1-S3 multiple times until the boundary feature coverage rate requirement is met. Figure 5 The corresponding single air combat data agent network boundary localization feature map is the result after executing S1-S4 once, while Figure 7 is the result after executing S1-S4 multiple times until the boundary feature coverage rate requirement is met.
[0124] For the game agent trained based on deep reinforcement learning, aiming at the problems of difficult intention understanding, difficult policy explanation, and difficult decision trust, the present invention proposes a method for locating the input-output boundary of the game agent by inverting time series data. Please refer to Figure 2 understand, the present invention is to give a decision network of deep reinforcement learning that has been trained completely, obtain the input situation data and output maneuver decision data of our agent in the actual or simulated environment, locate the performance boundary and network boundary of the agent simultaneously through the proposed maneuver input-output boundary method, and continuously analyze and iterate to improve the obtained input-output boundary of the agent, realizing a brand-new positioning scheme for the maneuver decision network boundary and performance boundary of the game agent.
[0125] This method uses feature-level unsupervised clustering interpretation methods and behavior-level real-time environmental situation awareness and discrimination, driven by time series data of actual or simulated environments, and inverts the boundary scene input of the agent from boundary features, thus achieving real-time boundary interpretability of the game agent at the level of maneuver strategy. Among them, based on high-fidelity game data, network boundary positioning realizes clustering analysis of features with poor generalization of decision networks, performance boundary positioning realizes the transformation of maneuver strategies made by decision networks into real-time situation judgments, and constructs a situation reversal discrimination method.
[0126] Specifically, the method of the present invention has the following beneficial effects:
[0127] 1. Deeper decision-making explanation capability: This invention can go deep into the internal feature level of the agent's decision through the feature-level unsupervised clustering explanation method. By solving the first-order derivative and second-order derivative of the cumulative reward function, it can accurately locate the interval of the agent's situation reversal and reveal the internal logic of the agent's decision-making under different battlefield situations. This in-depth explanation capability goes beyond the explanation of the action decision level that the existing technology can usually only provide, making the decision-making process more transparent and easier for operators to understand and trust.
[0128] 2. Higher decision fidelity and credibility: Existing technologies are usually based on task-level boundary definition. Under the same environment setting, the simulation results of each task of the agent may be inconsistent, resulting in low fidelity and credibility of boundary positioning. The present invention provides a new method for positioning the performance boundary and network boundary of game agents through temporal maneuver input and output boundary positioning, which can more accurately reflect the performance of the agent in the actual game and provide a more comprehensive battlefield situation awareness, thereby improving the fidelity and credibility of decision-making.
[0129] 3. Stronger strategy generalization and adaptability: This invention uses boundary feature inversion to identify the boundary scene input of the agent, and can identify the battlefield situations in which the agent performs poorly, so as to optimize the strategy in a targeted manner. This capability makes the agent's strategy effective not only in the training environment, but also in a wider range of actual game environments, improving the generalization and adaptability of the strategy.
[0130] 4. More efficient real-time response capability: The present invention can identify the input and output boundaries of the intelligent agent in real time based on battlefield situation perception, which is crucial for highly dynamic gaming environments. In contrast, existing technologies usually require post-analysis and cannot provide real-time decision support.
[0131] It should be noted that in the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more, unless otherwise specifically defined.
[0132] In the description of this specification, the description with reference to terms such as "an embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0133] The above is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A method for locating the input and output boundaries of a game agent based on time series data inversion, characterized in that: include: S1, constructing a game environment for the confrontation between enemy and friendly aircraft, obtaining the input situation data and output maneuver decision data of the enemy and friendly aircraft after the game task is completed, as well as the input features and output features of the decision network based on deep reinforcement learning in our intelligent agent; wherein our aircraft is an intelligent agent; the input features of the decision network are obtained according to the input situation data of our intelligent agent, and the output maneuver decision data of our intelligent agent are obtained according to the output features of the decision network; S2, according to the input situation data, calculate the cumulative reward value of the maneuvering decision of our agent at each moment during the entire game task execution process; S3, according to the cumulative reward value of the maneuvering decision of our intelligent agent at each moment, determine the performance boundary moment set of our situation change; perform inversion positioning at each moment according to the performance boundary moment set to obtain our performance boundary input situation data and performance boundary output maneuvering decision data; wherein, according to the cumulative reward value of the maneuvering decision of our intelligent agent at each moment, determine the performance boundary moment set of our situation change, including: according to the cumulative reward value of the maneuvering decision of our intelligent agent at each moment, determine the cumulative reward function; calculate the first-order derivative and the second-order derivative of the cumulative reward function, determine the reversal moment point when our situation changes from disadvantage to advantage according to the calculation result, obtain the first reversal moment set, and determine the reversal moment point when our situation changes from advantage to disadvantage, obtain the second reversal moment set; combine the first reversal moment set and the second reversal moment set to obtain the performance boundary moment set; S4, by performing unsupervised clustering on the output features, a network boundary time set corresponding to the boundary features of the decision network is obtained; inversion positioning is performed at each time according to the network boundary time set to obtain our network boundary input situation data and network boundary output maneuver decision data; S5, iteratively execute S1~S4, and calculate the boundary feature coverage rate based on the obtained performance boundary input situation data, performance boundary output maneuver decision data, network boundary input situation data and network boundary output maneuver decision data, and stop the iteration when the boundary feature coverage rate is greater than or equal to the preset feature coverage threshold to obtain our final intelligent agent input and output boundary.
2. The method for locating the input and output boundaries of a game agent based on time series data inversion according to claim 1, characterized in that: The input situation data includes the aircraft's altitude, longitude, latitude, heading angle, pitch angle, roll angle and speed; the output maneuver decision data includes the aircraft's aileron status, elevator status, rudder status, throttle status and whether to launch a missile.
3. The method for locating the input and output boundaries of a game agent based on time series data inversion according to claim 1, characterized in that: After obtaining the input situation data and output maneuver decision data of both the enemy and our party after the game task is completed, as well as the input features and output features of the decision network based on deep reinforcement learning in our party's intelligent agent, the method further includes: Determine the mapping relationship between the input features and the output features, as well as the corresponding relationship between the input features and the input situation data of our agent, and the corresponding relationship between the output maneuver decision data of our agent and the output features.
4. The method for locating the input and output boundaries of a game agent based on time series data inversion according to claim 1, characterized in that: According to the input situation data, the cumulative reward value of the maneuvering decision of our agent at each moment during the execution of the entire game task is calculated, including: According to the input situation data at each moment and the reward function preset by the decision network, the cumulative reward value of the mobile decision-making of our agent at each moment during the execution of the entire game task is calculated.
5. The method for locating the input and output boundaries of a game agent based on time series data inversion according to claim 4 is characterized in that: For a one-on-one game, the preset reward function is obtained by adding the distance reward, the posture advantage reward and the shootdown reward; wherein, The distance reward is expressed as: ; in, represents said distance reward; Indicates the position of our aircraft; Indicates the position of the enemy aircraft; the closer our aircraft is to the enemy aircraft, the greater the distance reward setting; The posture advantage bonus is expressed as: ; in, represents the said posture advantage bonus; Indicates the azimuth of our aircraft relative to the enemy aircraft; Indicates the azimuth of the enemy aircraft relative to our aircraft; when the sum of the azimuths of the enemy and our aircraft is 360°, it means that our agent is in a tail-chasing attack posture, and our angle posture advantage is the greatest at this time; conversely, when the sum of the azimuths of the enemy and our aircraft is 0°, it means that our agent is in a tail-chasing attack posture, and our angle posture advantage is the smallest at this time; The kill reward is It means that a positive reward is given if our aircraft wins, a negative reward is given if the enemy aircraft wins, and a zero reward is given if the game is a draw or the game is not over.
6. The method for locating the input and output boundaries of game agents by time series data inversion according to claim 1, characterized in that: After performing inversion positioning at each moment according to the performance boundary moment set to obtain our performance boundary input situation data and performance boundary output maneuver decision data, the method further includes: The performance boundary input situation data and the performance boundary output maneuver decision data are respectively extended at a given time according to the given performance boundary positioning visualization range to obtain the visualization input situation data and the visualization output situation data corresponding to the performance boundary; The visualized input situation data and the visualized output situation data corresponding to the performance boundary are reproduced through the game simulation platform used when constructing the game environment, and a performance boundary battlefield scene is generated.
7. The method for locating the input and output boundaries of a game agent by time series data inversion according to claim 1, characterized in that: By performing unsupervised clustering on the output features, a network boundary time set corresponding to the boundary features of the decision network is obtained, including: Using density clustering method on the output features to filter out a boundary feature set; A network boundary moment set is obtained according to the moments corresponding to the boundary features in the boundary feature set.
8. The method for locating the input and output boundaries of a gaming agent by time series data inversion according to claim 1 or 7, characterized in that: After performing inversion positioning at each moment according to the network boundary moment set to obtain our network boundary input situation data and network boundary output maneuver decision data, the method further includes: The network boundary input situation data and the network boundary output maneuver decision data are respectively extended at a given time according to the given network boundary positioning visualization range to obtain the visualization input situation data and visualization output situation data corresponding to the network boundary; The visualized input situation data and the visualized output situation data corresponding to the network boundary are reproduced through the game simulation platform used when constructing the game environment, and a network boundary battlefield scene is generated.
9. The method for locating the input and output boundaries of a game agent by time series data inversion according to claim 3, characterized in that: The step of calculating the boundary feature coverage rate based on the obtained performance boundary input situation data, performance boundary output maneuver decision data, network boundary input situation data, and network boundary output maneuver decision data includes: The performance boundary input situation data and the network boundary input situation data obtained in the current iteration are combined to obtain the combined input situation data of the current iteration; the performance boundary output situation data and the network boundary output situation data obtained in the current iteration are combined to obtain the combined output situation data of the current iteration; According to the determined correspondence between the input features and the input situation data of our agent, and the correspondence between the output maneuver decision data of our agent and the output features, the boundary input features corresponding to the combined input situation data of the current iteration and the boundary output features corresponding to the combined output situation data of the current iteration are calculated; The boundary feature coverage is obtained by calculating the intersection of the boundary input features and boundary output features obtained in the current iteration and the boundary input features and boundary output features accumulated before the current iteration; The iteration stops when the boundary feature coverage is greater than or equal to the preset feature coverage threshold, and our final agent input and output boundaries are obtained, including: Based on the boundary input features and boundary output features obtained at the corresponding iteration when the iteration is stopped, we get our final agent input and output boundaries.
Citation Information
Patent Citations
Method and device for determining maneuvering decision of unmanned aerial vehicle
CN118394127A
Primary helium fan fault evolution trend prediction method in combination with fault mechanism analysis
CN118934705A
3D morhphological or anatomical landmark detection method and device using deep reinforcement learning
EP4016470A1