Game strategy generation method and device, equipment and storage medium
By selecting game strategies with optimized win-loss ratios for reinforcement learning training, the problem of strategy imbalance in multi-role competitive games is solved, and the decision-making ability of NPCs and the utilization rate of computing resources are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NETEASE (HANGZHOU) NETWORK CO LTD
- Filing Date
- 2022-09-07
- Publication Date
- 2026-07-21
AI Technical Summary
In existing multiplayer competitive games, the training effect of NPC game strategies is not good, mainly because the strategies of different opponents cannot be effectively distinguished, resulting in wasted training resources and strategy imbalance.
Historical game strategies that meet the win-loss ratio criteria are selected from the game strategy pool as the opposing side. Training resources are obtained based on the win-loss ratio for reinforcement learning training, optimizing the diversity and size of the game strategy pool, and taking into account differences in the number of characters and team composition.
It improves the decision-making ability of NPCs, reduces training costs, enhances the effectiveness of game strategy models and the utilization of computing resources, and adapts to multi-role combat scenarios.
Smart Images

Figure CN117654052B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of game technology, and more specifically, to a game strategy generation method, apparatus, device, and storage medium. Background Technology
[0002] With the rapid development of game technology, multiplayer competitive games involving different game characters have become popular among players. In these games, in addition to the physical logic rules of the game itself, the behavior patterns and intelligence of non-player characters (NPCs) also determine the entertainment value and lifespan of the game.
[0003] In existing technologies, the game strategies adopted by NPCs in most multi-role adversarial games are obtained using self-game reinforcement learning techniques. Specifically, the agent selects opponent strategies from its own game strategy pool in a certain way for reinforcement learning training, and continuously improves in the process, eventually forming a strategy with powerful capabilities.
[0004] However, due to the imbalance of different opponent strategies in the game strategy pool, the above process did not distinguish between different opponent strategies, resulting in poor training results. Summary of the Invention
[0005] In view of this, embodiments of this application provide a game strategy generation method, apparatus, device, and storage medium to solve the problem in the prior art that the lack of differentiation between different opponent strategies leads to poor training results.
[0006] In a first aspect, embodiments of this application provide a game strategy generation method, including:
[0007] At least one historical game strategy is determined from the game strategy pool of the non-player character to be controlled, wherein the game strategy pool includes: multiple historical game strategies of the non-player character to be controlled, each historical game strategy being the game strategy adopted by the non-player character to be controlled in a historical game confrontation;
[0008] Based on the historical win / loss rate of the at least one historical game strategy, obtain the training resources for the at least one historical game strategy respectively;
[0009] Each historical game strategy is used as the opponent's competitive game strategy. The training resources of each historical game strategy are used to perform reinforcement learning training on the game strategy model of the non-player character to be controlled. The game strategy model is used to generate the game strategy of the non-player character to be controlled in the game competition.
[0010] Secondly, embodiments of this application also provide a game strategy generation apparatus, comprising:
[0011] A determining module is used to determine at least one historical game strategy from the game strategy pool of the non-player character to be controlled, wherein the game strategy pool includes: multiple historical game strategies of the non-player character to be controlled, each historical game strategy being the game strategy adopted by the non-player character to be controlled in a historical game confrontation;
[0012] The acquisition module is used to acquire the training resources of the at least one historical game strategy based on the historical game win rate of the at least one historical game strategy.
[0013] The training module is used to use each historical game strategy as the opponent's game strategy, and to use the training resources of each historical game strategy to perform reinforcement learning training on the game strategy model of the non-player character to be controlled. The game strategy model is used to generate the game strategy of the non-player character to be controlled in the game confrontation.
[0014] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the game strategy generation method described in any one of the first aspects.
[0015] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the game strategy generation method described in any one of the first aspects.
[0016] This application provides a game strategy generation method, apparatus, device, and storage medium. The method includes: determining at least one historical game strategy from a game strategy pool of a non-player character to be controlled; the game strategy pool includes multiple historical game strategies of the non-player character to be controlled, each historical game strategy being a game strategy adopted by the non-player character in historical game confrontations; acquiring training resources for at least one historical game strategy based on its historical win / loss rate; using each historical game strategy as the opponent's confrontation game strategy; and using the training resources of each historical game strategy to perform reinforcement learning training on a game strategy model of the non-player character to be controlled. The game strategy model is used to generate a game strategy for the non-player character to be controlled in game confrontations. In this application, considering the imbalance of game strategies, corresponding training resources are obtained based on the win / loss rate of the game strategies, and model training is performed based on the training resources of each strategy. The model training effect is excellent, improving the decision-making level of the non-player character.
[0017] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating the game strategy generation method provided in this application embodiment. Figure 1 ;
[0020] Figure 2 A flowchart illustrating the game strategy generation method provided in this application embodiment. Figure 2 ;
[0021] Figure 3 A schematic diagram of the game win rate curve obtained by fitting for an embodiment of this application;
[0022] Figure 4 A flowchart illustrating the game strategy generation method provided in this application embodiment. Figure 3 ;
[0023] Figure 5 A flowchart illustrating the game strategy generation method provided in this application embodiment. Figure 4 ;
[0024] Figure 6 A flowchart illustrating the game strategy generation method provided in this application embodiment. Figure 5 ;
[0025] Figure 7 A flowchart illustrating the game strategy generation method provided in this application embodiment. Figure 6 ;
[0026] Figure 8 A flowchart illustrating the game strategy generation method provided in this application embodiment. Figure 7 ;
[0027] Figure 9 This is a schematic diagram of the structure of the game strategy generation device provided in the embodiments of this application;
[0028] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0030] Self-play is currently the mainstream mode for applying reinforcement learning techniques to train high-intensity game agents. Its working principle is to obtain an ideal strategy through self-play. During its training process, the agent selects opponent strategies from its own game strategy pool in a certain way for reinforcement learning training, continuously improving in the process and ultimately forming a powerful strategy.
[0031] Reinforcement Learning (RL) is a machine learning method for solving sequential decision-making problems. Its main process involves an agent perceiving the current state in the environment, taking actions to interact with the environment, receiving feedback signals, and adjusting its policy based on these signals. This process of "perception, action, feedback, and optimization" is repeated continuously to maximize the accumulated reward signal. Theoretically, through continuous training, the agent can gradually develop an optimal behavioral policy for a given environment.
[0032] Various game agents developed by combining self-game theory and deep reinforcement learning (DRL) techniques have been widely applied and have achieved many remarkable results. In games, besides the physical and logical rules of the game itself, the behavioral patterns and intelligence levels of non-player characters (NPCs) also determine the game's entertainment value and lifespan. Currently, the game strategies adopted by NPCs in most multi-player competitive games are obtained using self-game reinforcement learning techniques or population-based methods. However, due to the imbalance of opponent strategies, the training results often fail to achieve ideal results.
[0033] Self-play reinforcement learning algorithms and population-based algorithms are simple to develop, highly modular, and supported by mature tools, making them widely used in various games. Representative self-play reinforcement learning algorithms include: fictitious self-play, heuristic fictitious self-play, and prioritized fictitious self-play. Among them, fictitious self-play models the training process as an extensive-form game, using reinforcement learning and supervised learning to approximate the best response strategy and the historical average strategy for historical opponent strategies, respectively. Heuristic self-play selects the latest opponent strategy and the historical average opponent strategy according to a certain proportion based on a predefined threshold (e.g., the probability of selecting the latest opponent strategy is 70%, and the remaining 30% of opponent strategies are sampled from historical strategies). Prioritized self-play assigns corresponding probabilities to different opponent strategies according to a certain sampling weight function, and prioritizes sampling according to these probabilities to train the current strategy.
[0034] However, while virtual self-game theory can theoretically converge to Nash equilibrium in two-player zero-sum games, in practical applications, the opponent strategy selection method is based on historical average strategies, failing to distinguish between opponents defeated with a near 100% win rate and those defeated with a 0% win rate. This not only wastes training resources but also, in unbalanced multi-role games, greatly increases the likelihood that the weaker side cannot learn, leading to passive combat. Furthermore, the weaker side's low-quality strategies can prevent the stronger side from further improvement, resulting in a "one-trick pony" strategy. In actual deployments against human players, this strategy is likely to be weak and fail to provide ideal feedback. Heuristic self-game theory struggles to overcome the catastrophic forgetting problem, i.e., the proportion of historical average opponent strategies selected is relatively small. Priority self-game theory needs further improvement to address the imbalances caused by unbalanced mechanisms or numerical settings for different game roles or professions.
[0035] Population-based agent training methods do not consider the imbalance in multi-player adversarial games, and require training multiple sets of policies simultaneously, which places high demands on computing resources.
[0036] In summary, the above methods have the following drawbacks: First, reinforcement learning agents in multi-role games with imbalance require simpler and more effective opponent policy selection methods. The opponent selection methods used in the above methods involve complex policy learning operations or rely on manual prior adjustments. Second, existing agent construction methods have high computational costs. For example, AlphaStar, based on a large-scale distributed computing architecture, uses a federation training method with a policy pool size of nearly one thousand. Both model training and deployment have high requirements for computing resources. A commercial multi-role game project team often does not have the computing resources required for large-scale distributed training, thus requiring a more streamlined opponent selection method and an efficient policy pool maintenance method. Third, existing methods cannot effectively solve the problem of further multi-agent expansion. In multi-role games, in addition to one-on-one confrontations, there are generally many many-to-many team scenarios. Therefore, direct self-play in the multi-agent environment in many-to-many team scenarios still needs to be addressed.
[0037] To address the aforementioned issues, this application provides a game strategy generation method. Considering the imbalance of game strategies, it extracts opponent strategies that are beneficial to its own progress from the game strategy pool for game competition, and maintains the game strategy pool to ensure its diversity. The diversity of the game strategy pool directly affects the effectiveness of opponent strategy selection and the reinforcement of the finally trained model, thus further improving the decision-making level of non-player characters. Furthermore, the size of the game strategy pool is optimized to reduce training costs. In addition, for multi-character competitive games, the differences in the number of multiple agents and the team formation between characters are taken into consideration, and reinforcement learning training is performed based on the game strategies of multiple characters.
[0038] The game strategy generation method provided in this application will be described below with reference to several specific embodiments.
[0039] Figure 1 A flowchart illustrating the game strategy generation method provided in this application embodiment. Figure 1 In this embodiment, the executing entity can be an electronic device, such as a mobile phone, tablet computer, laptop computer, or other device capable of data processing.
[0040] like Figure 1 As shown, the method includes:
[0041] S101. Determine at least one historical game strategy from the game strategy pool of the non-player character to be controlled.
[0042] The game strategy pool includes: multiple historical game strategies for the non-player character to be controlled, each historical game strategy being the game strategy adopted by the non-player character to be controlled in historical game confrontations.
[0043] During the process of a non-player character being controlled performing game operations in a historical game confrontation, the game operations of the non-player character being controlled in the historical game confrontation can be saved at certain time intervals. All the game operations of the non-player character being controlled in the historical game confrontation constitute the game strategy adopted by the non-player character being controlled in the historical game confrontation. Among them, game operations can include, for example, attack operations, movement operations, etc.
[0044] Among them, at least one historical game strategy can be a game strategy adopted by a non-player character to be controlled in the game strategy pool within a preset time period. The preset time period can be a preset time period with the current time as the end time. That is, based on the consideration of timeliness, the most recently used historical game strategy is taken as at least one historical game strategy.
[0045] In an optional implementation, determining at least one historical game strategy from a pool of game strategies for the non-player character to be controlled includes:
[0046] Based on the historical win rate of each historical game strategy, at least one historical game strategy is identified from the game strategy pool whose win rate meets the first preset condition.
[0047] The historical game strategy win rate is the win rate of the non-player character to be controlled when using the historical game strategy in historical game confrontation. For example, if there are 10 historical game confrontations, and the non-player character to be controlled uses the historical game strategy in all 10 games and wins 8 games, then the win rate is 80% and the loss rate is 20%.
[0048] Based on the historical win rate of each historical game strategy, at least one historical game strategy is determined from the game strategy pool whose win rate meets the first preset condition. The first preset condition may be that the win rate exceeds the first preset win rate threshold or the loss rate does not exceed the first preset loss rate threshold. Based on the historical win rate, the proportion of strategy training is determined from the game strategy pool, and opponent strategies that are beneficial to the improvement of the non-player character to be controlled are extracted for game playing.
[0049] The first preset win rate threshold and the first preset loss rate threshold can be selected according to the actual situation, and this embodiment does not impose any special restrictions on them.
[0050] S102. Based on the historical win / loss rate of at least one historical game strategy, obtain training resources for at least one historical game strategy.
[0051] Each historical game strategy has a historical win rate. Based on the historical win rate of each historical game strategy, training resources for each historical game strategy can be obtained. These training resources may include the number of training iterations.
[0052] There is a corresponding relationship between historical game win rate and training resources: the higher the historical game win rate, the more training resources are available; the lower the historical game win rate, the fewer training resources are available. Conversely, the lower the historical game loss rate, the more training resources are available; and the higher the historical game loss rate, the more training resources are available.
[0053] S103. Treat each historical game strategy as the opponent's adversarial game strategy, and use the training resources of each historical game strategy to perform reinforcement learning training on the game strategy model of the non-player character to be controlled.
[0054] Each historical game strategy is used as the adversary game strategy of the non-player character to be controlled. The training resources of each historical game strategy are used to perform self-game reinforcement learning on the game strategy model of the non-player character to be controlled until a preset iteration stopping condition is reached. The model corresponding to the maximum reward parameter when the preset iteration stopping condition is reached is used as the game strategy model of the non-player character to be controlled. The preset iteration stopping condition may include the number of training times for each historical game strategy.
[0055] The game strategy model is used to generate the game strategy for the non-player character to be controlled in the game confrontation. That is, after the game strategy model is generated, in the actual game confrontation of the non-player character to be controlled, the current state information of the non-player character to be controlled in the actual game confrontation is processed according to the game strategy model to obtain the next action information to be executed by the non-player character to be controlled in the actual game confrontation. All the next action information constitutes the game strategy of the non-player character to be controlled in the game confrontation.
[0056] Table 1 shows the win rate comparison results in actual game competition. As shown in Table 1, the win rate using the game strategy model of this application is 96% when facing the built-in rules, while the win rate using the heuristic self-game method is 36%. It can be seen that the game strategy model of this application is significantly better than the heuristic self-game method.
[0057] Game strategy model Heuristic self-game method Win rate against built-in rules 96% 36%
[0058] Table 1
[0059] In the game strategy generation method of this embodiment, at least one historical game strategy is determined from the game strategy pool of the non-player character to be controlled. The game strategy pool includes multiple historical game strategies of the non-player character to be controlled, each historical game strategy being a game strategy adopted by the non-player character in historical game confrontations. Based on the historical win-loss rate of at least one historical game strategy, training resources for each historical game strategy are obtained. Each historical game strategy is used as the opponent's confrontational game strategy. The training resources of each historical game strategy are used to train the game strategy model of the non-player character to be controlled through reinforcement learning. The game strategy model is used to generate the game strategy of the non-player character to be controlled in game confrontations. Determining training resources based on historical game win-loss rates directly considers the imbalance of historical strategies, is simple to implement, supports effective large-scale parallel training, and also supports training under limited computing resources.
[0060] Figure 2 A flowchart illustrating the game strategy generation method provided in this application embodiment. Figure 2 ,like Figure 2 As shown, based on the historical win / loss rates of at least one historical game strategy, training resources for at least one historical game strategy are obtained, including:
[0061] S201. Based on the historical win / loss rate of each historical game strategy, determine the first learning progress indicator for each historical game strategy.
[0062] The Absolute Learning Progress (ALP) metric is used to indicate the change in the win / loss ratio of each historical game strategy.
[0063] The first learning progress index for each historical game strategy can be calculated as follows: the learning progress derivative at the current time point (i.e., the difference quotient approximation), the Gaussian mixture model fitting algorithm, and the fitting algorithm that takes the difference between two exponential moving averages.
[0064] Taking the difference quotient approximation as an example, based on the historical game win rate of each historical game strategy at two historical times, the first learning progress index of each historical game strategy is determined. That is, the difference between the historical game win rates at two times is calculated, and the ratio of this difference to the historical time difference is used as the first learning progress index. In other words, for each historical game strategy, the historical game win rate at each historical time can be the win rate of the non-player character to be controlled using the historical game strategy in the historical game confrontation at that historical time. The non-player character to be controlled used the same historical game strategy in the historical game confrontation at both historical times.
[0065] For example, in the first historical time period, there are 10 historical game matches. If the non-player character being controlled uses the same historical game strategy in all 10 matches and wins 8 matches, the win rate is 80% and the loss rate is 20%. In the second historical time period, there are 6 historical game matches. If the non-player character being controlled uses the same historical game strategy in all 6 matches and wins 3 matches, the win rate is 50% and the loss rate is 50%. The first learning progress indicator is the ratio of -50% to the historical time difference, which can be the difference between the second historical time period and the first historical time period.
[0066] Figure 3 A schematic diagram of the game win rate curve obtained by fitting the data provided in the embodiments of this application, as shown below. Figure 3 As shown, the horizontal axis x represents time, and the vertical axis y represents the game win rate. The historical time difference Δx is calculated based on historical time B and historical time A, and the win rate change Δy is calculated based on the win rates of historical time B and historical time A. The first learning progress indicator is Δy / Δx.
[0067] S202. Obtain training resources for each historical game strategy based on the first learning progress indicator of each historical game strategy.
[0068] Different first learning progress indicators correspond to different training resources. Based on the first learning progress indicator of each historical game strategy, training resources for each historical game strategy can be obtained. The larger the first learning progress indicator, the more training resources are available, and the smaller the first learning progress indicator, the fewer training resources are available.
[0069] In the game strategy generation method of this embodiment, a first learning progress index is determined for each historical game strategy based on its historical win-loss rate. Training resources for each historical game strategy are then acquired based on this first learning progress index. Training resources are allocated to historical game strategies with higher learning progress, and strategies beneficial to one's own progress are selected from the game strategy pool as the opponent's game strategy. This method effectively considers the imbalance between different historical game strategies, is simple to implement, supports effective large-scale parallel training, and also supports training under limited computing resources.
[0070] First, combined Figures 4-6 This paper describes one implementation process of game strategies in the game strategy pool.
[0071] Figure 4 A flowchart illustrating the game strategy generation method provided in this application embodiment. Figure 3 ,like Figure 4 As shown, the method also includes:
[0072] S301: Obtain all historical game strategies adopted by the non-player character to be controlled at the preset historical time.
[0073] The preset historical time can be a preset time period before the current time. During the preset historical time, when the non-player character to be controlled performs game operations in the historical game confrontation, the game operations of the non-player character to be controlled in the historical game confrontation can be saved at certain time intervals. All the game operations of the non-player character to be controlled in the historical game confrontation constitute the game strategy adopted by the non-player character to be controlled in the historical game confrontation.
[0074] S302. Determine the target game strategy based on all historical game strategies.
[0075] S303. Add the target game strategy to the game strategy pool.
[0076] After obtaining all historical game strategies adopted by the non-player character to be controlled at the preset historical time, the target game strategy is determined based on all historical game strategies, and then added to the game strategy pool. The target game strategy can be a part of all historical game strategies, or it can be a game strategy evolved from all historical game strategies.
[0077] The following is combined Figure 5 This paper describes one implementation process of the target game strategy.
[0078] Figure 5 A flowchart illustrating the game strategy generation method provided in this application embodiment. Figure 4 ,like Figure 5 As shown, based on all historical game strategies, the target game strategy is determined, including:
[0079] S401. Based on the historical win / loss ratio of all historical game strategies, select game strategies from all historical game strategies whose historical win / loss ratio meets the second preset condition.
[0080] S402. Determine the first game strategy based on the game strategy that satisfies the second preset condition.
[0081] The historical game win / loss ratio meets the second preset condition, which means either the historical game win rate exceeds the second preset win rate threshold, or the historical game loss rate does not exceed the second preset loss rate threshold.
[0082] Based on the historical win-loss ratio of all historical game strategies, select game strategies whose historical win-loss ratio meets the second preset condition. Then, based on the game strategies that meet the second preset condition, determine the first game strategy. The target game strategy includes: the first game strategy.
[0083] The first preset win rate threshold and the first preset loss rate threshold can be selected according to the actual situation, and this embodiment does not impose any special restrictions on them.
[0084] In an optional implementation, determining a first game strategy based on a game strategy that satisfies a second preset condition includes:
[0085] Based on the historical win / loss rate of game strategies that meet the second preset conditions, a second learning progress indicator is determined for game strategies that meet the second preset conditions; based on the second learning progress indicator, a first game strategy is determined from game strategies that meet the second preset conditions.
[0086] The second learning progress indicator is used to indicate the change in the win rate of a game strategy that meets the second preset condition.
[0087] The second learning progress index of the game strategy that meets the second preset condition can be calculated in the following ways: the learning progress derivative at the current time point (i.e., the difference quotient approximation), the Gaussian mixture model fitting algorithm, and the fitting algorithm that takes the difference between two exponential moving averages.
[0088] Based on the historical win-loss ratio of game strategies that meet the second preset condition in two historical time periods, a second learning progress index for game strategies that meet the second preset condition is determined. Then, based on the second learning progress index, a first game strategy is selected from the game strategies that meet the second preset condition. The first game strategy can be a game strategy whose second learning progress index exceeds a preset learning progress threshold among the game strategies that meet the second preset condition, thereby avoiding the game strategy pool from being too large, affecting the opponent's strategy selection and occupying a large amount of memory.
[0089] It should be noted that the second learning progress index of the game strategy that meets the second preset condition can also be adjusted based on the timeliness information of each historical game strategy. The timeliness information of each historical game strategy is used to indicate the game confrontation time corresponding to that historical game strategy. The earlier the game confrontation time, the smaller the weight of the second learning progress index; the later the game confrontation time, the larger the weight of the second learning progress index. That is, newer historical game strategies are given higher weights to achieve better opponent strategy selection effect and a balance between exploration and utilization.
[0090] The following is combined Figure 6 An alternative implementation process for the target game strategy is described.
[0091] Figure 6 A flowchart illustrating the game strategy generation method provided in this application embodiment. Figure 5 ,like Figure 6 As shown, based on all historical game strategies, the target game strategy is determined, including:
[0092] S501. Based on the existing game strategies in the game strategy pool, evolve and update all historical game strategies to obtain a third game strategy, such that the similarity between the third game strategy and the existing game strategies does not meet the preset similarity condition.
[0093] The process involves determining the similarity between existing game strategies in the game strategy pool and each historical game strategy. This similarity is used as an evaluation criterion to evolve and update all historical game strategies, resulting in a third game strategy. The target game strategy includes the third game strategy. The similarity between the third game strategy and existing game strategies does not meet a preset similarity condition. The preset similarity condition can be that the similarity exceeds a preset similarity threshold (i.e., exceeds the population entropy threshold). In this way, the third game strategy is not similar to existing game strategies, thus ensuring the diversity of the game strategy pool when adding the third game strategy to it.
[0094] In an optional implementation, based on existing game strategies in the game strategy pool, all historical game strategies are evolved and updated to obtain a third game strategy, such that the similarity between the third game strategy and existing game strategies does not meet a preset similarity condition, including:
[0095] Based on the state information of the non-player character to be controlled when executing all historical game strategies, a preset strategy update model is used to evolve and update all historical game strategies to obtain the third game strategy.
[0096] Each state information corresponds to a historical game strategy. For all historical game strategies, the state information of the non-player character to be controlled when executing each historical game strategy is obtained. The state information may include, for example, the position information and behavioral characteristics of the non-player character to be controlled.
[0097] The preset strategy update model can be a reinforcement learning model. For each historical game strategy, the model input is the state information of the non-player character to be controlled when executing each historical game strategy, and the model output is the next action information to be executed. All the next action information constitutes the game strategy of the non-player character to be controlled. That is, according to the preset strategy update model, the non-player character to be controlled executes each historical game strategy and the evolution update is performed to obtain the third game strategy of the non-player character to be controlled. The target game strategy includes: the third game strategy.
[0098] The strategy update model is a model obtained by training the sample game strategies in the game strategy pool through reinforcement learning. The similarity between the trained sample game strategies and the existing game strategies does not meet the preset similarity conditions. In other words, the third game strategy obtained by evolution and update using this strategy update model is also not similar to the existing game strategies in the game strategy pool. In this way, the third game strategy is added to the game strategy pool to ensure the diversity of the game strategy pool.
[0099] During the training of the strategy update model, the similarity between the sample game strategy and the existing game strategy is used as an environmental reward, so that the strategy evolves in a direction that is dissimilar to the existing game strategies in the game strategy pool, thus avoiding the formation of many similar strategies in the game strategy pool.
[0100] In an optional implementation, before obtaining the third game strategy by updating all historical game strategies using a preset strategy update model based on the state information of the non-player character to be controlled when executing all historical game strategies, the method further includes:
[0101] Based on the game strategy pool, the sample game strategies are trained through multiple rounds of reinforcement learning. If the reinforcement learning training result of the current round meets the preset conditions, the preset similarity conditions are adjusted. In the next round, the sample game strategies are trained through reinforcement learning until the similarity between the trained sample game strategies and the existing game strategies no longer meets the adjusted preset similarity conditions. The model corresponding to the training sample game strategies that do not meet the adjusted preset similarity conditions is then used as the strategy update model.
[0102] Based on the game strategy pool, the sample game strategies are trained through multiple rounds of reinforcement learning. If the reinforcement learning training result of the current round meets the preset conditions, the preset similarity conditions are adjusted, that is, the population entropy threshold is reduced. Here, the reinforcement learning training result of the current round refers to the state information of the non-player character to be controlled when executing the game strategy generated by the strategy update model obtained from the current round of training, which meets the preset requirements. That is, the strategy update model generated during the model training process achieves the preset expected effect, and then the population entropy threshold is reduced.
[0103] Then, in the next round, reinforcement learning is performed on the sample game strategy until the similarity between the trained sample game strategy and the existing game strategy no longer meets the adjusted preset similarity condition. That is, during the model training process, the model iteration stopping condition is adjusted according to the training effect, and the model corresponding to the training sample game strategy and the existing game strategy no longer meeting the adjusted preset similarity condition is used as the strategy update model.
[0104] This is because: during model training, after the game strategy generated by the model reaches the preset expectation, in order to avoid continuing to train the model according to the original model iteration stopping condition, which would cause the strategy generated by the model to deviate from the preset expectation, the population entropy threshold is decayed. That is, the difference between the sample game strategy obtained by training and the existing game strategy is reduced, which helps the strategy to approach the direction of optimal reward.
[0105] Secondly, combine Figure 7 This paper describes another implementation process of game strategies in the game strategy pool.
[0106] Figure 7 A flowchart illustrating the game strategy generation method provided in this application embodiment. Figure 6 ,like Figure 7 As shown, the method also includes:
[0107] S601. Obtain the state information of the non-player character to be controlled when executing all historical game strategies.
[0108] S602. Based on the status information, determine the second game strategy from all historical game strategies.
[0109] S603. Add the second game strategy to the game strategy pool.
[0110] Each state information corresponds to a historical game strategy. For all historical game strategies, the state information of the non-player character to be controlled when executing each historical game strategy is obtained. The state information may include, for example, the position information and behavioral characteristics of the non-player character to be controlled. The historical game strategy is the game strategy adopted by the non-player character to be controlled at a preset historical time.
[0111] In this step, a specific application scenario can be defined. In this scenario, a second game strategy is determined from all historical game strategies based on the similarity between the state information corresponding to each historical game strategy and the state information corresponding to existing game strategies in the game strategy pool. The state information corresponding to existing game strategies in the game strategy pool refers to the state information of the non-player character to be controlled when executing the existing game strategy. This application scenario may include, for example, the presence of traps in the game scenario. That is, when traps exist in the game scenario, the state of the non-player character to be controlled is determined based on the state information of each historical game strategy and the state information of the existing game strategy, to determine whether the state of the non-player character to be controlled is similar, thereby determining whether each historical game strategy is similar to an existing game strategy.
[0112] The second game strategy can be any game strategy from all historical game strategies whose similarity to existing game strategies in the game strategy pool does not exceed a preset similarity threshold. The second game strategy is added to the game strategy pool, ensuring that the game strategies in the game strategy pool are not similar to each other and improving the diversity of the game strategy pool.
[0113] In an optional implementation, determining a second game strategy from all historical game strategies based on state information includes: determining the population entropy of all historical game strategies and existing game strategies in the game strategy pool based on state information, and determining historical game strategies whose population entropy does not exceed a preset population entropy threshold as the second game strategy.
[0114] Based on the state information, the population entropy of each historical game strategy and the existing game strategies in the game strategy pool is determined. The population entropy is used to indicate the similarity between the corresponding historical game strategy and the existing game strategies in the game strategy pool. Then, among all historical game strategies, the historical game strategies whose population entropy does not exceed the preset population entropy threshold are determined as the second game strategy. In this way, the second game strategy is dissimilar to the existing game strategies in the game strategy pool. This ensures the diversity of the game strategy pool when the second game strategy is added to the game strategy pool.
[0115] exist Figures 4 to 7 In the game strategy generation method, the multi-role strategy pool maintenance method optimizes the size and diversity of the game strategy pool for reinforcement learning algorithms to address the imbalance of multiple roles. This results in a game strategy pool with different styles of strategies, which can provide more guidance to the opponent's strategy selection algorithm and reinforcement learning algorithm. This makes it easier for the non-player characters to be controlled to learn better strategies, resulting in higher win rates and rewards in actual deployment. At the same time, it avoids the problem of the game strategy pool occupying too much memory and reduces training costs.
[0116] Figure 8 A flowchart illustrating the game strategy generation method provided in this application embodiment. Figure 7 ,like Figure 8 As shown, before using the training resources of each historical game strategy to perform reinforcement learning training on the game strategy model to be controlled by a non-player character, the method further includes:
[0117] S701: Obtain the historical game strategy of the first game character in the faction of the non-player character to be controlled, and the historical game strategy of the second game character in the faction of the opposing side.
[0118] The game strategy model for controlling a non-player character is trained using the training resources for each historical game strategy, including:
[0119] S702. Using the training resources of each historical game strategy, and based on each historical game strategy, the historical game strategy of the first game character, and the historical game strategy of the second game character, perform reinforcement learning training on the game strategy model of the non-player character to be controlled.
[0120] The system obtains the historical game strategies of the first game character in the faction of the non-player character to be controlled, and the historical game strategies of the second game character in the opposing faction. Each historical game strategy is then used as the opposing game strategy of the non-player character to be controlled. Based on each historical game strategy, the historical game strategies of the first game character, and the second game character, the system performs reinforcement learning training on the game strategy model of the non-player character to be controlled until a preset iteration stopping condition is reached. The model corresponding to the maximum reward parameter when the preset iteration stopping condition is reached is taken as the game strategy model of the non-player character to be controlled. The preset iteration stopping condition may include the number of training times for each historical game strategy.
[0121] In the game strategy generation method of this embodiment, the historical game strategies of the first game character in the faction to which the non-player character to be controlled belongs, and the historical game strategies of the second game character in the faction to which the opponent belongs, are obtained. Training resources for each historical game strategy are used, and reinforcement learning training is performed on the game strategy model of the non-player character to be controlled based on each historical game strategy, the historical game strategy of the first game character, and the historical game strategy of the second game character. By taking into account the difference in the number of characters and the team formation between characters, the application of asymmetric self-game reinforcement learning in multi-agent scenarios is expanded. This enables knowledge transfer between teams of different numbers of agents, forming a natural, gradual progression from easy to difficult. This alleviates the exponentially increasing exploration difficulty of general multi-role reinforcement learning algorithms as the number of characters increases, effectively improving sample utilization and reducing training time.
[0122] Figure 9 This is a schematic diagram of the structure of a game strategy generation device provided in an embodiment of this application. This device can be integrated into an electronic device. Figure 9 As shown, the device includes:
[0123] The determining module 801 is used to determine at least one historical game strategy from the game strategy pool of the non-player character to be controlled, wherein the game strategy pool includes: multiple historical game strategies of the non-player character to be controlled, each historical game strategy being the game strategy adopted by the non-player character to be controlled in a historical game confrontation;
[0124] The acquisition module 802 is used to acquire training resources for at least one historical game strategy based on the historical game win rate of at least one historical game strategy.
[0125] Training module 803 is used to treat each historical game strategy as the opponent's adversarial game strategy. It uses the training resources of each historical game strategy to perform reinforcement learning training on the game strategy model of the non-player character to be controlled. The game strategy model is used to generate the game strategy of the non-player character to be controlled in the game adversarial.
[0126] In an optional implementation, the acquisition module 802 is specifically used for:
[0127] Based on the historical win rate of each historical game strategy, a first learning progress index is determined for each historical game strategy. The first learning progress index is used to indicate the amount of change in the win rate of each historical game strategy.
[0128] Based on the first learning progress metric for each historical game strategy, obtain the training resources for each historical game strategy.
[0129] In an optional implementation, the determining module 801 is specifically used for:
[0130] Based on the historical win rate of each historical game strategy, at least one historical game strategy is identified from the game strategy pool whose win rate meets the first preset condition.
[0131] In an optional implementation, the acquisition module 802 is further configured to:
[0132] Obtain all historical game strategies employed by the non-player character to be controlled at a preset historical time.
[0133] The determination module 801 is also used to determine the target game strategy based on all historical game strategies;
[0134] Add module 804 to add the target game strategy to the game strategy pool.
[0135] In an optional implementation, the target game strategy includes: a first game strategy; the determining module 801 is further configured to:
[0136] Based on the historical win rate of all historical game strategies, select game strategies whose historical win rate meets the second preset condition.
[0137] The first game strategy is determined based on the game strategy that satisfies the second preset condition.
[0138] In an optional implementation, the determining module 801 is specifically used for:
[0139] Based on the historical win-loss rate of the game strategy that meets the second preset condition, determine the second learning progress index of the game strategy that meets the second preset condition.
[0140] Based on the second learning progress indicator, the first game strategy is determined from the game strategies that meet the second preset conditions.
[0141] In an optional implementation, the acquisition module 802 is further configured to:
[0142] Obtain the state information of the non-player character to be controlled when executing all historical game strategies;
[0143] The determination module 801 is also used to determine the second game strategy from all historical game strategies based on the state information;
[0144] Adding module 804 also adds a second game strategy to the game strategy pool.
[0145] In an optional implementation, the determining module 801 is specifically used for:
[0146] Based on the state information, the population entropy of all historical game strategies and existing game strategies in the game strategy pool is determined respectively. The population entropy is used to indicate the similarity between the corresponding historical game strategy and existing game strategies in the game strategy pool.
[0147] All historical game strategies in which the population entropy does not exceed a preset population entropy threshold are identified as the second game strategy.
[0148] In an optional implementation, the target game strategy includes: a third game strategy, and a determination module 801 is specifically used for:
[0149] Based on the existing game strategies in the game strategy pool, all historical game strategies are evolved and updated to obtain a third game strategy, such that the similarity between the third game strategy and the existing game strategies does not meet the preset similarity condition.
[0150] In an optional implementation, the determining module 801 is specifically used for:
[0151] Based on the state information of the non-player character to be controlled when executing all historical game strategies, a preset strategy update model is used to evolve and update all historical game strategies to obtain the third game strategy.
[0152] The strategy update model is a model obtained by training sample game strategies through reinforcement learning based on the game strategy pool, such that the similarity between the trained sample game strategies and existing game strategies does not meet the preset similarity conditions.
[0153] In an optional implementation, the training module 803 is further configured to:
[0154] Based on the game strategy pool, the sample game strategies are trained through multiple rounds of reinforcement learning.
[0155] If the reinforcement learning training results of the current round meet the preset conditions, then adjust the preset similarity conditions;
[0156] In the next round, reinforcement learning is performed on the sample game strategy until the similarity between the trained sample game strategy and the existing game strategy no longer meets the adjusted preset similarity condition. The model corresponding to the training sample game strategy that does not meet the adjusted preset similarity condition is then used as the strategy update model.
[0157] In an optional implementation, the acquisition module 802 is further configured to:
[0158] Obtain the historical game strategies of the first character in the faction to which the non-player character to be controlled belongs, and the historical game strategies of the second character in the faction to which the opposing side belongs;
[0159] Training module 803 is specifically used for:
[0160] Using training resources for each historical game strategy, reinforcement learning training is performed on the game strategy model of the non-player character to be controlled, based on each historical game strategy, the historical game strategy of the first game character, and the historical game strategy of the second game character.
[0161] In the game strategy generation apparatus of this embodiment, a determining module is used to determine at least one historical game strategy from the game strategy pool of the non-player character to be controlled. The game strategy pool includes multiple historical game strategies of the non-player character to be controlled, each historical game strategy being a game strategy adopted by the non-player character in historical game confrontations. An acquiring module is used to acquire training resources for at least one historical game strategy based on its historical win / loss rate. A training module is used to use each historical game strategy as the opponent's confrontation game strategy, and to perform reinforcement learning training on the game strategy model of the non-player character to be controlled using the training resources of each historical game strategy. The game strategy model is used to generate the game strategy of the non-player character to be controlled in game confrontations. Considering the imbalance of game strategies, acquiring corresponding training resources based on the win / loss rate of game strategies and training the model based on the training resources of each strategy results in good model training performance and improves the decision-making level of the non-player character.
[0162] Figure 10 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application, such as... Figure 10 As shown, the device includes a processor 901, a memory 902, and a bus 903. The memory 902 stores machine-readable instructions executable by the processor 901. When the electronic device is running, the processor 901 communicates with the memory 902 via the bus 903. The processor 901 executes the machine-readable instructions to perform the following steps:
[0163] Determine at least one historical game strategy from the game strategy pool of the non-player character to be controlled, wherein the game strategy pool includes: multiple historical game strategies of the non-player character to be controlled, each historical game strategy being the game strategy adopted by the non-player character to be controlled in a historical game confrontation;
[0164] Based on the historical win / loss rate of at least one historical game strategy, obtain training resources for at least one historical game strategy.
[0165] Each historical game strategy is used as the opponent's competitive game strategy. The training resources of each historical game strategy are used to train the game strategy model of the non-player character to be controlled through reinforcement learning. The game strategy model is used to generate the game strategy of the non-player character to be controlled in the game competition.
[0166] In an optional implementation, when the processor 901 performs the task of acquiring training resources for at least one historical game strategy based on the historical game win-loss rates of at least one historical game strategy, it is specifically used for:
[0167] Based on the historical win rate of each historical game strategy, a first learning progress index is determined for each historical game strategy. The first learning progress index is used to indicate the amount of change in the win rate of each historical game strategy.
[0168] Based on the first learning progress metric for each historical game strategy, obtain the training resources for each historical game strategy.
[0169] In an optional implementation, when the processor 901 determines at least one historical game strategy from the game strategy pool of the non-player character to be controlled, it specifically performs the following:
[0170] Based on the historical win rate of each historical game strategy, at least one historical game strategy is identified from the game strategy pool whose win rate meets the first preset condition.
[0171] In an optional implementation, the processor 901 is further configured to:
[0172] Obtain all historical game strategies employed by the non-player character to be controlled at a preset historical time.
[0173] Determine the target game strategy based on all historical game strategies;
[0174] Add the target game strategy to the game strategy pool.
[0175] In an optional implementation, the target game strategy includes: a first game strategy, which the processor 901 specifically uses when determining the target game strategy based on all historical game strategies:
[0176] Based on the historical win rate of all historical game strategies, select game strategies whose historical win rate meets the second preset condition.
[0177] The first game strategy is determined based on the game strategy that satisfies the second preset condition.
[0178] In an optional implementation, when the processor 901 executes a game strategy to determine a first game strategy based on a game strategy that satisfies a second preset condition, it is specifically used to:
[0179] Based on the historical win-loss rate of the game strategy that meets the second preset condition, determine the second learning progress index of the game strategy that meets the second preset condition.
[0180] Based on the second learning progress indicator, the first game strategy is determined from the game strategies that meet the second preset conditions.
[0181] In an optional implementation, the processor 901 is further configured to:
[0182] Obtain the state information of the non-player character to be controlled when executing all historical game strategies;
[0183] Based on the state information, determine the second game strategy from all historical game strategies;
[0184] Add the second game strategy to the game strategy pool.
[0185] In an optional implementation, when the processor 901 determines the second game strategy from all historical game strategies based on state information, it specifically performs the following:
[0186] Based on the state information, the population entropy of all historical game strategies and existing game strategies in the game strategy pool is determined respectively. The population entropy is used to indicate the similarity between the corresponding historical game strategy and existing game strategies in the game strategy pool.
[0187] All historical game strategies in which the population entropy does not exceed a preset population entropy threshold are identified as the second game strategy.
[0188] In an optional implementation, the target game strategy includes: a third game strategy, which the processor 901 specifically uses when determining the target game strategy based on all historical game strategies:
[0189] Based on the existing game strategies in the game strategy pool, all historical game strategies are evolved and updated to obtain a third game strategy, such that the similarity between the third game strategy and the existing game strategies does not meet the preset similarity condition.
[0190] In an optional implementation, when the processor 901 performs evolutionary updates on all historical game strategies based on existing game strategies in the game strategy pool to obtain a third game strategy, such that the similarity between the third game strategy and existing game strategies does not meet a preset similarity condition, it specifically performs the following:
[0191] Based on the state information of the non-player character to be controlled when executing all historical game strategies, a preset strategy update model is used to evolve and update all historical game strategies to obtain the third game strategy.
[0192] The strategy update model is a model obtained by training sample game strategies through reinforcement learning based on the game strategy pool, such that the similarity between the trained sample game strategies and existing game strategies does not meet the preset similarity conditions.
[0193] In an optional implementation, the processor 901 is further configured to:
[0194] Based on the game strategy pool, the sample game strategies are trained through multiple rounds of reinforcement learning.
[0195] If the reinforcement learning training results of the current round meet the preset conditions, then adjust the preset similarity conditions;
[0196] In the next round, reinforcement learning is performed on the sample game strategy until the similarity between the trained sample game strategy and the existing game strategy no longer meets the adjusted preset similarity condition. The model corresponding to the training sample game strategy that does not meet the adjusted preset similarity condition is then used as the strategy update model.
[0197] In an optional implementation, the processor 901 is further configured to:
[0198] Obtain the historical game strategies of the first character in the faction to which the non-player character to be controlled belongs, and the historical game strategies of the second character in the faction to which the opposing side belongs;
[0199] When processor 901 performs reinforcement learning training on the game strategy model for the non-player character being controlled using training resources for each historical game strategy, it is specifically used for:
[0200] Using training resources for each historical game strategy, reinforcement learning training is performed on the game strategy model of the non-player character to be controlled, based on each historical game strategy, the historical game strategy of the first game character, and the historical game strategy of the second game character.
[0201] In this manner, the processor determines at least one historical game strategy from the game strategy pool of the non-player character to be controlled. Based on the historical win-loss rate of each historical game strategy, it acquires training resources for each strategy. Each historical game strategy is used as the adversary's strategy. The training resources for each strategy are then used to train the game strategy model of the non-player character to be controlled through reinforcement learning. This model is then used to generate the game strategy for the non-player character in the game. In this application, considering the imbalance of game strategies, training resources are acquired based on the win-loss rate of each strategy. Model training is then performed based on these training resources, resulting in excellent model training performance and improved decision-making ability for the non-player character.
[0202] This application embodiment also provides a computer-readable storage medium storing a computer program, which is executed by a processor, and the processor performs the following steps:
[0203] Determine at least one historical game strategy from the game strategy pool of the non-player character to be controlled, wherein the game strategy pool includes: multiple historical game strategies of the non-player character to be controlled, each historical game strategy being the game strategy adopted by the non-player character to be controlled in a historical game confrontation;
[0204] Based on the historical win / loss rate of at least one historical game strategy, obtain training resources for at least one historical game strategy.
[0205] Each historical game strategy is used as the opponent's competitive game strategy. The training resources of each historical game strategy are used to train the game strategy model of the non-player character to be controlled through reinforcement learning. The game strategy model is used to generate the game strategy of the non-player character to be controlled in the game competition.
[0206] In an optional implementation, when the processor acquires training resources for at least one historical game strategy based on the historical win-loss ratio of at least one historical game strategy, it specifically performs the following:
[0207] Based on the historical win rate of each historical game strategy, a first learning progress index is determined for each historical game strategy. The first learning progress index is used to indicate the amount of change in the win rate of each historical game strategy.
[0208] Based on the first learning progress metric for each historical game strategy, obtain the training resources for each historical game strategy.
[0209] In an optional implementation, when the processor determines at least one historical game strategy from the game strategy pool of the non-player character to be controlled, it specifically performs the following:
[0210] Based on the historical win rate of each historical game strategy, at least one historical game strategy is identified from the game strategy pool whose win rate meets the first preset condition.
[0211] In an alternative implementation, the processor is further configured to:
[0212] Obtain all historical game strategies employed by the non-player character to be controlled at a preset historical time.
[0213] Determine the target game strategy based on all historical game strategies;
[0214] Add the target game strategy to the game strategy pool.
[0215] In an optional implementation, the target game strategy includes: a first game strategy, which the processor, when determining the target game strategy based on all historical game strategies, specifically uses to:
[0216] Based on the historical win rate of all historical game strategies, select game strategies whose historical win rate meets the second preset condition.
[0217] The first game strategy is determined based on the game strategy that satisfies the second preset condition.
[0218] In an optional implementation, when the processor determines the first game strategy based on a game strategy that satisfies a second preset condition, it specifically performs the following:
[0219] Based on the historical win-loss rate of the game strategy that meets the second preset condition, determine the second learning progress index of the game strategy that meets the second preset condition.
[0220] Based on the second learning progress indicator, the first game strategy is determined from the game strategies that meet the second preset conditions.
[0221] In an alternative implementation, the processor is further configured to:
[0222] Obtain the state information of the non-player character to be controlled when executing all historical game strategies;
[0223] Based on the state information, determine the second game strategy from all historical game strategies;
[0224] Add the second game strategy to the game strategy pool.
[0225] In an optional implementation, when the processor determines the second game strategy from all historical game strategies based on state information, it specifically performs the following:
[0226] Based on the state information, the population entropy of all historical game strategies and existing game strategies in the game strategy pool is determined respectively. The population entropy is used to indicate the similarity between the corresponding historical game strategy and existing game strategies in the game strategy pool.
[0227] All historical game strategies in which the population entropy does not exceed a preset population entropy threshold are identified as the second game strategy.
[0228] In an optional implementation, the target game strategy includes: a third game strategy, which the processor, when determining the target game strategy based on all historical game strategies, specifically uses to:
[0229] Based on the existing game strategies in the game strategy pool, all historical game strategies are evolved and updated to obtain a third game strategy, such that the similarity between the third game strategy and the existing game strategies does not meet the preset similarity condition.
[0230] In an optional implementation, when the processor performs evolutionary updates on all historical game strategies based on existing game strategies in the game strategy pool to obtain a third game strategy, such that the similarity between the third game strategy and existing game strategies does not meet a preset similarity condition, it specifically performs the following:
[0231] Based on the state information of the non-player character to be controlled when executing all historical game strategies, a preset strategy update model is used to evolve and update all historical game strategies to obtain the third game strategy.
[0232] The strategy update model is a model obtained by training sample game strategies through reinforcement learning based on the game strategy pool, such that the similarity between the trained sample game strategies and existing game strategies does not meet the preset similarity conditions.
[0233] In an alternative implementation, the processor is further configured to:
[0234] Based on the game strategy pool, the sample game strategies are trained through multiple rounds of reinforcement learning.
[0235] If the reinforcement learning training results of the current round meet the preset conditions, then adjust the preset similarity conditions;
[0236] In the next round, reinforcement learning is performed on the sample game strategy until the similarity between the trained sample game strategy and the existing game strategy no longer meets the adjusted preset similarity condition. The model corresponding to the training sample game strategy that does not meet the adjusted preset similarity condition is then used as the strategy update model.
[0237] In an alternative implementation, the processor is further configured to:
[0238] Obtain the historical game strategies of the first character in the faction to which the non-player character to be controlled belongs, and the historical game strategies of the second character in the faction to which the opposing side belongs;
[0239] When the processor executes reinforcement learning training on the game strategy model for the non-player character under control using training resources for each historical game strategy, it is specifically used for:
[0240] Using training resources for each historical game strategy, reinforcement learning training is performed on the game strategy model of the non-player character to be controlled, based on each historical game strategy, the historical game strategy of the first game character, and the historical game strategy of the second game character.
[0241] In this manner, the processor determines at least one historical game strategy from the game strategy pool of the non-player character to be controlled. Based on the historical win-loss rate of each historical game strategy, it acquires training resources for each strategy. Each historical game strategy is used as the adversary's strategy. The training resources for each strategy are then used to train the game strategy model of the non-player character to be controlled through reinforcement learning. This model is then used to generate the game strategy for the non-player character in the game. In this application, considering the imbalance of game strategies, training resources are acquired based on the win-loss rate of each strategy. Model training is then performed based on these training resources, resulting in excellent model training performance and improved decision-making ability for the non-player character.
[0242] In this embodiment, the computer program, when run by the processor, can also execute other machine-readable instructions to perform other methods as described in the embodiments. For details on the specific execution steps and principles, please refer to the description of the embodiments, which will not be repeated here.
[0243] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the shown or discussed mutual couplings or direct couplings or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0244] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0245] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0246] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0247] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0248] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.
Claims
1. A method for generating game strategies, characterized in that, include: At least one historical game strategy is determined from the game strategy pool of the non-player character to be controlled, wherein the game strategy pool includes: multiple historical game strategies of the non-player character to be controlled, each historical game strategy being the game strategy adopted by the non-player character to be controlled in a historical game confrontation; Based on the historical win / loss rate of the at least one historical game strategy, obtain the training resources for the at least one historical game strategy respectively; Each historical game strategy is used as the opponent's game strategy. The training resources of each historical game strategy are used to perform reinforcement learning training on the game strategy model of the non-player character to be controlled. The game strategy model is used to generate the game strategy of the non-player character to be controlled in the game confrontation. The step of obtaining training resources for the at least one historical game strategy based on its historical game win / loss rate includes: Based on the historical win rate of each historical game strategy, a first learning progress index is determined for each historical game strategy, and the first learning progress index is used to indicate the amount of change in the win rate of each historical game strategy. Based on the first learning progress index of each historical game strategy, obtain the training resources for each historical game strategy.
2. The method according to claim 1, characterized in that, The step of determining at least one historical game strategy from the game strategy pool of the non-player character to be controlled includes: Based on the historical win rate of each historical game strategy, the historical game strategy whose win rate meets the first preset condition is determined from the game strategy pool as the at least one historical game strategy.
3. The method according to claim 1, characterized in that, The method further includes: Obtain all historical game strategies adopted by the non-player character to be controlled at the preset historical time. Based on all the historical game strategies, determine the target game strategy; Add the target game strategy to the game strategy pool.
4. The method according to claim 3, characterized in that, The target game strategy includes: a first game strategy; determining the target game strategy based on all historical game strategies includes: Based on the historical win-loss ratio of all historical game strategies, select game strategies whose historical win-loss ratio meets the second preset condition from all historical game strategies; The first game strategy is determined based on the game strategy that satisfies the second preset condition.
5. The method according to claim 4, characterized in that, The step of determining the first game strategy based on a game strategy that satisfies the second preset condition includes: Based on the historical win-loss rate of the game strategy that meets the second preset condition, a second learning progress indicator for the game strategy that meets the second preset condition is determined. Based on the second learning progress indicator, the first game strategy is determined from the game strategies that meet the second preset conditions.
6. The method according to claim 1, characterized in that, The method further includes: Obtain the state information of the non-player character to be controlled when executing all historical game strategies; Based on the state information, determine the second game strategy from all historical game strategies; Add the second game strategy to the game strategy pool.
7. The method according to claim 6, characterized in that, The step of determining the second game strategy from all historical game strategies based on the state information includes: Based on the state information, the population entropy of all historical game strategies and the existing game strategies in the game strategy pool is determined respectively. The population entropy is used to indicate the similarity between the corresponding historical game strategy and the existing game strategies in the game strategy pool. The historical game strategy whose population entropy does not exceed a preset population entropy threshold among all historical game strategies is determined as the second game strategy.
8. The method according to claim 3, characterized in that, The target game strategy includes: a third game strategy, wherein determining the target game strategy based on all historical game strategies includes: Based on the existing game strategies in the game strategy pool, all historical game strategies are evolved and updated to obtain the third game strategy, such that the similarity between the third game strategy and the existing game strategies does not meet the preset similarity condition.
9. The method according to claim 8, characterized in that, The step of evolving and updating all historical game strategies based on existing game strategies in the game strategy pool to obtain the third game strategy, such that the similarity between the third game strategy and the existing game strategies does not meet a preset similarity condition, includes: Based on the state information of the non-player character to be controlled when executing all the historical game strategies, a preset strategy update model is used to evolve and update all the historical game strategies to obtain the third game strategy. The strategy update model is obtained by training the sample game strategies with reinforcement learning based on the game strategy pool, such that the similarity between the trained sample game strategies and the existing game strategies does not meet the preset similarity condition.
10. The method according to claim 9, characterized in that, Before obtaining the third game strategy by updating all historical game strategies using a preset strategy update model based on the state information of the non-player character to be controlled when executing all historical game strategies, the method further includes: Based on the game strategy pool, the sample game strategy is trained through multiple rounds of reinforcement learning. If the reinforcement learning training result of the current round meets the preset conditions, then the preset similarity conditions are adjusted. In the next round, the sample game strategy is trained using reinforcement learning until the similarity between the trained sample game strategy and the existing game strategy no longer meets the adjusted preset similarity condition. The model corresponding to the training sample game strategy and the existing game strategy not meeting the adjusted preset similarity condition is then used as the strategy update model.
11. The method according to claim 1, characterized in that, Before performing reinforcement learning training on the game strategy model of the non-player character to be controlled using the training resources of each historical game strategy, the method further includes: Obtain the historical game strategy of the first game character in the faction to which the non-player character to be controlled belongs, and the historical game strategy of the second game character in the faction to which the opposing side belongs; The step of using the training resources of each historical game strategy to perform reinforcement learning training on the game strategy model of the non-player character to be controlled includes: Using the training resources of each historical game strategy, reinforcement learning training is performed on the game strategy model of the non-player character to be controlled, based on each historical game strategy, the historical game strategy of the first game character, and the historical game strategy of the second game character.
12. A game strategy generation device, characterized in that, include: A determining module is used to determine at least one historical game strategy from the game strategy pool of the non-player character to be controlled, wherein the game strategy pool includes: multiple historical game strategies of the non-player character to be controlled, each historical game strategy being a game strategy adopted by the non-player character to be controlled in a historical game confrontation; The acquisition module is used to acquire the training resources of the at least one historical game strategy based on the historical game win rate of the at least one historical game strategy. The training module is used to use each historical game strategy as the opponent's game strategy, and to use the training resources of each historical game strategy to perform reinforcement learning training on the game strategy model of the non-player character to be controlled. The game strategy model is used to generate the game strategy of the non-player character to be controlled in the game confrontation. The acquisition module is specifically used for: Based on the historical win rate of each historical game strategy, a first learning progress index is determined for each historical game strategy, and the first learning progress index is used to indicate the amount of change in the win rate of each historical game strategy. Based on the first learning progress index of each historical game strategy, obtain the training resources for each historical game strategy.
13. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the game strategy generation method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, performs the game strategy generation method as described in any one of claims 1 to 11.