Data processing method and device of asymmetric game, electronic equipment and storage medium
Through training and simulation game methods, the game data changes in asymmetric competitive games are analyzed, and the problem of difficulty in optimizing game balance in the existing technology is solved, and efficient and objective balance optimization is achieved.
Patent Information
- Application Number
- CN202311459285.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2043-11-03
AI Technical Summary
The existing technology is difficult to effectively optimize the balance of asymmetric competitive games in practice, mainly due to the problems of low data quality, high algorithm complexity and strong subjectivity of user feedback.
By training non-player characters, iteratively update their behavior Q value table, simulate the game to obtain game data, analyze the game data change information to determine the game state, and adjust the game parameters according to the status to achieve balanced optimization.
It achieves a rapid and objective evaluation of game balance, improves the efficiency and effect of balance optimization, and reduces testing costs and time.
Smart Images

Figure CN119925936A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a data processing method, device, electronic device and storage medium for an asymmetric game. Background Art
[0002] Asymmetric competitive games are a type of multiplayer game in which different players play different roles with different skills, goals, and abilities. The goals and abilities of these players are different, so the balance and difficulty of the game often need to be carefully designed to be guaranteed.
[0003] In the prior art, machine learning has been used as a solution for automatically optimizing the balance of asymmetric competitive games. By collecting game data after the game is launched, machine learning technology is used to analyze the game data, predict balance issues in the game, and propose corresponding optimization solutions.
[0004] However, in the prior art, the quality of game data has a great impact on the accuracy and reliability of the analysis results. There may be noise and errors in the data collection process, and the preferences of players may also cause data deviation. In addition, the algorithm is relatively complex, so it is difficult to optimize the balance of the game in practice. Summary of the invention
[0005] The purpose of this application is to provide a data processing method, device, electronic device and storage medium for an asymmetric game in order to address the deficiencies in the above-mentioned prior art and to solve the problem of optimizing the balance of the game in practice in the prior art.
[0006] To achieve the above objectives, the technical solutions adopted in this application are as follows:
[0007] In a first aspect, the present application provides a data processing method for an asymmetric game, the method comprising:
[0008] Iteratively updating a behavior Q value table of the non-player character based on a game match result of the non-player character in a training match of the asymmetric game until the trained non-player character is obtained when an iteration termination condition is satisfied; the behavior Q value table is used to represent the Q value corresponding to the state action of the non-player character;
[0009] Acquire first game data of the trained non-player character in a first simulated game of the asymmetric game; the first simulated game includes first game parameters before parameter modification;
[0010] Acquire second game data of the trained non-player character in a second simulated game of the asymmetric game, wherein the second simulated game includes second game parameters obtained by modifying the first game parameters;
[0011] Determining game data change information of the non-player character according to the second game data and the first game data;
[0012] Based on the game data change information, a game state of the asymmetric game is determined to determine whether to modify the second game parameter based on the game state; the game state includes a balanced state and an unbalanced state.
[0013] In a second aspect, the present application provides a data processing device for an asymmetric game, the device comprising:
[0014] A training module, configured to iteratively update a behavior Q value table of the non-player character based on a game match result of the non-player character in a training match of the asymmetric game, until the trained non-player character is obtained when an iteration termination condition is satisfied; the behavior Q value table is used to represent a Q value corresponding to a state action of the non-player character;
[0015] A first acquisition module is used to acquire first game data of the trained non-player character in a first simulated game of the asymmetric game; the first simulated game includes first game parameters before parameter modification;
[0016] A second acquisition module is used to acquire second game data of the trained non-player character in a second simulated game of the asymmetric game, wherein the second simulated game includes second game parameters obtained by modifying the first game parameters;
[0017] a change determination module, configured to determine game data change information of the non-player character according to the second game data and the first game data;
[0018] A state determination module is used to determine the game state of the asymmetric game based on the game data change information, so as to determine whether to modify the second game parameter based on the game state; the game state includes a balanced state and an unbalanced state.
[0019] In a third aspect, an embodiment of the present application further provides an electronic device, comprising: a processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of a data processing method for an asymmetric game as described in any one of the first aspects.
[0020] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a data processing method for an asymmetric game as described in any one of the first aspects are executed.
[0021] The beneficial effects of the present application are as follows: by training non-player characters in training games, and updating the behavior Q value table of the non-player characters according to the game operation results of the non-player characters in the training games, non-player characters of all levels and roles that are closer to the level of real players are obtained; after the non-player characters are obtained, the trained non-player characters are used to perform simulated games, which can not only simulate the game process of players of different levels, but also obtain a large amount of real and objective game data in a short time, and quickly discover problems and defects in the game in a short time, so as to facilitate timely repair and optimization; after a large amount of game data is obtained by using non-player characters to perform simulated games, the balance of the game can be quickly measured based on these game data and the data before modifying the parameters, thereby improving the efficiency and effect of game balance optimization.
[0022] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0024] Figure 1 A schematic diagram showing an application scenario provided by an embodiment of the present application is shown;
[0025] Figure 2 A flow chart showing a data processing method for an asymmetric game provided by an embodiment of the present application;
[0026] Figure 3 A flowchart of updating a behavior Q value table provided in an embodiment of the present application is shown;
[0027] Figure 4 A flowchart for determining an action to be performed provided by an embodiment of the present application is shown;
[0028] Figure 5 A flowchart of determining an updated Q value of an action to be performed provided by an embodiment of the present application is shown;
[0029] Figure 6 A flowchart for determining whether a behavior Q value table is effective is shown in an embodiment of the present application;
[0030] Figure 7 A flowchart of determining game data change information provided by an embodiment of the present application is shown;
[0031] Figure 8 A flowchart of another method of determining game data change information provided by an embodiment of the present application is shown;
[0032] Fig. 9 A schematic diagram showing the structure of a data processing device for an asymmetric game provided by an embodiment of the present application is shown;
[0033] Fig.10 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0034] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0035] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.
[0036] Asymmetric competitive games are a type of multiplayer game in which different players play different roles with different skills, goals, and abilities. The goals and abilities of these players are different, so the balance and difficulty of the game often need to be carefully designed to be guaranteed.
[0037] In asymmetric competitive games, there are complex and subtle relationships between each role, rule, scene design and other factors. Only by ensuring that various parameters are coordinated and balanced can players enjoy a better experience.
[0038] However, it is not easy to achieve a good balance in practice. In asymmetric competitive games, each player has his or her own unique style and strategy, which makes data collection and analysis difficult. In addition, there are many elements in the game, and different role combinations will also lead to sparse data. In addition, the product iteration speed is fast, and new functions will also affect the overall balance. In addition, player needs are obviously differentiated. Therefore, optimizing the balance in asymmetric competitive games has become an urgent problem to be solved.
[0039] Existing solutions for optimizing the balance of asymmetric competitive games include machine learning, data mining, and user feedback.
[0040] Among them, machine learning uses machine learning technology to analyze game data, predict balance issues in the game, and propose corresponding optimization solutions.
[0041] However, machine learning algorithms are highly complex and require reliable data support. However, game data has low data quality due to players' personal preferences and noise during the collection process. Therefore, using machine learning to optimize balance is still insufficient.
[0042] Data mining refers to game planners using data mining technology to discover rules and patterns in game data, discover balance issues in the game, and optimize the balance through version fine-tuning.
[0043] However, this method relies on the subjective experience of game planners. If planners do not fully understand the game's mechanics and players' needs, they may incorrectly modify certain values, resulting in deviations in the final balance optimization results.
[0044] User feedback refers to collecting user feedback and opinions through a questionnaire system, organizing data to analyze user behavior patterns and needs, and using this to solve balance problems in the game, planning to propose corresponding optimization plans based on the problems found.
[0045] However, user feedback is highly subjective, and the quality and quantity of user feedback may be problematic, which may lead to deviations in the results of game balance optimization.
[0046] In summary, the existing solutions for optimizing the balance of asymmetric competitive games have problems such as low data quality, high algorithm complexity, and excessive subjectivity in user feedback and game planning.
[0047] Based on the above problems, the present application proposes a data processing method for asymmetric games. After the game version is iteratively updated, the robot is trained to learn different characters to play the game, ensuring that the robot has a reasonable winning rate at different levels of the characters and will not have abnormalities due to different operations of the players. After the training is completed, the robot is used to simulate the game, and based on the change information of the game data, the impact of the iterative update of the game version on the game balance is verified and optimized.
[0048] First, the basic process of optimizing balance in this application is explained. After a new character or a new game map is released, it is generally necessary to iterate the game version. After the game version is iterated, the balance of the game can generally be ensured by adjusting some parameters in the game. After adjusting the parameters, if the game data of the same character at the same rank does not change significantly before and after the parameter modification, then it can be considered that the game is in a balanced state. Otherwise, it is considered that the balance of the game is broken, and the game parameters need to be readjusted at this time.
[0049] like Figure 1 As shown, it is a brief flow chart of the method of the present application. After the game version is iterated, a training game between a non-player character and a real player can be conducted on an electronic device based on the method of the present application to train a non-player character of all levels and all roles, wherein the non-player character can be a game robot, and then game data of different combinations are obtained through robot simulation games, so as to determine the balance state of the current game based on the game data, and adjust the game parameters to achieve balance optimization.
[0050] Compared with the existing technology, the algorithm of the present application has low complexity and comprehensive and objective data, which can provide more objective and realistic game evaluation results. In addition, the game robot can save game simulation time, save a lot of testing costs, and improve testing efficiency and effectiveness.
[0051] Next, the data processing method of the asymmetric game of the present application is described in conjunction with a specific embodiment. The execution subject of the method may be an electronic device, such as Figure 2 As shown, the method includes:
[0052] S201: Iteratively updating a behavior Q value table of the non-player character based on a game match result of the non-player character in a training match of the asymmetric game until a trained non-player character is obtained when an iteration termination condition is satisfied; the behavior Q value table is used to characterize the Q value corresponding to the state action of the non-player character.
[0053] Optionally, the non-player character can be a program script running on the electronic device for simulating the real operation of the player. The tester can create multiple processes on the electronic device, and each process runs a non-player character of a different role.
[0054] As another possible implementation, the tester may also create multiple processes on different electronic devices to run non-player characters of different roles.
[0055] Optionally, the electronic device can create or participate in different training games, and use non-player characters of different roles and ranks to learn simulated games in the training games, where the training games include non-player characters and real players.
[0056] Among them, non-player characters of the same rank and the same role share the same storage space, and the game operation results generated in the training game can all be stored in the storage space.
[0057] Optionally, the game operation result may be game data generated by the non-player character in the training game, such as a non-player character performing a certain action in a certain state in the game scene, and the result of the action.
[0058] The status of the non-player character may include the position of the non-player character in the game scene, the life information of the non-player character, and the prop information currently possessed by the non-player character, etc. The life information may include, for example, health points and skill cooldown time, and the prop information may include, for example, the prop currently used by the non-player character, etc.
[0059] The actions of the non-player character may be game operations that can be performed by the non-player character in the current state. For example, at a certain map position, the actions that can be taken are up, down, left, right, attack, release skills, etc.
[0060] Optionally, for each character of different ranks, the non-player character can initialize a behavior Q value table. The values of each item in the initial behavior Q value table can all be 0 or random numbers. During the training game, the initial behavior Q value table can be updated based on the game operation results of the non-player character to obtain the behavior Q value table of the character of the current rank of the non-player character in the training game.
[0061] As shown in Table 1, this is an example of a behavior Q value table given in the present application, wherein S1 represents a state, which includes 7 actions, and the numbers under each action represent the Q value of the action in the state of S1.
[0062] Table 1. Example of Q value table
[0063]
[0064] The Q value indicates the degree of influence of executing the action in the S1 state on the result. The larger the Q value, the more likely it is that executing the action in the S1 state will lead to a game victory. The smaller the Q value, the more likely it is that executing the action in the S1 state will lead to a game failure.
[0065] It is worth noting that in training games, non-player characters can be arranged to play against real players, and the proportion of non-player characters in the training games may not exceed a preset proportion of the total number of players, such as 10%, to ensure that the non-player characters can learn game strategies of different characters in a game that is close to the real thing.
[0066] In the present application, the non-player character can iteratively perform multiple training games, iteratively update the behavior Q value table, and when the behavior Q value table meets the iteration termination condition, save the behavior Q value table of the non-player character obtained by the last training.
[0067] S202: Acquire first game data of a trained non-player character in a first simulated game of an asymmetric game; the first simulated game includes first game parameters before parameter modification.
[0068] After obtaining trained non-player characters of all ranks and roles, simulated games with different character combinations can be created, and game robots can be triggered to perform game operations in the simulated games.
[0069] Among them, the participants in the first simulated game can all be non-player characters, the roles and ranks of the non-player characters can be randomly selected, and the electronic device can select different arrangements and combinations to conduct multiple simulated games to obtain more comprehensive game data.
[0070] Optionally, the first game data represents the game data before the game parameters are modified. The first game parameters may include: skill cooldown time of the character in the game, attack damage value of the character, and other configuration information that can affect the strength of the character.
[0071] S203: Acquire second game data of the trained non-player character in a second simulated game of the asymmetric game, where the second simulated game includes second game parameters obtained by modifying the first game parameters.
[0072] After obtaining the first simulated game data, the game parameters of the asymmetric game can be modified, and the first game parameters can be modified to second game parameters, such as adjusting the skill cooldown time of the character in the game, the character's attack damage value, etc.
[0073] Among them, all participants in the second simulated game are non-player characters, and the second game data represents the game data after the game parameters are modified.
[0074] It should be noted that the strategy for the non-player character to select the next action to be performed in the first simulated game and the second simulated game may be the same as the strategy for the non-player character to select the next action to be performed in the training game.
[0075] S204: Determine game data change information of the non-player character according to the second game data and the first game data.
[0076] When a non-player character is playing a simulated game, the game operation results of the non-player character in the simulated game can be transmitted to the database in real time, and the database can automatically perform data statistics to obtain the game data of each character in the simulated game at different levels.
[0077] Before conducting the second simulated game, the developer can first modify the game parameters. After the second simulated game is completed, the game data change information of the non-player character can be determined based on the second game data of the second simulated game and the first game data of the first simulated game of the non-player character. If the game data changes significantly before and after the modification, it means that the parameter modification has affected the balance of the game, and the game parameters can be modified again at this time.
[0078] Among them, the game data change information may include the change in win rate and ranking of the same character at the same rank before and after the game parameters are modified.
[0079] S205: Determine the game state of the asymmetric game based on the game data change information, and determine whether to modify the second game parameter based on the game state; the game state includes a balanced state and an unbalanced state.
[0080] Among them, when the game is in a balanced state, on the premise of fully understanding the rules of the game, n participants of equal level use various characters of the same rank to participate in the game against each other, and the difference in the winning rate of each participant is less than the preset threshold.
[0081] In a game where n participants of equal skill levels fully understand the rules of the game and use various characters of the same level to compete against each other, if the difference in the winning rate of at least one participant is greater than a preset threshold, the game can be considered to be in an unbalanced state.
[0082] Optionally, based on the game data change information, the developer can learn the impact of the modified parameters on the game characters and determine the game state of the current asymmetric game. When the asymmetric game is in an unbalanced state, continue to modify the game parameters and execute the above S203-S204 steps again until the game is in a balanced state.
[0083] As a possible implementation, the game parameters of a game character may be modified based on the game data change information of the game character in all simulated games to ensure that the character does not affect the balance of the game.
[0084] It is worth noting that the above steps S202-S205 can be iterated multiple times, ultimately achieving the dual goals of modifying game parameters without affecting the overall balance of the game, thereby optimizing the balance of the game.
[0085] In the embodiment of the present application, a non-player character is trained through a training match to obtain a non-player character of all ranks and roles. After modifying the game parameters, the trained non-player character is used to perform a first simulated match to obtain first match data, and then the game parameters are modified and a second simulated match is performed to obtain second match data. Finally, match data change information is determined based on the first match data and the second match data, thereby determining the game state of the asymmetric game based on the match data change information and determining whether to modify the game parameters. The present application achieves the dual goals of modifying the game parameters without affecting the overall balance of the game, thereby optimizing the balance of the game.
[0086] By training non-player characters in training games and obtaining a behavior Q-value table for the non-player characters based on the game operation results of the non-player characters in the training games, non-player characters of all ranks and all characters that are closer to the level of real players can be obtained. The behavior Q-value tables of non-player characters between different versions are similar, so they can be reused to a certain extent.
[0087] By using trained non-player characters to simulate games, not only can the game process of players of different levels be simulated, but also a large amount of real and objective game data can be obtained in a short period of time, and problems and flaws in the game can be quickly discovered in a short period of time, so as to facilitate timely repair and optimization. After using non-player characters to simulate games to obtain a large amount of game data, the balance of the game can be quickly measured based on these game data and the data before modifying the parameters, thereby improving the efficiency and effectiveness of game balance optimization.
[0088] Next, the process of obtaining the behavior Q value table of the non-player character according to the game operation results of the non-player character in each training game is described. Figure 3 As shown, the above step S201 includes:
[0089] S301: Determine an action to be performed of the non-player character according to a current state of the non-player character in a current training game and action source information, wherein the action source information includes: a current behavior Q value table or an action selection strategy, and the current behavior Q value table is used to characterize the Q value corresponding to the state action of the non-player character in the current training game.
[0090] Optionally, the non-player character may select one action to be executed from multiple actions in the current state based on the current behavior Q-value table or action selection strategy.
[0091] The current behavior Q value table includes each state, each action corresponding to each state, and the Q value of each action in the current training game.
[0092] The action selection strategy can be to randomly select an action from various actions in the current state as the action to be executed with a certain probability.
[0093] It should be noted that during the training process, in order to more comprehensively explore new states and actions, non-player characters can randomly determine action source information with a certain probability, that is, non-player characters can randomly select the actions to be performed by the non-player characters according to the current behavior Q value table or according to the action selection strategy.
[0094] S302: Triggering the non-player character to perform the action to be performed.
[0095] In the training game, after the electronic device determines the action to be executed of the non-player character, the non-player character can execute the action to be executed and obtain the next state after executing the action to be executed.
[0096] S303: Determine an updated Q value corresponding to the action to be executed according to the result of the non-player character executing the action to be executed, and update the current behavior Q value table according to the updated Q value corresponding to the action to be executed.
[0097] Optionally, the result of the non-player character performing the action to be performed may be the state of the non-player character after performing the action to be performed, including the position of the non-player character in the game map after performing the action to be performed, the survival status of the character, the health of the character, and the skill cooldown time of the character.
[0098] According to the result of the non-player character executing the action to be executed, the Q value of the action to be executed can be determined. For example, referring to Table 1, assuming that after the non-player character executes the action "Move Up" in the current state, the survival state of the character changes from alive to dead, which means that the Q value of the action to be executed in the current state is low. At this time, the updated Q value of the action to be executed can be determined based on the current state and the result of executing the action to be executed, and the updated Q value is used as the Q value of the action "Move Up" in the S1 state of the current behavior Q value table.
[0099] S304: When the current training game ends, an updated current behavior Q value table is obtained, and based on the updated current behavior Q value table, a current behavior Q value table for the next training game is obtained.
[0100] In the training game, each time an action is performed, the above step S303 can be executed to update the current behavior Q value table. At the end of the current training game, the current behavior Q value table of each non-player character in the training game can be obtained.
[0101] It is worth noting that the current behavior Q value table of the non-player character refers to the current behavior Q value table of the character used by the non-player character at the current level.
[0102] It should be understood that the results of the same action in the same state may be random. Therefore, in this application, the same character at the same level can be trained multiple times, and the current behavior Q value table can be iterated to obtain the behavior Q value table of the character at the level.
[0103] Exemplarily, after the first training game is over, the behavior Q value table obtained from the first training game can be used as the current behavior Q value table for the second training game, and after the second training game is over, the updated behavior Q value table can be used as the current behavior Q value table for the third training game.
[0104] S305: Determine whether the current behavior Q value table meets the iteration termination condition. If so, use the updated current behavior Q value table obtained from the last training game as the behavior Q value table of the non-player character.
[0105] Optionally, the current behavior Q value table satisfies the iteration termination condition that the Q value in the behavior Q value table satisfies a preset convergence condition, or the number of training games reaches a preset number of iterations. At this time, the updated current behavior Q value table obtained from the last training game can be used as the behavior Q value table of the non-player character, and the behavior Q value table can be stored in the corresponding storage space of the character rank currently used by the non-player character.
[0106] The preset convergence condition may be that after multiple iterations of updating, the change in the Q value is less than a preset change value.
[0107] In the embodiment of the present application, by determining and executing actions based on the behavior Q value table or action selection strategy in the training game, it is possible to fully explore each action in each state, thereby achieving a more comprehensive training effect. By continuously iteratively optimizing the training game, the authenticity and reliability of the data can be improved, and the problem of incomplete data caused by a small number of iterations can be avoided.
[0108] The following is a further explanation of the above-mentioned determination of the pending actions of the non-player character based on the current state of the non-player character in the current training game and the action source information, such as Figure 4 As shown, the above step S301 includes:
[0109] S401: If the action source information is the current behavior Q value table, the Q value of each action in the current state is read from the current behavior Q value table.
[0110] S402: Determine the actions to be performed by the non-player character according to the Q value of each action in the current state.
[0111] As a possible implementation, if the non-player character determines the action to be executed according to the current behavior Q-value table, the non-player character can read the Q-value of each action in the current state from the current behavior Q-value table, and select the optimal action in the current state as the action to be executed in each state, so that the robot can learn the game strategies of different characters.
[0112] Exemplarily, the action with the highest Q value in the behavior Q value table in the current state may be used as the action to be executed.
[0113] After the action to be executed is determined, the selected action to be executed can be executed, and the Q value of the action to be executed in the current state in the behavior Q value table is updated.
[0114] It should be noted that if the selection of the actions to be executed is performed only according to the above steps S401-S402, it may happen that only one action with the highest Q value is selected each time, resulting in incomplete exploration. Therefore, as another possible implementation method, the present application can also determine the actions to be executed of non-player characters based on the action selection strategy. The selection of actions to be executed based on the action strategy and the selection of actions to be executed based on the behavior Q value table can be performed alternately or randomly, and the present application does not impose any restrictions on this.
[0115] When the action source information is an action selection strategy, the above step S301 includes:
[0116] If the action source information is an action selection strategy, the action to be performed by the non-player character is determined based on the action selection strategy.
[0117] Optionally, the action selection strategy may be an ε-greedy strategy, that is, randomly selecting one of the actions corresponding to the current state as the action to be executed with a certain probability, so as to explore new states and actions.
[0118] After determining the action to be executed, the updated Q value corresponding to the action to be executed can be determined based on the result of the non-player character executing the action to be executed, such as Figure 5 As shown, the above step S303 includes:
[0119] S501: Determine the next state of the non-player character and the reward value of the next state according to the result of the non-player character performing the action to be performed.
[0120] After the non-player character executes the action to be executed, the position of the non-player character after execution, the blood volume, survival status and skill cooldown of the non-player character can be obtained, that is, the next state of the robot and each executable action of the next state can be obtained.
[0121] Optionally, the reward value of the next state of the non-player character may be a reward value of the action to be performed calculated based on a reward function after the non-player character performs the action to be performed.
[0122] The reward function may determine the reward value of the action to be performed based on the next state caused by the action to be performed.
[0123] S502: Determine an updated Q value corresponding to the action to be executed based on the next state of the non-player character, the reward value of the next state, the Q value of the action to be executed in the next state in the current behavior Q value table, and the Q value of the action to be executed in the current state in the current behavior Q value table.
[0124] After executing the action to be executed and determining the next state of the non-player character, the action to be executed in the next state can be determined among the various actions in the next state. The method for determining the action to be executed in the next state can be the same as the above-mentioned step S301, which is not repeated in this application.
[0125] It should be noted that the behavior Q value table in the present application is continuously updated during the training game of the non-player character. Therefore, after multiple training games, the Q value of the action in each state in the behavior Q value table has a historical value, and the Q value of the action to be executed in the next state in the current behavior Q value table can be the historical value of the action to be executed in the next state.
[0126] Exemplarily, for each time step t, the robot observes the current state S_t, selects an action to be executed A_t according to the current behavior Q value table and the ε-greedy strategy, executes the action to be executed and determines the next state S_t+1 and reward R_t+1, and then determines the updated Q value Q'(S_t, A_t) corresponding to the action to be executed according to the next state and the current behavior Q value table, that is:
[0127] Q'(S_t,A_t)=Q(S_t,A_t)+α*[R_t+1+γ*Q(S_t+1,A_t+1)-Q(S_t,A_t)]
[0128] Among them, α is the learning rate, γ is the discount factor, which controls the decay rate of future rewards. Q(S_t+1,A_t+1) represents the Q value of the action to be executed in the next state in the current behavior Q value table, R_t+1 represents the reward value of the next state, and Q(S_t,A_t) represents the Q value of the action to be executed in the current state in the current behavior Q value table.
[0129] By taking the Q value of the action to be executed in the next state in the current behavior Q value table as a calculation factor for the Q value of the action to be executed, it can be ensured that the action to be executed in the next state also affects the Q value of the action to be executed in the current state, thereby ensuring that the Q value of the action to be executed is optimal in combination with the current state and the next state.
[0130] It is worth noting that the above steps S501-S502 are instructions for determining an action to be executed and updating the behavior Q-value table. It should be understood that after updating the behavior Q-value table, the next state can be used as the new current state, and the above steps S501-S502 can be repeated until the end of the training game, and the behavior Q-value table last updated at the end will be used as the behavior Q-value table for the training game.
[0131] The training game includes non-player characters and real players. After the training game is over, in order to avoid a large difference between the training results and the player's needs, the application can also judge the pros and cons of the training game based on the player's feedback information on the training game after the training game is over.
[0132] like Figure 6 As shown, the above step S304 obtains the current behavior Q value table of the next training game according to the updated current behavior Q value table, and also includes:
[0133] S601: Receive training feedback information from real players of the current training game regarding the current training game.
[0134] After the current training game ends, the electronic device can receive training feedback information from the real player of the current training game. The training feedback information can indicate whether the operation of the non-player character in the current training game is reasonable, the difficulty of the current training game, etc.
[0135] S602: Determine whether the updated current behavior Q value table is effective according to the training feedback information.
[0136] After collecting the training feedback information of the real players for the current training game, the electronic device can analyze the training feedback information to determine whether to discard the result of the current training game.
[0137] S603: If yes, the updated current behavior Q value table is used as the current behavior Q value table for the next training game.
[0138] As a possible implementation, if the training feedback information indicates that the difficulty of the current training game is moderate and the non-player character has no unreasonable operations in the current training game, the current behavior Q-value table updated for the current training game can be used as the current behavior Q-value table for the next training game. During the training of the next training game, the action to be executed can be selected based on the behavior Q-value table, and the Q value of the action to be executed can be calculated.
[0139] Exemplarily, unreasonable operations of non-player characters may include: the actual game participation time is shorter than the normal game time, the actions performed by the non-player characters are contrary to common sense, etc.
[0140] S604: If not, the current behavior Q value table is used as the current behavior Q value table for the next training game.
[0141] As another possible implementation, if the training feedback information indicates that the difficulty of the current training game is too easy or too difficult, and the non-player character has unreasonable operations in the current training game, the electronic device can discard the training results of the current training game, that is, instead of using the updated behavior Q-value table of the current training game as the current behavior Q-value table of the next training game, the current behavior Q-value table of the current training game, that is, the updated behavior Q-value table of the previous training game of the current training game, is used as the current behavior Q-value table of the next training game.
[0142] It is worth noting that in order to prevent non-player characters from being too strong or too weak and affecting the player's gaming experience and the outcome of the game, the winning rate of non-player characters in the game can also be limited to be within a preset winning rate range. For example, if the preset winning rate range is 40%-55%, the game data of the non-player characters can be dynamically adjusted based on the game performance of real players, such as the critical hit rate, or the probability of the appearance of hidden props, so that the current game matches the player's gaming level.
[0143] In the embodiment of the present application, by using the training feedback information of the user for the training game as an influencing factor for training the non-player characters, the trained non-player characters can be made closer to real players and better meet the real game needs of the players.
[0144] The following is a further explanation of the above-mentioned determination of the game data change information of the non-player character based on the game operation results of the non-player character in the simulated game and the historical game data of the non-player character. It should be noted that before conducting a simulated game, the developer can first modify the game parameters. For example, if the purpose is to adjust the strength of the new character after the new character is released to ensure the balance of the game, the configuration information of the new character can be modified, such as modifying the damage value, health volume, skill refresh frequency, etc. of the new character.
[0145] In order to improve the reliability of the game data, the non-player character can participate in multiple simulated games, and the game data of the non-player character in each simulated game can be obtained based on the game operation results of the non-player character in each simulated game. The game data of each simulated game is counted to obtain the simulated game data of the character of the rank used by the non-player character in the simulated game.
[0146] The first game data may be statistical information of game data of the non-player character in a first simulated game before the game parameters are modified, including the winning rate, ranking, etc. of the non-player character.
[0147] The second game data may be game data statistics of a second simulated game of a character used by the non-player character in a corresponding rank after the game parameters are modified, including the winning rate and ranking of the character in the rank.
[0148] The following is a further explanation of determining the game data change information of the non-player character based on the first game data and the second game data. Figure 7 As shown, the above step S204 includes:
[0149] S701: Determine a win rate difference between a win rate of the first match data of each non-player character and a win rate of the second match data of each non-player character.
[0150] Optionally, in the present application, the average of the first game data win rate and the average of the second game data win rate of each non-player character can be calculated respectively, and the difference between the average win rate of each non-player character is used as the win rate difference of the non-player character.
[0151] S702: Determine a ranking difference between the ranking of the first game data of each non-player character and the ranking of the second game data of each non-player character.
[0152] Optionally, in the present application, the average of the first game data ranking and the average of the second game data ranking of each non-player character may be calculated respectively, and the difference between the average rankings of each non-player character may be used as the ranking difference of the non-player character.
[0153] S703: Using the win rate difference and / or ranking difference as the game data change information of the non-player character.
[0154] In the present application, the first game data and the second game data of the same character at the same rank can be compared to obtain the win rate difference and ranking difference of the character at the corresponding rank, and the win rate difference and ranking difference can be used as the game data change information of the non-player character.
[0155] It is worth noting that in the present application, the game data change information of non-player characters can be counted for each character in each rank, and the balance of the game can be verified based on the game data change information of all non-player characters.
[0156] It is worth noting that whether the balance of the game is affected by parameter modifications can also be determined through other indicators in the game data, that is, the game data change information can also include change information of other data. This application only gives a possible implementation method using winning rate and ranking as an example, and should not be limited to this.
[0157] After determining the game data change information, the game state of the asymmetric game can be determined based on the game data change information, and it can be determined whether the game parameters need to be modified. The above step S205 includes:
[0158] If the win rate difference is greater than or equal to the preset win rate change threshold, and / or the ranking difference is greater than or equal to the preset ranking change threshold, it is determined that the asymmetric game is in an unbalanced state, and the second game parameters are modified based on the unbalanced state to obtain the third game parameters.
[0159] If the win rate difference is greater than or equal to the preset win rate change threshold, or the ranking difference is greater than or equal to the preset ranking change threshold, or the win rate difference is greater than or equal to the preset win rate change threshold and the ranking difference is greater than or equal to the preset ranking change threshold, then it is determined that the asymmetric game is in an unbalanced state. At this time, the designer can continue to modify the second game parameters to obtain the third game parameters.
[0160] For example, assuming that under the second game parameters, the win rate difference of game character A in the second game parameters is greater than the win rate change difference, the second game parameters can be adjusted or the parameters of game character A can be adjusted separately, such as reducing the single damage value of character A.
[0161] like Figure 8 As shown, after the third game parameter is modified, the method of the present application further includes:
[0162] S801: Obtain third game data of a trained non-player character in a third simulated game of an asymmetric game, where the third simulated game includes third game parameters obtained by modifying the second game parameters.
[0163] Under the third game parameters, a third simulated game consisting of non-player characters may be created, and third game data of the third simulated game may be obtained.
[0164] S802: Determine game data change information of the non-player character according to the third game data and the second game data.
[0165] S803: Determine the game state of the asymmetric game based on the game data change information, and determine whether to modify the third game parameter based on the game state.
[0166] In the present application, the win rate difference and ranking difference in the second game data and the third game data can be calculated to obtain the game data change information of the third simulated game compared with the second simulated game, and then the game state of the asymmetric game under the third game parameters can be determined based on the game data change information.
[0167] If the game state of the asymmetric game is an unbalanced state, the electronic device can output prompt information to prompt the designer to modify the third game parameter, obtain the fourth game parameter, and conduct a fourth simulated game under the fourth game parameter, and then determine the game data change information according to the fourth game data of the fourth simulated game and the third game data, and continue to determine the game state of the asymmetric game based on the game data change information, until the game state of the asymmetric game is a balanced state, and the game parameters after the last modification are determined as the final game parameters of the asymmetric game.
[0168] As another possible implementation, the step of determining the game state of the asymmetric game in S205 above, and determining whether to modify the second game parameter based on the game state further includes:
[0169] If the win rate difference is less than a preset win rate change threshold, and / or the ranking difference is less than a preset ranking change threshold, it is determined that the asymmetric game is in a balanced state, and the second game parameters are not modified based on the balanced state.
[0170] If the win rate difference is less than the preset win rate change threshold, or the ranking difference is less than the preset ranking change threshold, or the win rate difference is less than the preset win rate change threshold and the ranking difference is less than the preset ranking change threshold, then it can be determined that this parameter modification has not affected the balance of the game, and the asymmetric game is in a balanced state. At this time, there is no need to modify the game parameters.
[0171] When the asymmetric game is in a balanced state, the last modified game parameters can be used as official game parameters and run when the game goes online.
[0172] In an embodiment of the present application, a non-player character is first trained through a human-computer training game to obtain a non-player character of all ranks and roles. After the game parameters are modified, a simulated game is performed using the trained non-player character, and game data change information is determined based on the second game data after the parameters are modified and the first game data before the parameters are modified. Therefore, based on the game data change information, it is determined whether the game is in a balanced state and whether the game parameters need to be modified, thereby achieving the dual goals of modifying the game parameters without affecting the overall balance of the game, thereby optimizing the game balance.
[0173] Based on the same inventive concept, the embodiment of the present application also provides an asymmetric game data processing device corresponding to the asymmetric game data processing method. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the above-mentioned asymmetric game data processing method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0174] Fig. 9 A schematic diagram of the structure of a data processing device for an asymmetric game provided in an embodiment of the present application is shown.
[0175] The training module 901 is used to iteratively update the behavior Q value table of the non-player character based on the game results of the non-player character in the training game of the asymmetric game, until the trained non-player character is obtained when the iteration termination condition is met; the behavior Q value table is used to represent the Q value corresponding to the state action of the non-player character;
[0176] A first acquisition module 902 is used to acquire first game data of a trained non-player character in a first simulated game of an asymmetric game; the first simulated game includes first game parameters before parameter modification;
[0177] A second acquisition module 903 is used to acquire second game data of the trained non-player character in a second simulated game of the asymmetric game, where the second simulated game includes second game parameters obtained by modifying the first game parameters;
[0178] A change determination module 904, configured to determine game data change information of a non-player character according to the second game data and the first game data;
[0179] The state determination module 905 is used to determine the game state of the asymmetric game based on the game data change information, so as to determine whether to modify the second game parameter based on the game state; the game state includes a balanced state and an unbalanced state.
[0180] In a feasible implementation scheme, the training module 901 is specifically used for:
[0181] Determine the action to be performed of the non-player character according to the current state of the non-player character in the current training game and the action source information, wherein the action source information includes: a current behavior Q value table or an action selection strategy, the current behavior Q value table is used to characterize the Q value corresponding to the state action of the non-player character in the current training game;
[0182] Trigger non-player characters to perform pending actions;
[0183] Determine an updated Q value corresponding to the action to be executed according to the result of the non-player character executing the action to be executed, and update the current behavior Q value table according to the updated Q value corresponding to the action to be executed;
[0184] At the end of the current training game, an updated current behavior Q value table is obtained, and based on the updated current behavior Q value table, a current behavior Q value table for the next training game is obtained;
[0185] Determine whether the current behavior Q value table meets the iteration termination condition. If so, use the updated current behavior Q value table obtained from the last training game as the behavior Q value table of the non-player character.
[0186] In a feasible implementation scheme, the training module 901 is specifically used for:
[0187] If the Q value in the current behavior Q value table converges, or the number of iterations of the current behavior Q value table is equal to the preset number of iterations, the current behavior Q value table meets the iteration termination condition, and the updated current behavior Q value table obtained from the last training game is used as the behavior Q value table for non-player characters.
[0188] In a feasible implementation scheme, the training module 901 is specifically used for:
[0189] If the action source information is the current behavior Q value table, then read the Q value of each action in the current state from the current behavior Q value table;
[0190] Determine the actions to be performed by the non-player character based on the Q value of each action in the current state.
[0191] In a feasible implementation scheme, the training module 901 is specifically used for:
[0192] If the action source information is an action selection strategy, the action to be performed by the non-player character is determined based on the action selection strategy.
[0193] In a feasible implementation scheme, the training module 901 is specifically used for:
[0194] Determining a next state of the non-player character and a reward value of the next state according to a result of the non-player character performing the action to be performed;
[0195] Determine the updated Q value corresponding to the action to be executed based on the next state of the non-player character, the reward value of the next state, the Q value of the action to be executed in the next state in the current behavior Q value table, and the Q value of the action to be executed in the current state in the current behavior Q value table.
[0196] In a feasible implementation scheme, the training module 901 is specifically used for:
[0197] Receive training feedback information from real players of the current training game regarding the current training game;
[0198] Determine whether the updated current behavior Q value table is effective based on the training feedback information;
[0199] If yes, the updated current behavior Q value table is used as the current behavior Q value table for the next training game;
[0200] If not, the current behavior Q value table is used as the current behavior Q value table for the next training game.
[0201] In a feasible implementation manner, both the first game data and the second game data include a winning rate of the non-player character and / or a ranking of the non-player character.
[0202] In a feasible implementation, the change determination module 904 is specifically used to:
[0203] Determine a win rate difference between a win rate of the first match data of each non-player character and a win rate of the second match data of each non-player character;
[0204] Determine a ranking difference between a ranking of the first match data of each non-player character and a ranking of the second match data of each non-player character;
[0205] The win rate difference and / or ranking difference is used as the game data change information of the non-player character.
[0206] In a feasible implementation scheme, the state determination module 905 is specifically used for:
[0207] If the win rate difference is greater than or equal to the preset win rate change threshold, and / or the ranking difference is greater than or equal to the preset ranking change threshold, it is determined that the asymmetric game is in an unbalanced state, and the second game parameters are modified based on the unbalanced state to obtain the third game parameters.
[0208] In a feasible implementation scheme, the state determination module 905 is specifically used for:
[0209] Acquire third game data of the trained non-player character in a third simulated game of the asymmetric game, where the third simulated game includes third game parameters after modifying the second game parameters;
[0210] Determine the game data change information of the non-player character according to the third game data and the second game data;
[0211] Based on the game data change information, the game state of the asymmetric game is determined to determine whether to modify the third game parameter based on the game state.
[0212] In a feasible implementation scheme, the state determination module 905 is specifically used for:
[0213] If the win rate difference is less than a preset win rate change threshold, and / or the ranking difference is less than a preset ranking change threshold, it is determined that the asymmetric game is in a balanced state, and the second game parameters are not modified based on the balanced state.
[0214] By training the game robot in training games and obtaining the game robot's behavioral Q value table based on the game operation results of the game robot in the training games, a game robot with full ranks and full roles that is closer to the level of real players is obtained. After obtaining the game robot, the trained game robot is used for simulated games, which can not only simulate the game process of players of different levels, but also obtain a large amount of real and objective game data in a short time, and quickly discover problems and defects in the game in a short time, so as to facilitate timely repair and optimization. After using the game robot to simulate the game to obtain a large amount of game data, the balance of the game can be quickly measured based on these game data and the data before modifying the parameters, thereby improving the efficiency and effect of game balance optimization.
[0215] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown, including: a processor 1001, a storage medium 1002 and a bus 1003, wherein the storage medium 1002 stores machine-readable instructions executable by the processor 1001. When the electronic device runs a data processing method for an asymmetric game in the embodiment, the processor 1001 communicates with the storage medium 1002 via the bus 1003, and the processor 1001 executes the machine-readable instructions. The processor 1001 performs the preamble of the method item to perform the following steps:
[0216] Based on the game results of the non-player character in the training game of the asymmetric game, iteratively updating the behavior Q value table of the non-player character until the trained non-player character is obtained when the iteration termination condition is met; the behavior Q value table is used to represent the Q value corresponding to the state action of the non-player character;
[0217] Acquire first game data of a trained non-player character in a first simulated game of an asymmetric game; the first simulated game includes first game parameters before the parameters are modified;
[0218] Acquire second game data of the trained non-player character in a second simulated game of the asymmetric game, where the second simulated game includes second game parameters after modifying the first game parameters;
[0219] Determine the game data change information of the non-player character according to the second game data and the first game data;
[0220] Based on the game data change information, the game state of the asymmetric game is determined to determine whether to modify the second game parameter based on the game state; the game state includes a balanced state and an unbalanced state.
[0221] In a feasible implementation, the processor 1001 is specifically configured to iteratively update the behavior Q value table of the non-player character based on the game game results of the non-player character in the training game of the asymmetric game until the trained non-player character is obtained when the iteration termination condition is met:
[0222] Determine the action to be performed of the non-player character according to the current state of the non-player character in the current training game and the action source information, wherein the action source information includes: a current behavior Q value table or an action selection strategy, the current behavior Q value table is used to characterize the Q value corresponding to the state action of the non-player character in the current training game;
[0223] Trigger non-player characters to perform pending actions;
[0224] Determine an updated Q value corresponding to the action to be executed according to the result of the non-player character executing the action to be executed, and update the current behavior Q value table according to the updated Q value corresponding to the action to be executed;
[0225] At the end of the current training game, an updated current behavior Q value table is obtained, and based on the updated current behavior Q value table, a current behavior Q value table for the next training game is obtained;
[0226] Determine whether the current behavior Q value table meets the iteration termination condition. If so, use the updated current behavior Q value table obtained from the last training game as the behavior Q value table of the non-player character.
[0227] In a feasible implementation, when the processor 1001 determines whether the current behavior Q value table satisfies the iteration termination condition, and if so, uses the updated current behavior Q value table obtained in the last training game as the behavior Q value table of the non-player character, it is specifically configured to:
[0228] If the Q value in the current behavior Q value table converges, or the number of iterations of the current behavior Q value table is equal to the preset number of iterations, the current behavior Q value table meets the iteration termination condition, and the updated current behavior Q value table obtained from the last training game is used as the behavior Q value table for non-player characters.
[0229] In a feasible implementation, when the processor 1001 determines the action to be performed by the non-player character according to the current state of the non-player character in the current training game and the action source information, it is specifically configured to:
[0230] If the action source information is the current behavior Q value table, then read the Q value of each action in the current state from the current behavior Q value table;
[0231] Determine the actions to be performed by the non-player character based on the Q value of each action in the current state.
[0232] In a feasible implementation, when the processor 1001 determines the action to be performed by the non-player character according to the current state of the non-player character in the current training game and the action source information, it is specifically configured to:
[0233] If the action source information is an action selection strategy, the action to be performed by the non-player character is determined based on the action selection strategy.
[0234] In a feasible implementation, when the processor 1001 determines the updated Q value corresponding to the to-be-executed action according to the result of the non-player character executing the to-be-executed action, it is specifically configured to:
[0235] Determining a next state of the non-player character and a reward value of the next state according to a result of the non-player character performing the action to be performed;
[0236] According to the next state of the non-player character, the reward value of the next state, the Q value of the action to be executed in the next state in the current behavior Q value table, and the Q value of the action to be executed in the current state in the current behavior Q value table, the updated Q value corresponding to the action to be executed is determined.
[0237] In a feasible implementation manner, when the processor 1001 obtains the current behavior Q value table of the next training game according to the updated current behavior Q value table, it is specifically configured to:
[0238] Receive training feedback information from real players of the current training game regarding the current training game;
[0239] Determine whether the updated current behavior Q value table is effective based on the training feedback information;
[0240] If yes, the updated current behavior Q value table is used as the current behavior Q value table for the next training game;
[0241] If not, the current behavior Q value table is used as the current behavior Q value table for the next training game.
[0242] In a feasible implementation manner, when the processor 1001 determines the game data change information of the non-player character according to the second game data and the first game data, it is specifically configured to:
[0243] Determine a win rate difference between a win rate of the first match data of each non-player character and a win rate of the second match data of each non-player character;
[0244] Determine a ranking difference between a ranking of the first match data of each non-player character and a ranking of the second match data of each non-player character;
[0245] The win rate difference and / or ranking difference is used as the game data change information of the non-player character.
[0246] In a feasible implementation, when the processor 1001 determines the game state of the asymmetric game based on the game data change information, and determines whether to modify the second game parameter based on the game state, it is specifically configured to:
[0247] If the win rate difference is greater than or equal to the preset win rate change threshold, and / or the ranking difference is greater than or equal to the preset ranking change threshold, it is determined that the asymmetric game is in an unbalanced state, and the second game parameters are modified based on the unbalanced state to obtain the third game parameters.
[0248] In one feasible implementation, the processor 1001 is further configured to:
[0249] Acquire third game data of the trained non-player character in a third simulated game of the asymmetric game, where the third simulated game includes third game parameters after modifying the second game parameters;
[0250] Determine the game data change information of the non-player character according to the third game data and the second game data;
[0251] Based on the game data change information, the game state of the asymmetric game is determined to determine whether to modify the third game parameter based on the game state.
[0252] In a feasible implementation, when the processor 1001 determines the game state of the asymmetric game based on the game data change information, and determines whether to modify the second game parameter based on the game state, it is specifically configured to:
[0253] If the win rate difference is less than a preset win rate change threshold, and / or the ranking difference is less than a preset ranking change threshold, it is determined that the asymmetric game is in a balanced state, and the second game parameters are not modified based on the balanced state.
[0254] By training non-player characters in training games and updating the behavior Q value table of the non-player characters according to the game operation results of the non-player characters in the training games, non-player characters of all levels and roles that are closer to the level of real players are obtained. After obtaining the non-player characters, the trained non-player characters are used to perform simulated games, which can not only simulate the game process of players of different levels, but also obtain a large amount of real and objective game data in a short time, and quickly discover problems and defects in the game in a short time, so as to facilitate timely repair and optimization. After using non-player characters to perform simulated games to obtain a large amount of game data, the balance of the game can be quickly measured based on these game data and the data before modifying the parameters, thereby improving the efficiency and effect of game balance optimization.
[0255] The embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. The computer program is executed when a processor is running, and the processor performs the following steps:
[0256] Based on the game results of the non-player character in the training game of the asymmetric game, iteratively updating the behavior Q value table of the non-player character until the trained non-player character is obtained when the iteration termination condition is met; the behavior Q value table is used to represent the Q value corresponding to the state action of the non-player character;
[0257] Acquire first game data of a trained non-player character in a first simulated game of an asymmetric game; the first simulated game includes first game parameters before the parameters are modified;
[0258] Acquire second game data of the trained non-player character in a second simulated game of the asymmetric game, where the second simulated game includes second game parameters after modifying the first game parameters;
[0259] Determine the game data change information of the non-player character according to the second game data and the first game data;
[0260] Based on the game data change information, the game state of the asymmetric game is determined to determine whether to modify the second game parameter based on the game state; the game state includes a balanced state and an unbalanced state.
[0261] In a feasible implementation, when the processor executes the game game results of the non-player character in the training game of the asymmetric game, iteratively updates the behavior Q value table of the non-player character until the trained non-player character is obtained when the iteration termination condition is met, it is specifically configured to:
[0262] Determine the action to be performed of the non-player character according to the current state of the non-player character in the current training game and the action source information, wherein the action source information includes: a current behavior Q value table or an action selection strategy, the current behavior Q value table is used to characterize the Q value corresponding to the state action of the non-player character in the current training game;
[0263] Trigger non-player characters to perform pending actions;
[0264] Determine an updated Q value corresponding to the action to be executed according to the result of the non-player character executing the action to be executed, and update the current behavior Q value table according to the updated Q value corresponding to the action to be executed;
[0265] At the end of the current training game, an updated current behavior Q value table is obtained, and based on the updated current behavior Q value table, a current behavior Q value table for the next training game is obtained;
[0266] Determine whether the current behavior Q value table meets the iteration termination condition. If so, use the updated current behavior Q value table obtained from the last training game as the behavior Q value table of the non-player character.
[0267] In a feasible implementation scheme, when the processor determines whether the current behavior Q value table satisfies the iteration termination condition, and if so, uses the updated current behavior Q value table obtained from the last training game as the behavior Q value table of the non-player character, it is specifically used to:
[0268] If the Q value in the current behavior Q value table converges, or the number of iterations of the current behavior Q value table is equal to the preset number of iterations, the current behavior Q value table meets the iteration termination condition, and the updated current behavior Q value table obtained from the last training game is used as the behavior Q value table for non-player characters.
[0269] In a feasible implementation, when the processor determines the action to be performed by the non-player character according to the current state of the non-player character in the current training game and the action source information, it is specifically configured to:
[0270] If the action source information is the current behavior Q value table, then read the Q value of each action in the current state from the current behavior Q value table;
[0271] Determine the actions to be performed by the non-player character based on the Q value of each action in the current state.
[0272] In a feasible implementation, when the processor determines the action to be performed by the non-player character according to the current state of the non-player character in the current training game and the action source information, it is specifically configured to:
[0273] If the action source information is an action selection strategy, the action to be performed by the non-player character is determined based on the action selection strategy.
[0274] In a feasible implementation scheme, when the processor determines the updated Q value corresponding to the to-be-executed action according to the result of the non-player character executing the to-be-executed action, it is specifically configured to:
[0275] Determining a next state of the non-player character and a reward value of the next state according to a result of the non-player character performing the action to be performed;
[0276] Determine the updated Q value corresponding to the action to be executed based on the next state of the non-player character, the reward value of the next state, the Q value of the action to be executed in the next state in the current behavior Q value table, and the Q value of the action to be executed in the current state in the current behavior Q value table.
[0277] In a feasible implementation scheme, when the processor executes the process of obtaining the current behavior Q value table of the next training game according to the updated current behavior Q value table, the processor is specifically configured to:
[0278] Receive training feedback information from real players of the current training game regarding the current training game;
[0279] Determine whether the updated current behavior Q value table is effective based on the training feedback information;
[0280] If yes, the updated current behavior Q value table is used as the current behavior Q value table for the next training game;
[0281] If not, the current behavior Q value table is used as the current behavior Q value table for the next training game.
[0282] In a feasible implementation manner, when the processor determines the game data change information of the non-player character according to the second game data and the first game data, it is specifically configured to:
[0283] Determine a win rate difference between a win rate of the first match data of each non-player character and a win rate of the second match data of each non-player character;
[0284] Determine a ranking difference between a ranking of the first match data of each non-player character and a ranking of the second match data of each non-player character;
[0285] The win rate difference and / or ranking difference is used as the game data change information of the non-player character.
[0286] In a feasible implementation, when the processor determines the game state of the asymmetric game based on the game data change information, and determines whether to modify the second game parameter based on the game state, the processor is specifically configured to:
[0287] If the win rate difference is greater than or equal to the preset win rate change threshold, and / or the ranking difference is greater than or equal to the preset ranking change threshold, it is determined that the asymmetric game is in an unbalanced state, and the second game parameters are modified based on the unbalanced state to obtain the third game parameters.
[0288] In one possible embodiment, the processor is further configured to:
[0289] Acquire third game data of the trained non-player character in a third simulated game of the asymmetric game, where the third simulated game includes third game parameters after modifying the second game parameters;
[0290] Determine the game data change information of the non-player character according to the third game data and the second game data;
[0291] Based on the game data change information, the game state of the asymmetric game is determined to determine whether to modify the third game parameter based on the game state.
[0292] In a feasible implementation, when the processor determines the game state of the asymmetric game based on the game data change information, and determines whether to modify the second game parameter based on the game state, the processor is specifically configured to:
[0293] If the win rate difference is less than a preset win rate change threshold, and / or the ranking difference is less than a preset ranking change threshold, it is determined that the asymmetric game is in a balanced state, and the second game parameters are not modified based on the balanced state.
[0294] By training non-player characters in training games, and obtaining a behavioral Q value table of the non-player characters based on the game operation results of the non-player characters in the training games, non-player characters of all levels and roles that are closer to the level of real players are obtained. After obtaining the non-player characters, the trained non-player characters are used to perform simulated games, which can not only simulate the game process of players of different levels, but also obtain a large amount of real and objective game data in a short time, and quickly discover problems and defects in the game in a short time, so as to facilitate timely repair and optimization. After obtaining a large amount of game data through simulated games using non-player characters, the balance of the game can be quickly measured based on these game data and the data before modifying the parameters, thereby improving the efficiency and effect of game balance optimization.
[0295] In the embodiment of the present application, the computer program can also execute other machine-readable instructions when run by the processor to execute other methods described in the embodiment. For the specific execution method steps and principles, please refer to the description of the embodiment, which will not be repeated here.
[0296] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0297] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0298] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0299] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0300] It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance.
[0301] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The protection scope of the present application is not limited thereto. Although the present application is described in detail with reference to the above-mentioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-mentioned embodiments within the technical scope disclosed in the present application, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application. They should all be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A data processing method for an asymmetric game, characterized in that: include: Iteratively updating a behavior Q value table of the non-player character based on a game match result of the non-player character in a training match of the asymmetric game until the trained non-player character is obtained when an iteration termination condition is satisfied; the behavior Q value table is used to represent the Q value corresponding to the state action of the non-player character; Acquire first game data of the trained non-player character in a first simulated game of the asymmetric game; The first simulated game includes first game parameters before the parameters are modified; Acquire second game data of the trained non-player character in a second simulated game of the asymmetric game, wherein the second simulated game includes second game parameters obtained by modifying the first game parameters; Determining game data change information of the non-player character according to the second game data and the first game data; Determining a game state of the asymmetric game based on the game data change information, so as to determine whether to modify the second game parameter based on the game state; The game state includes a balanced state and an unbalanced state.
2. The method according to claim 1, characterized in that The iterative updating of the behavior Q value table of the non-player character based on the game match result of the non-player character in the training match of the asymmetric game until the trained non-player character is obtained when an iteration termination condition is satisfied includes: Determining the action to be performed of the non-player character according to the current state of the non-player character in the current training game and action source information, wherein the action source information includes: a current behavior Q value table or an action selection strategy, wherein the current behavior Q value table is used to characterize the Q value corresponding to the state action of the non-player character in the current training game; Triggering the non-player character to perform the action to be performed; Determine, according to the result of the non-player character executing the action to be executed, an updated Q value corresponding to the action to be executed, and update the current behavior Q value table according to the updated Q value corresponding to the action to be executed; When the current training game ends, an updated current behavior Q value table is obtained, and according to the updated current behavior Q value table, a current behavior Q value table for the next training game is obtained; Determine whether the current behavior Q value table meets the iteration termination condition. If so, use the updated current behavior Q value table obtained in the last training game as the behavior Q value table of the non-player character.
3. The method according to claim 2, characterized in that The step of determining whether the current behavior Q value table satisfies the iteration termination condition includes: If the Q value in the current behavior Q value table meets the preset convergence condition, or the number of iteration updates is equal to the preset number of iterations, it is determined that the current behavior Q value table meets the iteration termination condition.
4. The method according to claim 2, characterized in that: The step of determining the action to be performed of the non-player character according to the current state of the non-player character in the current training game and the action source information includes: If the action source information is the current behavior Q value table, then read the Q value of each action in the current state from the current behavior Q value table; The actions to be performed by the non-player character are determined according to the Q values of the actions in the current state.
5. The method according to claim 2, characterized in that: The step of determining the action to be performed of the non-player character according to the current state of the non-player character in the current training game and the action source information includes: If the action source information is the action selection strategy, the action to be performed by the non-player character is determined based on the action selection strategy.
6. The method according to claim 2, characterized in that The step of determining, according to the result of the non-player character executing the action to be executed, an updated Q value corresponding to the action to be executed, comprises: Determining a next state of the non-player character and a reward value of the next state according to a result of the non-player character performing the to-be-performed action; Determine the updated Q value corresponding to the action to be executed according to the next state of the non-player character, the reward value of the next state, the Q value of the action to be executed in the current behavior Q value table in the next state, and the Q value of the action to be executed in the current behavior Q value table in the current state.
7. The method according to claim 2, characterized in that The step of obtaining the current behavior Q value table for the next training game according to the updated current behavior Q value table includes: Receiving training feedback information from real players of the current training game regarding the current training game; Determine whether the updated current behavior Q value table is effective according to the training feedback information; If yes, the updated current behavior Q value table is used as the current behavior Q value table for the next training game; If not, the current behavior Q value table is used as the current behavior Q value table for the next training game.
8. The method according to any one of claims 1 to 7, characterized in that: The first game data and the second game data both include a winning rate of the non-player character and / or a ranking of the non-player character.
9. The method according to claim 8, characterized in that The determining the game data change information of the non-player character according to the second game data and the first game data includes: Determine a win rate difference between a win rate of the first match data of each non-player character and a win rate of the second match data of each non-player character; Determine a ranking difference between a ranking of the first game data of each non-player character and a ranking of the second game data of each non-player character; The win rate difference and / or the ranking difference is used as the game data change information of the non-player character.
10. The method according to claim 9, characterized in that The determining the game state of the asymmetric game based on the game data change information, and determining whether to modify the second game parameter based on the game state, includes: If the win rate difference is greater than or equal to a preset win rate change threshold, and / or the ranking difference is greater than or equal to a preset ranking change threshold, it is determined that the asymmetric game is in the unbalanced state, and the second game parameter is modified based on the unbalanced state to obtain a third game parameter; The method further comprises: Acquire third game data of the trained non-player character in a third simulated game of the asymmetric game, wherein the third simulated game includes third game parameters obtained by modifying the second game parameters; Determining game data change information of the non-player character according to the third game data and the second game data; Based on the game data change information, a game state of the asymmetric game is determined, so as to determine whether to modify the third game parameter based on the game state.
11. The method according to claim 9, characterized in that The determining the game state of the asymmetric game based on the game data change information, and determining whether to modify the second game parameter based on the game state, includes: If the winning rate difference is less than the preset winning rate change threshold, and / or the ranking difference is less than the preset ranking change threshold, it is determined that the asymmetric game is in the equilibrium state, and the second game parameter is not modified based on the equilibrium state.
12. A data processing device for an asymmetric game, characterized in that: include: A training module, configured to iteratively update a behavior Q value table of the non-player character based on a game match result of the non-player character in a training match of the asymmetric game, until the trained non-player character is obtained when an iteration termination condition is satisfied; the behavior Q value table is used to represent a Q value corresponding to a state action of the non-player character; A first acquisition module is used to acquire first game data of the trained non-player character in a first simulated game of the asymmetric game; the first simulated game includes first game parameters before parameter modification; A second acquisition module is used to acquire second game data of the trained non-player character in a second simulated game of the asymmetric game, wherein the second simulated game includes second game parameters obtained by modifying the first game parameters; a change determination module, configured to determine game data change information of the non-player character according to the second game data and the first game data; A state determination module, configured to determine a game state of the asymmetric game based on the game data change information, so as to determine whether to modify the second game parameter based on the game state; The game state includes a balanced state and an unbalanced state.
13. An electronic device, characterized in that: include: A processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the storage medium communicate via the bus, and the processor executes the machine-readable instructions to perform the steps of a data processing method for an asymmetric game as described in any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the data processing method for an asymmetric game as claimed in any one of claims 1 to 11 are executed.
Citation Information
Patent Citations
Method for learning non-player character combat strategies on basis of deep Q-learning networks
CN108211362A
Game map balance test method, device, equipment and storage medium
CN110489340A
Game balance test method and device, electronic equipment and storage medium
CN115463415A
Game artificial intelligence control method and device, equipment and storage medium
CN116943220A
Systems and methods for test deployment of computational code on virtual servers
KR102282499B1