Parameter adjustment method, device and equipment for asymmetric competitive game and storage medium

By applying reinforcement learning and Q-Learning algorithms in asymmetric competitive games and adjusting game balance parameters, the problem of poor balance optimization effects in the existing technology is solved, and more efficient game balance optimization and player needs are achieved.

CN120094210APending Publication Date: 2025-06-06NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311669902.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art has problems such as data quality, algorithm complexity, user feedback and game planning subjectivity when automating the optimization of balance in asymmetric competitive games, resulting in poor balance optimization effects.

Method used

The reinforcement learning method is adopted and combined with the Q-Learning algorithm to adjust the game balance parameters of asymmetric competitive games, and quickly optimize the game parameters through a large amount of game game data to ensure the differentiated needs of players and improve the effect of balance optimization.

Benefits of technology

Through the combination of reinforcement learning and Q-Learning algorithm, game parameters can be quickly optimized, balanced optimization effects can be improved, players' differentiated needs can be met, and subjective intervention by game planners can be reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120094210A_ABST
    Figure CN120094210A_ABST
Patent Text Reader

Abstract

The invention provides a parameter adjustment method, device and equipment for an asymmetric competitive game and a storage medium, and the method comprises the steps: adjusting a game balance parameter of the asymmetric competitive game based on a Q-Learning algorithm, and obtaining a game log of at least one game match performed by a player based on the adjusted game balance parameter; determining a balance coefficient corresponding to the adjusted game balance parameter based on the game log of at least one game match, and updating the Q value table based on the balance coefficient; in response to the condition of continuing to adjust the game balance parameter, adjusting the game balance parameter of the asymmetric competitive game based on a Q-Learning algorithm; and determining a final game balance parameter based on the updated Q value table in response to the condition that the condition of continuing to adjust the game balance parameter is not met. The game parameters are rapidly optimized in a reinforcement learning mode in combination with a large number of game games, the differentiated requirements of players can be guaranteed, and the balance optimization effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic game technology, and in particular to a method, device, equipment and storage medium for adjusting parameters of an asymmetric competitive game. Background Art

[0002] Asymmetric competitive games are a type of multiplayer game in which different players play different roles with different skills, goals, and abilities. The goals and abilities of these players are different, so the balance and difficulty of the game often need to be carefully designed to be guaranteed.

[0003] At present, the technical solutions for automatic optimization of balance in asymmetric competitive games mainly include machine learning, data mining, simulation and user feedback, etc. Different solutions can be used in combination to achieve better optimization effects.

[0004] However, although the existing technical solutions can be used for automatic optimization of balance in asymmetric competitive games, they still have some shortcomings:

[0005] (1) Data quality issues: Data quality has a significant impact on the accuracy and reliability of analysis results. Noise and errors may exist during data collection and processing. Different preferences of players may cause data deviation, which will affect the accuracy of analysis results.

[0006] (2) Algorithm complexity problem: Some machine learning and data mining algorithms require a lot of computing resources and time, which makes the algorithms too complex and difficult to apply in actual games.

[0007] (3) User feedback issues: User feedback is an important source for optimizing balance, but the quality and quantity of user feedback may be problematic, and user feedback may be affected by the user’s personal bias and subjectivity.

[0008] (4) Subjective issues of game planning: It can be seen that many ways to adjust the balance of the game require the participation and decision-making of game planners. If the planners do not understand the game mechanism and the needs of the players, they may modify certain values ​​​​in the wrong way, and ultimately the optimization results may be biased.

[0009] In summary, the above technical solutions have the problem of poor balance optimization effect due to various factors such as data quality, algorithm complexity, user feedback and game planning. Summary of the invention

[0010] In view of this, the present application provides a parameter adjustment method, device, equipment and storage medium for an asymmetric competitive game, which can quickly optimize game parameters by combining a large number of game matches through reinforcement learning, thereby ensuring the differentiated needs of players and improving the effect of balance optimization.

[0011] In a first aspect, an embodiment of the present invention provides a parameter adjustment method for an asymmetric competitive game, the method comprising: adjusting the game balance parameters of the asymmetric competitive game based on a Q-Learning algorithm, and obtaining a game log of a player playing at least one game based on the adjusted game balance parameters; wherein the Q-Learning algorithm includes an adjustment strategy for at least one game balance parameter of the asymmetric competitive game; determining a balance coefficient corresponding to the adjusted game balance parameter based on the game log of at least one game, and updating a Q value table based on the balance coefficient; wherein the Q value table records the adjustment strategy for each adjustment and the degree of balance of the game under the adjusted game balance parameters, and the degree of balance is determined based on the balance coefficient; in response to meeting the condition for continuing to adjust the game balance parameters, adjusting the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm; in response to not meeting the condition for continuing to adjust the game balance parameters, determining the final game balance parameters based on the updated Q value table.

[0012] In a second aspect, an embodiment of the present invention further provides a parameter adjustment device for an asymmetric competitive game, the device comprising: a parameter adjustment module, for adjusting the game balance parameters of the asymmetric competitive game based on a Q-Learning algorithm, and obtaining a game log of a player playing at least one game based on the adjusted game balance parameters; wherein the Q-Learning algorithm includes an adjustment strategy for at least one game balance parameter of the asymmetric competitive game; a Q value table update module, for determining a balance coefficient corresponding to the adjusted game balance parameter based on the game log of at least one game, and updating the Q value table based on the balance coefficient; wherein the Q value table records the adjustment strategy for each adjustment and the degree of balance of the game under the adjusted game balance parameters, and the degree of balance is determined based on the balance coefficient; a parameter continuous adjustment module, for adjusting the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm in response to meeting the conditions for continuing to adjust the game balance parameters; a parameter determination module, for determining the final game balance parameters based on the updated Q value table in response to not meeting the conditions for continuing to adjust the game balance parameters.

[0013] In a third aspect, an embodiment of the present invention further provides an electronic device, including a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of the above-mentioned method for adjusting parameters of an asymmetric competitive game.

[0014] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the steps of the above-mentioned parameter adjustment method for an asymmetric competitive game.

[0015] The embodiments of the present invention bring the following beneficial effects:

[0016] The embodiment of the present invention provides a parameter adjustment method, device, equipment and storage medium for an asymmetric competitive game, which adjusts the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm, obtains the game log of the player playing at least one game based on the adjusted game balance parameters; determines the balance coefficient corresponding to the adjusted game balance parameters based on the game log of at least one game, and updates the Q value table based on the balance coefficient; in response to the condition of continuing to adjust the game balance parameters being met, adjusts the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm; in response to the condition of not continuing to adjust the game balance parameters being met, determines the final game balance parameters based on the updated Q value table. In this method, the game parameters are quickly optimized by combining a large number of game matches through reinforcement learning, which can ensure the differentiated needs of players and improve the effect of balance optimization.

[0017] Other features and advantages of the present disclosure will be set forth in the following description, or some features and advantages may be inferred or unambiguously determined from the description, or may be learned by implementing the above-mentioned technology of the present disclosure.

[0018] In order to make the above-mentioned objectives, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 A flow chart of a method for adjusting parameters of an asymmetric competitive game provided by an embodiment of the present invention;

[0021] Figure 2 A flowchart of another method for adjusting parameters of an asymmetric competitive game provided by an embodiment of the present invention;

[0022] Figure 3A schematic diagram of a method for adjusting parameters of an asymmetric competitive game provided by an embodiment of the present invention;

[0023] Figure 4 A schematic diagram of the structure of a parameter adjustment device for an asymmetric competitive game provided by an embodiment of the present invention;

[0024] Figure 5 A schematic diagram of the structure of another device for adjusting parameters of an asymmetric competitive game provided by an embodiment of the present invention;

[0025] Figure 6 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0027] Asymmetric competitive games are a type of multiplayer game in which different players play different roles with different skills, goals, and abilities. The goals and abilities of these players are different, so the balance and difficulty of the game often need to be carefully designed to be guaranteed.

[0028] For example, in an asymmetric competitive game, players play different roles in the game and fight against monsters played by a single player on the other side. Due to the differences in skills between different characters and monsters, the balance problem caused by the skill combination causes great trouble to the planner. How to ensure the strength of the new game characters while balancing the fairness of the overall game is a very important issue.

[0029] There are already some technical solutions for automatically optimizing the balance of asymmetric competitive games, mainly including the following:

[0030] (1) Machine Learning: Use machine learning techniques to analyze game data, predict balance issues in the game, and propose corresponding optimization solutions. For example, reinforcement learning algorithms can be used to train intelligent agents to automatically optimize game balance.

[0031] (2) Data mining: Game planners use data mining technology to discover patterns and rules in game data, identify balance issues in the game, and optimize the balance through version fine-tuning.

[0032] (3) Simulation: By simulating the game process, different characters and skills are tested and evaluated to identify balance issues in the game and propose corresponding optimization solutions.

[0033] (4) User feedback: Collect user feedback and opinions through a questionnaire system, organize data to analyze user behavior patterns and needs, and use this to find balance issues in the game. Plan to propose corresponding optimization plans based on the problems found.

[0034] In summary, the technical solutions for automatically optimizing the balance of asymmetric competitive games mainly include machine learning, data mining, simulation, and user feedback. Different solutions can be used in combination to achieve better optimization effects.

[0035] However, although the existing technical solutions can be used for automatic optimization of balance in asymmetric competitive games, they still have some shortcomings:

[0036] (1) Data quality issues: Data quality has a significant impact on the accuracy and reliability of analysis results. Noise and errors may exist during data collection and processing. Different preferences of players may cause data deviation, which will affect the accuracy of analysis results.

[0037] For example: The above noise and error can be:

[0038] 1. Errors caused by abnormal log records, log duplication, and log loss.

[0039] 2. Player preferences lead to data imbalance. Specifically, when players play on the official server, there are often popular and unpopular characters. There is a lot of data for popular characters, but very little data for unpopular characters. The smaller amount of data can lead to large differences due to data sparsity and different player habits, and is not of reference value.

[0040] 3. At the same time, since the combinations of different characters may not be complete in the official server, this part of the data is missing.

[0041] (2) Algorithm complexity problem: Some machine learning and data mining algorithms require a lot of computing resources and time, which makes the algorithms too complex and difficult to apply in actual games.

[0042] For example, some machine learning and data mining algorithms are very complex and require a lot of computing resources and time to complete the calculation. These algorithms may need to use more complex models such as multi-layer neural networks and support vector machines, which makes the algorithms themselves very complex. Many parameters lead to slow convergence, which often takes a long time to converge, and ultimately is not suitable for the rapid iteration and update of the game, so it is also important to choose the right machine learning algorithm.

[0043] (3) User feedback issues: User feedback is an important source for optimizing balance, but the quality and quantity of user feedback may be problematic, and user feedback may be affected by the user’s personal bias and subjectivity.

[0044] (4) Subjective issues of game planning: It can be seen that many ways to adjust the balance of the game require the participation and decision-making of game planners. If the planners do not understand the game mechanism and the needs of the players, they may modify certain values ​​​​in the wrong way, and ultimately the optimization results may be biased.

[0045] In summary, existing technical solutions have problems such as data quality, algorithm complexity, user feedback and game planning, and it is necessary to comprehensively consider various factors to optimize the balance.

[0046] Based on this, the embodiments of the present invention provide a parameter adjustment method, device, equipment and storage medium for an asymmetric competitive game, and specifically provide a system design for automatically optimizing the balance of an asymmetric competitive game based on a Q-learning algorithm, which can be applied to asymmetric competitive games.

[0047] In asymmetric competitive games, balance is very important because these games are usually multiplayer online competitive games. If one side is too advantageous or disadvantageous, then players will not be able to obtain a fair competition environment and a pleasant gaming experience in the game. In addition, in asymmetric competitive games, there is a complex and subtle relationship between each role, rule, scene design and other factors. Only by ensuring that various parameters are coordinated and balanced can players enjoy a better experience.

[0048] However, getting the balance right in practice is not easy. Here are a few reasons:

[0049] 1. Variability: Since each player has his or her own unique style and strategy, it is difficult to predict how they will use different types of characters, which makes data collection and analysis difficult.

[0050] 2. Complexity: In asymmetric competitive games, the design of relatively large maps with multiple elements (such as character attributes, equipment systems, etc.) requires developers to consider many details and ensure that all elements match each other reasonably.

[0051] 3. Continuous updates and iterations: Asymmetric competitive products are continuously updated and iterated at a fast speed, and new features may affect old features and even the overall balance.

[0052] 4. Differentiation of player demands: In a large user base, player demands are very differentiated, and developers must meet the demands of players at different levels based on market feedback.

[0053] This embodiment can use reinforcement learning ideas instead of simply relying on the planning experience to adjust the game balance. It uses a large amount of data support to ensure the differentiated needs of players, and combines the Q-Learning algorithm to quickly iterate and obtain the parameter modification configuration that achieves the fastest game balance.

[0054] Specifically, this embodiment designs a random gameplay for the test server, which can make the player data distribution even and reasonable. It uses the Q-Learning algorithm idea in combination with massive player data, takes different parameter changes as execution actions, and measures the balance of corresponding parameters through the data results of real players in actual games. This is used as a reward for training to update the Q value table. Through the iteration of the Q value table update, the minimum parameter modification path to achieve balance is finally obtained.

[0055] To facilitate understanding of this embodiment, a parameter adjustment method for an asymmetric competitive game disclosed in an embodiment of the present invention is first introduced in detail.

[0056] This embodiment provides a method for adjusting parameters of an asymmetric competitive game. The machine learning model in this embodiment can adjust the game balance parameters by means of reinforcement learning.

[0057] Reinforcement learning is a machine learning method based on the interaction between an agent and the environment, which aims to enable the agent to learn the optimal behavior strategy through continuous interaction with the environment. In reinforcement learning, the agent will continuously adjust its behavior strategy based on its own behavior and environmental feedback to achieve the goal of maximizing the cumulative reward. The above-mentioned agent is the machine learning model of this embodiment, and the machine learning model can continuously decide which game balance parameters to modify.

[0058] See also Figure 1 The flowchart of a method for adjusting parameters of an asymmetric competitive game is shown, and the method for adjusting parameters of an asymmetric competitive game comprises the following steps:

[0059] Step S102, adjusting the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm, and obtaining a game log of at least one game match played by the player based on the adjusted game balance parameters.

[0060] The asymmetric competitive game in this embodiment is a multiplayer game type, in which different players control different virtual characters, each character has different skills, goals and abilities, and different players can control the virtual characters to play the game.

[0061] The game balance parameter of an asymmetric competitive game is a parameter that can affect the balance of the asymmetric competitive game. The balance of the asymmetric competitive game can be characterized by the winning rate of different virtual characters. For example, if the change of the control time of the skill of virtual character A affects the winning rate of virtual character A, the control time can be a game balance parameter.

[0062] Among them, the Q-Learning algorithm includes at least one game balance parameter adjustment strategy for asymmetric competitive games. In this embodiment, a Q-Learning algorithm including at least one game balance parameter adjustment strategy can be set, and each time the game balance parameters of the asymmetric competitive game are adjusted, an adjustment strategy for the game balance parameters can be selected, and the game balance parameters can be updated based on the selected game balance parameters.

[0063] After updating the game balance parameters, players can play multiple games based on the adjusted game balance parameters, and can obtain game logs of the game games after the game games.

[0064] Step S104, determining a balance coefficient corresponding to the adjusted game balance parameter based on a game log of at least one game match, and updating the Q value table based on the balance coefficient.

[0065] The Q value table records the adjustment strategy for each adjustment and the balance degree of the game under the adjusted game balance parameters, and the balance degree is determined based on the balance coefficient.

[0066] Referring to a Q value table shown in Table 1, a Q value table can be provided in advance in this embodiment, which records the adjustment strategy of each adjustment (i.e., strategy 1-strategy 3, for example: the adjustment strategy of strategy 3 for the S1th adjustment can be to increase the game balance parameter A by 0.1) and the balance degree of the game under the adjusted game balance parameters (i.e., h1-h3).

[0067] Table 1

[0068] Strategy 1 Strategy 2 Strategy 3 S1 h1 S2 h2 S3 h3

[0069] After obtaining the game logs of the game, the game logs of multiple game matches can be analyzed to determine whether the adjusted game is balanced. The balance coefficient can be used to characterize the degree of balance of the adjusted game. For example, a positive balance coefficient can indicate that the degree of balance after adjustment is increased compared to before adjustment, and a negative balance coefficient can indicate that the degree of balance after adjustment is decreased compared to before adjustment.

[0070] After determining the balance coefficient, the balance degree in the Q value table can be updated according to the balance coefficient. As shown in Table 1, the balance degree adjusted for the S1th time is h1 and written into the corresponding position of the column to which strategy 3 belongs in the Q value table.

[0071] Step S106, in response to satisfying the condition for continuing to adjust the game balance parameters, adjusting the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm.

[0072] In this embodiment, after adjusting the game balance parameters once and updating the Q value table, it can be determined whether the conditions for continuing to adjust the game balance parameters are met. In particular, it can be determined whether the conditions for continuing to adjust the game balance parameters are met based on some game data such as the winning rate in the game, the winning rate of each game character, the number of times used, etc.

[0073] If the conditions for continuing to adjust the game balance parameters are met, it means that the balance level of the adjusted game is poor, and the game balance parameters of the asymmetric competitive game can continue to be adjusted based on the Q-Learning algorithm.

[0074] Step S108, in response to the condition for continuing to adjust the game balance parameters not being met, determining the final game balance parameters based on the updated Q value table.

[0075] If the conditions for continuing to adjust the game balance parameters are not met, it means that the balance level of the adjusted game is better. You can stop adjusting the game balance parameters of the asymmetric competitive game, and determine the final game balance parameters based on the updated Q value table, thereby completing the game balance parameter adjustment process.

[0076] The embodiment of the present invention provides a parameter adjustment device for an asymmetric competitive game, which adjusts the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm, obtains the game log of the player playing at least one game based on the adjusted game balance parameters; determines the balance coefficient corresponding to the adjusted game balance parameters based on the game log of at least one game, and updates the Q value table based on the balance coefficient; in response to the condition of continuing to adjust the game balance parameters being met, adjusts the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm; in response to the condition of not continuing to adjust the game balance parameters being met, determines the final game balance parameters based on the updated Q value table. In this method, the game parameters are quickly optimized by combining a large number of game matches through reinforcement learning, which can ensure the differentiated needs of players and improve the effect of balance optimization.

[0077] like Figure 2 The flowchart of another method for adjusting parameters of an asymmetric competitive game in an optional embodiment is shown. The method for adjusting parameters of an asymmetric competitive game in the optional embodiment includes the following steps:

[0078] Step S202: adjusting the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm, and obtaining a game log of at least one game match played by the player based on the adjusted game balance parameters.

[0079] In this embodiment, the adjustment strategy of at least one game balance parameter of the asymmetric competitive game included in the Q-Learning algorithm can be determined first. The above adjustment strategy can be written into the initialized Q value table, for example: specifying several indicators that affect the game balance according to the actual content of the game, such as skill CD and attack backswing. The minimum adjustment benchmark of these two indicators is obtained as a single Q value action change, for example, the minimum adjustment benchmark of skill CD and attack backswing time is 0.1s.

[0080] Refer to a Q value table shown in Table 2, which sets 4 adjustment strategies: skill CD increases by 0.1, skill CD decreases by 0.1, attack backswing time increases by 0.1, and attack backswing time decreases by 0.1. The initialized Q value table can default to a balance degree of 0.

[0081] Table 2

[0082]

[0083] In some embodiments, a target adjustment strategy may be selected from at least one game balance parameter adjustment strategy included in the Q-Learning algorithm; and the game balance parameter may be adjusted based on the target adjustment strategy.

[0084] In this embodiment, at least one adjustment strategy in the initialized Q value table can be selected as a target adjustment strategy, for example, "reducing the attack backswing time by 0.1" in Table 2 can be used as the target adjustment strategy. Then, the game balance parameters can be adjusted according to the target adjustment strategy.

[0085] In this embodiment, the target adjustment strategy can be selected according to a greedy strategy or a random strategy. In some embodiments, the target adjustment strategy (i.e., a random strategy) can be randomly selected from the adjustment strategies of at least one game balance parameter included in the Q-Learning algorithm; or, the target adjustment strategy (i.e., a greedy strategy) can be selected from the adjustment strategies of at least one game balance parameter included in the Q-Learning algorithm based on the balance degree recorded in the Q value table.

[0086] In the Q-learning algorithm, you can also set the probability of using a random strategy or a greedy strategy for each adjustment, and determine whether to select a random strategy or a greedy strategy for each adjustment based on the above probability.

[0087] The greedy strategy is an action selection strategy that selects the action that looks optimal in the current state. The greedy strategy does not consider possible future rewards, but only focuses on the current rewards. The reward in this embodiment is the degree of balance. Therefore, this embodiment can implement the greedy strategy by selecting a target adjustment strategy from the adjustment strategies of at least one game balance parameter based on the degree of balance in the Q value table.

[0088] A random strategy is an action selection strategy that randomly selects one of the feasible actions. A random strategy randomly selects between each feasible action. It does not pay attention to the reward of each action, but tries multiple random selections to obtain better results in the long run. Therefore, this embodiment can implement a random strategy by randomly selecting a target adjustment strategy.

[0089] In this embodiment, the behavior can be explored by different choices each time, and the optimal behavior can be found by updating the Q value table through result feedback.

[0090] In some embodiments, the adjusted game balance parameters can be set in a test server so that the player can play at least one game match based on the test server; and obtain a game log of at least one game match.

[0091] Since each player has his or her own unique style, strategy, and operating habits, it is difficult to predict how they will use different types of characters to operate, resulting in an imbalance in data collection. In this embodiment, the adjusted game balance parameters can be set in the test server. Players can randomly use different game characters in each game match on the test server to ensure that data under different game character combinations can be taken into account. Rewards are used to guide players to play positive games and provide effective data.

[0092] Step S204, determining a balance coefficient corresponding to the adjusted game balance parameter based on a game log of at least one game match, and updating the Q value table based on the balance coefficient.

[0093] In some embodiments, the win rate variance and / or average ranking variance of each game character can be determined based on the game log of at least one game match; and the balance coefficient corresponding to the adjusted game balance parameter can be determined based on the win rate variance and / or average ranking variance of each game character.

[0094] For asymmetric competitive games, the win rate and / or average ranking of different game characters can be used as a balance consideration. Therefore, the strength difference of each game character can be judged according to the variance of the win rate and / or the variance of the average ranking of each game character, so as to determine the balance coefficient corresponding to the adjusted game balance parameter. Among them, the larger the variance, the greater the difference in the strength of each game character. Among them, in this embodiment, the win rate variance or the average ranking variance of the game character can be used as the balance coefficient, and the win rate variance and the average ranking variance of the game character can also be weighted as the balance coefficient.

[0095] This embodiment can obtain game logs of multiple game matches, determine the win rate variance and / or average ranking variance of each game character in the game match based on the game logs to calculate the balance coefficient, and use the balance coefficient to characterize the balance change of the adjusted game balance parameters, thereby giving the algorithm rewards or penalties.

[0096] In some embodiments, the balance degree of the currently adjusted target adjustment strategy can be determined based on the change in the balance coefficient after adjustment and the balance coefficient before adjustment; the target adjustment strategy and the balance degree corresponding to the target adjustment strategy are added to the Q value table.

[0097] After the balance coefficient is calculated, the present embodiment can determine the balance degree of the currently adjusted target adjustment strategy based on the change in the balance coefficient after adjustment and the balance coefficient before adjustment. For example, if the balance coefficient decreases, the balance degree is positive, and if the balance coefficient increases, the balance degree is negative.

[0098] Referring to a Q value table shown in Table 3, after "reducing the attack backswing time by 0.1" is used as the target adjustment strategy, the balance degree becomes larger, indicating that it is more unbalanced, and the balance degree can be -1.

[0099] Table 3

[0100]

[0101] Among them, -1 in Table 3 refers to the updated value obtained after the reward and estimated value are calculated after the state changes. If it is a negative value, it means that the return benefit of the behavior is negative, and it can be updated according to the difference between the current value and the estimated value. S1 and S2 in Table 3 can be the states after each iteration, S1 can be regarded as the original state, and S2 is the state of the original state of S1 after "the attack backswing time is reduced by 0.1". In addition, the second and third rows in Table 3 are generally not merged. Each row represents a single behavior, and the following is the value corresponding to the result obtained by the single behavior in the corresponding environment S. After implementing a behavior, the state moves down one row.

[0102] In addition, it may also be determined whether to cancel the adjustment of the game balance parameters based on the game logs of multiple game matches.

[0103] After a parameter adjustment, if the balance degree indicates that the adjusted game has extremely poor balance, the current adjustment can be canceled, and the game balance parameters can be retraced to the game balance parameters before the adjustment and then readjusted using a different adjustment strategy.

[0104] Step S206, in response to satisfying the condition for continuing to adjust the game balance parameters, adjusting the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm.

[0105] In some embodiments, the conditions for continuing to adjust the game balance parameters include: the currently adjusted balance coefficient is less than a preset threshold; and / or the average value of the balance coefficients adjusted multiple times is less than a preset threshold.

[0106] In this embodiment, the game balance parameter can be judged based on the currently adjusted balance coefficient whether it reaches the balance state experienced by the player. If the currently adjusted balance coefficient is less than a preset threshold; and / or the average value of the balance coefficient adjusted multiple times is less than the preset threshold, it means that the game balance parameter needs to be further adjusted.

[0107] Step S208, in response to the condition for continuing to adjust the game balance parameters not being met, determining the final game balance parameters based on the updated Q value table.

[0108] If the conditions for continuing to adjust the game balance parameters are not met, it means that the game balance parameters converge within a reasonable modification range, and ultimately the game balance is optimized and a balanced state of player experience is achieved.

[0109] In addition, in some embodiments, after each game match ends, the evaluation of the players participating in the game match can be obtained; and whether to continue adjusting the game balance parameters can be determined based on the player's evaluation.

[0110] In this embodiment, players can provide feedback and evaluations, and the player evaluations can be collected after each game. This embodiment can also determine whether a balance state of player experience is reached based on the player evaluations.

[0111] With multiple iterations, the Q value table will be updated until the balance state experienced by the player is reached. In some embodiments, the adjustment strategy for each adjustment is determined based on the updated Q value table; and the final game balance parameter is determined based on the adjustment strategy for each adjustment.

[0112] For example, the target adjustment strategy for the 1st to 3rd adjustments is to reduce the skill CD by 0.1, and the target adjustment strategy for the 4th to 8th adjustments is to increase the attack backswing time by 0.1. Therefore, the skill CD can be reduced by 0.3 and the attack backswing time can be increased by 0.5 to obtain the final game balance parameters.

[0113] Step S210, setting the final game balance parameters in the official server.

[0114] The final game balance parameters generally mean a better degree of balance, so the final game balance parameters can be set in the official server, and players can play games in the official server.

[0115] In some embodiments, it may also be determined whether to continue adjusting the game balance parameters based on the game match on the official server.

[0116] In addition, this embodiment can also perform long-term monitoring and adjustment of the official server, mainly including collecting in-game log data, performing character and environment data analysis to observe the overall game experience and winning rate. Through data feedback, a series of parameters obtained by the balance optimization system are manually fine-tuned.

[0117] In this embodiment, various data in the game can also be collected based on the game logs of multiple game matches, and processed and analyzed. These data include but are not limited to the character selected by the player, attributes, winning rate, cause of death, etc.

[0118] The above method provided by the embodiment of the present invention can be performed by opening a random balance exploration gameplay on the test server before the official server of each game version is launched, using the Q-learning algorithm to dynamically adjust various parameters in the game, taking the game's win-loss balance as the balance degree, and performing corresponding operations and updating the Q value table for all initialized parameter values ​​(such as skill CD, back-swing hard value, etc.) according to the greedy strategy (i.e., selecting the action with the maximum Q value) or the random strategy (i.e., selecting according to the probability distribution), and repeating this process until the parameters are basically converged. The final game balance parameters are launched on the official server and can be fine-tuned, and finally the final balance optimization of the game is achieved.

[0119] For the overall process of the above method provided by the embodiment of the present invention, please refer to Figure 3 Schematic diagram of a method for adjusting parameters of an asymmetric competitive game. Figure 3 As shown, you can first initialize the Q value table, select the target adjustment strategy, and update the test server. After the player plays multiple games on the test server, calculate the balance coefficient and update the Q value table. Determine whether the player experience is balanced. If the player experience is balanced, determine the final game balance parameters and update the official server. If the player experience is not balanced, reselect the target adjustment strategy.

[0120] The above method provided by the embodiment of the present invention is based on real-time data, and makes minimal modifications to the game balance parameters to achieve maximum fairness. This advantage can help game developers reduce interference with players while ensuring game balance, thereby improving the fairness and credibility of the game, and has the following specific advantages:

[0121] (1) Fast iteration speed can meet the differentiated needs of different players. This advantage can help game developers understand players’ needs more timely and adjust the game, thereby improving the user experience and satisfaction of the game.

[0122] (2) Controllable complexity: This advantage can help game developers reduce the cost of game development and maintenance while ensuring the quality and stability of the game.

[0123] (3) Through machine learning algorithms and a large amount of data, it can meet the differentiated needs of players. This advantage can help game developers understand players’ needs and behaviors more accurately, thereby providing players with a more personalized and high-quality gaming experience.

[0124] (4) Prevent the imperfection of manual annotation by game planners. This advantage can reduce the subjective annotation and errors of game planners through machine learning algorithms and data analysis, thereby improving the balance and fairness of the game. At the same time, it can also help game developers understand the balance and playability of the game more comprehensively, thereby providing more powerful support for game optimization.

[0125] Corresponding to the above method embodiment, the embodiment of the present invention provides a parameter adjustment device for an asymmetric competitive game, such as Figure 4 The schematic diagram of the structure of a parameter adjustment device for an asymmetric competitive game is shown, and the parameter adjustment device for an asymmetric competitive game includes:

[0126] The parameter adjustment module 41 is used to adjust the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm, and obtain the game log of the player playing at least one game based on the adjusted game balance parameters; wherein the Q-Learning algorithm includes an adjustment strategy for at least one game balance parameter of the asymmetric competitive game;

[0127] A Q value table updating module 42 is used to determine a balance coefficient corresponding to the adjusted game balance parameter based on the game log of at least one game match, and update the Q value table based on the balance coefficient; wherein the Q value table records the adjustment strategy of each adjustment and the balance degree of the game under the adjusted game balance parameter, and the balance degree is determined based on the balance coefficient;

[0128] A parameter continuous adjustment module 43, configured to adjust the game balance parameters of the asymmetric competitive game based on a Q-Learning algorithm in response to satisfying the condition for continuing to adjust the game balance parameters;

[0129] The parameter determination module 44 is configured to determine the final game balance parameters based on the updated Q value table in response to the condition for continuing to adjust the game balance parameters not being met.

[0130] The embodiment of the present invention provides a parameter adjustment device for an asymmetric competitive game, which adjusts the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm, obtains the game log of the player playing at least one game based on the adjusted game balance parameters; determines the balance coefficient corresponding to the adjusted game balance parameters based on the game log of at least one game, and updates the Q value table based on the balance coefficient; in response to the condition of continuing to adjust the game balance parameters being met, adjusts the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm; in response to the condition of not continuing to adjust the game balance parameters being met, determines the final game balance parameters based on the updated Q value table. In this method, the game parameters are quickly optimized by combining a large number of game matches through reinforcement learning, which can ensure the differentiated needs of players and improve the effect of balance optimization.

[0131] In an optional embodiment of the present invention, the parameter adjustment module is used to select a target adjustment strategy from at least one game balance parameter adjustment strategy included in the Q-Learning algorithm; and adjust the game balance parameter based on the target adjustment strategy.

[0132] In an optional embodiment of the present invention, the above-mentioned parameter adjustment module is used to randomly select a target adjustment strategy from the adjustment strategy of at least one game balance parameter included in the Q-Learning algorithm; or, to select a target adjustment strategy from the adjustment strategy of at least one game balance parameter included in the Q-Learning algorithm based on the balance degree recorded in the Q value table.

[0133] In an optional embodiment of the present invention, the adjusted game balance parameters are set in a test server so that players can play at least one game match based on the test server; and a game log of at least one game match is obtained.

[0134] In an optional embodiment of the present invention, the above-mentioned Q value table updating module is used to determine the win rate variance and / or average ranking variance of each game character based on the game log of at least one game match; and determine the balance coefficient corresponding to the adjusted game balance parameter based on the win rate variance and / or average ranking variance of each game character.

[0135] In an optional embodiment of the present invention, the above-mentioned conditions for continuing to adjust the game balance parameters include: the currently adjusted balance coefficient is less than a preset threshold; and / or the average value of the balance coefficients adjusted multiple times is less than a preset threshold.

[0136] In an optional embodiment of the present invention, the above-mentioned Q value table updating module is used to determine the balance degree of the currently adjusted target adjustment strategy based on the change in the balance coefficient after adjustment and the balance coefficient before adjustment; and add the target adjustment strategy and the balance degree corresponding to the target adjustment strategy in the Q value table.

[0137] In an optional embodiment of the present invention, the parameter determination module is used to determine an adjustment strategy for each adjustment based on the updated Q value table; and determine a final game balance parameter based on the adjustment strategy for each adjustment.

[0138] In an alternative embodiment of the present invention, see Figure 5 The schematic diagram of the structure of another parameter adjustment device for an asymmetric competitive game is shown, the device comprises: a parameter online module 45 connected to a parameter determination module 44; the parameter online module 45 is used to set the final game balance parameters in the official server.

[0139] In an optional embodiment of the present invention, the above-mentioned parameter online module is also used to determine whether to continue adjusting the game balance parameters based on the game match of the official server.

[0140] The parameter adjustment device for an asymmetric competitive game provided in an embodiment of the present invention has the same technical features as the parameter adjustment method for an asymmetric competitive game provided in the above-mentioned embodiment, and therefore can also solve the same technical problems and achieve the same technical effects.

[0141] The embodiment of the present invention also provides an electronic device for running the parameter adjustment method of the above-mentioned asymmetric competitive game; see Figure 6 The structure diagram of an electronic device shown in FIG. 1 includes a memory 100 and a processor 101, wherein the memory 100 is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor 101 to perform the following steps:

[0142] Adjusting game balance parameters of an asymmetric competitive game based on a Q-Learning algorithm, obtaining a game log of at least one game match played by a player based on the adjusted game balance parameters; wherein the Q-Learning algorithm includes an adjustment strategy for at least one game balance parameter of the asymmetric competitive game; determining a balance coefficient corresponding to the adjusted game balance parameter based on the game log of at least one game match, and updating a Q value table based on the balance coefficient; wherein the Q value table records the adjustment strategy for each adjustment and the degree of balance of the game under the adjusted game balance parameters, and the degree of balance is determined based on the balance coefficient; in response to satisfying a condition for continuing to adjust the game balance parameters, adjusting the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm; in response to not satisfying the condition for continuing to adjust the game balance parameters, determining the final game balance parameters based on the updated Q value table.

[0143] In an optional embodiment of the present invention, the above-mentioned step of adjusting the game balance parameters of an asymmetric competitive game based on the Q-Learning algorithm includes: selecting a target adjustment strategy from at least one game balance parameter adjustment strategy included in the Q-Learning algorithm; and adjusting the game balance parameters based on the target adjustment strategy.

[0144] In an optional embodiment of the present invention, the step of selecting a target adjustment strategy from the adjustment strategy of at least one game balance parameter included in the Q-Learning algorithm includes: randomly selecting a target adjustment strategy from the adjustment strategy of at least one game balance parameter included in the Q-Learning algorithm; or, selecting a target adjustment strategy from the adjustment strategy of at least one game balance parameter included in the Q-Learning algorithm based on the degree of balance recorded in the Q value table.

[0145] In an optional embodiment of the present invention, the step of obtaining a game log of at least one game match played by the player based on the adjusted game balance parameters includes: setting the adjusted game balance parameters in a test server so that the player plays at least one game match based on the test server; and obtaining a game log of at least one game match.

[0146] In an optional embodiment of the present invention, the step of determining the balance coefficient corresponding to the adjusted game balance parameter based on the game log of at least one game match includes:

[0147] Based on the game log of at least one game match, the win rate variance and / or average ranking variance of each game character is determined; based on the win rate variance and / or average ranking variance of each game character, the balance coefficient corresponding to the adjusted game balance parameter is determined.

[0148] In an optional embodiment of the present invention, the above-mentioned conditions for continuing to adjust the game balance parameters include: the currently adjusted balance coefficient is less than a preset threshold; and / or the average value of the balance coefficients adjusted multiple times is less than a preset threshold.

[0149] In an optional embodiment of the present invention, the above-mentioned step of updating the Q value table based on the balance coefficient includes: determining the balance degree of the currently adjusted target adjustment strategy based on the change in the balance coefficient after adjustment and the balance coefficient before adjustment; adding the target adjustment strategy and the balance degree corresponding to the target adjustment strategy in the Q value table.

[0150] In an optional embodiment of the present invention, the above-mentioned step of determining the final game balance parameters based on the updated Q value table includes: determining an adjustment strategy for each adjustment based on the updated Q value table; and determining the final game balance parameters based on the adjustment strategy for each adjustment.

[0151] In an optional embodiment of the present invention, after the above step of determining the final game balance parameters based on the updated Q value table, the method further includes: setting the final game balance parameters in the official server.

[0152] In an optional embodiment of the present invention, after the above step of setting the final game balance parameters in the official server, the method further includes: judging whether to continue adjusting the game balance parameters based on the game match on the official server.

[0153] The embodiment of the present invention can adjust the game balance parameters of an asymmetric competitive game based on the Q-Learning algorithm, obtain the game log of a player playing at least one game based on the adjusted game balance parameters; determine the balance coefficient corresponding to the adjusted game balance parameters based on the game log of at least one game, and update the Q value table based on the balance coefficient; in response to the conditions for continuing to adjust the game balance parameters being met, adjust the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm; in response to the conditions for continuing to adjust the game balance parameters not being met, determine the final game balance parameters based on the updated Q value table. In this method, the game parameters are quickly optimized by combining a large number of game matches through reinforcement learning, which can ensure the differentiated needs of players and improve the effect of balance optimization.

[0154] Further, Figure 6 The electronic device shown further includes a bus 102 and a communication interface 103 , and the processor 101 , the communication interface 103 and the memory 100 are connected via the bus 102 .

[0155] The memory 100 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 103 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used. The bus 102 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0156] The processor 101 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 101. The above processor 101 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present invention can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module may be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 100, and the processor 101 reads the information in the memory 100 and completes the steps of the method of the above embodiment in combination with its hardware.

[0157] The embodiment of the present invention further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the parameter adjustment method of the asymmetric competitive game, and the following steps may be performed:

[0158] Adjusting game balance parameters of an asymmetric competitive game based on a Q-Learning algorithm, obtaining a game log of at least one game match played by a player based on the adjusted game balance parameters; wherein the Q-Learning algorithm includes an adjustment strategy for at least one game balance parameter of the asymmetric competitive game; determining a balance coefficient corresponding to the adjusted game balance parameter based on the game log of at least one game match, and updating a Q value table based on the balance coefficient; wherein the Q value table records the adjustment strategy for each adjustment and the degree of balance of the game under the adjusted game balance parameters, and the degree of balance is determined based on the balance coefficient; in response to satisfying a condition for continuing to adjust the game balance parameters, adjusting the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm; in response to not satisfying the condition for continuing to adjust the game balance parameters, determining the final game balance parameters based on the updated Q value table.

[0159] In an optional embodiment of the present invention, the above-mentioned step of adjusting the game balance parameters of an asymmetric competitive game based on the Q-Learning algorithm includes: selecting a target adjustment strategy from at least one game balance parameter adjustment strategy included in the Q-Learning algorithm; and adjusting the game balance parameters based on the target adjustment strategy.

[0160] In an optional embodiment of the present invention, the step of selecting a target adjustment strategy from the adjustment strategy of at least one game balance parameter included in the Q-Learning algorithm includes: randomly selecting a target adjustment strategy from the adjustment strategy of at least one game balance parameter included in the Q-Learning algorithm; or, selecting a target adjustment strategy from the adjustment strategy of at least one game balance parameter included in the Q-Learning algorithm based on the degree of balance recorded in the Q value table.

[0161] In an optional embodiment of the present invention, the step of obtaining a game log of at least one game match played by the player based on the adjusted game balance parameters includes: setting the adjusted game balance parameters in a test server so that the player plays at least one game match based on the test server; and obtaining a game log of at least one game match.

[0162] In an optional embodiment of the present invention, the step of determining the balance coefficient corresponding to the adjusted game balance parameter based on the game log of at least one game match includes:

[0163] Based on the game log of at least one game match, the win rate variance and / or average ranking variance of each game character is determined; based on the win rate variance and / or average ranking variance of each game character, the balance coefficient corresponding to the adjusted game balance parameter is determined.

[0164] In an optional embodiment of the present invention, the above-mentioned conditions for continuing to adjust the game balance parameters include: the currently adjusted balance coefficient is less than a preset threshold; and / or the average value of the balance coefficients adjusted multiple times is less than a preset threshold.

[0165] In an optional embodiment of the present invention, the above-mentioned step of updating the Q value table based on the balance coefficient includes: determining the balance degree of the currently adjusted target adjustment strategy based on the change in the balance coefficient after adjustment and the balance coefficient before adjustment; adding the target adjustment strategy and the balance degree corresponding to the target adjustment strategy in the Q value table.

[0166] In an optional embodiment of the present invention, the above-mentioned step of determining the final game balance parameters based on the updated Q value table includes: determining an adjustment strategy for each adjustment based on the updated Q value table; and determining the final game balance parameters based on the adjustment strategy for each adjustment.

[0167] In an optional embodiment of the present invention, after the above step of determining the final game balance parameters based on the updated Q value table, the method further includes: setting the final game balance parameters in the official server.

[0168] In an optional embodiment of the present invention, after the above step of setting the final game balance parameters in the official server, the method further includes: judging whether to continue adjusting the game balance parameters based on the game match on the official server.

[0169] The embodiment of the present invention can adjust the game balance parameters of an asymmetric competitive game based on the Q-Learning algorithm, obtain the game log of a player playing at least one game based on the adjusted game balance parameters; determine the balance coefficient corresponding to the adjusted game balance parameters based on the game log of at least one game, and update the Q value table based on the balance coefficient; in response to the conditions for continuing to adjust the game balance parameters being met, adjust the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm; in response to the conditions for continuing to adjust the game balance parameters not being met, determine the final game balance parameters based on the updated Q value table. In this method, the game parameters are quickly optimized by combining a large number of game matches through reinforcement learning, which can ensure the differentiated needs of players and improve the effect of balance optimization.

[0170] The computer program product of the parameter adjustment method, device, equipment and storage medium for asymmetric competitive games provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method in the previous method embodiment. The specific implementation can be found in the method embodiment, which will not be repeated here.

[0171] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and / or device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0172] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0173] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, electronic device, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0174] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.

[0175] Finally, it should be noted that the above embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A parameter adjustment method for an asymmetric competitive game. It is characterized in that The method comprises: Adjusting the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm, and obtaining a game log of at least one game match played by the player based on the adjusted game balance parameters; wherein the Q-Learning algorithm includes an adjustment strategy for at least one game balance parameter of the asymmetric competitive game; Determine a balance coefficient corresponding to the adjusted game balance parameter based on the game log of the at least one game match, and update a Q value table based on the balance coefficient; wherein the Q value table records an adjustment strategy for each adjustment and a balance degree of the game under the adjusted game balance parameter, and the balance degree is determined based on the balance coefficient; In response to satisfying a condition for continuing to adjust the game balance parameter, adjusting the game balance parameter of the asymmetric competitive game based on a Q-Learning algorithm; In response to the condition for continuing to adjust the game balance parameter not being met, determining the final game balance parameter based on the updated Q value table.

2. The method according to claim 1, It is characterized in that The step of adjusting the game balance parameters of the asymmetric competitive game based on the Q-Learning algorithm includes: Selecting a target adjustment strategy from at least one game balance parameter adjustment strategy included in the Q-Learning algorithm; The game balance parameter is adjusted based on the target adjustment strategy.

3. The method according to claim 2, It is characterized in that The step of selecting a target adjustment strategy from at least one game balance parameter adjustment strategy included in the Q-Learning algorithm comprises: Randomly selecting a target adjustment strategy from at least one game balance parameter adjustment strategy included in the Q-Learning algorithm; Or, the target adjustment strategy is selected from the adjustment strategies of at least one game balance parameter included in the Q-Learning algorithm based on the balance degree recorded in the Q-value table.

4. The method according to claim 1, It is characterized in that The step of obtaining a game log of at least one game match played by a player based on the adjusted game balance parameter comprises: Setting the adjusted game balance parameters in a test server so that the player can play at least one game based on the test server; Obtain a game log of the at least one game match.

5. The method according to claim 1, It is characterized in that The step of determining the balance coefficient corresponding to the adjusted game balance parameter based on the game log of the at least one game match comprises: Determine the win rate variance and / or average ranking variance of each game character based on the game log of the at least one game match; The balance coefficient corresponding to the adjusted game balance parameter is determined based on the win rate variance and / or average ranking variance of each of the game characters.

6. The method according to claim 1, It is characterized in that The conditions for continuing to adjust the game balance parameters include: The currently adjusted balance coefficient is less than a preset threshold; And / or, an average value of the balance coefficient adjusted multiple times is less than a preset threshold.

7. The method according to claim 2, It is characterized in that The step of updating the Q value table based on the balance coefficient comprises: Determining the balance degree of the currently adjusted target adjustment strategy based on the change amount between the balance coefficient after adjustment and the balance coefficient before adjustment; The target adjustment strategy and the balance degree corresponding to the target adjustment strategy are added to the Q value table.

8. The method according to claim 1, It is characterized in that The step of determining the final game balance parameter based on the updated Q value table includes: Determine an adjustment strategy for each adjustment based on the updated Q value table; The final game balance parameter is determined based on the adjustment strategy of each adjustment.

9. The method according to claim 1, It is characterized in that After the step of determining the final game balance parameter based on the updated Q value table, the method further includes: Set the final game balance parameters on the official server.

10. The method according to claim 9, It is characterized in that After the step of setting the final game balance parameters in the official server, the method further includes: Determine whether to continue adjusting the game balance parameters based on the game match on the official server.

11. A parameter adjustment device for an asymmetric competitive game, It is characterized in that The device comprises: A parameter adjustment module, configured to adjust the game balance parameters of the asymmetric competitive game based on a Q-Learning algorithm, and obtain a game log of a player playing at least one game based on the adjusted game balance parameters; wherein the Q-Learning algorithm includes an adjustment strategy for at least one game balance parameter of the asymmetric competitive game; A Q value table updating module, used to determine the balance coefficient corresponding to the adjusted game balance parameter based on the game log of the at least one game match, and update the Q value table based on the balance coefficient; wherein the Q value table records the adjustment strategy of each adjustment and the balance degree of the game under the adjusted game balance parameter, and the balance degree is determined based on the balance coefficient; A parameter continuous adjustment module, configured to adjust the game balance parameter of the asymmetric competitive game based on a Q-Learning algorithm in response to satisfying a condition for continuing to adjust the game balance parameter; A parameter determination module is used to determine the final game balance parameter based on the updated Q value table in response to the condition for continuing to adjust the game balance parameter not being met.

12. An electronic device, It is characterized in that It includes a processor and a memory, the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the parameter adjustment method of the asymmetric competitive game described in any one of claims 1-10.

13. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions prompt the processor to implement the parameter adjustment method for the asymmetric competitive game described in any one of claims 1-10.