Game agent design method and system based on deep reinforcement learning
By grouping and training game server users and driving iterations with feedback, and combining relevant game data, a general intelligent agent that adapts to various user behaviors and habits is generated. This solves the problem of low training efficiency in existing technologies and achieves the stability and robustness of the intelligent agent in different environments.
Patent Information
- Application Number
- CN202511246653.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-02
AI Technical Summary
Existing game agent design methods based on deep reinforcement learning cannot effectively analyze users' game thinking, leading to repetition or interference during training, low training efficiency, and an inability to generate general agents that adapt to various user behaviors and habits.
By collecting and analyzing user data and behaviors on game servers, grouping and training users to form a set of users with high similarity, conducting cross-set integration training, iterating training based on user feedback, identifying relevant game data for data fusion, and finally generating a general intelligent agent.
It improves training efficiency, enhances the adaptability and generalization ability of the agent, enabling it to adapt to different user groups and game environments, and improves user experience.
Smart Images

Figure CN121093992A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of game intelligent design, in particular to a game intelligent agent design method and system based on deep reinforcement learning. BACKGROUND
[0002] As an important form of entertainment, the complexity and diversity of electronic games provide a natural research platform for the development of artificial intelligence (AI). In recent years, with the rapid improvement of computing power and continuous innovation of algorithms, the application of artificial intelligence in the game field has made remarkable progress, especially the rise of deep reinforcement learning (DRL), which enables AI to learn and master high-level strategies from complex interactive environments, thereby performing better than humans or even professional players in various games. Traditional game AI design, such as rule-based systems, search algorithms (such as Minimax), and finite state machines, while performing well in some simple games or specific situations, is not sufficient to face the growing complexity of modern games, massive state spaces, and dynamic uncertain environments. These methods often require a lot of manual design and tuning and are difficult to generalize to new games or different game scenarios. Deep reinforcement learning provides a powerful solution to the above problems. DRL combines the powerful feature extraction capability of deep learning and the decision-making learning framework of reinforcement learning, enabling it to learn optimal strategies directly from raw game pixels or high-dimensional state spaces. Through extensive interaction with the game environment, the agent can continuously try and error and adjust its internal neural network parameters based on the reward signals obtained, ultimately learning complex behaviors that maximize cumulative rewards. Existing game agent design methods and systems based on deep reinforcement learning cannot analyze the game thinking of game users and organize users with similar game thinking into the same set. They cannot train each user set first and then train the overall data. They cannot group users of similar games based on their game thinking and train the agent according to the idea of grouping training first and then overall training, which makes the agent prone to repeated or interfering data during training, slowing down the training process and reducing training efficiency. The practicality has certain limitations. SUMMARY
[0003] The present application provides a game agent design method and system based on deep reinforcement learning, which is used to promote the solution to the problems in the background art.
[0004] The present application provides the following technical solutions: a game agent design method based on deep reinforcement learning, comprising: collecting data of all users in the game server; User behavior analysis is performed on all users in the game server; Users in the game server are grouped and grouped training is performed accordingly; According to the grouping training, cross-set integration training is performed; Collect user feedback on the trained intelligent agent, and form feedback-driven iterative training based on user feedback; Identify other games related to the current game and perform data collection and analysis; According to the data collected and analyzed on other games related to the current game, cross-game user grouping and training are performed; All training data of the current game and related games are fused and finally trained to generate a general intelligent agent that can adapt to various user behaviors and habits.
[0005] As the game intelligent agent design method based on deep reinforcement learning described in the application, wherein: data collection is performed on all users in the game server, specifically: For each user , extract its original behavior data from the game server, denoted as : ; Use feature extraction function to convert the original behavior data into a low-dimensional feature vector : ; Collect user 's direct feedback data, denoted as : ; Normalize the indirect feedback data to obtain and : ; ; Integrate direct feedback and normalized indirect feedback to form comprehensive feedback data, denoted as : ; Splice behavior features and feedback features to obtain comprehensive features : .
[0006] As the game intelligent agent design method based on deep reinforcement learning described in the application, wherein: user behavior analysis is performed on all users in the game server, specifically: Obtain the feature vector of each user in the game server ; Cluster the behavior features of all users to identify different behavior patterns: ; Analyze each behavior pattern to mine the logic and rules behind user behavior: .
[0007] As the game agent design method based on deep reinforcement learning of the application, wherein: the users in the game server are grouped and corresponding grouping training is carried out, specifically: S1, extract a user from the game server at random, define it as an analysis user, denoted as ; Extract another user from the game server at random, define it as a comparison user, denoted as ; S2, obtain the feature vectors of the analysis user and the comparison user , denoted as and ; S3, calculate the cosine similarity between the analysis user and the comparison user : ; S4, repeat steps S1-S3 until there is a cosine similarity value between all other users in the game server and the analysis user; S5, repeat steps S1-S4 until all users in the game server have been analyzed as analysis users; S6, for each analysis user , separate it as a set; S7, extract all comparison users corresponding to the analysis user ; S8, extract the similarity matrix between the analysis user and all comparison users : ; S9, define a screening function, for each analysis user , find all comparison users with high similarity: ; If If yes, it is determined that the analysis user is similar to the comparison user . If yes , it is determined that the analysis user is similar to the comparison user . S10, all comparison users similar to the analysis user and the comparison user are merged into the same set . S11, repeat steps S1-S10 until the similarity comparison of all users in the game server is completed, and multiple user sets are formed . S12, initialize an agent for each set : . . S13, for each set , collect the game behavior data and feedback data of all users in the set to form a training data set : . S14, train each agent independently . S15, maximize the performance of the agent in the user set environment . S16, optimize the objective function and update the parameters of the policy network . S17, repeat steps S14-S15 until the performance of the agent in the user set environment reaches convergence.
[0008] As the game agent design method based on deep reinforcement learning of the present application, wherein: according to the grouping training, cross-set integration training is performed, specifically: Obtain the corresponding training data set of each set , and integrate to form a comprehensive data set . Initialize a global agent : . Perform environment interaction After the game ends, use GAE to calculate the advantage function : . Update the policy network by maximizing the objective function of PPO ; Update the parameters of the policy network using gradient ascent: ; Update the parameters of the value function using gradient descent: ; Update the parameters of the value function using gradient descent: .
[0009] As the game agent design method based on deep reinforcement learning, wherein: collecting user feedback on the trained agent, and based on the user feedback to form a feedback-driven iterative training, specifically: Select a group of users from all users in the game server Conduct experience testing: ; For users Provide a game environment consistent with the training environment, and deploy the trained agent Into the game environment; Users Interact with the agent Generate experience data : ; For each user In the user Collect their direct feedback : ; Collect indirect feedback data from users And normalize the data: ; ; Integrate indirect feedback data : ; Splice the direct feedback And indirect feedback Into a comprehensive feedback vector : ; Form a comprehensive feedback dataset : ; Perform sentiment analysis on the comments Of each user In the user Identify positive and negative feedback, and get sentiment scores : ; Calculate the importance score of each behavior feature : : ; Cluster the comprehensive feedback dataset , obtaining a clustering result : ; According to the feedback analysis result, define the improvement target : ; Filter out data related to the improvement target from the original training data : ; Use the improved dataset to train the agent and update the policy network and value function : ; ; Update the improved policy network and value function to the global agent ; ; Repeat the overall steps for each user in the user : , collect their direct feedback to update the improved policy network and value function to the global agent until the performance of the agent reaches a satisfactory level.
[0010] As the game agent design method based on deep reinforcement learning described in the present application, wherein: other games related to the current game are identified and data collection and analysis are performed, specifically: Analyze the type, play and user group of the current game, and identify other games related to the current game ; For each related game , respectively obtain user data ; Collect data and analyze user behavior for each user in each related game , obtaining and ; Integrate the user data in the related games into the user data of the current game to form the final training dataset .
[0011] As a kind of game agent design method based on deep reinforcement learning described in the application, wherein: the training data of current game and related game are fused, and final training is carried out, to generate a general agent capable of adapting to various user behaviors and habits, specifically: Obtain final training data set ; The final training data set Carry out cross-game user grouping and training.
[0012] As the application also discloses a system for executing a game agent design method based on deep reinforcement learning, wherein: Data acquisition module: collect the game data of users of current game and related games, collect the feedback of users on the game and the trained agent; Data analysis module: analyze the game data of users of current game and related games, classify users, and classify users with similar game operation thinking into the same set for data training; Data training module: according to the classified set of users of current game or related games, first train the agent for the data in each set, and then train the agent for the whole set according to the training condition of the set; Data integration module: integrate the training condition of the agent for current game and the training condition of the agent for related game, and train the agent again; Data generation module: generate or update the agent according to the training condition of the agent.
[0013] The application has the following beneficial effects: 1. The game agent design method and system based on deep reinforcement learning extracts all user training data from the game server, sets a feedback mechanism in the game to collect direct feedback on the performance of the agent, records indirect feedback indicators such as user game duration, frequency, and retention rate, pre-processes and extracts the behavior data of each user, extracts key features that reflect the user's thinking style and habits, uses clustering algorithms to analyze the behavior characteristics of users, identifies different behavior patterns, further mines the logic and rules behind user behavior through decision tree, random forest and other algorithms, defines a similarity measurement method between users, calculates the similarity between users according to their behavior characteristics and patterns, groups users according to similarity, forms multiple user sets, and each set has similar thinking style and game habits. Collect comprehensive user behavior and feedback data to provide a rich information base for subsequent analysis and training, covering different types of data to ensure that the agent can learn the diversity of user behavior, understand the user's thinking style and game habits through feature extraction and behavior pattern recognition, and provide the basis for subsequent personalized training. The agent can better adapt to the needs of different users, reduce noise in the training process by grouping similar users, improve training efficiency, and optimize the agent for different user groups to improve user experience.
[0014] 2、The game agent design method and system based on deep reinforcement learning independently trains the data in each user set, uses the deep reinforcement learning algorithm to train the agent, enables it to adapt to the behavior characteristics of the set users, adjusts the agent's strategy in the training process to improve its performance in the user set environment, integrates all the training data of the user sets to form a comprehensive data set, uses the integrated data set to globally train the agent, enables it to adapt to the behavior characteristics of different user sets, enables the user to experience the trained agent, observes the agent's performance in the game process, collects user feedback, including direct feedback (such as score, comment) and indirect feedback (such as game behavior change), analyzes user feedback, identifies the advantages and disadvantages of the agent, according to the feedback result, trains the agent specifically, optimizes its performance in specific scenarios or behavior patterns, can learn the behavior characteristics of specific user sets, provides personalized game experience, through independent training and strategy optimization, improves the performance of the agent in specific user sets, integrates the data of multiple user sets, improves the generalization ability of the agent, enables it to adapt to a wider user group, enhances the stability and robustness of the agent in different environments, through the real experience and feedback of the user, obtains the objective evaluation of the agent's performance, provides direct improvement suggestions for subsequent iterative training, forms a closed-loop optimization, trains the agent specifically according to the user feedback, solves specific problems, improves user satisfaction, continuously optimizes the performance of the agent, ensures its competitiveness under the changing user demand.
[0015] 3、The game agent design method and system based on deep reinforcement learning identifies other games related to the current game, obtains user data from these games, analyzes the user data of other games, extracts the behavior characteristics and patterns of users, integrates all the training data of the current game and related games to form the final training data set, uses the integrated data set to finally train the agent, generates a general agent that can adapt to various user behaviors and habits, expands the training data set by collecting data from related games, improves the adaptability of the agent, enables the agent to adapt to user behaviors in different game environments, enhances its generalization ability, integrates all the data for final training, ensures that the agent performs evenly in all aspects, generates a general agent that can adapt to various user behaviors and game scenarios, improves the overall user experience. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 The game agent design method based on deep reinforcement learning of the present application is a flowchart; Figure 2 The system block diagram of the game agent design method based on deep reinforcement learning of the present application is executed. DETAILED DESCRIPTION
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Example 1: A method for designing game agents based on deep reinforcement learning (see [reference]). Figure 1 ,include: Data is collected from all users on the game server. Perform user behavior analysis on all users on the game server; Users on the game server are grouped and trained accordingly. Based on the results of group training, perform cross-set integration training; Collect user feedback on the trained agent and form feedback-driven iterative training based on the user feedback; Identify other games related to the current game and collect and analyze the data; Based on data collected and analyzed from other games related to the current game, cross-game user grouping and training are conducted; By fusing all training data from the current game and related games and performing final training, a general intelligent agent capable of adapting to various user behaviors and habits is generated.
[0019] Example 2 is an improvement upon Example 1. This deep reinforcement learning-based game agent design method collects data from all users on the game server, specifically as follows: For each user Extract its raw behavioral data from the game server, denoted as The raw behavioral data refers to the user's specific operation records in the game, such as button clicks and movement trajectories. ; in, This represents the user's behavior across different dimensions, which are the results of classifying the raw data according to dimensions such as operation, strategy, social, and achievement. Using feature extraction functions Transform raw behavioral data into low-dimensional feature vectors : ; in, is the weight matrix of the first principal components, used to project the original data into a low-dimensional subspace defined by the first principal components, denotes the transpose of the weight matrix of the first principal components, denotes the standardized original behavior data, the standardization process can make the mean of the data to be 0 and the standard deviation to be 1, so as to eliminate the dimensional difference between different features, specifically: ; wherein, is the mean of each feature dimension , is the standard deviation of each feature dimension , the feature dimension is each feature in the original behavior data; collect the direct feedback data of the user , denoted as , the direct feedback data is the feedback provided by the user, such as satisfaction score, questionnaire opinion, etc.: ; normalize the indirect feedback data to obtain and , the indirect feedback data is obtained indirectly by analyzing the user behavior, for example, game duration, retention rate, task completion rate, etc.: ; ; wherein, is the normalized game duration, is the normalized retention rate, and are the maximum and minimum values of the game duration of all users respectively, and are the maximum and minimum values of the retention rate of all users respectively, and are the game duration and retention rate before normalization, after normalization, the game duration and retention rate will be normalized to the range of [0, 1], the retention rate refers to the proportion of users who use or participate in a certain service, game or product again within a certain period of time, in the game field, it is usually used to measure the frequency of players continuing to visit the game after registration or first playing, and is an important index for evaluating the attraction of the game and the user stickiness; integrate the direct feedback and the normalized indirect feedback to form comprehensive feedback data, denoted as : ; the behavior features and feedback features splicing, obtaining comprehensive features : .
[0020] The embodiment also provides user behavior analysis on all users in the game server, specifically: obtaining a feature vector of each user in the game server ; using a clustering algorithm to cluster the behavior features of all users, and identifying different behavior patterns: ; wherein, is a set of behavior feature vectors of all users, denoted as , is a clustering result, denoted as , each represents a behavior pattern; using a decision tree or random forest algorithm to analyze each behavior pattern , and mining the logic and rules behind the user behavior: ; wherein, is a set of feature vectors corresponding to the behavior pattern , is a corresponding behavior pattern label, is a constructed behavior pattern model, used to understand the logic and rules of user behavior in the behavior pattern ; wherein, the behavior features of all users are clustered to identify different behavior patterns, specifically, a K-Means algorithm is used for clustering: randomly selecting feature vectors as initial cluster centers, denoted as ; for each feature vector , calculating the distance between it and each cluster center, and assigning it to the nearest cluster: ; setting the cluster center as the current cluster center; for each cluster , recalculating the cluster center as the average value of all feature vectors in the cluster: ; setting the recalculated cluster center as the re-calculated cluster center; calculating the Euclidean distance between the current cluster center and the re-calculated cluster center, and the calculated value is the cluster center change: ; wherein, is the current cluster center, is the cluster center, is the dimension of the feature, i.e., the number of features in the feature vector of each user in the game server; If the change of the cluster center is less than a predetermined threshold, the algorithm is determined to converge, and the iteration is stopped; If the change of the cluster center is greater than or equal to the predetermined threshold, return to step 1, i.e., for each feature vector , calculate its distance from each cluster center, and assign it to the nearest cluster, and repeat the steps; where, for each behavior pattern , the logic and rules behind user behavior are mined, specifically: extract the behavior feature vector of the user from the clustering result ; prepare the label corresponding to each feature vector , which can be the behavior category of the user, game result, etc.; use and , through algorithms such as decision tree, random forest, and logistic regression, to build a model such as a classification model or a regression model; extract rules from the built model, which can explain the logic behind user behavior, for example, the branches of the decision tree can represent different behavior decision rules; evaluate the performance of the model using a validation dataset, and optimize the model according to performance indicators such as accuracy and recall, where the validation dataset is a portion of data divided from the original dataset, used to evaluate the performance of the model during training, it is mainly used to adjust the hyperparameters of the model, select the best model version, and avoid overfitting, the validation dataset does not participate in the training process of the model, so it can provide a relatively objective performance evaluation.
[0021] The embodiment also provides grouping and corresponding grouping training for users in the game server, specifically: S1, randomly extract a user from the game server as an analysis user, denoted as ; randomly extract another user from the game server as a comparison user, denoted as , where the other user is a user in the game server except the analysis user; S2, obtain the feature vectors of the analysis user and the comparison user , denoted as and ; S3, calculate the distance between the analysis user Compared with users Cosine similarity measures the angular difference between two user behavior feature vectors: ; in, For vectors The modulus length represents the user being analyzed. The overall strength of behavioral traits For vectors The modulus length indicates the comparison user The overall strength of behavioral characteristics; S4. Repeat steps S1-S3 until all other users in the game server have a cosine similarity value with the analyzed user. S5. Repeat steps S1-S4 until all users in the game server have been analyzed as users. S6. For each analytics user Treat it as a separate set; S7, Extract and analyze users All corresponding comparison users ; S8, Extract and analyze users Compared with all users Similarity matrix between : ; S9. Define a filtering function for each analysis user. Find all users with high similarity to it. : ; in, The similarity threshold is a threshold set based on business needs or data analysis results to find users with high similarity. like Then determine the user analysis Compared with this user High similarity; like Then determine the user analysis Compared with this user Low similarity; S10, Analyze the user's judgment Compared with this user All users with high similarity Merge into the same set; S11. Repeat steps S1-S10 until the similarity comparison of all users on the game server is completed, forming multiple user sets: ; Each set Users in the game share similar ways of thinking and gaming habits; S12, for each set Initialize an agent : ; in, A policy network is a mapping from a state space to an action space, providing guidance to an agent on how to act in a specific situation. Its goal is to learn the optimal policy, that is, the policy that maximizes cumulative reward in long-term operation. The value function is used to evaluate the quality of an agent in a certain state. It reflects the expected cumulative reward that the agent can obtain in the future if it follows a specific policy, starting from the current state. It helps the agent judge the quality of different states, thereby guiding the learning and optimization of the policy network. S13. For each set Collect game behavior data and feedback data from all users within this set to form a training dataset. : ; in, It is a comprehensive feature, including behavioral features. and feedback features ; S14. For each agent Conduct independent training; S15. Optimize the objective function using the PPO algorithm to maximize the agent's performance in this user set environment: ; in, These are the parameters of the policy network, defining the probability distribution of the agent's action selection under different states. Is the policy network in state Select action The probability represents the agent's preference for actions under the current policy. Under the old strategy, i.e., the strategy before the last update, select the action. The probability is used to compare changes between the old and new strategies. It is the dominance function, representing the state. Select action Advantages, measuring the choice of action Compared to the average action in the state The additional rewards that can be obtained guide the agent to choose more advantageous actions. is a clipping coefficient, used to control the maximum allowed change between the old and new policies when updating PPO, limiting the magnitude of policy updates, preventing policy updates from being too large and causing unstable training, is a policy ratio, representing the preference ratio of the new and old policies for actions , is a clipping operation, ensuring that the policy update does not exceed a predetermined range; S16, optimize the objective function by gradient ascent method, update the parameters of the policy network: , is the learning rate, is the gradient of the objective function with respect to the parameters of the policy network; S17, repeat steps S14-S15 until the performance of the agent in the user set environment reaches convergence, i.e. the average reward value of multiple rounds within a certain number of rounds fluctuates less than a set threshold, or the change in policy parameters is less than a set threshold, or the change in mean square error (MSE) of the value function (i.e. loss value) is less than a set threshold; , each agent is trained independently, specifically: perform environment interaction; if it is determined that the game is over, use generalized advantage estimation (GAE) to calculate the advantage function , is a discount factor, used to calculate the present value of future rewards, with a value range of [0, 1), close to 1, the agent pays more attention to long-term rewards, close to 0, the agent pays more attention to immediate rewards, is a smoothing parameter for GAE, used to control the smoothing degree of GAE, with a value range of [0, 1], used to balance bias and variance, is a time difference (TD) error, representing the TD error at time step t+k, used to calculate the advantage function, specifically , is the state value estimate at time step t+k+1, is the state value estimate at time step t+k, is the estimated value of the advantage function, used to measure the advantage of choosing action , over the average action value in state ;update the policy network by maximizing the objective function of PPO: , is the clipping coefficient for PPO, used to control the maximum amplitude of policy update, preventing the update from being too large and causing instability in training, is the parameter of the policy network, used to define the policy network, and the parameters are adjusted by optimization to improve the policy of the agent, is the state , the policy network selects the action with the probability, which represents the agent's preference for the action under the current policy, is the probability of selecting the action under the old policy, i.e. the policy before the last update, used to compare the changes between the new and old policies, helping to control the amplitude of policy update, is the policy ratio, representing the preference ratio of the new and old policies for the action , is the clipping operation to ensure that the policy update does not exceed the preset range, is the advantage function, used to guide the agent to select more advantageous actions; The parameters of the policy network are updated using the gradient ascent method: ; where is the learning rate, used to control the step size of policy network parameter update, affecting the convergence speed and stability of training, is the PPO objective function , the gradient of the policy network parameters , used to indicate the direction of parameter update to maximize the objective function; The value function is updated by minimizing the mean square error: ; where is the parameter of the value function network, used to define the value function network , used to estimate the value of the state, is the estimated value of the value function network in state , representing the agent's estimate of the value of the state , used to evaluate the expected return in that state, is the estimated value of the value function, representing the value function estimate at time step t, used to update the value function network, specifically , is the reward obtained at time step t, is the discount factor, used to balance the importance of immediate rewards and future rewards, is the state value estimate at time step t+1; The parameters of the value function are updated using the gradient descent method: ; where is the learning rate of the value function, used to control the step size of the value function network parameter update, affecting the convergence speed and stability of the training, is the value function loss is the gradient of the value function network parameters , used to indicate the direction of parameter update, to minimize the loss function; Repeat the overall steps of independent training for each agent , until the performance of the agent in the user set environment reaches convergence, that is, within a certain number of rounds, the fluctuation of the average reward value of multiple rounds is less than the set threshold, or the change of the policy parameter is less than the set threshold, or the change of the mean square error (MSE) of the value function (i.e. the loss value) is less than the set threshold.
[0022] The embodiment also provides that, according to the grouping training, cross-set integrated training is performed, specifically: Obtain the corresponding training data set of each set ; Integrate to form a comprehensive data set: ; Among them, contains the comprehensive feature vector of all user sets ; Initialize a global agent to train based on the comprehensive features of the entire user group: ; Among them, is the policy network, which is a mapping from the state space to the action space, providing guidance for the agent on how to act in a particular situation, and its goal is to learn the optimal policy, that is, the policy that maximizes the cumulative reward in the long run, is the value function, used to evaluate the pros and cons of the agent in a certain state, reflecting the expected cumulative reward that the agent can obtain in the future if it follows a certain policy from the current state, helping the agent to judge the pros and cons of different states, and guiding the learning and optimization of the policy network; Perform environment interaction; If it is determined that the game is over, use GAE to calculate the advantage function : ; Among them, is the discount factor, used to calculate the present value of future rewards, with a value range of [0, 1), close to 1, the agent pays more attention to long-term rewards, close to 0, the agent pays more attention to immediate rewards, This is the smoothing parameter for GAE, used to control the smoothness of GAE. Its value ranges from [0,1] and is used to balance bias and variance. The time difference (TD) error represents the TD error at time step t+k, used to calculate the dominance function. , For the state value estimation at time step t+k+1, For the state value estimation at time step t+k, This is an estimate of the advantage function, used to measure performance in state... Select action The advantage of this, relative to the value of average movement; Update the policy network by maximizing the objective function of PPO: ; in, This is the shearing factor for PPO, used to control the maximum magnitude of policy updates and prevent excessively large updates from causing training instability. These are the parameters of the policy network, used to define the policy network. By optimizing and adjusting these parameters, the agent's policy can be improved. In the state Below, policy network Select Action The probability represents the agent's preference for actions under the current policy. To select an action under the old strategy, i.e., the strategy before the last update. The probability is used to compare changes between the old and new strategies, helping to control the magnitude of strategy updates. The strategy ratio represents the ratio of the old and new strategies to actions. The preference ratio, This is a cut operation to ensure that policy updates do not exceed a preset range. This is the advantage function, used to guide the agent to choose a more advantageous action; Update the parameters of the policy network using gradient ascent: ; in, The learning rate controls the step size for updating the policy network parameters, affecting the convergence speed and stability of training. PPO objective function Policy network parameters The gradient is used to indicate the direction of parameter updates in order to maximize the objective function; Update the value function by minimizing the mean squared error: ; in, The parameters of the value function network are used to define the value function network. , for estimating the value of a state, is the estimated value of the value function network at state , represents the agent's estimate of the value of state for evaluating the expected return in that state, is the estimated value of the value function, represents the value function estimate at time step t, for updating the value function network, specifically , is the reward obtained at time step t, is the discount factor, for balancing the importance of immediate rewards and future rewards, is the state value estimate at time step t+1; update the parameters of the value function using gradient descent: ; where, is the learning rate of the value function, for controlling the step size of the value function network parameter updates, affecting the convergence speed and stability of training, is the gradient of the value function loss with respect to the value function network parameters , for indicating the direction of parameter updates to minimize the loss function; repeat the overall steps of environment interaction until the agent's performance in the user set environment reaches convergence, i.e., within a certain number of rounds, the fluctuation of the average reward value of multiple rounds is less than the set threshold, or the change in policy parameters is less than the set threshold, or the change in the mean square error (MSE) of the value function (i.e., the loss value) is less than the set threshold; the execution of environment interaction, specifically: reset the game environment to obtain the initial state ; the agent selects an action according to the current policy network in state : ; where, is the current state, representing the agent's current state in the environment, for decision-making to select actions, is the current action, represents the action selected by the agent in state , is the probability of the policy network selecting action in state , to define the agent's behavior strategy, for selecting actions; execute the action in the environment, to obtain the next state and rewards ; wherein, is a reward value, representing an immediate reward obtained by the agent after performing the action ; storing trajectory data into a buffer ; wherein, is the next state, representing a new state to which the environment transitions after performing the action , is a buffer for storing trajectory data, to save trajectory data of the agent interacting with the environment, for subsequent policy update; defining conditions for ending the game, such as the player reaching a specific game goal (e.g., scoring more than a certain value), the player failing (e.g., running out of lives, running out of time), reaching a preset game duration or number of rounds, etc. after each game step, checking whether the current game state satisfies any termination condition; if the termination condition is satisfied, determining that the game is over; if the termination condition is not satisfied, determining that the game is not over, returning the state , and repeating the above steps of environment interaction.
[0023] The embodiment also provides collecting feedback of the user on the trained agent, and forming feedback-driven iterative training based on the feedback of the user, specifically: selecting a group of users from all users in the game server for experience testing: ; wherein, is a set of all users, is a set of users selected for experience testing, is a function of selecting a group of users from all users in the game server according to predetermined conditions, such as selecting the most active users in the game server according to a set selection number, or selecting a corresponding number of users in each user group according to a proportion in the corresponding set according to a set selection number; providing a game environment consistent with the training environment for the user , and deploying the trained agent into the game environment; the user interacts with the agent , generating experience data : ; wherein, To simulate the game interaction process between the user and the global intelligent agent and to evaluate the performance of the intelligent agent, a function is developed. Through actual game interaction, interaction data between the user and the intelligent agent is collected to provide raw materials for subsequent analysis of the performance of the intelligent agent and user satisfaction. For users Each user in Collect their direct feedback : ; in, A function to obtain direct user feedback on the performance of the intelligent agent, so as to directly understand the user's needs and dissatisfactions; used to collect feedback information directly from users, so as to quickly identify the strengths and weaknesses of the intelligent agent and make targeted improvements. Collect users The indirect feedback data is then normalized. ; ; Integrating indirect feedback data : ; Direct feedback and indirect feedback Concatenate into a comprehensive feedback vector : ; Forming a comprehensive feedback dataset : ; For users Each user in Comments Sentiment analysis is performed to identify positive and negative feedback and obtain a sentiment score. : ; in, To analyze the sentiment of user reviews and quantify the user's satisfaction or dissatisfaction with the performance of the agent, this function helps development teams quickly understand the sentiment in a large number of user reviews, thereby identifying potential problems in the agent design and providing direction for further optimization. Calculate each behavioral feature Importance score : ; in, To evaluate the importance of different features in user feedback, help the team determine which features have the greatest impact on user experience, so that the function of optimizing these features can be prioritized, for quantifying the importance of each feature, helping the development team to determine the priority of improvement, concentrating resources to optimize those features that have the greatest impact on user experience, improving the overall performance of the agent and user satisfaction; Clustering the comprehensive feedback dataset to obtain a clustering result : ; According to the feedback analysis result, define the improvement target : ; wherein each improvement point represents the shortcomings of the agent in a specific scenario or behavior pattern; Filtering data related to the improvement target from the original training data : ; wherein, is a function for determining whether the data is related to the improvement target , for filtering out data useful for the improvement target; Using the improved dataset to train the agent and update the policy network and the value function : ; wherein, is a function for training using the PPO algorithm, for training and returning the improved policy network and value function according to the provided data and initial model; Update the improved policy network and value function to the global agent: ; ; Repeat the process for each user in the user , collect their direct feedback The overall steps between updating the improved policy network and value function into the global agent until the performance of the agent reaches a satisfactory level, i.e., continuously training and optimizing the agent until its performance meets the predetermined performance standards or goals, which can be specifically that the cumulative reward obtained by the agent reaches a predetermined threshold (for example, the score of the agent in the game stabilizes at a high level), in a competitive game, the win rate of the agent reaches a satisfactory level (such as a win rate of more than 70%), the performance of the agent remains stable for multiple evaluation periods without large fluctuations, the change of the strategy parameters of the agent is small, the change (i.e., loss value) of the loss function in the training process is less than a set threshold, etc.
[0024] Embodiment three is an improvement based on embodiment two. In this embodiment, other games related to the current game are identified and data collection and analysis are performed. Specifically: The type, gameplay, and user group of the current game are analyzed to identify other games related to the current game: ; Among them, is a function for identifying other games related to the current game, which is used to find other games with similar characteristics by analyzing the type, gameplay, user group, etc. of the current game, to expand the data source and provide more diverse data for model training, to improve the generalization ability of the agent, is the current game, is the set of identified related games; For each related game , user data is obtained: ; Among them, is the user data obtained from the related game , is a function for collecting user data from a specified related game , which is used to obtain user data in related games, including but not limited to user game behavior, operation record, decision selection, etc., to enrich the user behavior data set and help the agent better understand user behavior patterns in different game environments; Data collection and user behavior analysis are performed on each user in each related game , to obtain and , i.e., the steps in claim 2 and claim 3 are performed on each user in each related game ; The user data in the related games is integrated into the user data of the current game to form the final training data set: ; Among them, is the extended final training data set.
[0025] The embodiment also provides that all training data of the current game and related games are fused and finally trained to generate a general intelligent agent that can adapt to various user behaviors and habits, specifically: obtaining the final training data set ; training the final training data set performing cross-game user grouping and training, that is, performing the contents in claims 4-6 on the final training data set, specifically: comparing users of other games with users of the current game, and integrating users with high similarity into the same set; jointly training the data of each set (containing users of the current game and related games) to enable the intelligent agent to adapt to more extensive user behavior characteristics.
[0026] Embodiment four, the embodiment also discloses a system for executing a game intelligent agent design method based on deep reinforcement learning, referring to Figure 2 , comprising: a data acquisition module: acquiring game data of users of the current game and related games, that is, the game situation of the users, and acquiring feedback of the users on the game and the trained intelligent agent; a data analysis module: analyzing the game data of the users of the current game and related games, classifying the users, and classifying users with similar game operation ideas into the same set for data training; a data training module: training the intelligent agent according to the classified sets of users of the current game or related games, then training the intelligent agent according to the training of the sets, and finally training the intelligent agent according to the whole; a data integration module: integrating the training of the intelligent agent for the current game and the training of the intelligent agent for related games, and training the intelligent agent again; a data generation module: generating or updating the intelligent agent according to the training of the intelligent agent.
Claims
1. A method for designing game intelligent agents based on deep reinforcement learning, characterized in that: include: Data is collected from all users on the game server. Perform user behavior analysis on all users on the game server; Users on the game server are grouped and trained accordingly. Based on the results of group training, perform cross-set integration training; Collect user feedback on the trained agent and form feedback-driven iterative training based on the user feedback; Identify other games related to the current game and collect and analyze the data; Based on data collected and analyzed from other games related to the current game, cross-game user grouping and training are conducted; By fusing all training data from the current game and related games and performing final training, a general intelligent agent capable of adapting to various user behaviors and habits is generated.
2. The method for designing a game agent based on deep reinforcement learning according to claim 1, characterized in that: Data collection is performed on all users on the game server, specifically as follows: For each user Extract its raw behavioral data from the game server, denoted as : ; Using feature extraction functions Transform raw behavioral data into low-dimensional feature vectors : ; Collect users Direct feedback data, denoted as : ; The indirect feedback data is normalized to obtain and : ; ; Integrating direct feedback and normalized indirect feedback, a comprehensive feedback data is formed, denoted as... : ; behavioral characteristics and feedback features By splicing together, we obtain the comprehensive features. : .
3. The method for designing a game agent based on deep reinforcement learning according to claim 1, characterized in that: User behavior analysis is performed on all users on the game server, specifically as follows: Obtain the feature vector of each user in the game server ; Clustering the behavioral characteristics of all users to identify different behavioral patterns: ; For each behavioral pattern Analyze and uncover the logic and rules behind user behavior: 。 4. The method for designing a game agent based on deep reinforcement learning according to claim 1, characterized in that: Users on the game server are grouped and trained accordingly, specifically as follows: S1. Randomly extract a user from the game server, designate them as the analysis user, and denot them as... ; Randomly select another user from the game server and designate them as the comparison user, denoted as _____. ; S2, Obtain and analyze users respectively Compared with users The eigenvectors of are denoted as . and ; S3, Calculation and Analysis User Compared with users Cosine similarity: ; S4. Repeat steps S1-S3 until all other users in the game server have a cosine similarity value with the analyzed user. S5. Repeat steps S1-S4 until all users in the game server have been analyzed as users. S6. For each analytics user Treat it as a separate set; S7, Extract and analyze users All corresponding comparison users ; S8, Extract and analyze users Compared with all users Similarity matrix between : ; S9. Define a filtering function for each analysis user. Find all users with high similarity to it. : ; like Then determine the user analysis Compared with this user High similarity; like Then determine the user analysis Compared with this user Low similarity; S10, Analyze the user's judgment Compared with this user All users with high similarity Merge into the same set; S11. Repeat steps S1-S10 until the similarity comparison of all users on the game server is completed, forming multiple user sets: ; S12, for each set Initialize an agent : ; S13. For each set Collect game behavior data and feedback data from all users within this set to form a training dataset. : ; S14. For each agent Conduct independent training; S15. Maximize the agent's performance in this user set environment: ; S16. Optimize the objective function and update the parameters of the policy network: ; S17. Repeat steps S14-S15 until the agent's performance in the user set environment converges.
5. The method for designing a game agent based on deep reinforcement learning according to claim 1, characterized in that: Based on the group training results, cross-set integration training is performed, specifically as follows: Get each collection Corresponding training dataset This will be integrated into a comprehensive dataset. ; Initialize a global agent : ; Execution environment interaction; After the game ends, use GAE to calculate the advantage function. : ; Update the policy network by maximizing the objective function of PPO: ; Update the parameters of the policy network using gradient ascent: ; Update the value function by minimizing the mean squared error: ; Update the parameters of the value function using gradient descent: .
6. The method for designing a game agent based on deep reinforcement learning according to claim 1, characterized in that: Collect user feedback on the trained agent and use this feedback to generate feedback-driven iterative training, specifically: Select a group of users from all users on the game server. Conduct experience testing: ; For users Provide a game environment consistent with the training environment, and then train the agent. Deploy to the game environment; user With intelligent agents Interaction, generating experience data : ; For users Each user in Collect their direct feedback : ; Collect users The indirect feedback data is then normalized. ; ; Integrating indirect feedback data : ; Direct feedback and indirect feedback Concatenate into a comprehensive feedback vector : ; Forming a comprehensive feedback dataset : ; For users Each user in Comments Sentiment analysis is performed to identify positive and negative feedback and obtain a sentiment score. : ; Calculate each behavioral feature Importance score : ; For the comprehensive feedback dataset Perform clustering to obtain clustering results : ; Based on the feedback analysis results, define improvement goals. : ; Select data relevant to the improvement goals from the original training data. : ; Using the improved dataset Targeted training of the agent and updating of the policy network and value function : ; Update the improved policy network and value function to the global agent: ; ; Repeated execution for users Each user in Collect their direct feedback The overall steps from updating the improved policy network and value function to the global agent are continued until the agent's performance reaches a satisfactory level.
7. The method for designing a game agent based on deep reinforcement learning according to claim 1, characterized in that: Identify other games related to the current game and collect and analyze the data, specifically: Analyze the current game's genre, gameplay, and user base to identify other games related to it: ; For each relevant game To obtain user data separately: ; For each relevant game Data is collected and user behavior is analyzed for each user in the process, resulting in... and ; The user data from related games is integrated into the user data of the current game to form the final training dataset: .
8. The game agent design method based on deep h-reinforcement learning according to claim 1, characterized in that: The training data from the current game and related games are merged and then trained to generate a general-purpose intelligent agent capable of adapting to various user behaviors and habits. Specifically: Obtain the final training dataset ; For the final training dataset Conduct cross-game user grouping and training.
9. A system for implementing the game agent design method based on deep reinforcement learning as described in claim 1, characterized in that: include: Data acquisition module: Collects game data from users of the current game and related games, and collects user feedback on the game and the trained agent; Data Analysis Module: Analyzes game data of users of the current game and related games, classifies users, and groups users with similar game operation strategies into the same set for data training; Data training module: Based on the user classification set of the current game or related games, first train the agent on the data within each set, and then train the agent as a whole according to the training status of the sets; Data integration module: Integrates the training data of the agent for the current game and the training data of the agent for related games, and then trains the agent again; Data generation module: Generates or updates the agent based on the training results.
Citation Information
Patent Citations
Cloud game engine intelligent optimization method and device based on reinforcement learning
CN112121439A
Game type-based data management method and device
CN117258305A
An online evolutionary algorithm and training platform for intelligent agents based on self-game
CN119783778A
Game dynamic adjustment method, system and equipment based on emotion prediction and medium
CN120437636A
Dynamic role customization intelligent customer service system driven by multi-agent user portraits
CN120543179A