Role personalized behavior generation method based on reinforcement learning
By introducing reinforcement learning, diffusion transformers and positive and negative spiral whale search algorithms in the generation of personalized behaviors of roles, combined with environmental feedback and large language model evaluation, the stability and diversity problems in the generation of personalized behaviors are solved, and a more natural and intelligent role behavior is achieved.
Patent Information
- Application Number
- CN202510557948.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has problems in the generation of personalized role behaviors, insufficient personal characteristics modeling, poor behavior generation stability, and insufficient diversity in behavior optimization.
A reinforcement learning-based method is adopted, combined with the diffusion transformer and the forward and reverse spiral whale search algorithm, and optimized behavior strategies are evaluated through environmental feedback and large language models, and the character personality characteristics parameters are dynamically adjusted to generate a personalized behavior sequence.
It improves the diversity and stability of role behavior, ensures the authenticity, nature and personality matching of behavior, adapts to changes in complex environments, and avoids local optimal problems.
Smart Images

Figure CN120106129A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of reinforcement learning, and in particular to a method for generating personalized role behaviors based on reinforcement learning. Background Art
[0002] With the development of artificial intelligence technology, personalized character behavior modeling has been widely used in intelligent dialogue systems, game AI, virtual character creation and other fields. Existing character behavior generation methods mainly rely on rule-driven or statistical learning-based strategies. Rule-based methods drive the character's behavioral response through predefined behavioral logic and state transition rules, but such methods have low flexibility and are difficult to adapt to complex and changing interactive environments. At the same time, statistical learning-based methods, such as traditional Markov decision processes (MDPs) and hidden Markov models (HMMs), although they can simulate the character's behavior patterns to a certain extent, still lack deep modeling of personalized features, making it difficult to ensure the naturalness and consistency of character behavior.
[0003] In recent years, the progress of deep learning and reinforcement learning technologies has provided new ideas for the personalized behavior modeling of characters. Through the nonlinear mapping ability of neural networks, deep learning can learn complex behavior patterns and generate behavior sequences that match the personality of the character. However, traditional deep learning methods usually rely on large-scale data for training and lack effective personalized feature control mechanisms, resulting in the lack of stability of the generated behaviors. In addition, reinforcement learning, as a learning method based on reward feedback, can optimize behavioral strategies during the interaction process so that the character behavior can continuously adapt to environmental changes. However, existing reinforcement learning models still face great challenges in personalized modeling, mainly manifested in the balance between the stability of personality characteristics and the diversity of behavior and consistency of personality.
[0004] Existing technologies still have many deficiencies in generating personalized behaviors. First, the modeling of personality characteristics is not sophisticated enough. Traditional methods mainly rely on fixed rules or simple feature mappings, and cannot dynamically adjust the character personality parameters, which limits the adaptability and plasticity of the character. Secondly, the stability of behavior generation is poor. Existing generation models are usually based on end-to-end training methods, which fail to fully consider the long-term consistency of behavior sequences, resulting in the generated behavior may be offset during long-term interactions, affecting the user experience. In addition, in the process of optimizing personalized behaviors, existing methods usually adopt strategies based on gradient optimization, but these methods are often prone to falling into local optimality and lack global search capabilities, resulting in insufficient diversity in behavior optimization and difficulty in meeting the needs of character personalization in different scenarios.
[0005] Therefore, how to provide a method for generating personalized character behaviors based on reinforcement learning is an urgent problem to be solved by those skilled in the art. Summary of the invention
[0006] One purpose of the present invention is to propose a method for generating personalized character behaviors based on reinforcement learning. The present invention combines reinforcement learning, diffusion transformers, and forward and reverse spiral whale search algorithms to achieve personalized character behavior generation, and optimizes behavior strategies through environmental feedback and large language model evaluation. This method can improve the diversity and stability of behavior generation while ensuring the authenticity, naturalness, and personality matching of character behavior. By dynamically adjusting personality feature parameters, character behavior can adapt to long-term interactions and complex environmental changes, avoiding the defects of traditional methods that are prone to fall into local optimality, single or unstable behavior patterns, and is widely applicable to the fields of intelligent interactive systems, game AI, and virtual character creation.
[0007] A method for generating personalized role behaviors based on reinforcement learning according to an embodiment of the present invention includes the following steps: S1. Collect the historical behavior data of the role, perform data cleaning and feature extraction, and then construct the character personality feature vector. Use Transformer to embed the character personality feature vector into a behavior sequence to generate a behavior pattern sequence. S2, using the behavior pattern sequence to train the diffusion transformer, establish a probability distribution model of the role's personalized behavior, and generate an initial personalized behavior sequence by adding noise to the behavior pattern sequence and gradually denoising it; S3, using forward and reverse spiral whale search algorithms to optimize the initial personalized behavior sequence, wherein the forward spiral search is used to generate a variety of candidate behavior sequences, and the reverse spiral search is used to adjust the character personality characteristic parameters to stabilize the character behavior pattern, and output the optimized personalized behavior sequence; S4, using the optimized personalized behavior sequence and the feedback information of the character's environment, the character's personality characteristic parameters are adjusted dynamically for a second time to generate a preliminary personalized behavior strategy; S5. Use a large language model to evaluate personalized behavior strategies, obtain evaluation results of the authenticity, naturalness, and personality matching of the behavior strategies, optimize the personalized behavior strategies based on the evaluation results and human feedback information, and output the final personalized behavior strategies.
[0008] Optionally, the character's historical behavior data includes dialogue content, action information, and emotional reactions.
[0009] Optionally, the S2 specifically includes: S21. Obtaining behavior pattern sequence , define the time step ,in , Represents the maximum time step and constructs the composite diffusion coefficient sequence: ; in, represents the composite diffusion coefficient sequence, represents the exponential function, represents the reference noise adjustment constant, represents the exponential decay adjustment parameter, represents the offset adjustment parameter, represents the integral variable; S22, based on the composite diffusion coefficient sequence , generate a noise injection behavior sequence: ; in, Indicates that at time step The noise injection behavior sequence, Indicates that at time step Random noise sampled from a multivariate normal distribution; S23, construct a denoising mapping network, and define a denoising mapping function of the denoising mapping network: ; in, Indicates that at time step Denoising mapping function for denoising noise-injected behavior sequences, represents the total number of layers of the denoising mapping network, Indicates The time decay parameter of the layer, Indicates The layer has a pre-set feature mapping function; S24, using the denoising mapping function to implement a multi-step denoising inverse process on the noise injection behavior sequence, starting from the maximum time step Decrease to 0, and gradually restore the original behavior pattern characteristics; S25, in the inverse denoising process, dynamically correct the behavior sequence output at each stage according to the recorded noise injection information, and filter out abnormal noise interference; S26. Integrate the inverse process outputs of each time step to generate an initial personalized behavior sequence after denoising and restoration.
[0010] Optionally, the S3 specifically includes: S31, receiving an initial personalized behavior sequence as an input to the forward and reverse spiral whale search algorithm optimization process; S32, using forward spiral search to generate candidate behavior sequences, the definition formula is: ; in, Indicates The number is forward searched for candidate behavior sequences, represents the initial personalized behavior sequence, Indicates the candidate behavior sequence number, represents the exponential function, represents the diffusion factor, represents the exponential factor, represents the pitch parameter, represents the amplitude coefficient; S33. Candidate behavior sequence generated by forward search Perform normalization and sorting to form a preliminary candidate set; S34, using reverse spiral search to adjust the character personality characteristic parameters of each candidate behavior sequence in the preliminary candidate set, the definition formula is: ; in, Indicates Number reverse search candidate behavior sequence, represents the stability factor, represents the total number of iteration steps, represents the convergence index, represents the compensation constant, represents the attenuation parameter; S35, the candidate behavior sequence adjusted by reverse search Perform multiple rounds of iterative fusion to generate optimized personalized behavior sequences: ; in, represents the optimized personalized behavior sequence, Indicates Round weighting coefficient, Indicates Round Number reverse search candidate behavior sequence.
[0011] Optionally, the S4 specifically includes: S41. Receive optimized personalized behavior sequence , extract the character's current environment state and construct the environment feedback matrix, where the environment feedback matrix is defined as , represents the total number of environmental factors, Indicates environmental factors; S42. Calculate the matching degree based on the optimized personalized behavior sequence and the environmental feedback matrix: ; in, It represents the matching degree of the optimized personalized behavior sequence under the current environmental feedback. represents the optimized personalized behavior sequence length, Indicates The weight of a behavior unit, Indicates Behavior Unit In the environmental feedback matrix The following adaptability scores: ; in, represents the hyperbolic tangent function, Indicates The weight of the impact of environmental factors on behavior matching, Indicates Environmental factors affect the amplitude adjustment parameters. represents the offset term, Indicates Behavior unit and The characteristic distance of each environmental factor; S43. Based on the matching degree of the optimized personalized behavior sequence under the current environmental feedback, the adjustment gradient of the character personality characteristic parameters is calculated, and a dynamic step size update strategy is used to perform a secondary dynamic adjustment of the character personality characteristic parameters: ; in, Indicates The individual parameters of the iteration, Indicates The individual parameters of the iteration, represents the dynamic learning rate, Represents the adjustment gradient of the character's personality trait parameters; S44, performing boundary constraints on the character personality characteristic parameters after the secondary dynamic adjustment to ensure that they are within the character personality setting range: ; in, Indicates the lower limit of the character's personality characteristic parameters, Indicates the upper limit of the character's personality characteristic parameters. represents the minimum function, represents the maximum value function; S45, recalculate the personalized behavior sequence using the character personality characteristic parameters after the secondary dynamic adjustment, and and the environmental feedback matrix Perform secondary fitness matching, if The increment is below the threshold Then the optimization is stopped, otherwise the process returns to step S43 to continue adjusting and generate a preliminary personalized behavior strategy.
[0012] Optionally, the S5 specifically includes: S51, obtaining a preliminary personalized behavior strategy, and inputting the personalized behavior strategy into a large language model for semantic parsing and context matching, to generate a preliminary evaluation result for authenticity, naturalness, and personality matching; S52, converting the evaluation results of the large language model into quantifiable evaluation data, and establishing an evaluation indicator set corresponding to the role characteristics, marking and recording each indicator; S53, combining human feedback information, screening and summarizing the evaluation index set, filtering abnormal or invalid evaluation items, forming comprehensive evaluation data, and recording the score or grade of the role behavior strategy under each evaluation index; S54. Based on the comprehensive evaluation data, delete the behavior content that is obviously in conflict with the role setting, optimize the personalized behavior strategy, and output the final personalized behavior strategy.
[0013] The beneficial effects of the present invention are: First, the present invention realizes the dynamic adjustment of character behavior by introducing reinforcement learning and personalized behavior modeling technology, so that it can adapt to different personality settings and interaction scenarios. Compared with traditional rule-based or statistical learning methods, the present invention uses Transformer to extract features and embed behavior sequences from historical behavior data of characters, so that the long-term consistency of character behavior can be maintained. At the same time, the introduction of diffusion transformers enables the generation of personalized character behaviors based on probability distribution models, which not only improves the generation quality of behaviors, but also ensures the controllability and stability of behavior sequences.
[0014] Secondly, the present invention uses forward and reverse spiral whale search algorithms to optimize the initial personalized behavior sequence, while improving the diversity of behavior generation and ensuring the stability of the character's personality characteristics. The forward spiral search is used to discover new behavior patterns, thereby expanding the character's behavior selection space, while the reverse spiral search is used to adjust personality parameters to prevent drastic changes in behavior patterns and ensure that the behavior conforms to the characteristics set for the character. This optimization strategy overcomes the defects of existing methods that are prone to falling into local optimality and single behavior patterns during the behavior optimization process, making the character behavior richer and more layered.
[0015] In addition, the present invention combines environmental feedback information to dynamically adjust the character personality characteristic parameters, so that the character behavior is not only generated according to the historical behavior pattern, but also can be adaptively optimized according to the current interactive environment. This adjustment mechanism ensures that the character behavior remains stable during long-term interaction and can be optimized as the environment changes, thereby avoiding the problem of behavioral deviation or behavior distortion caused by environmental changes in traditional methods. At the same time, the personality characteristic parameter update mechanism based on reinforcement learning allows the character personality to evolve over time, ensuring that the character's behavior in different scenarios always conforms to the established personality settings.
[0016] Finally, in the process of evaluating and optimizing personalized behavior strategies, the present invention uses a large language model to conduct multiple rounds of evaluation on the behavior strategy to ensure that the generated behavior meets high quality standards in terms of authenticity, naturalness, and personality matching. Through the deep semantic analysis and context understanding capabilities of the large language model, it is possible to accurately evaluate whether the character behavior is consistent with the personality setting, and optimize it in combination with human feedback information to reduce behaviors that do not meet interaction expectations or are unnatural. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 This is an overall flow chart of a method for generating personalized role behaviors based on reinforcement learning proposed by the present invention. DETAILED DESCRIPTION
[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.
[0019] refer to Figure 1 , a method for generating personalized role behaviors based on reinforcement learning, comprising the following steps: S1. Collect the historical behavior data of the role, perform data cleaning and feature extraction, and then construct the character personality feature vector. Use Transformer to embed the character personality feature vector into a behavior sequence to generate a behavior pattern sequence. S2, using the behavior pattern sequence to train the diffusion transformer, establish a probability distribution model of the role's personalized behavior, and generate an initial personalized behavior sequence by adding noise to the behavior pattern sequence and gradually denoising it; S3, using forward and reverse spiral whale search algorithms to optimize the initial personalized behavior sequence, wherein the forward spiral search is used to generate a variety of candidate behavior sequences, and the reverse spiral search is used to adjust the character personality characteristic parameters to stabilize the character behavior pattern, and output the optimized personalized behavior sequence; S4, using the optimized personalized behavior sequence and the feedback information of the character's environment, the character's personality characteristic parameters are adjusted dynamically for a second time to generate a preliminary personalized behavior strategy; S5. Use a large language model to evaluate personalized behavior strategies, obtain evaluation results of the authenticity, naturalness, and personality matching of the behavior strategies, optimize the personalized behavior strategies based on the evaluation results and human feedback information, and output the final personalized behavior strategies.
[0020] In this embodiment, the character's historical behavior data includes dialogue content, action information, and emotional response.
[0021] In this implementation, S2 specifically includes: S21. Obtaining behavior pattern sequence , define the time step ,in , Represents the maximum time step and constructs the composite diffusion coefficient sequence: ; in, represents the composite diffusion coefficient sequence, represents the exponential function, represents the reference noise adjustment constant, represents the exponential decay adjustment parameter, represents the offset adjustment parameter, represents the integral variable; S22, based on the composite diffusion coefficient sequence , generate a noise injection behavior sequence: ; in, Indicates that at time step The noise injection behavior sequence, Indicates that at time step Random noise sampled from a multivariate normal distribution; S23, construct a denoising mapping network, and define a denoising mapping function of the denoising mapping network: ; in, Indicates that at time step Denoising mapping function for denoising noise-injected behavior sequences, represents the total number of layers of the denoising mapping network, Indicates The time decay parameter of the layer, Indicates The layer has a pre-set feature mapping function; S24, using the denoising mapping function to implement a multi-step denoising inverse process on the noise injection behavior sequence, starting from the maximum time step Decrease to 0, and gradually restore the original behavior pattern characteristics; S25, in the inverse denoising process, dynamically correct the behavior sequence output at each stage according to the recorded noise injection information, and filter out abnormal noise interference; S26. Integrate the inverse process outputs of each time step to generate an initial personalized behavior sequence after denoising and restoration.
[0022] In this implementation, S3 specifically includes: S31, receiving an initial personalized behavior sequence as an input to the forward and reverse spiral whale search algorithm optimization process; S32, using forward spiral search to generate candidate behavior sequences, the definition formula is: ; in, Indicates The number is forward searched for candidate behavior sequences, represents the initial personalized behavior sequence, Indicates the candidate behavior sequence number, represents the exponential function, represents the diffusion factor, represents the exponential factor, represents the pitch parameter, represents the amplitude coefficient; S33. Candidate behavior sequence generated by forward search Perform normalization and sorting to form a preliminary candidate set; S34, using reverse spiral search to adjust the character personality characteristic parameters of each candidate behavior sequence in the preliminary candidate set, the definition formula is: ; in, Indicates Number reverse search candidate behavior sequence, represents the stability factor, represents the total number of iteration steps, represents the convergence index, represents the compensation constant, represents the attenuation parameter; S35, the candidate behavior sequence adjusted by reverse search Perform multiple rounds of iterative fusion to generate optimized personalized behavior sequences: ; in, represents the optimized personalized behavior sequence, Indicates Round weighting coefficient, Indicates Round Number reverse search candidate behavior sequence.
[0023] In this implementation manner, the S4 specifically includes: S41. Receive optimized personalized behavior sequence , extract the character's current environment state and construct the environment feedback matrix, where the environment feedback matrix is defined as , represents the total number of environmental factors, Indicates environmental factors; S42. Calculate the matching degree based on the optimized personalized behavior sequence and the environmental feedback matrix: ; in, It represents the matching degree of the optimized personalized behavior sequence under the current environmental feedback. represents the optimized personalized behavior sequence length, Indicates The weight of a behavior unit, Indicates Behavior Unit In the environmental feedback matrix The following adaptability scores: ; in, represents the hyperbolic tangent function, Indicates The weight of the impact of environmental factors on behavior matching, Indicates Environmental factors affect the amplitude adjustment parameters. represents the offset term, Indicates Behavior unit and The characteristic distance of each environmental factor; S43. Based on the matching degree of the optimized personalized behavior sequence under the current environmental feedback, the adjustment gradient of the character personality characteristic parameters is calculated, and a dynamic step size update strategy is used to perform a secondary dynamic adjustment of the character personality characteristic parameters: ; in, Indicates The individual parameters of the iteration, Indicates The individual parameters of the iteration, represents the dynamic learning rate, Represents the adjustment gradient of the character's personality trait parameters; S44, performing boundary constraints on the character personality characteristic parameters after the secondary dynamic adjustment to ensure that they are within the character personality setting range: ; in, Indicates the lower limit of the character's personality characteristic parameters, Indicates the upper limit of the character's personality characteristic parameters. represents the minimum function, represents the maximum value function; S45, recalculate the personalized behavior sequence using the character personality characteristic parameters after the secondary dynamic adjustment, and and the environmental feedback matrix Perform secondary fitness matching, if The increment is below the threshold Then the optimization is stopped, otherwise the process returns to step S43 to continue adjusting and generate a preliminary personalized behavior strategy.
[0024] In this implementation manner, S5 specifically includes: S51, obtaining a preliminary personalized behavior strategy, and inputting the personalized behavior strategy into a large language model for semantic parsing and context matching, to generate a preliminary evaluation result for authenticity, naturalness, and personality matching; S52, converting the evaluation results of the large language model into quantifiable evaluation data, and establishing an evaluation indicator set corresponding to the role characteristics, marking and recording each indicator; S53, combining human feedback information, screening and summarizing the evaluation index set, filtering abnormal or invalid evaluation items, forming comprehensive evaluation data, and recording the score or grade of the role behavior strategy under each evaluation index; S54. Based on the comprehensive evaluation data, delete the behavior content that is obviously in conflict with the role setting, optimize the personalized behavior strategy, and output the final personalized behavior strategy.
[0025] Embodiment 1: In order to verify the feasibility of the present invention in implementation, the present invention is applied to the personalized behavior generation task of virtual characters in a certain intelligent interactive system, which is used for personalized interaction of intelligent customer service, game NPC and intelligent assistant. Traditional character behavior generation methods usually adopt preset rules or statistical learning-based methods, but there are problems such as single behavior pattern, lack of personalization, and difficulty in adapting to environmental changes. For example, in the intelligent customer service scenario, different users have different communication methods, emotional expressions and needs, and traditional intelligent customer service systems can often only interact based on fixed reply templates or simple machine learning algorithms, resulting in poor user experience and lack of naturalness and personalized matching. In addition, game NPCs should have different behavioral feedback under the operation of different players, while traditional NPC behaviors are often preset. No matter how the player's behavior changes, the NPC's feedback lacks flexibility and cannot adapt to personalized interaction needs. The present invention constructs a method for generating personalized character behaviors through technologies such as reinforcement learning, diffusion transformers, and positive and negative spiral whale search algorithms to improve the naturalness, personality matching and environmental adaptability of character behaviors.
[0026] In the intelligent customer service system, the application process of the present invention includes four stages: data collection, behavior modeling, optimization and behavior evaluation. First, the system collects the user's historical conversation data, including text communication content, user emotion tags and interaction patterns, and embeds the behavior sequence through Transformer to generate the character personality feature vector and behavior pattern sequence. Subsequently, the probability distribution model of the character behavior is trained using the diffusion transformer, and the initial personalized behavior sequence is generated through the noise injection and denoising process. In order to enhance the diversity and stability of the behavior, the system uses the forward and reverse spiral whale search algorithm to optimize the behavior strategy, in which the forward spiral search explores new behavior patterns, and the reverse spiral search adjusts the personality parameters to maintain the stability of the behavior. After the optimization of the role behavior is completed, the system combines the environmental feedback to make a secondary dynamic adjustment to the personality parameters, so that the role behavior can adapt to the real-time needs of the user, and adjust the personality feature parameters according to the interaction habits of different users. Finally, the large language model is used to evaluate the authenticity, naturalness and personality matching of the behavior strategy, and the personalized behavior strategy is optimized in combination with human feedback to ensure that the final output behavior meets the user's expectations and improve the interactive experience.
[0027] In order to verify the effectiveness of the present invention in different scenarios, we conducted experiments in a certain intelligent customer service system in Beijing and a certain game company in Shanghai from March to June 2024. In the intelligent customer service experiment, we selected 5,000 users for A / B testing and randomly divided the users into two groups, one using the traditional intelligent customer service system (control group) and the other using the personalized behavior generation system based on the present invention (experimental group). During the three-month experimental period, we tracked core indicators such as user satisfaction, naturalness of responses, problem solving rate, and average number of conversation turns. In the game NPC behavior test, we deployed two groups of NPC behavior models in a role-playing game (RPG), one using the traditional rule-based behavior control model, and the other using the method of the present invention to generate NPC interaction behaviors through reinforcement learning and personalized behavior optimization, and compared and analyzed indicators such as player immersion, NPC behavior diversity, and in-game interaction experience.
[0028] The experimental results show that in the intelligent customer service scenario, the user satisfaction of the experimental group increased by 18.7%, the response naturalness score increased from 3.8 points to 4.5 points (out of 5 points), the problem solving rate increased from 71.3% to 86.5%, and the average number of conversation rounds decreased by 27.2%, indicating that users can get satisfactory responses faster. In the game NPC behavior test, the game players who adopted the NPC interactive behavior model of the present invention had a game immersion score that was 22.5% higher than that of the control group, the NPC behavior diversity score increased by 31.4%, and the average game time of the players increased from 72.5 minutes to 89.3 minutes. These data show that the present invention can not only improve the authenticity and naturalness of character behavior, but also effectively improve user experience and interaction satisfaction, making the intelligent system more intelligent and personalized.
[0029] The application of the present invention is not only applicable to intelligent customer service and game AI, but also has broad application prospects in the fields of intelligent assistants, virtual human interaction, educational AI, etc. Through reinforcement learning and personalized behavior optimization, the present invention can solve the problems of single behavior, poor adaptability, and insufficient personalization in traditional intelligent role behavior modeling, and provide a more natural and intelligent personalized behavior modeling method for intelligent interactive systems.
[0030] Table 1 Comparison of test data of intelligent customer service and game NPC interaction
[0031] From the experimental data in Table 1 above, the present invention has shown obvious advantages in both intelligent customer service and game NPC behavior optimization. In the intelligent customer service scenario, user satisfaction increased from 75.2% of the traditional method to 89.3%, an increase of 18.7%. This result shows that after adopting the method of the present invention, intelligent customer service can understand user needs more accurately and provide responses that are more in line with personalized expectations, so that users can have a better experience in the process of interacting with customer service. At the same time, the naturalness score of the response increased from 3.8 points to 4.5 points (out of 5 points), an increase of 18.4%, indicating that the intelligent customer service system using the method of the present invention is closer to human communication habits in terms of language fluency and expression, reducing mechanized or templated responses, and making the conversation more real and natural. In addition, the problem solving rate increased from 71.3% to 86.5%, an increase of 21.3%, indicating that intelligent customer service can not only understand user problems more accurately, but also provide effective solutions more quickly, reducing the situation of repeated consultation by users. In terms of interaction efficiency, after adopting the method of the present invention, the average number of conversation rounds was reduced from 6.2 to 4.5, a decrease of 27.2%. This means that users can obtain effective information in a shorter conversation process, reduce unnecessary communication costs, and improve customer service response efficiency.
[0032] In the game NPC behavior optimization test, the present invention also showed significant improvement. The player immersion score increased from 7.1 points to 8.7 points (out of 10 points), an increase of 22.5%, which shows that the NPC behavior using the method of the present invention is more realistic, and can produce scene-compliant reactions according to different player operations, allowing players to feel a stronger sense of reality and interaction during the game. At the same time, the NPC behavior diversity score increased from 6.5 points to 8.5 points, an increase of 31.4%. This data shows that after adopting the method of the present invention, the behavior of the game NPC is no longer a single fixed mode, but can be personalized according to the player's behavior, making the game world richer and more vivid, and improving the playability of the game and the depth of character creation. In addition, in terms of player game experience, the method of the present invention increases the average game time of players from 72.5 minutes to 89.3 minutes, an increase of 23.1%. This result shows that because the personalized behavior of NPC enhances the player's interactive experience, players are more willing to invest more time in the game, which has important commercial value for game developers.
[0033] On the whole, the present invention shows outstanding advantages in both intelligent customer service and game NPC behavior optimization. In the intelligent customer service scenario, the present invention significantly improves the intelligence of the customer service system, enabling it to understand user intentions more accurately, improve the naturalness of conversations, and optimize service efficiency. In terms of game NPC behavior optimization, the present invention successfully improves the intelligence of NPCs, giving them higher behavioral diversity and personality matching, thereby enhancing the player's immersion and gaming experience. These experimental data verify the feasibility and superiority of the present invention in generating personalized character behaviors, providing technical support for the development of future intelligent interactive systems, and also providing a more intelligent and personalized solution for fields such as intelligent customer service and game AI.
[0034] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A method for generating personalized role behaviors based on reinforcement learning, characterized in that: The steps include: S1. Collect the historical behavior data of the role, perform data cleaning and feature extraction, and then construct the character personality feature vector. Use Transformer to embed the character personality feature vector into a behavior sequence to generate a behavior pattern sequence. S2, using the behavior pattern sequence to train the diffusion transformer, establish a probability distribution model of the role's personalized behavior, and generate an initial personalized behavior sequence by adding noise to the behavior pattern sequence and gradually denoising it; S3, using forward and reverse spiral whale search algorithms to optimize the initial personalized behavior sequence, wherein the forward spiral search is used to generate a variety of candidate behavior sequences, and the reverse spiral search is used to adjust the character personality characteristic parameters to stabilize the character behavior pattern, and output the optimized personalized behavior sequence; S4, using the optimized personalized behavior sequence and the feedback information of the character's environment, the character's personality characteristic parameters are adjusted dynamically for a second time to generate a preliminary personalized behavior strategy; S5. Use a large language model to evaluate personalized behavior strategies, obtain evaluation results of the authenticity, naturalness, and personality matching of the behavior strategies, optimize the personalized behavior strategies based on the evaluation results and human feedback information, and output the final personalized behavior strategies.
2. The method for generating personalized role behaviors based on reinforcement learning according to claim 1, characterized in that: The character's historical behavior data includes dialogue content, action information, and emotional reactions.
3. The method for generating personalized role behaviors based on reinforcement learning according to claim 1, characterized in that: The S2 specifically includes: S21. Obtaining behavior pattern sequence , define the time step ,in , Represents the maximum time step and constructs the composite diffusion coefficient sequence: ; in, represents the composite diffusion coefficient sequence, represents the exponential function, represents the reference noise adjustment constant, represents the exponential decay adjustment parameter, represents the offset adjustment parameter, represents the integral variable; S22, based on the composite diffusion coefficient sequence , generate a noise injection behavior sequence: ; in, Indicates that at time step The noise injection behavior sequence, Indicates that at time step Random noise sampled from a multivariate normal distribution; S23, construct a denoising mapping network, and define a denoising mapping function of the denoising mapping network: ; in, Indicates that at time step Denoising mapping function for denoising noise-injected behavior sequences, represents the total number of layers of the denoising mapping network, Indicates The time decay parameter of the layer, Indicates The layer has a pre-set feature mapping function; S24, using the denoising mapping function to implement a multi-step denoising inverse process on the noise injection behavior sequence, starting from the maximum time step Decrease to 0, and gradually restore the original behavior pattern characteristics; S25, in the inverse denoising process, dynamically correct the behavior sequence output at each stage according to the recorded noise injection information, and filter out abnormal noise interference; S26. Integrate the inverse process outputs of each time step to generate an initial personalized behavior sequence after denoising and restoration.
4. The method for generating personalized role behaviors based on reinforcement learning according to claim 1, characterized in that: The S3 specifically includes: S31, receiving an initial personalized behavior sequence as an input to the forward and reverse spiral whale search algorithm optimization process; S32, using forward spiral search to generate candidate behavior sequences, the definition formula is: ; in, Indicates The number is forward searched for candidate behavior sequences, represents the initial personalized behavior sequence, Indicates the candidate behavior sequence number, represents the exponential function, represents the diffusion factor, represents the exponential factor, represents the pitch parameter, represents the amplitude coefficient; S33. Candidate behavior sequence generated by forward search Perform normalization and sorting to form a preliminary candidate set; S34, using reverse spiral search to adjust the character personality characteristic parameters of each candidate behavior sequence in the preliminary candidate set, the definition formula is: ; in, Indicates Number reverse search candidate behavior sequence, represents the stability factor, represents the total number of iteration steps, represents the convergence index, represents the compensation constant, represents the attenuation parameter; S35, the candidate behavior sequence adjusted by reverse search Perform multiple rounds of iterative fusion to generate optimized personalized behavior sequences: ; in, represents the optimized personalized behavior sequence, Indicates Round weighting coefficient, Indicates Round Number reverse search candidate behavior sequence.
5. The method for generating personalized role behaviors based on reinforcement learning according to claim 1, characterized in that: The S4 specifically includes: S41. Receive optimized personalized behavior sequence , extract the character's current environment state and construct the environment feedback matrix, where the environment feedback matrix is defined as , represents the total number of environmental factors, Indicates environmental factors; S42. Calculate the matching degree based on the optimized personalized behavior sequence and the environmental feedback matrix: ; in, It represents the matching degree of the optimized personalized behavior sequence under the current environmental feedback. represents the optimized personalized behavior sequence length, Indicates The weight of a behavior unit, Indicates Behavior Unit In the environmental feedback matrix The following adaptability scores: ; in, represents the hyperbolic tangent function, Indicates The weight of the impact of environmental factors on behavior matching, Indicates Environmental factors affect the amplitude adjustment parameters. represents the offset term, Indicates Behavior unit and The characteristic distance of each environmental factor; S43. Based on the matching degree of the optimized personalized behavior sequence under the current environmental feedback, the adjustment gradient of the character personality characteristic parameters is calculated, and a dynamic step size update strategy is used to perform a secondary dynamic adjustment of the character personality characteristic parameters: ; in, Indicates The individual parameters of the iteration, Indicates The individual parameters of the iteration, represents the dynamic learning rate, Represents the adjustment gradient of the character's personality trait parameters; S44, performing boundary constraints on the character personality characteristic parameters after the secondary dynamic adjustment to ensure that they are within the character personality setting range: ; in, Indicates the lower limit of the character's personality characteristic parameters, Indicates the upper limit of the character's personality characteristic parameters. represents the minimum function, represents the maximum value function; S45, recalculate the personalized behavior sequence using the character personality characteristic parameters after the secondary dynamic adjustment, and and the environmental feedback matrix Perform secondary fitness matching, if The increment is below the threshold Then the optimization is stopped, otherwise the process returns to step S43 to continue adjusting and generate a preliminary personalized behavior strategy.
6. The method for generating personalized role behaviors based on reinforcement learning according to claim 1, characterized in that: The S5 specifically includes: S51, obtaining a preliminary personalized behavior strategy, and inputting the personalized behavior strategy into a large language model for semantic parsing and context matching, to generate a preliminary evaluation result for authenticity, naturalness, and personality matching; S52, converting the evaluation results of the large language model into quantifiable evaluation data, and establishing an evaluation indicator set corresponding to the role characteristics, marking and recording each indicator; S53, combining human feedback information, screening and summarizing the evaluation index set, filtering abnormal or invalid evaluation items, forming comprehensive evaluation data, and recording the score or grade of the role behavior strategy under each evaluation index; S54. Based on the comprehensive evaluation data, delete the behavior content that is obviously in conflict with the role setting, optimize the personalized behavior strategy, and output the final personalized behavior strategy.
Citation Information
Cited By
Multi-module cooperative intelligent role playing system
CN120542581A