Dialogue generation method and system

By obtaining scene information and the initial personality vector of the target character, generating candidate dialogues and conducting multi-dimensional evaluations, the problems of stereotyped personality simulation and poor scene adaptability in multi-character communication are solved, high-quality dialogue generation is achieved, and the user interaction experience and character simulation authenticity are improved.

CN120725142APending Publication Date: 2025-09-30SHANDONG CHAOYUE DATA CONTROL ELECTRONICS CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510868683.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing technologies make it difficult to achieve multi-dimensional evaluation and dynamic adjustment in multi-person communication, resulting in stereotyped character simulation and poor scene adaptability, which cannot meet complex and changing communication needs.

Method used

By obtaining scene information and the initial personality vector of the target character, candidate dialogues are generated, and the comprehensive reward score of the candidate dialogues is evaluated based on a preset multi-dimensional reward function. Combined with the initial personality vector and the target's historical interaction information, the personality performance is dynamically adjusted to ensure that the dialogue conforms to the personality setting and scene adaptation.

Benefits of technology

It improves the accuracy and effectiveness of dialogue generation, enhances the user interaction experience, and improves the realism and scene adaptability of character simulation. It is suitable for fields such as virtual social networking, game NPCs, and film and television script generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725142A_ABST
    Figure CN120725142A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and natural language processing, and discloses a dialogue generation method and system. The method comprises the following steps: acquiring scene information and an initial character vector of a target role, and acquiring target historical interaction information based on the scene information; generating a plurality of candidate conversations based on the initial character vector, the target historical interaction information and a preset large language model; evaluating a comprehensive reward score of each candidate dialogue based on a preset multi-dimensional reward function; and sorting all the candidate conversations according to the comprehensive reward score to obtain a candidate sequence, and taking a first preset number of candidate conversations in the candidate sequence as target conversations. Through the scheme of the invention, the dialogue generation method with multi-character character dynamic adjustment, multi-dimensional evaluation and good scene adaptability can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence and natural language processing, and in particular to a dialogue generation method and system. Background Art

[0002] Currently, large-scale AI models face numerous challenges in simulating multi-person communication across diverse scenarios. Traditional character simulation approaches rely heavily on static personality templates, such as the Myers-Briggs (MBTI) personality categorization. These fixed templates lack dynamic adaptability, making it difficult to flexibly adapt character expressions to changing interaction scenarios. Consequently, the simulated characters exhibit a rigid and monotonous appearance, making them incapable of meeting complex and ever-changing communication needs. Furthermore, existing reinforcement learning methods often focus on a single objective, such as conversational fluency, when evaluating character communication effectiveness, while neglecting multi-dimensional criteria such as personality consistency and contextual adaptability. This one-sided evaluation approach prevents the model from effectively optimizing the fidelity of character simulation. Furthermore, when faced with cross-scenario situations, such as transitioning from a business negotiation to a family conflict, the model struggles to maintain the logical coherence of character behavior, exhibiting poor generalization and inability to adapt effectively to the communication needs of diverse scenarios. This leads to logical discontinuities in character behavior during scenario transitions, severely impacting simulation effectiveness.

[0003] Although some related technologies currently use multi-dimensional feedback reinforcement learning to perform multi-dimensional evaluation of responses, there is still room for improvement in the dynamic optimization of multi-character interactions and cross-scenario character simulation.

[0004] Therefore, developing a dialogue generation method and system that can achieve dynamic adjustment of multiple character personalities, multi-dimensional evaluation, and good scene adaptability has become an urgent problem to be solved in this field. Summary of the Invention

[0005] In view of this, the present invention proposes a dialogue generation method and system, which can realize dynamic adjustment of multiple character personalities, multi-dimensional evaluation and good scene adaptability.

[0006] Based on the above objectives, an embodiment of the present invention provides a method for generating a dialogue, which specifically includes the following steps: Obtaining scene information and an initial personality vector of a target character, and obtaining target historical interaction information based on the scene information; generating a plurality of candidate dialogues based on the initial personality vector, the target historical interaction information, and a preset large language model; Evaluate the comprehensive reward score of each candidate dialogue based on a preset multi-dimensional reward function; All the candidate dialogues are sorted according to the comprehensive reward scores to obtain a candidate sequence, and the first preset number of candidate dialogues in the candidate sequence are used as target dialogues.

[0007] In some implementations, the step of generating a plurality of candidate dialogues based on the initial personality vector, the target historical interaction information, and a preset large language model includes: generating an initial query vector and an initial key vector based on the target historical interaction information; The initial personality vector, the initial query vector, and the initial key vector are used as input parameters of the preset large language model to generate a preset number of candidate dialogues.

[0008] In some implementations, the step of using the initial personality vector, the initial query vector, and the initial key vector as input parameters of the preset large language model to generate a preset number of candidate dialogues includes: Determining a dimension weight corresponding to each personality dimension based on the scenario information; Obtaining a target personality vector based on the initial personality vector and all the dimension weights; Based on the large language model, the target personality vector is respectively fused with the initial query vector and the initial key vector to obtain a target query vector and a target key vector; determining an attention score based on the target query vector and the target key vector; Based on the attention scores, the large language model generates a preset number of candidate dialogues.

[0009] In some implementations, the step of evaluating the comprehensive reward score of each candidate dialogue based on a preset multi-dimensional reward function includes: Obtaining the feature values ​​corresponding to each candidate dialogue in all the personality dimensions, and determining a first reward score for each candidate dialogue based on the initial personality vector, all the dimension weights, and all the feature values ​​corresponding to each candidate dialogue; Performing a preset text analysis on each candidate dialogue to obtain an analysis result for each candidate dialogue, and determining a second reward score for each candidate dialogue based on the analysis result and a scenario goal corresponding to the scenario information; Obtaining a third reward score for each candidate conversation according to the user's rating of each candidate conversation; Determining, according to the scenario type corresponding to the scenario information, a first reward weight corresponding to the first reward score, a second reward weight corresponding to the second reward score, and a third reward weight corresponding to the third reward score; A comprehensive reward score is obtained based on the preset multi-dimensional reward function, the first reward score, the first reward weight, the second reward score, the second reward weight, the third reward score and the third reward weight.

[0010] In some implementations, the step of sorting all the candidate dialogues according to the comprehensive reward scores to obtain a candidate sequence includes: Sorting all the candidate dialogues according to the comprehensive reward scores from high to low to obtain a first sequence; For each candidate dialogue, obtaining a personality score of the candidate dialogue on each personality dimension according to all corresponding feature values; Obtaining a score threshold corresponding to each personality dimension according to the scenario information, and determining whether the candidate dialogue has a personality score on at least one personality dimension that is less than the corresponding score threshold; If so, lowering the priority of the candidate dialogue in the first sequence; The adjusted first sequence is used as the candidate sequence.

[0011] In some implementations, the dialog generation method further includes: Calculate the vector gradient value based on the comprehensive reward scores corresponding to all the target dialogues; Based on the vector gradient value and the initial character vector, a new character vector is obtained and used as the new initial character vector; Obtaining a dimension threshold corresponding to each of the personality dimensions according to the scene information, and determining whether there is a personality dimension in the new initial personality vector whose value is less than the corresponding dimension threshold; If so, increase the dimension weight corresponding to the personality dimension, and return to the step of generating several candidate dialogues based on the initial personality vector, the target historical interaction information and the preset large language model.

[0012] In some embodiments, the step of obtaining scene information and an initial personality vector of a target character, and obtaining target historical interaction information based on the scene information, includes: generating the initial personality vector based on a preset personality trait library and all historical interaction information corresponding to the target character; Receiving a scene type, a character relationship, and a scene target input by a user, and converting the scene type, the character relationship, and the scene target based on a preset data format to obtain the scene information; The target historical interaction information is obtained by querying a preset global memory database according to the scenario information.

[0013] In some implementations, the dialog generation method further includes: storing the first second preset number of candidate dialogues in the candidate sequence into a preset buffer; Synthesize virtual dialogue data based on all data in the preset buffer, annotate scene information of the virtual dialogue data, and store the virtual dialogue data and its scene information in the preset global memory database.

[0014] In some implementations, the step of constructing the preset large language model includes: Initialize the large language model; An input interface of the personality vector is added to the decoding layer of the initialized large language model and pre-training is performed to obtain the preset large language model.

[0015] Another aspect of the present invention provides a dialogue generation system, which includes: a data acquisition module configured to acquire scene information and an initial personality vector of a target character, and acquire target historical interaction information based on the scene information; a candidate dialogue generation module configured to generate a plurality of candidate dialogues based on the initial personality vector, the target historical interaction information, and a preset large language model; An evaluation module configured to evaluate a comprehensive reward score of each candidate dialogue based on a preset multi-dimensional reward function; The target dialogue selection module is configured to sort all the candidate dialogues according to the comprehensive reward score to obtain a candidate sequence, and select the first preset number of candidate dialogues in the candidate sequence as target dialogues.

[0016] The present invention has at least the following beneficial technical effects: The dialogue generation method of the present invention, by obtaining scene information and the initial personality vector of the target character, and retrieving the target historical interaction information based on the scene information, can enable the model to fully grasp the key elements of the current situation and the character's past behavior logic before generating a dialogue, laying the foundation for subsequent dialogue generation. Under this premise, combining the initial personality vector, the target historical interaction information and the preset large language model to generate several candidate dialogues can ensure that the candidate dialogues are consistent with the character personality settings and can be associated with effective information in historical interactions, thereby improving the coherence and rationality of the dialogue. The comprehensive reward score of each candidate dialogue is evaluated based on a preset multi-dimensional reward function, and a comprehensive consideration is made from multiple dimensions such as personality consistency, scene adaptability, and user preferences to avoid the one-sidedness of a single standard, so that the evaluation results are closer to real needs. Finally, the candidate dialogues are sorted from high to low based on the comprehensive reward score, and the first preset number of candidate dialogues are selected as the target dialogues corresponding to the target character. This screening mechanism can accurately lock in high-quality dialogues, effectively improve the accuracy and effectiveness of dialogue generation, make the target dialogues more in line with actual needs in complex multi-character interaction scenarios, and enhance the user interaction experience. At the same time, through multi-dimensional evaluation and screening, it also helps the preset large language model to continuously improve the authenticity of character simulation and scene adaptability during continuous optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 A block diagram of an embodiment of the dialogue generation method provided by the present invention; Figure 2 A schematic diagram of an embodiment of the core architecture of the dialogue generation method provided by the present invention; Figure 3 A schematic diagram of an embodiment of the hierarchical reinforcement learning framework provided by the present invention; Figure 4 A schematic diagram of an embodiment of an optimization process for generating a dialogue provided by the present invention; Figure 5 A schematic diagram of an embodiment of the reinforcement learning process provided by the present invention; Figure 6 A schematic diagram of an embodiment of a dialogue generation system provided by the present invention; Figure 7 A schematic structural diagram of an embodiment of a computer device provided by the present invention; Figure 8 This is a schematic structural diagram of an embodiment of a computer-readable storage medium provided by the present invention. DETAILED DESCRIPTION

[0019] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the embodiments of the present invention are further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.

[0020] It should be noted that all expressions using "first" and "second" in the embodiments of the present invention are for distinguishing two non-identical entities with the same name or non-identical parameters. It can be seen that "first" and "second" are only for the convenience of expression and should not be understood as limitations on the embodiments of the present invention. Subsequent embodiments will not explain this one by one.

[0021] Based on the above objectives, the first aspect of the embodiment of the present invention provides a method for generating a dialogue. Figure 1 As shown, it includes the following steps: Step S100, obtaining scene information and an initial personality vector of a target character, and obtaining target historical interaction information based on the scene information; Step S200 , generating a number of candidate dialogues based on the initial personality vector, the target's historical interaction information, and a preset large language model; Step S300, evaluating the comprehensive reward score of each candidate dialogue based on a preset multi-dimensional reward function; Step S400: sort all candidate dialogues according to the comprehensive reward scores to obtain a candidate sequence, and use the first preset number of candidate dialogues in the candidate sequence as target dialogues.

[0022] In some embodiments, as Figure 2As shown, the dialogue generation method of the present invention includes four parts: character modeling, scenario generation, reinforcement learning, and dialogue generation. The character modeling part dynamically assigns an initial personality vector to each character. The initial personality vector uses a multidimensional parameterized representation, including multiple personality dimensions such as openness, responsibility, aggressiveness, and empathy. Each personality dimension corresponds to a different numerical value, forming a two-dimensional array. The initial value of each personality dimension in the initial personality vector can be preset by the user or generated through data mining. The data mining-based generation method uses natural language processing technology to analyze large amounts of text data such as film and television scripts, social platform conversations, and game conversations, extract personality-related features, construct a personality feature library, and then use data mining methods such as clustering algorithms to generate an initial personality vector for each character. Furthermore, this initial personality vector is embedded in the attention layer of a preset large language model. During the attention mechanism calculation process, the initial personality vector is fused with the query vector, key vector, and value vector. Specifically, when calculating the attention score, the initial personality vector is used as an additional input parameter. It is linearly transformed with the query vector and key vector, and then multiplied together to obtain an attention score that incorporates personality information. This allows the personality vector to play a role in the dialogue generation process, enabling differentiated expressions for characters with different personalities. For example, for a highly "aggressive" character in a conflict scenario, the aggressiveness parameter in their personality vector will increase the attention mechanism's focus on adversarial vocabulary and expressions, thereby generating more adversarial responses.

[0023] In some embodiments, the scenario generation component is configured to receive scenario information input by the user and, based on this scenario information, extract historical interaction information related to the scenario information from a global memory database, using this information as target historical interaction information. Scenario information includes information such as scenario type, scenario goal, and character relationships. Scenario types include, but are not limited to, "workplace competition" and "family dinner," character relationships include, but are not limited to, "superior-subordinate" and "sibling," and scenario goals include, but are not limited to, "persuasion" and "reconciliation." Scenario information is represented in a structured data format, such as JSON or XML, to facilitate model parsing and processing of relationships between different scenario elements. Scenario information can be used to understand the basic context of the interaction, clarify the social relationships between characters, and identify the core tasks of the interaction. The scenario generation component, in conjunction with a global memory mechanism, leverages the 64KB long-chain reasoning capabilities of SenseNova V6 (a large model system) to record and process historical interaction information between characters across different scenarios. This is achieved by establishing a global memory database to store information such as conversation content, behavioral performance, and personality vector changes for each character in each scenario. When the scene changes, the pre-set large language model extracts historical interaction information related to the current scene from the global memory database based on the converted scene information. This information is then integrated into the dialogue generation process of the current scene through the attention mechanism, maintaining the consistency of character behavior across scenes and ensuring the logical coherence of the character's behavior when switching between scenes. For example, a character who exhibits strong logic and decision-making ability in a business negotiation scenario will continue to exhibit these personality traits when transitioning to a project management scenario.

[0024] In some embodiments, during the dialogue generation process, the scene information and the target character's initial personality vector are first obtained. Based on the scene information, the target's historical interaction information is then retrieved from a global memory database. The scene information uses structured data, including the scene type, character relationships, and scene objectives. The character's initial personality vector is a vector containing multidimensional personality traits. Next, based on the initial personality vector, the target's historical interaction information, and a preset large language model, several candidate dialogues are generated using the large language model and the embedded personality vector. During the generation process, the attention mechanism fully considers the personality vector and scene information to ensure that the generated dialogues are consistent with the character's personality and the scene settings. For example, when constructing simulated communication scenarios and interaction content, multiple rounds of dialogues are generated that are tailored to the characters and scenes with different personalities.

[0025] The dialogue generation method of the present invention, by obtaining scene information and the initial personality vector of the target character, and retrieving the target historical interaction information based on the scene information, can enable the model to fully grasp the key elements of the current situation and the character's past behavior logic before generating a dialogue, laying the foundation for subsequent dialogue generation. Under this premise, a number of candidate dialogues are generated by combining the initial personality vector, the target historical interaction information and the preset large language model to ensure that the candidate dialogues are consistent with the character personality settings and can be associated with effective information in historical interactions, thereby improving the coherence and rationality of the dialogue. The comprehensive reward score of each candidate dialogue is evaluated based on a preset multi-dimensional reward function, and a comprehensive consideration is made from multiple dimensions such as personality consistency, scene adaptability, and user preferences to avoid the one-sidedness of a single standard, so that the evaluation results are closer to real needs. Finally, the candidate dialogues are sorted from high to low based on the comprehensive reward score, and the first preset number of candidate dialogues are selected as the target dialogues corresponding to the target character. This screening mechanism can accurately lock in high-quality dialogues, effectively improve the accuracy and effectiveness of dialogue generation, make the target dialogues more in line with actual needs in complex multi-character interaction scenarios, and enhance the user interaction experience. At the same time, through multi-dimensional evaluation and screening, it also helps the preset large language model to continuously improve the authenticity of character simulation and scene adaptation capabilities during continuous optimization. The present invention optimizes the authenticity of character behavior through scenario-based interactive feedback and can be applied to fields such as virtual social networking, game NPCs, and film and television script generation.

[0026] In some embodiments, the step of generating a number of candidate dialogues based on the initial personality vector, the target historical interaction information, and a preset large language model includes: generating an initial query vector and an initial key vector based on the target historical interaction information; and using the initial personality vector, the initial query vector, and the initial key vector as input parameters of the preset large language model to generate a preset number of candidate dialogues.

[0027] In some embodiments, the target historical interaction information is vectorized, and the text content in the historical interaction is segmented, tagged with parts of speech, and semantically encoded using natural language processing technology, and converted into a numerical vector form that can be processed by a computer. Then, the encoded historical interaction vector is mapped using a preset linear transformation matrix. Specifically, two different learnable linear transformation matrices are used to process the historical interaction vectors separately. One matrix is ​​used to generate an initial query vector, and the other matrix is ​​used to generate an initial key vector. That is, through matrix multiplication, the historical interaction vector is multiplied by the corresponding linear transformation matrix to obtain the initial query vector and the initial key vector.

[0028] In some embodiments, the initial personality vector, the initial query vector, and the initial key vector are used as input parameters of a preset large language model to generate a preset number of candidate dialogues, including: determining the dimension weight corresponding to each personality dimension based on scenario information; obtaining a target personality vector based on the initial personality vector and all dimension weights; fusing the target personality vector with the initial query vector and the initial key vector based on the large language model to obtain a target query vector and a target key vector; determining an attention score based on the target query vector and the target key vector; and generating a preset number of candidate dialogues by the large language model based on the attention score.

[0029] In some embodiments, the dimension weight corresponding to each personality dimension is first determined based on the scenario information. The weights of all dimensions corresponding to each scenario information can be pre-configured by the user, or the scenario information can be analyzed through a neural network model to assign corresponding weight values ​​to each personality dimension such as openness, responsibility, aggressiveness, empathy, etc. Based on the initial personality vector and all dimension weights, a target personality vector is obtained. Specifically, each dimension value in the initial personality vector is weighted with the corresponding dimension weight, and through element-level multiplication, each dimension value of the initial personality vector is multiplied by the corresponding weight and then accumulated to obtain an adjusted target personality vector to make it more suitable for the current scenario requirements.

[0030] In some embodiments, a target personality vector is fused with an initial query vector and an initial key vector, respectively, based on a large language model to produce a target query vector and a target key vector. During the fusion process, the target personality vector is first mapped to the same dimensional space as the initial query vector and initial key vector using a learnable linear transformation matrix. It is then multiplied with the initial query vector and initial key vector, respectively, to produce a target query vector and a target key vector containing personality information. An attention score is then determined based on the target query vector and the target key vector. An attention weight distribution is obtained by calculating the dot product of the target query vector and the target key vector, scaling them, and performing softmax (an activation function) normalization. This distribution reflects the importance of historical interaction information to be considered when generating dialogue. For example, for a highly "aggressive" character in a conflict scenario, the aggressiveness-related dimensions in the target query vector and target key vector will increase the attention weight for adversarial vocabulary and expressions. Finally, based on the attention scores, the large language model generates a preset number of candidate dialogues, such as 10. This preset number is determined based on the actual application and is not specifically limited here. The decoder layer of the large language model performs a weighted aggregation of the target's historical interaction information based on the attention score, and combines it with the multidimensional personality traits in the target's personality vector to generate multiple candidate dialogues that match the character's personality and the scenario. For example, during the generation process, the model adjusts its response to the user's emotional expressions based on the weight of the empathy dimension in the target's personality vector, generating candidate responses that are more empathetic or professional.

[0031] In some embodiments, in the learning reinforcement part, a hybrid training strategy combining offline and online stages is used to optimize the above process of the preset large language model. In the offline stage, historical interaction data is used for pre-training. The historical interaction data includes multi-scene data such as film and television scripts, social platform conversations, game conversations, etc., and the data is labeled to clarify information such as scene type, character relationship and personality label. The imitation learning method is adopted to allow the preset large language model to learn the dialogue generation and personality expression patterns of the characters in the historical data, and master the basic dialogue rules and personality expression skills. In the online stage, the GRPO (an online learning algorithm) algorithm is adopted to dynamically adjust the personality vector through a generative replay mechanism. Furthermore, HRL (Hierarchical Reinforcement Learning) is used to optimize personality parameters. The hierarchical reinforcement learning framework includes high-level strategies and low-level strategies. The hierarchical reinforcement learning framework is such as Figure 3As shown in the figure, the high-level strategy is responsible for adjusting the dimension weights of personality dimensions, dynamically adjusting the importance of each dimension based on different scenarios and interaction requirements. For example, in conflict scenarios, the weights of dimensions such as aggressiveness and empathy are increased, while in cooperative scenarios, the weights of dimensions such as responsibility and openness are increased. The high-level strategy determines the optimal weight distribution for each personality dimension by analyzing scenario information and reward signals. The low-level strategy optimizes the generation of specific dialogues. Based on the dimension weights of personality dimensions determined by the high-level strategy and the current scenario information, it generates dialogue content that meets the character's personality and scenario requirements. The low-level strategy leverages the generative capabilities of large language models, combined with the reward signals of reinforcement learning, to continuously adjust the dialogue generation strategy to ensure that the dialogue content is more in line with the requirements, achieving comprehensive optimization from the overall to the details.

[0032] In some embodiments, when simulating multiple rounds of dialogue in a virtual environment, for the current speaking role, the scenario information and the current personality vector are input, and 10 candidate replies are generated through the large language model, each of which corresponds to a different personality expression tendency. For example, in a simulated workplace competition scenario, when the scenario information of "project report-colleague competition-striving for resources" and a personality vector with high "responsibility" and high "openness" are input, the large language model will generate multiple candidate replies. Some of the replies will focus on demonstrating professional ability and innovative ideas, reflecting openness, while other replies will emphasize the responsible attitude towards the project and teamwork spirit, reflecting responsibility. Through the optimization of hierarchical reinforcement learning, high-quality candidate dialogues that meet the scenario and personality requirements are ultimately output.

[0033] The dialogue generation method of the present invention effectively captures the semantic associations and temporal features of historical dialogues by generating an initial query vector and initial key vector based on target historical interaction information, providing a context-dependent foundation for dialogue generation. The initial personality vector, initial query vector, and initial key vector are used as input parameters for a pre-set large language model. The initial personality vector, initial query vector, and initial key vector are combined with scene information to determine the dimension weights of each personality dimension and generate a target personality vector. This allows the model to generate candidate dialogues while respecting the inherent personality traits of the character and dynamically adjusting the intensity of personality expression based on scene requirements. The target personality vector is fused with the initial query vector and initial key vector using the pre-set large language model to generate the target query vector and target key vector. This further guides the attention mechanism to focus on semantic information related to personality and scene, ensuring that the generated candidate dialogues strike a balance between personality consistency and scene adaptability. Generating a preset number of candidate dialogues based on attention scores not only enhances the diversity of dialogue content but also accurately matches user preferences with scene objectives through a multi-dimensional reward evaluation and screening mechanism. Ultimately, this optimizes the entire process from historical information understanding, personality trait modulation, to scene-adaptive generation, significantly enhancing the authenticity of character simulation and interactive effectiveness of the dialogue system in complex scenarios.

[0034] In some embodiments, the step of evaluating the comprehensive reward score of each candidate dialogue based on a preset multi-dimensional reward function includes: obtaining the eigenvalues ​​corresponding to each candidate dialogue in all personality dimensions, and determining the first reward score of each candidate dialogue based on the initial personality vector, all dimension weights and all eigenvalues ​​corresponding to each candidate dialogue; performing a preset text analysis on each candidate dialogue to obtain the analysis results of each candidate dialogue, and determining the second reward score of each candidate dialogue based on the analysis results and the scenario goals corresponding to the scenario information; obtaining a third reward score for each candidate dialogue based on the user's rating of each candidate dialogue; determining a first reward weight corresponding to the first reward score, a second reward weight corresponding to the second reward score, and a third reward weight corresponding to the third reward score based on the scenario type corresponding to the scenario information; and obtaining a comprehensive reward score based on the preset multi-dimensional reward function, the first reward score, the first reward weight, the second reward score, the second reward weight, the third reward score, and the third reward weight.

[0035] In some embodiments, in the reinforcement learning part, a multi-dimensional scoring is performed on the dialogue generated by the preset large language model, and reinforcement learning is performed on the preset large language model based on the multi-dimensional scoring results, such as Figure 4 As shown. Obtain the eigenvalues ​​corresponding to each candidate dialogue in all personality dimensions, and determine the first reward score for each candidate dialogue based on the initial personality vector, all dimension weights, and all eigenvalues ​​corresponding to each candidate dialogue. Specifically, for each candidate dialogue, analyze the dialogue text through natural language processing technology to extract personality-related features. For example, calculate the eigenvalue of the "extraversion" dimension from indicators such as the number of active speeches and word style in the dialogue, and calculate the eigenvalue of the "logic" dimension from the degree of logical rigor. Compare these eigenvalues ​​with the set values ​​of each dimension in the initial personality vector, and combine them with the weights of the corresponding dimensions, and calculate according to Formula 1 to obtain the first reward score reflecting personality consistency to ensure that the dialogue conforms to the character setting. Formula 1 is as follows: ; in, Score for the first reward, is the dimension weight corresponding to each personality dimension, is the characteristic value of the i-th personality dimension reflected in the candidate dialogue, is the value of the i-th personality dimension in the initial personality vector, and n is the total number of personality dimensions.

[0036] In some implementations, a preset text analysis is performed on each candidate conversation to obtain an analysis result. A secondary reward score is then determined for each candidate conversation based on the analysis result and the scenario goal corresponding to the scenario information. Specifically, an automatic evaluation model, such as LLaVA-Critic (a multimodal large-scale universal evaluator), is employed to first analyze the scenario goal and extract key elements, such as "reaching a cooperation" in business negotiations or "resolving customer issues" in customer service scenarios. Semantic and sentiment analysis are then performed on the candidate conversation to determine whether the conversation content contributes to achieving the scenario goal. For example, in a business negotiation scenario, if a conversation contains an effective compromise or a balance of interests, a higher secondary reward score is awarded based on its contribution to the scenario goal.

[0037] In some implementations, a third reward score for each candidate conversation is derived based on the user's rating of each candidate conversation. Specifically, the third reward score reflecting user preferences is generated by automatically evaluating the candidate conversations through direct user ratings and feedback, or by leveraging an AI (Artificial Intelligence) referee model that learns user evaluation patterns, such as a model trained using the PPO-max (an algorithm based on proximal policy optimization) optimization strategy of FudanNLP (an open source natural language processing toolkit).

[0038] In some embodiments, the first reward weight corresponding to the first reward score, the second reward weight corresponding to the second reward score, and the third reward weight corresponding to the third reward score are determined based on the scenario type corresponding to the scenario information. The first reward weight, the second reward weight, and the third reward weight are all obtained in advance by analyzing the scenario type by the neural network model. For example, in a conflict scenario, setting the first reward weight to a higher weight value can emphasize the importance of the character's behavior being consistent with the personality setting, setting the second reward weight to a higher weight value can emphasize the importance of the conversation's contribution to achieving the scenario goal, and setting the second reward weight to a higher weight value can emphasize the importance of user preferences. Finally, based on the preset multi-dimensional reward function, combined with the first reward score, first reward weight, second reward score, second reward weight, third reward score, and third reward weight obtained above, the comprehensive reward score for each candidate conversation is calculated. Among them, the preset multi-dimensional reward function is shown in the following formula 2: ; in, is the comprehensive reward score of the k-th candidate dialogue, is the first reward score of the k-th candidate dialogue, is the second reward score of the k-th candidate dialogue, is the third reward score of the k-th candidate dialogue, is the first reward weight, is the second reward weight, is the third reward weight. Dynamically adjusted by the scenario type. For example, when the scenario type changes from workplace competition scenario to conflict mediation scenario, the first reward weight Increased to 40%.

[0039] The dialogue generation method of the present invention obtains the characteristic values ​​corresponding to each personality dimension of the candidate dialogue, and determines the first reward score in combination with the initial personality vector and the dimension weight. It can accurately measure the degree of fit between the dialogue and the character personality setting, and ensure that the character image is always consistent in the dialogue. The candidate dialogue is subjected to preset text analysis, and the second reward score is determined according to the scene goal. It can effectively evaluate the role of the dialogue in promoting the current scene task, so that the content of the dialogue is closely linked to the core goal of the interaction. The third reward score is obtained with the help of user ratings, and the user's subjective preferences are fully incorporated to make the dialogue generation more in line with actual usage needs. The weight of each reward score is dynamically determined according to the scene type, which can be flexibly adapted to different interaction scenarios. The comprehensive reward score comprehensively and dynamically evaluates the quality of the dialogue from multiple key angles such as personality, scene, user preference, etc., avoids the one-sidedness of single-dimensional evaluation, provides accurate and comprehensive guidance signals for model optimization, and significantly improves the accuracy, practicality and user satisfaction of dialogue generation.

[0040] In some embodiments, the step of sorting all candidate conversations according to the comprehensive reward score to obtain a candidate sequence includes: sorting all candidate conversations according to the comprehensive reward score from high to low to obtain a first sequence; for each candidate conversation, obtaining the personality score of the candidate conversation on each personality dimension according to all its corresponding feature values; obtaining the score threshold corresponding to each personality dimension according to the scene information, and determining whether the candidate conversation has a personality score on at least one personality dimension that is less than the corresponding score threshold; if so, lowering the priority of the candidate conversation in the first sequence; and using the adjusted first sequence as the candidate sequence.

[0041] In some embodiments, the personality score of each personality dimension is calculated using Formula 3, which is shown as follows: ; in, The personality score of the candidate dialogue on the i-th personality dimension.

[0042] In some embodiments, the scene information adopts a structured data format, including fields such as scene type, character relationship and scene goal. The scene information is analyzed by a neural network model, and a corresponding score threshold is assigned to each personality dimension. For example, in a family conflict mediation scenario, the score threshold of the "empathy" dimension will be set higher to ensure that the generated dialogue can reflect sufficient empathy. If there is a situation where the personality score on a certain personality dimension is less than the corresponding score threshold, the priority of the candidate dialogue in the first sequence is reduced. The specific method of reducing the priority can be to multiply the comprehensive reward score of the candidate dialogue by an adjustment coefficient less than 1, or to move its position in the first sequence backward by a certain number of digits. Finally, the first sequence after priority adjustment is used as the candidate sequence.

[0043] The dialogue generation method of the present invention, when sorting candidate dialogues, first sorts them from high to low according to the comprehensive reward score to obtain a first sequence, and performs a preliminary screening of the dialogue quality as a whole to ensure that high-scoring dialogues are presented first. For each candidate dialogue, the personality score of each personality dimension is calculated based on its characteristic value to achieve a refined assessment of the dialogue personality performance. The score threshold is used to determine whether the candidate dialogue has any personality performance that does not meet the standard. If so, its priority in the first sequence is reduced. This mechanism can strictly control the consistency of the character's personality while ensuring the overall quality of the dialogue, avoiding the appearance of dialogue content that is inconsistent with the scene and character setting. The final candidate sequence not only guarantees the high quality and diversity of dialogue generation, but also ensures the stable presentation of the character's personality in different scenarios, making the dialogue content more in line with actual interaction needs, significantly improving the user's interaction experience and the practicality of the dialogue system.

[0044] In some embodiments, the dialogue generation method of the present invention also includes: calculating a vector gradient value based on the comprehensive reward scores corresponding to all the target dialogues; obtaining a new personality vector based on the vector gradient value and the initial personality vector and using it as the new initial personality vector; obtaining the dimension threshold corresponding to each personality dimension according to the scene information, and judging whether there is a personality dimension in the new initial personality vector whose value is less than the corresponding dimension threshold; if so, increasing the dimension weight corresponding to the personality dimension, and returning to the step of generating several candidate dialogues based on the initial personality vector, the target historical interaction information and the preset large language model.

[0045] In some embodiments, the GRPO algorithm is used in the online stage to dynamically update the personality vector using the candidate sequence, such as Figure 5As shown. Truncated Importance Sampling is used to select the first preset number of candidate dialogues in the candidate sequence as target dialogues, such as 3. The first preset number is set according to the application and is not specifically limited here. For all target dialogues, the vector gradient value is calculated based on Formula 4. Formula 4 is as follows: ; in, is the kth candidate dialogue, is the vector gradient value, is the filtered target dialogue set, Indicates the strategy in state s Generate candidate dialogues The probability of is the comprehensive reward score of the kth candidate dialogue.

[0046] Based on the above calculation process, the contribution of each candidate dialogue to the update direction of the personality vector is clarified. Based on the calculated vector gradient value and the initial personality vector, the personality vector is updated using the Adam (AdaptiveMomentEstimation) optimizer, a gradient descent-based optimization algorithm, to obtain a new personality vector, which is used as the new initial personality vector. In practice, if analysis reveals that the new personality vector's contribution to high-reward dialogues in terms of empathy is insufficient, the weight coefficient of this dimension in the attention mechanism is specifically increased. This allows the preset large language model to focus more on the semantics and expressions that reflect empathy during subsequent dialogue generation.

[0047] In some embodiments, the dimension threshold corresponding to each personality dimension is obtained based on the scenario information. For different scenario information, the corresponding dimension threshold is pre-configured for each personality dimension. It is determined whether there is a personality dimension in the new initial personality vector whose value is less than the corresponding dimension threshold. For example, in the "psychological counseling" scenario, the "empathy" dimension threshold is set relatively high. If the dimension value in the new personality vector is lower than the threshold, the dimension weight adjustment mechanism is triggered. If there is a situation that triggers the dimension weight adjustment mechanism, the dimension weight of the corresponding personality dimension is increased, and the process returns to the step of generating several candidate dialogues based on the initial personality vector, the target historical interaction information and the preset large language model, and the cycle of candidate dialogue generation, reward evaluation, and strategy optimization is repeated.

[0048] Through the above iterative optimization, the dialogue generation method of the present invention can more efficiently process high-dimensional action spaces in multi-character parallel decision-making scenarios compared to traditional algorithms, significantly improving the diversity of dialogue generation and the accuracy of scene adaptation, so that the generated dialogue not only conforms to the character personality settings, but also better meets the scene target requirements.

[0049] In some embodiments, the steps of obtaining scene information and an initial personality vector of a target character, and obtaining target historical interaction information based on the scene information, include: generating the initial personality vector based on a preset personality trait library and all historical interaction information corresponding to the target character; receiving scene type, character relationship, and scene target input by a user, converting the scene type, character relationship, and scene target based on a preset data format, and using the conversion result as the scene information; and querying a preset global memory database according to the scene information to obtain target historical interaction information.

[0050] In some implementations, the implementation process also involves data preparation. Multi-scenario dialogue datasets are collected from sources such as film and television scripts, social media platform conversations, and game dialogues to ensure coverage of different scenarios and character relationships. The collected data is annotated to clarify information such as scene type, character relationships, and personality tags. The annotation process utilizes a combination of manual and automatic annotation to improve annotation efficiency and accuracy. This provides foundational data for the construction of a pre-set personality trait library and a global memory database, thereby providing high-quality data support for subsequent model training and dialogue generation.

[0051] In some embodiments, the conversation generation method of the present invention is characterized in that it also includes: storing the first second preset number of candidate conversations in the candidate sequence into a preset buffer; synthesizing virtual conversation data based on all data in the preset buffer, marking the scene information of the virtual conversation data, and storing the virtual conversation data and its scene information into the preset global memory database.

[0052] In some implementations, during each round of interaction, the first second-predetermined number of candidate dialogues and their personality vector states in the candidate sequence are stored in a generative replay buffer. For example, for the top 50% of candidate dialogues, 100,000 virtual dialogues are synthesized weekly using StyleGAN-NLP to alleviate data sparsity during the cold start phase. These virtual dialogues are annotated to identify scene types, character relationships, and personality labels, and then stored in a pre-set global memory database.

[0053] The dialogue generation method of the present invention establishes a generative replay buffer to store historical high-reward dialogues and their corresponding personality vector states, and regularly synthesizes virtual interaction data to alleviate the data sparsity problem of online learning and improve the efficiency of strategy optimization.

[0054] In some embodiments, the step of constructing a preset large language model includes: initializing the large language model; adding an input interface of the personality vector to the decoding layer of the initialized large language model and performing pre-training to obtain the preset large language model.

[0055] In some implementations, a large language model (LLM), such as DeepSeek-R1, is first initialized, and the personality vector is embedded into the decoder layer to construct a pre-defined large language model. The specific structure of the LLM's decoder layer is adjusted accordingly, adding a personality vector input interface to enable the model to consider character traits when generating dialogue. Then, using RFT (reinforcement learning-based fine-tuning) technology, a small amount of high-quality annotated data is selected as training samples to optimize the reward model. During fine-tuning, model parameters are adjusted based on reward signals to improve the model's sensitivity to reward signals and judgment accuracy, enabling the reward model to more accurately assess dialogue effectiveness.

[0056] The dialogue generation method of the present invention, during the construction of a pre-set large language model, initializes the large language model to establish a basic language processing framework, ensuring the model possesses basic semantic understanding and text generation capabilities. Subsequently, an input interface for personality vectors is added to the decoding layer, enabling the model to receive and process character personality information during dialogue generation, providing structural support for personalized dialogue generation.

[0057] In some embodiments, when simulating multiple rounds of interactions between multiple characters across multiple scenarios in a virtual environment using the method of the present invention, the preset large language model will execute the following steps in each round of interaction: (1) Candidate dialogue generation: The current speaking character in this round of interaction is used as the target character, the last updated personality vector of the target character is obtained as the initial personality vector, and the scene information of this round of interaction input by the user (such as "conflict scene-brother-sister-reconciliation goal") is obtained. 8 candidate dialogues are generated based on the initial personality vector and scene information. Among them, each candidate dialogue generated by the preset large language model corresponds to the same or different personality expression tendencies. For example, candidate dialogue 1 focuses on expressing the target character's aggressiveness, and candidate dialogue 2 focuses on expressing the target character's empathy. (2) Multi-dimensional scoring: The 8 generated dialogues are scored separately from the three dimensions of personality consistency, scene adaptability, and user preference through a multimodal evaluator. Finally, the scores of each dimension are combined to generate corresponding reward signals or penalty signals, providing a basis for the optimization of the preset large language model. For example, when scoring the scene adaptability dimension, if it is detected that in the scene type of "family dinner", the personality score of a candidate dialogue in the "empathy" personality dimension is lower than the score threshold corresponding to the "empathy" personality dimension, then the priority of the candidate dialogue in the ranking of all candidate dialogues will be reduced; (3) GRPO strategy update: The top three candidate dialogues are selected from all candidate dialogues to calculate the gradient direction of the personality vector, and the personality vector of the target character is updated through the Adam optimizer. If it is found that the dimension contribution of a personality dimension (such as "empathy") in the updated personality vector is insufficient in the current scenario type (such as "family dinner"), the dimension weight of the personality dimension in the attention mechanism is increased, so that the subsequent generation strategy is shifted to the high reward area; (4) Replay buffer update: The top 50% candidate dialogues and their personality vector states in this round of interaction are stored in the generative replay buffer. 100,000 virtual dialogue data are synthesized through StyleGAN-NLP every preset period (e.g., every week) to alleviate the data sparsity problem in the cold start phase.

[0058] The dialogue generation method of the present invention, in order to achieve flexible adjustment of the personalities of multiple characters in interactive scenarios, defines a dynamic personality vector containing multi-dimensional personality parameters, and updates the personality vector in combination with a reinforcement learning optimization mechanism, so that the personality of the character can be represented in a dynamic parameterized form, thereby naturally showing a personality performance that conforms to the situation in different scenarios. At the level of scene interaction effect evaluation, a multi-dimensional reward mechanism is constructed, and key dimensions such as personality consistency, scene adaptability, and user preferences are incorporated into the comprehensive evaluation system. By designing a scientific and reasonable reward function, a comprehensive and accurate guidance signal is provided for model optimization to ensure an all-round improvement in the interaction effect. In terms of model optimization and capability improvement, the personality parameters are continuously optimized with the help of reinforcement learning technology. At the same time, a hierarchical reinforcement learning framework and a hybrid training strategy are adopted to perform hierarchical optimization on the personality parameter adjustment and dialogue generation processes, significantly enhancing the model's processing capabilities in complex multi-character interaction scenarios, and further improving the authenticity and scene adaptability of the character simulation.

[0059] Based on the same inventive concept, according to another aspect of the present invention, Figure 6 As shown, an embodiment of the present invention further provides a dialogue generation system, characterized by comprising: The data acquisition module 110 is configured to acquire scene information and an initial personality vector of a target character, and acquire target historical interaction information based on the scene information; a candidate dialogue generation module 120 configured to generate a plurality of candidate dialogues based on the initial personality vector, the target historical interaction information, and a preset large language model; An evaluation module 130 is configured to evaluate a comprehensive reward score of each candidate dialogue based on a preset multi-dimensional reward function; The target dialogue selection module 140 is configured to sort all candidate dialogues according to the comprehensive reward scores to obtain a candidate sequence, and select the first preset number of candidate dialogues in the candidate sequence as target dialogues.

[0060] The dialogue generation system of the present invention, by obtaining scene information and the initial personality vector of the target character, and calling the target historical interaction information based on the scene information, can enable the preset large language model to fully grasp the key elements of the current situation and the character's past behavior logic before generating a dialogue, laying the foundation for subsequent dialogue generation. Under this premise, combining the initial personality vector, the target historical interaction information and the preset large language model to generate several candidate dialogues can ensure that the candidate dialogues are consistent with the character personality settings and can be associated with effective information in historical interactions, thereby improving the coherence and rationality of the dialogue. The comprehensive reward score of each candidate dialogue is evaluated based on a preset multi-dimensional reward function, and a comprehensive consideration is made from multiple dimensions such as personality consistency, scene adaptability, and user preferences to avoid the one-sidedness of a single standard, so that the evaluation results are closer to real needs. Finally, the candidate dialogues are sorted from high to low based on the comprehensive reward score, and the first preset number of candidate dialogues are selected as the target dialogues corresponding to the target character. This screening mechanism can accurately lock in high-quality dialogues, effectively improve the accuracy and effectiveness of dialogue generation, make the target dialogues more in line with actual needs in complex multi-character interaction scenarios, and enhance the user interaction experience. At the same time, through multi-dimensional evaluation and screening, it also helps the preset large language model to continuously improve the authenticity of character simulation and scene adaptability during continuous optimization.

[0061] Based on the same inventive concept, according to another aspect of the present invention, Figure 7 As shown, an embodiment of the present invention further provides a computer device 30, which includes a processor 310 and a memory 320. The memory 320 stores a computer program 321 that can be run on the processor. When the processor 310 executes the program, the steps of the above method are performed.

[0062] Based on the same inventive concept, according to another aspect of the present invention, Figure 8 As shown, an embodiment of the present invention further provides a computer-readable storage medium 40 , which stores a computer program 410 for executing the above method when executed by a processor.

[0063] Finally, it should be noted that those skilled in the art will understand that all or part of the processes in the above-described method embodiments can be implemented using a computer program to instruct the relevant hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes in the above-described method embodiments. The program storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM). The above-described computer program embodiments can achieve the same or similar effects as any of the corresponding aforementioned method embodiments.

[0064] It will also be appreciated by those skilled in the art that the various exemplary logic blocks, modules, circuits and algorithmic steps described in conjunction with the disclosure herein can be implemented as electronic hardware, computer software or a combination of the two. In order to clearly illustrate this interchangeability of hardware and software, a general description has been given of the functions of various schematic components, blocks, modules, circuits and steps. Whether this function is implemented as software or hardware depends on specific applications and the design constraints imposed on the entire system. Those skilled in the art can implement the function in various ways for each specific application, but this implementation decision should not be interpreted as causing a departure from the disclosed scope of the embodiments of the present invention.

[0065] The above are exemplary embodiments disclosed in the present invention, but it should be noted that various changes and modifications can be made without departing from the scope of the disclosure of the embodiments of the present invention as defined in the claims. The functions, steps and / or actions of the method claims according to the disclosed embodiments described herein do not need to be performed in any particular order. The serial numbers of the embodiments disclosed in the above embodiments of the present invention are for description only and do not represent the advantages and disadvantages of the embodiments. In addition, although the elements disclosed in the embodiments of the present invention can be described or required in individual form, they can also be understood as multiple unless expressly limited to the singular.

[0066] It should be understood that, as used herein, the singular forms "a" and "an" are intended to include the plural forms as well, unless the context clearly supports an exception. It should also be understood that, as used herein, "and / or" is intended to include any and all possible combinations of one or more of the associated listed items.

[0067] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to limit the scope of the disclosure of the present invention (including the claims) to these examples. Within the spirit of the present invention, the technical features of the above embodiments or different embodiments may be combined, and many other variations exist in different aspects of the above embodiments, which are not provided in detail for the sake of clarity. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for generating a dialogue, characterized in that: include: Obtaining scene information and an initial personality vector of a target character, and obtaining target historical interaction information based on the scene information; generating a plurality of candidate dialogues based on the initial personality vector, the target historical interaction information, and a preset large language model; Evaluate the comprehensive reward score of each candidate dialogue based on a preset multi-dimensional reward function; All the candidate dialogues are sorted according to the comprehensive reward scores to obtain a candidate sequence, and the first preset number of candidate dialogues in the candidate sequence are used as target dialogues.

2. The method for generating a dialogue according to claim 1, wherein: The step of generating a plurality of candidate dialogues based on the initial personality vector, the target historical interaction information, and a preset large language model includes: generating an initial query vector and an initial key vector based on the target historical interaction information; The initial personality vector, the initial query vector, and the initial key vector are used as input parameters of the preset large language model to generate a preset number of candidate dialogues.

3. The method for generating a dialogue according to claim 2, wherein: The step of using the initial personality vector, the initial query vector, and the initial key vector as input parameters of the preset large language model to generate a preset number of candidate dialogues includes: Determining a dimension weight corresponding to each personality dimension based on the scenario information; Obtaining a target personality vector based on the initial personality vector and all the dimension weights; Based on the large language model, the target personality vector is respectively fused with the initial query vector and the initial key vector to obtain a target query vector and a target key vector; determining an attention score based on the target query vector and the target key vector; Based on the attention scores, the large language model generates a preset number of candidate dialogues.

4. The method for generating a dialogue according to claim 3, wherein: The step of evaluating the comprehensive reward score of each candidate dialogue based on a preset multi-dimensional reward function includes: Obtaining the feature values ​​corresponding to each candidate dialogue in all the personality dimensions, and determining a first reward score for each candidate dialogue based on the initial personality vector, all the dimension weights, and all the feature values ​​corresponding to each candidate dialogue; Performing a preset text analysis on each candidate dialogue to obtain an analysis result for each candidate dialogue, and determining a second reward score for each candidate dialogue based on the analysis result and a scenario goal corresponding to the scenario information; Obtaining a third reward score for each candidate conversation according to the user's rating of each candidate conversation; Determining, according to the scenario type corresponding to the scenario information, a first reward weight corresponding to the first reward score, a second reward weight corresponding to the second reward score, and a third reward weight corresponding to the third reward score; A comprehensive reward score is obtained based on the preset multi-dimensional reward function, the first reward score, the first reward weight, the second reward score, the second reward weight, the third reward score and the third reward weight.

5. The method for generating a dialogue according to claim 4, wherein: The step of sorting all the candidate dialogues according to the comprehensive reward scores to obtain a candidate sequence includes: Sorting all the candidate dialogues according to the comprehensive reward scores from high to low to obtain a first sequence; For each candidate dialogue, obtaining a personality score of the candidate dialogue on each personality dimension according to all corresponding feature values; Obtaining a score threshold corresponding to each personality dimension according to the scenario information, and determining whether the candidate dialogue has a personality score on at least one personality dimension that is less than the corresponding score threshold; If so, lowering the priority of the candidate dialogue in the first sequence; The adjusted first sequence is used as the candidate sequence.

6. The method for generating a dialogue according to claim 3, wherein: Also includes: Calculate the vector gradient value based on the comprehensive reward scores corresponding to all the target dialogues; Based on the vector gradient value and the initial character vector, a new character vector is obtained and used as the new initial character vector; Obtaining a dimension threshold corresponding to each of the personality dimensions according to the scene information, and determining whether there is a personality dimension in the new initial personality vector whose value is less than the corresponding dimension threshold; If so, increase the dimension weight corresponding to the personality dimension, and return to the step of generating several candidate dialogues based on the initial personality vector, the target historical interaction information and the preset large language model.

7. The method for generating a dialogue according to claim 1, wherein: The step of obtaining scene information and an initial personality vector of a target character, and obtaining target historical interaction information based on the scene information, includes: generating the initial personality vector based on a preset personality trait library and all historical interaction information corresponding to the target character; Receiving a scene type, a character relationship, and a scene target input by a user, and converting the scene type, the character relationship, and the scene target based on a preset data format to obtain the scene information; The target historical interaction information is obtained by querying a preset global memory database according to the scenario information.

8. The method for generating a dialogue according to claim 7, wherein: Also includes: storing the first second preset number of candidate dialogues in the candidate sequence into a preset buffer; Synthesize virtual dialogue data based on all data in the preset buffer, annotate scene information of the virtual dialogue data, and store the virtual dialogue data and its scene information in the preset global memory database.

9. The method for generating a dialogue according to claim 1, wherein: The steps of constructing the preset large language model include: Initialize the large language model; An input interface of the personality vector is added to the decoding layer of the initialized large language model and pre-training is performed to obtain the preset large language model.

10. A dialogue generation system, characterized in that: include: a data acquisition module configured to acquire scene information and an initial personality vector of a target character, and acquire target historical interaction information based on the scene information; a candidate dialogue generation module configured to generate a plurality of candidate dialogues based on the initial personality vector, the target historical interaction information, and a preset large language model; An evaluation module configured to evaluate a comprehensive reward score of each candidate dialogue based on a preset multi-dimensional reward function; The target dialogue selection module is configured to sort all the candidate dialogues according to the comprehensive reward score to obtain a candidate sequence, and select the first preset number of candidate dialogues in the candidate sequence as target dialogues.

Citation Information

Cited By

  • Interaction control method and device, electronic equipment and computer readable storage medium

    CN121028654A

  • Multi-agent collaborative role large model same-environment personality generation method and related product

    CN121745150A

  • Method for aligning and tuning personality consistency of large reinforcement learning role model and related product

    CN121745210A

  • Multi-context session large language model system

    CN122021876A