Chat robot style control method based on external knowledge guidance
By obtaining background information of virtual characters and generating vector libraries, chatbots can more accurately simulate character language style and emotions, solving the problems of poor character consistency and single dialogue style in the existing technology, and achieving a more natural and diverse dialogue.
Patent Information
- Application Number
- CN202510050544.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-13
AI Technical Summary
When existing chatbots simulate specific virtual characters to talk to users, they cannot accurately reflect the unique language style and complex psychology of the characters, resulting in poor character consistency and single dialogue style.
By obtaining the background information of the virtual role, generating structured prompt words and sample dialogue data, collecting training dialogue data related to the virtual role, generating vector libraries, using vectorization technology to extract text and emotional vectors, and combining the basic information of the role to generate output vectors to achieve more natural and consistent role-playing.
The chatbot is able to more accurately simulate character language style, personality traits and emotional changes, ensure character consistency and dialogue diversity, and improve user satisfaction.
Smart Images

Figure CN119988545A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human-computer chat interaction, and in particular to a chat robot style control method based on external knowledge guidance. Background Art
[0002] At present, chatbots based on deep learning, especially dialogue generation technology based on large language models, have made significant progress. However, existing technologies still have certain limitations in accurately simulating the personality, style and background of specific virtual characters in conversations. Especially in virtual role-playing, existing chat dialogue models often have problems such as poor role consistency and a single dialogue style.
[0003] Since existing models usually rely on large-scale unsupervised training data, but this data cannot fully cover the deep personality and emotional characteristics of specific characters, even if the training data contains some character dialogues, the answers generated by the model cannot accurately reflect the unique style of the character and cannot maintain the consistency of the character. At the same time, although the existing models can generate relevant answers through prompt words, due to the lack of in-depth background knowledge and character settings, the system is often unable to accurately express the complex psychology and emotions of the characters in multiple rounds of dialogue, resulting in the relationship between the characters in the dialogue not being consistent with the original setting. In addition, the existing models are relatively single in style control and cannot fully capture the unique personality traits of the characters. During the dialogue, it is impossible to flexibly adjust the character's tone to match the character's unique style.
[0004] To address these problems, the present invention proposes a chatbot style control method based on external knowledge guidance. By introducing detailed role background information and dynamic dialogue examples, the chatbot can play a specific role more accurately and naturally, ensuring the consistency of the role, the diversity of style and the richness of background knowledge. Summary of the invention
[0005] In view of this, the present invention provides a chat robot style control method based on external knowledge guidance, which is used to solve the technical problems that when existing chat robots simulate a specific virtual character to have a conversation with a user, they fail to fully capture the character's personality traits, resulting in the inability to accurately reflect the character's unique language style, and the character's tone cannot be flexibly adjusted during the conversation to match the character's unique style.
[0006] In order to achieve the above technical objectives, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a chat robot style control method based on external knowledge guidance, comprising:
[0008] Obtaining background information of the virtual character, determining structured prompt words of the background information, and generating basic character information and sample dialogue data based on the structured prompt words and the background information;
[0009] Collect training dialogue data related to the virtual character, generate dialogue example data based on the training dialogue data and sample dialogue data, extract text vectors and sentiment vectors of the dialogue example data through vectorization technology, and generate a vector library;
[0010] Receive the user's chat content, determine the conversation context based on the text vector of the current chat content, search for conversation examples related to the current context from the vector library based on the conversation context, determine the output vector based on the basic role information, and generate reply information for the current chat based on the output vector.
[0011] Furthermore, the text vector and sentiment vector of the dialogue example data are extracted through vectorization technology to generate a vector library, including:
[0012] Label each conversation example with emotional style;
[0013] Use the first language model to perform text vectorization on each conversation example data to obtain a conversation example text vector;
[0014] The text vector is input into the second language model. The second language model clusters the text vectors under each emotional style category based on the emotional style label, and uses the vector of the cluster center as the direction vector of the emotional style of this category to obtain the emotional vector.
[0015] The text vectors and sentiment vectors are stored to generate a vector library.
[0016] Furthermore, the conversation context is determined according to the text vector of the current chat content, and based on the conversation context, a conversation example related to the current context is searched from the vector library, and an output vector is determined in combination with the basic information of the role, and the reply information of the current chat is generated according to the output vector, including:
[0017] Perform text vectorization on the chat content input by the user to obtain a text input vector;
[0018] Recalling multiple dialogue examples in corresponding dialogue situations in the vector library according to the text input vector, and determining the emotion vectors in the multiple dialogue examples;
[0019] Get the output vector based on the text input vector, the sentiment vector and the basic information of the role;
[0020] The output vector is input into the preset conversation model to generate the reply information of the current chat.
[0021] Furthermore, the output vector is obtained according to the text input vector, the emotion vector and the basic information of the role, including:
[0022] V out =V L +λ·V em
[0023] Among them, V L Represents the input text vector, V em represents the emotion vector, and λ represents the adjustment weight, which is determined according to the basic information of the role and the preset experience parameters.
[0024] Furthermore, the method further comprises:
[0025] When having multiple rounds of chats with a user, the tone and content of each round of reply information generated are adjusted according to the current user input and historical chat content, so that the character's current conversation style is consistent with the historical conversation style.
[0026] Furthermore, each round of reply information generation is adjusted in tone and content according to the current user input and historical chat content, including:
[0027] Adjust the decoder parameters of the conversation generation model according to the current user input content, and generate initial reply content through the conversation model according to the adjusted decoder parameters;
[0028] Calculate the similarity between the sentiment vector in the initial reply content and the sentiment vector in the previous round of reply information, and adjust the initial reply content according to the similarity;
[0029] When the similarity is within the preset threshold range, the reply content of this round of dialogue is output.
[0030] Furthermore, the method further comprises:
[0031] Obtain the state parameters, reward parameters and action parameters of each round of dialogue; the state parameters are high-dimensional vectors, including user feedback, sentiment changes and context information; the reward parameters include the user's real-time score and satisfaction feedback; the action parameters include the text and sentiment vector generated by the decoder of the conversation generation model;
[0032] Iteratively adjust the action parameters according to the state parameters and reward parameters to optimize the reply content of this round of dialogue.
[0033] Furthermore, the action parameters are iteratively adjusted according to the state parameters and reward parameters, including:
[0034] The Q-learning algorithm is used to iteratively update the decoder parameters of the conversation model and adjust the sentiment weight parameters.
[0035] In a second aspect, the present invention further provides a chat robot style control system based on external knowledge guidance, comprising:
[0036] A background information acquisition module, used to acquire background information of a virtual character, determine structured prompt words of the background information, and generate basic character information and sample dialogue data based on the structured prompt words and the background information;
[0037] A vector library generation module is used to collect training dialogue data related to the virtual character, generate dialogue example data based on the training dialogue data and sample dialogue data, extract text vectors and sentiment vectors of the dialogue example data through vectorization technology, and generate a vector library;
[0038] The dialogue reply module is used to receive the user's chat content, determine the dialogue context based on the text vector of the current chat content, search for dialogue examples related to the current context from the vector library based on the dialogue context, determine the output vector based on the basic role information, and generate reply information for the current chat based on the output vector.
[0039] In a third aspect, the present invention also provides an electronic device, comprising a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the chat robot style control method based on external knowledge guidance as described in any of the above technical solutions is implemented.
[0040] Compared with the prior art, the present invention provides the following advantages:
[0041] (1) Through detailed information about the character’s background, personality traits, and historical events, the chatbot can more accurately simulate the character’s language style, personality traits, and emotional changes, thereby achieving more natural role-playing.
[0042] (2) By dynamically introducing dialogue examples and context matching, the model can flexibly adjust the tone, attitude, and wording according to different dialogue contexts, making the conversation more natural and reducing the repetitiveness and monotony of language generation.
[0043] The method of the present invention can not only optimize the character's dialogue content according to the set personality, emotion and background, but also make dynamic adjustments according to the user's interaction mode, so as to control the style and personalized performance of the chat robot, and conduct chat dialogues with users in language and emotion that are more in line with the character's style, thereby achieving more accurate role-playing and improving user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 A flow chart of a chat robot style control method based on external knowledge guidance provided by the present invention;
[0045] Figure 2 A schematic diagram of the structure of a chatbot style control system based on external knowledge guidance provided by the present invention;
[0046] Figure 3 The present invention provides a schematic structural diagram of an electronic device. DETAILED DESCRIPTION
[0047] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.
[0048] Example 1
[0049] See also Figure 1 This embodiment provides a chat robot style control method based on external knowledge guidance, including:
[0050] Step S101: obtaining background information of a virtual character, determining structured prompt words according to the background information, and generating sample dialogue data based on the structured prompt words and the background information;
[0051] Step S102: Collect training dialogue data related to the virtual character, generate dialogue sample data according to the training dialogue data and sample dialogue data, extract text vectors and sentiment vectors of the dialogue sample data by vectorization technology, and generate a vector library;
[0052] Step S103: Receive the user's chat content, determine the conversation context based on the text vector of the current chat content, search for conversation examples related to the current context from the vector library based on the conversation context, determine the output vector in combination with the basic role information, and generate reply information for the current chat based on the output vector.
[0053] The method of this embodiment, first, through the background setting, personality characteristics and detailed information of historical events of the character, the chat robot can more accurately simulate the language style, personality characteristics and emotional changes of the character, thereby achieving more natural role-playing. By dynamically introducing dialogue examples and context matching, the model can flexibly adjust the tone, attitude and wording according to different dialogue contexts, making the dialogue more natural and reducing the repetitiveness and monotony of language generation. Through the method of this embodiment, the character's dialogue content can not only be optimized according to the set personality, emotion and background, but also dynamically adjusted according to the user's interaction mode, thereby controlling the style and personalized performance of the chat robot, and chatting with the user in language and emotion that is more in line with the character's style, achieving more accurate role-playing, thereby improving user satisfaction.
[0054] The method of the present invention is described in detail below through a specific embodiment. Assume that we want to generate a chat robot with Vegeta in Dragon Ball as the character.
[0055] The first step is to set Vegeta's detailed background information in the system's input prompt, including multi-dimensional information such as his origin, personality traits, growth story, and interpersonal relationships. Specifically:
[0056] 1. Character origin, for example: Vegeta was born on the Saiyan planet. He is the prince of the Saiyan Kingdom. Saiyans are a race that lives on fighting and aggression. They are born with strong strength and transformation abilities. However, the Saiyan planet was eventually destroyed by the evil Frieza, and Vegeta became one of the few surviving Saiyans.
[0057] 2. Personality traits, for example: Vegeta's personality is complex and changeable. He is arrogant, ruthless, and extremely confident in his own strength and status. In the many battles with Earth Warriors such as Son Goku, Vegeta gradually showed more humanity, including loyalty to his family and concern for his friends.
[0058] 3. Growth story, for example: He came to Earth as Frieza's subordinate to find the Dragon Balls to realize his wish of immortality. Through continuous training and fighting, he learned the Super Saiyan transformation and continued to improve his strength in subsequent stories.
[0059] 4. Interpersonal relationships, for example: Vegeta and Son Goku's relationship gradually changed from enemies to rivals and friends who respected each other. Although he and Bulma had very different personalities, they unexpectedly developed a deep relationship and had two children.
[0060] Based on the collected background information, structured prompt words are determined. The purpose of determining structured prompt words is to extract character features, abstract plots, and classify situations for the virtual character background. It can help the dialogue generation model accurately generate dialogue content that matches the character features and background information. Based on the above background information of Vegeta, the following structured prompt words can be divided:
[0061] Emotional labels: cold, proud, proud;
[0062] Character Background: He has a fierce competition with Sun Wukong and desires to surpass him;
[0063] A Deep Display of Responsibility: He is protective and affectionate towards his family, especially showing his human side in front of Bulma and the children.
[0064] Based on the above structured prompt words, we can first determine the basic emotional type of the character, and then determine the different emotions to complete the dialogue process when different contexts or character relationship prompt words are used. When the model generates dialogues later, it can guide the virtual character in the current dialogue to generate corresponding dialogue emotions based on the acquired background information. For example, if information related to Sun Wukong is received in the dialogue scene, the weight of hostility and respect in the emotion can be increased to show the complex emotions towards the competitor; if the dialogue scene contains the lover Bulma, adjust the weight of deep feeling in the emotion.
[0065] Generate basic character information and sample dialogue data based on structured prompt words and background information. The generated dialogue samples can be more consistent with the character settings and the characteristic background of the virtual character, for example:
[0066] Vegeta: "You, a monkey from Earth, will never surpass me. Although you can always catch up, I am the true prince of the Saiyans!"
[0067] Sun Wukong: "Haha, Vegeta, your pride is really unbearable."
[0068] Vegeta: "You will never be able to surpass me."
[0069] The second step is to collect Vegeta’s typical dialogues in Dragon Ball and perform vectorization processing to establish a vector database.
[0070] Since there is still a certain gap between the conversations generated by structured prompt words and background information and the chats with real users, it is also necessary to collect sample data of conversations between virtual characters and users, such as:
[0071] Example 1: User asks: "Who are you?" Vegeta replies: "I am Vegeta, the Saiyan Prince. What do you want?"
[0072] Example 2: User asks: "Is Bulma your wife?" Vegeta replies: "My private life is none of your business. Tell me what you want, or leave."
[0073] These conversations are stored in the database through text vectorization so that they can be called in subsequent conversations. The specific steps are as follows:
[0074] Step S21: Use the encoder of the first language model to vectorize the example text of the conversation; the first language model here can be BERT, GPT, etc., to vectorize Vegeta's conversation, and each conversation will be converted into a vector of fixed length to capture the semantic information of the text. As in Example 1 above, after being processed by the encoder, Vegeta's answer will be converted into a vector of fixed dimension. Assume that the output is a long vector representing the semantic information in the sentence, text vector: [0.21, -0.14, 0.53, 0.12, ...], each number in the vector corresponds to the position of the text in a high-dimensional space. The model will learn the semantic relationship between each word, phrase, sentence, etc. of the text and other texts in a given context, and map these relationships to a high-dimensional space to represent the semantics, grammar, and contextual relationships of the text.
[0075] Step S22: Use the second language model (emotion style encoder) to convert the text into an emotion vector; emotion labels such as positive, cold, angry, and threatening. According to the text vector, each conversation will be passed through the emotion encoder to obtain an emotion vector representation, such as: emotion vector [0.45, -0.28, 0.31, -0.15, ...], etc. The second language model here can be an emotion embedding model such as DeepMoji, EmoBERT, etc. that can generate multi-dimensional emotion representation, or it can be a deep learning model based on LSTM, GRU, or BERT.
[0076] Step S23: Calculate the cluster center vector v_emotion_i of each emotion category;
[0077] After the emotion vectorization process, we can perform cluster analysis. By clustering the emotion vectors of multiple conversations, we can get the cluster center vector of each emotion category. The cluster center vector represents the typical emotion style of the emotion category.
[0078] Step S24: Store the text vector and v_emotion_i in the vector library respectively. In future conversations, the system can retrieve similar conversation content as needed, or adjust the emotional direction of the emotional style conversation. For example:
[0079] Dialogue 1:
[0080] Text vector: [0.21,-0.14,0.53,0.12,...]
[0081] Sentiment vector: [0.45,-0.28,0.31,-0.15,...]
[0082] Cluster center vector (threat): [0.47,-0.29,0.33,-0.17,...]
[0083] The third step is to dynamically introduce and match the situation.
[0084] When a user asks a question, the system recalls relevant conversation examples from the vector database based on the context of the chat content and generates appropriate answers based on the current conversation situation. For example, when a user asks a question about Vegeta's family, the system selects conversation examples related to the family. Suppose the vector database contains the following conversation examples:
[0085] Round 1: User asks: "Who are you?" Vegeta replies: "I am Vegeta, Prince of Saiyan. What do you want?" Round 2: User asks: "Is Bulma your wife?" Vegeta replies: "My private life is none of your business. Tell me what you want, or leave."
[0086] The tone of the virtual character is adjusted appropriately according to Vegeta's character setting and the situation mentioned in the current context. The adjusted answer is as follows:
[0087] "You ask too many questions. This is my personal privacy."
[0088] The specific processing process is:
[0089] Step S31: vectorize the text input by the user "Who are you?": input_vector = [0.45, 0.32, ..., 0.78];
[0090] Step S32: Select the target emotional style (such as arrogance): v_emotion_i;
[0091] Step S33: Calculate the adjusted vector new_vector and express it as:
[0092] new_vector=input_vector+0.7*v_emotion_i
[0093] Among them, 0.7 is an adjustment coefficient, which is obtained based on the basic information of the character, the experience parameter library (manually set for each virtual character) and reinforcement learning (the process of reinforcement learning is described in detail in Example 4).
[0094] Step S34: Recall the top three similar conversation examples from the vector library as references for generating answers, and output the reply content through the preset conversation model.
[0095] The preset conversation model here can be the first language model in the second step, or it can be a third language model based on deep learning and NLP technology, such as the Transformer model, which combines contextual information and the user's emotional needs to output natural, coherent and character-style responses.
[0096] Example 2
[0097] In multiple rounds of conversation, the overall conversation style needs to be adjusted to maintain the characteristics of the characters.
[0098] When having multiple rounds of chats with a user, the tone and content of each round of reply information generated are adjusted according to the current user input and historical chat content, so that the character's current conversation style is consistent with the historical conversation style.
[0099] As a preferred embodiment, the generation of each round of reply information adjusts the reply tone and content according to the current user input content and the historical chat content, including:
[0100] Adjust the decoder parameters of the conversation generation model according to the current user input content, and generate initial reply content through the conversation model according to the adjusted decoder parameters;
[0101] Calculate the similarity between the sentiment vector in the initial reply content and the sentiment vector in the previous round of reply information, and adjust the initial reply content according to the similarity;
[0102] When the similarity is within the preset threshold range, the reply content of this round of dialogue is output.
[0103] Through the above method, the tone and content of the answer are adjusted according to the changes in the content of the user's question, combined with the character's personality and emotional changes, so that each round of dialogue is consistent with the character's background setting and personality.
[0104] This embodiment combines the decoding intervention method to modify the probability distribution generated by the decoder of the conversation generation model in real time to influence the vocabulary selection, ensuring that the style, emotion and tone of the generated text are consistent. Which tokens in the decoder are modified in probability depends on the recalled dialogue examples. The tokens corresponding to the words that best reflect the emotion or style in the recalled dialogue examples are modified in probability (usually increasing the probability of this token and synonymous tokens), thereby affecting the probability of the conversation generation model generating text with similar style.
[0105] In addition, the conversation model can also be guided by self-feedback. During the generation process, it continuously evaluates the current output token (estimates the similarity between the vector of the text it currently generates and the recalled emotion vector v_emotion_i), and adjusts the responses that do not conform to the character characteristics. In other words, it re-decodes the decoder of the conversation generation model to optimize its style and tone.
[0106] In some embodiments, during multiple rounds of conversation, the user may also add "requirements" information reflecting the user's personalized needs at the end of the current last round of information, for example:
[0107] User: Why are you here?
[0108] Vegeta: Humph! This is my territory. Do you think you can understand why I'm here? Go ask someone else.
[0109] User: Are you the strongest?
[0110] [Requirement] (user input): Tone: cold and high, example: "Enough to make you disappear in an instant".
[0111] At this point, based on the user's request, the current conversation content, the context information, and the background information previously entered, the system-generated answer may be: Vegeta: Of course! In this universe, only I deserve to be called the strongest! You can't even compare to my toes.
[0112] Specifically, the steps of adjusting the overall conversation style according to user requirements, current conversation content, and context information include:
[0113] Step S41: construct personalized requirements, for example: required tone: based on the emotional style of the previous round of dialogue, required examples: based on the TOP-3 similar dialogue examples of the previous round of dialogue.
[0114] Step S42: Decoder intervention: increase the generation probability of feature words such as "disappear", "moment", and "enough".
[0115] Step S43: self-feedback evaluation, generating similarity calculation between the text vector and the target emotion vector, for example, the similarity between the sexy vector of "Of course! In this universe, only I deserve to be called the strongest! You can't even compare to my toes." and the "arrogant" emotion style vector v_emotion_i.
[0116] If the similarity is lower than the threshold, regeneration is triggered, and the generation probability of the feature words in the decoder is continuously adjusted until a response with a similarity within the preset threshold is generated.
[0117] Adjustments based on historical conversations and user input ensure that the character's emotions and tone remain consistent, making the conversation more coherent. Each round of responses not only reflects the character's characteristics, but also forms a close connection with the previous conversation content, making it less likely for users to feel abrupt or unnatural. In multiple rounds of conversations, the model can continue to maintain the character's unique style, handle complex emotions and relationships between characters based on the character's background knowledge, and make the user's interaction with the virtual character more realistic and coherent.
[0118] Example 3
[0119] In order to optimize the chat output while maintaining the virtual character's self-style to adapt to the user's personalized needs and improve the quality of conversation, this method also designs a personalized optimization and feedback mechanism (reinforcement learning) to achieve the function of continuously optimizing the character's tone and style based on user feedback.
[0120] For example: If the user enjoys the challenge of interacting with Vegeta, then you need to continue to strengthen the tone.
[0121] User: Are you the strongest?
[0122] Vegeta: Humph, of course! I, Vegeta, the Saiyan Prince, who can rival me?
[0123] User: I want to challenge you!
[0124] Vegeta: Oh? Then you try it!
[0125] At this point, get the user's feedback on this round of conversation: likes to challenge Vegeta, Vegeta needs to be more provocative when responding. (Score: 3 points)
[0126] The system performs reinforcement learning based on the feedback and generates the following response again:
[0127] Ha! You dare to challenge me? You are so ignorant. Come on, give it a try. Let me see how strong you are and whether you can withstand the force of my blow. However, you'd better prepare yourself for the most painful battle of your life!
[0128] Through the process of reinforcement learning, it is possible to accurately simulate the style of the character Vegeta and keep the dialogue consistent and personalized.
[0129] As a specific embodiment, the steps of reinforcement learning optimization are as follows:
[0130] Step S51: Construct an initial dialogue state vector:
[0131] User history conversation vector: ["Are you the best?"] → [0.75, 0.62, ..., 0.83];
[0132] Current sentiment style vector: Arrogant style → [0.90, 0.85, ..., 0.78];
[0133] Dialogue context vector: [dialogue turn = 1, topic = strength discussion, ...];
[0134] Step S52: The system responds for the first time. The action information is: sentiment weight λ = 0.6. Decoder parameters: increase the generation probability of words such as "powerful" and "prince". The generated initial response result is:
[0135] "Humph, of course! I, Vegeta, the Saiyan Prince, who can rival me?";
[0136] Step S53: User subsequent input and feedback: New input: "I want to challenge you!", system reply: "Oh? Then you try it."
[0137] The user feedback in this round is as follows: Text feedback: "I like to challenge Vegeta, and I am more provocative when Vegeta needs a reply", score: 3 points (moderate positive feedback);
[0138] Step S54: Update parameters through Q-learning using the current state and user rewards:
[0139] Current state: [dialogue history vector, challenging topic identifier, current sentiment vector];
[0140] Select action: Increase the sentiment weight λ to 0.8 to increase the probability of generating provocative words;
[0141] Calculate the reward: The basic score is 3 points, which is converted into a reward value of r = 0.6; emotional feedback bonus: if the user expresses "like", the reward value is additional +0.2; the total reward value is r = 0.8;
[0142] Parameter update: learning rate α = 0.1 discount factor γ = 0.9;
[0143] Step S55: Output the optimized response to generate sentiment weight λ=0.8 (enhanced), and the provocative word weight is increased by 30%. Based on the above parameters, the conversation model generates a new reply as follows:
[0144] "Ha! You dare to challenge me? You simply don't know your own limitations. Come on, give it a try. Let me see how strong you are and whether you can withstand the force of my blow. However, you'd better be prepared for the most painful battle of your life!";
[0145] Step S56: Strategy storage: store the current successful state-action pair into the experience pool for reference in similar scenarios in the future: {"Scenario":"Strength Challenge","Emotional Tendency":"Strong Provocation","Optimal λ":0.8,"Lexical Preference":["Ha!","You Dare","Try It"],"Reward Value":0.8}.
[0146] Through reinforcement learning, the chatbot can adjust its dialogue strategy based on real-time feedback after each interaction with the user, ensuring that the dialogue style and emotional expression are more in line with user needs. It is particularly suitable for application scenarios such as games, animation, virtual assistants, etc. that require a high degree of role-playing.
[0147] It should be noted that the states, actions, and rewards in the above reinforcement learning are defined as follows:
[0148] State definition: The state of each conversation includes the user's historical conversation content vector, the currently generated text vector, the current conversation emotional style vector, etc. The state is represented as a high-dimensional vector that covers user feedback, emotional changes, conversation context, and other information.
[0149] Action space: The style adjustments of the responses generated by the system (e.g., decoder-generated text, weight λ of the sentiment style vector, etc.).
[0150] Reward signal: Provide rewards to chatbots based on real-time user feedback, such as ratings (user ratings of a specific reply text) and emotional feedback (such as whether the user is satisfied, grateful, happy, etc.). If the user is satisfied, the reward is positive; if the user is dissatisfied, a negative reward is given; if there is no feedback, a zero reward is given.
[0151] After collecting the chatbot's response to the user (i.e. specific response style adjustment measures) and the user's feedback on the robot's response (scoring, input text), it will continuously adjust its strategy through the algorithm (Q-learning method) and iterate the model parameters to increase the probability of selecting the best answer in the future.
[0152] Through continuous feedback optimization, the chatbot can fine-tune its style based on user demand feedback to achieve a highly personalized conversation experience, which is suitable for application scenarios such as games, animation, and virtual assistants.
[0153] Example 4
[0154] This embodiment also provides a chat robot style control system based on external knowledge guidance, such as Figure 2 As shown, the chatbot style control system 200 based on external knowledge guidance includes:
[0155] A background information acquisition module 201 is used to acquire background information of a virtual character, determine structured prompt words of the background information, and generate basic character information and sample dialogue data based on the structured prompt words and the background information;
[0156] A vector library generation module 202 is used to collect training dialogue data related to the virtual character, generate dialogue example data according to the training dialogue data and sample dialogue data, extract text vectors and sentiment vectors of the dialogue example data by vectorization technology, and generate a vector library;
[0157] The conversation reply module 203 is used to receive the user's chat content, determine the conversation context based on the text vector of the current chat content, search for conversation examples related to the current context from the vector library based on the conversation context, determine the output vector in combination with the basic role information, and generate reply information for the current chat based on the output vector.
[0158] Example 5
[0159] like Figure 3 As shown, the chatbot style control method based on external knowledge guidance is described above. The present invention also provides an electronic device 300, which can be a computing device such as a mobile terminal, a desktop computer, a notebook, a palm computer, and a server. The electronic device includes a processor 301, a memory 302, and a display 303.
[0160] In some embodiments, the memory 302 may be an internal storage unit of a computer device, such as a hard disk or memory of a computer device. In other embodiments, the memory 302 may also be an external storage device of a computer device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Further, the memory 302 may also include both an internal storage unit of a computer device and an external storage device. The memory 302 is used to store application software and various types of data installed on the computer device, such as program codes for installing the computer device. The memory 302 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a chat robot style control method program 304 based on external knowledge guidance is stored on the memory 302, and the chat robot style control method program 304 based on external knowledge guidance can be executed by the processor 301, thereby realizing a chat robot style control method based on external knowledge guidance in each embodiment of the present invention.
[0161] In some embodiments, the processor 301 can be a central processing unit (CPU), a microprocessor or other data processing chip, used to run the program code or process data stored in the memory 302, such as executing a chat robot style control method program based on external knowledge guidance.
[0162] In some embodiments, the display 303 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 303 is used to display information on the computer device and to display a visual user interface. The components 301-303 of the computer device communicate with each other through a system bus.
[0163] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A chatbot style control method based on external knowledge guidance, characterized in that: include: Obtaining background information of the virtual character, determining structured prompt words of the background information, and generating basic character information and sample dialogue data based on the structured prompt words and the background information; Collect training dialogue data related to the virtual character, generate dialogue example data based on the training dialogue data and sample dialogue data, extract text vectors and sentiment vectors of the dialogue example data through vectorization technology, and generate a vector library; Receive the user's chat content, determine the conversation context based on the text vector of the current chat content, search for conversation examples related to the current context from the vector library based on the conversation context, determine the output vector based on the basic role information, and generate reply information for the current chat based on the output vector.
2. The chat robot style control method based on external knowledge guidance according to claim 1 is characterized in that: The text vector and sentiment vector of the conversation example data are extracted through vectorization technology to generate a vector library, including: Label each conversation example with emotional style; Use the first language model to perform text vectorization on each conversation example data to obtain a conversation example text vector; The text vector is input into the second language model. The second language model clusters the text vectors under each emotional style category based on the emotional style label, and uses the vector of the cluster center as the direction vector of the emotional style of this category to obtain the emotional vector. The text vectors and sentiment vectors are stored to generate a vector library.
3. The chat robot style control method based on external knowledge guidance according to claim 1 is characterized in that: Determine the conversation context based on the text vector of the current chat content, search for conversation examples related to the current context from the vector library based on the conversation context, determine the output vector based on the basic role information, and generate the reply information of the current chat based on the output vector, including: Perform text vectorization on the chat content input by the user to obtain a text input vector; Recalling multiple dialogue examples in corresponding dialogue situations in the vector library according to the text input vector, and determining the emotion vectors in the multiple dialogue examples; Get the output vector based on the text input vector, the sentiment vector and the basic information of the role; The output vector is input into the preset conversation model to generate the reply information of the current chat.
4. The chat robot style control method based on external knowledge guidance according to claim 3 is characterized in that: The output vector is obtained based on the text input vector, emotion vector and character basic information, including: V out =V L +λ·V em Among them, V L Represents the input text vector, V em represents the emotion vector, and λ represents the adjustment weight, which is determined according to the basic information of the role and the preset experience parameters.
5. The chat robot style control method based on external knowledge guidance according to claim 1 is characterized in that: Also includes: When having multiple rounds of chats with a user, the tone and content of each round of reply information generated are adjusted according to the current user input and historical chat content, so that the character's current conversation style is consistent with the historical conversation style.
6. The chat robot style control method based on external knowledge guidance according to claim 5 is characterized in that: Each round of reply information is generated based on the current user input and historical chat content, and the tone and content of the reply are adjusted, including: Adjust the decoder parameters of the conversation generation model according to the current user input content, and generate initial reply content through the conversation model according to the adjusted decoder parameters; Calculate the similarity between the sentiment vector in the initial reply content and the sentiment vector in the previous round of reply information, and adjust the initial reply content according to the similarity; When the similarity is within the preset threshold range, the reply content of this round of dialogue is output.
7. The chat robot style control method based on external knowledge guidance according to claim 5 is characterized in that: Also includes: Obtain the state parameters, reward parameters and action parameters of each round of dialogue; the state parameters are high-dimensional vectors, including user feedback, sentiment changes and context information; the reward parameters include the user's real-time score and satisfaction feedback; the action parameters include the text and sentiment vector generated by the decoder of the conversation generation model; Iteratively adjust the action parameters according to the state parameters and reward parameters to optimize the reply content of this round of dialogue.
8. The chat robot style control method based on external knowledge guidance according to claim 7 is characterized in that: Iteratively adjust the action parameters according to the state parameters and reward parameters, including: The Q-learning algorithm is used to iteratively update the decoder parameters of the conversation model and adjust the sentiment weight parameters.
9. A chatbot style control system based on external knowledge guidance, characterized in that: A background information acquisition module, used to acquire background information of a virtual character, determine structured prompt words of the background information, and generate basic character information and sample dialogue data based on the structured prompt words and the background information; A vector library generation module is used to collect training dialogue data related to the virtual character, generate dialogue example data based on the training dialogue data and sample dialogue data, extract text vectors and sentiment vectors of the dialogue example data through vectorization technology, and generate a vector library; The dialogue reply module is used to receive the user's chat content, determine the dialogue context based on the text vector of the current chat content, search for dialogue examples related to the current context from the vector library based on the dialogue context, determine the output vector based on the basic role information, and generate reply information for the current chat based on the output vector.
10. An electronic device, characterized in that: It includes a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the chat robot style control method based on external knowledge guidance as described in any one of claims 1-8 is implemented.
Citation Information
Cited By
Chat method and device based on artificial intelligence, and storage medium
CN120296136A
Children AI dialogue generation method and system based on IP role personalization, and storage medium
CN120492596A
Dynamic interaction method based on multi-modal dynamic fusion large model and intelligent agent collaboration
CN120560517A
Dynamic interaction method based on multimodal dynamic fusion large model and intelligent agent collaboration
CN120560517B