Emotional accompanying interaction method and related device
By recording the emotional changes of virtual characters in a human-computer interaction system and generating dialogue responses that align with the user's emotions, the problem of stiff dialogue responses from virtual characters is solved, the naturalness and fluency of the dialogue and the richness of emotional expression are improved, and the realism of the virtual characters is enhanced.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- IFLYTEK CO LTD
- Filing Date
- 2025-12-02
- Publication Date
- 2026-04-17
AI Technical Summary
In existing human-computer interaction systems based on emotional companionship, the emotional expression in the dialogue responses of virtual characters is stiff and lacks richness, affecting the realism of the virtual characters.
By acquiring the character description information of the virtual character and the current dialogue information of the target user, the target memory is retrieved from the memory storage unit, a dialogue response prompt instruction is generated, and the instruction is input into the large dialogue model to generate dialogue response information aligned with the user's emotions. The emotional change information of the virtual character is recorded so that subsequent reasoning can generate more natural dialogue responses.
It improves the natural fluency of dialogue responses and the richness of emotional expression, enhances the realism of virtual characters, and achieves a deeper emotional exchange between virtual characters and users.
Smart Images

Figure CN121880503A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of human-computer interaction technology, and in particular to an emotional companionship interaction method and related device. Background Technology
[0002] With the rapid development of artificial intelligence technology, human-computer interaction systems based on emotional companionship have been widely used in many fields, such as education, healthcare, and customer service.
[0003] Currently, most human-computer interaction systems based on emotional companionship generate dialogue responses based on the user's input of the current dialogue. However, the emotional expression in these responses is often stiff and lacks richness, affecting the realism of the virtual character. Summary of the Invention
[0004] In view of the above problems, this application provides an emotional companionship interaction method and related device to achieve richer and more natural and fluent emotional expression in dialogue with users, thereby enhancing the realism of virtual characters. The specific solution is as follows:
[0005] The first aspect of this application provides an emotional companionship interaction method, including:
[0006] Obtain the character description information of the virtual character and the current dialogue information input by the target user, wherein the current dialogue information is used to communicate with the virtual character;
[0007] Retrieve target memories related to the current dialogue information from the memory storage unit. The target memories include the emotional change information of the virtual character, and the changed emotional state in the emotional change information is aligned with the actual emotional state of the target user.
[0008] Generate dialogue response prompts based on the current dialogue information, the character description information, and the target memory;
[0009] The dialogue response prompt instruction is input into the configured dialogue model to obtain the dialogue response information output by the model, which serves as the virtual character's dialogue response information to the current dialogue information.
[0010] In one possible implementation, the process of determining the changed emotional state in the emotional change information includes:
[0011] Obtain a first dialogue context, which includes user dialogue information at the time of emotion refresh and / or at least one round of dialogue history information prior to the user dialogue information;
[0012] Personality quantification is performed based on the character description information to obtain the personality characteristic parameters of the virtual character;
[0013] An emotion recognition prompt instruction is generated based on the first dialogue context and the personality feature parameters, and then input into the large dialogue model to obtain the changed emotional state output by the model.
[0014] One possible implementation also includes:
[0015] Based on the personality trait parameters, the expression state information of the virtual character is determined. The expression state information includes the latest emotional state and / or the target language style expression paradigm. The latest emotional state is the emotional state after the most recent emotional refresh.
[0016] The expressed state information is injected into the dialogue response generation process of the large dialogue model to guide the large dialogue model to generate dialogue response information that conforms to the expressed state information.
[0017] In one possible implementation, the process of determining the target language style expression paradigm includes:
[0018] For each dimension included in the personality trait parameters, generate a scene description corresponding to that dimension and a language expression paradigm under that scene description;
[0019] Each dimension of the personality trait parameters, the scene description corresponding to the dimension, and the language expression paradigm under the scene description constitute a grammatical information set.
[0020] Using the current dialogue information and / or at least one round of dialogue history information preceding the current dialogue information as the second dialogue context, at least one piece of grammatical information is filtered from the grammatical information set based on the second dialogue context and / or the target memory.
[0021] The target language style expression paradigm is obtained based on at least one piece of grammatical information.
[0022] In one possible implementation, the memories in the memory storage unit include short-term memory and long-term memory;
[0023] The short-term memory includes information on the virtual character's emotional changes and information on behavioral events extracted from dialogue history. The long-term memory includes personality-related parameters of the virtual character and information summarized and extracted from one or more short-term memories.
[0024] In one possible implementation, the generation process of any of the aforementioned short-term memories includes:
[0025] Acquire memory content, wherein the memory content includes at least one of the behavioral event information, the dialogue history information, and the emotion change information;
[0026] Identify the first entity in the memory content, and determine the entity identifier of the second entity that is most similar to the first entity from a predefined entity table;
[0027] The first entity in the memory content is replaced with the entity identifier of the second entity, and the short-term memory is generated based on the memory content after the entity identifier is replaced.
[0028] In one possible implementation, retrieving the target memory related to the current dialogue information from the memory storage unit includes:
[0029] Generate a query vector corresponding to the current dialogue information;
[0030] Obtain a set of memory vectors consisting of memory vectors corresponding to all memories in the memory storage unit, as well as the importance and timeliness values corresponding to all memories;
[0031] Calculate the similarity between the memory vector in the memory vector set and the query vector, and use it as the similarity of the memory corresponding to the memory vector, so as to obtain the similarity of each memory.
[0032] Based on the similarity, importance, and timeliness values corresponding to all the memories, the target memory related to the current dialogue information is retrieved from the memory storage unit.
[0033] One possible implementation also includes:
[0034] When the conditions for triggering an active dialogue are met, one or more of the most recently generated memories are retrieved from the memory storage unit;
[0035] Based on one or more memories, an active dialogue is initiated with the target user, wherein the virtual character has the same dialogue identifier for both active and passive dialogues with the same user.
[0036] One possible implementation also includes:
[0037] If the large dialogue model generates the same dialogue response information for different user dialogue information under the same dialogue identifier, then the target parameters of the large dialogue model are adjusted.
[0038] And / or,
[0039] If the same user dialogue information is obtained in multiple consecutive rounds of dialogue under the same dialogue identifier, then the dialogue response information output by the large dialogue model in the first round of dialogue in the multiple consecutive rounds of dialogue will be used as the dialogue response information in subsequent rounds.
[0040] One possible implementation also includes:
[0041] When the conditions for generating a blocking message are met, a reply template is selected from a preset reply template library to maintain the semantic coherence of the dialogue. The dialogue reply information is generated based on the selected reply template. The conditions for generating the blocking message include that the current dialogue information contains preset sensitive words.
[0042] A second aspect of this application provides a computer program product, including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement the emotional companionship interaction method described in the first aspect or any implementation thereof.
[0043] A third aspect of this application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:
[0044] The memory is used to store computer programs;
[0045] The processor is used to execute the computer program so that the electronic device can implement the emotional companionship interaction method of the first aspect or any implementation thereof.
[0046] The fourth aspect of this application provides a computer storage medium carrying one or more computer programs, which, when executed by an electronic device, enable the electronic device to implement the emotional companionship interaction method described in the first aspect or any implementation thereof.
[0047] By employing the aforementioned technical solution, the emotional companionship interaction method provided in this application acquires the character description information of a virtual character and the current dialogue information input by the target user. It retrieves target memories related to the current dialogue information from a memory storage unit. These target memories include information about the virtual character's emotional changes. Based on the current dialogue information, character description information, and target memories, it generates dialogue response prompts. These prompts are then input into a configured large-scale dialogue model to obtain the dialogue response information output by the model, which serves as the virtual character's response to the current dialogue information. Therefore, this application enables the virtual character's emotions to be variable during the dialogue between the virtual character and the target user. Furthermore, the information about the virtual character's emotional changes is stored in the memory storage unit in memory form, achieving continuous recording of the virtual character's emotions. Subsequently, through memory retrieval, the information about the virtual character's past emotional changes can be input into the large-scale dialogue model in memory form. This allows the large-scale dialogue model to generate dialogue response information that is more adapted to the dialogue scenario based on the dynamic evolution of the virtual character's emotions, improving the natural fluency of the dialogue response information and thus enhancing the realism of the virtual character.
[0048] Furthermore, this application aligns the changed emotional state of the virtual character with the real emotional state of the target user, which can guide the large dialogue model to generate dialogue response information that is aligned with the target user's emotions. This deepens the emotional exchange between the virtual character and the target user, and the variable emotional state of the virtual character makes the emotional expression of the dialogue response information richer and more diverse, further enhancing the realism of the virtual character. Attached Figure Description
[0049] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0050] Figure 1 A schematic diagram of a system architecture provided for this application;
[0051] Figure 2 A flowchart illustrating an emotional companionship interaction method provided in this application;
[0052] Figure 3 A schematic diagram of the structure of an emotional companionship interaction device provided in this application;
[0053] Figure 4 This is a schematic diagram of the structure of an electronic device provided in this application. Detailed Implementation
[0054] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.
[0055] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0056] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0057] As described in the background section, most human-computer interaction systems based on emotional companionship generate dialogue responses based on the user's input of the current dialogue. However, the emotional expression in these responses is often stiff and lacks richness, affecting the realism of the virtual character. In-depth research reveals one possible reason: the dialogue responses are generated based on a pre-set static emotional state for the virtual character, which does not change with each round of dialogue. This results in stiff and unrich emotional expression in the responses, ultimately leading to a low level of realism in the virtual character.
[0058] To address this issue, this application provides an emotional companionship interaction method and related apparatus, enabling the virtual character's emotions to follow the target user's real emotional changes, and storing the information of the virtual character's past emotional changes as memory. This allows the dialogue response information generated based on memory to be more realistically and naturally integrated into the dialogue scene, thereby improving the realism of the virtual character.
[0059] Optionally, the emotional companionship interaction method provided in this application can be applied to, for example... Figure 1 The system architecture shown includes a terminal 100 and a server 200. The server 200 may include one or more servers (…). Figure 1 (This example uses a server as an illustration).
[0060] Either terminal 100 or server 200 can be used independently to execute the emotional companionship interaction method provided in the embodiments of this application. Alternatively, terminal 100 and server 200 can also be used collaboratively to execute the emotional companionship interaction method provided in the embodiments of this application.
[0061] The following description Figure 1 The product form of the mid-terminal 100;
[0062] The terminal 100 in this application embodiment can be a mobile phone, tablet computer, wearable device, vehicle device, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc., and this application embodiment does not impose any restrictions on it.
[0063] To enable those skilled in the art to better understand this application, the emotional companionship interaction method of the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0064] Reference Figure 2 , Figure 2 This is a flowchart illustrating an emotional companionship interaction method provided in an embodiment of this application, as shown below. Figure 2 As shown, this emotional companionship interaction method may include:
[0065] Step S101: Obtain the character description information of the virtual character and the current dialogue information input by the target user.
[0066] Among them, the character description information of virtual characters refers to the descriptive information that describes the character attributes such as personality, occupation, hobbies, catchphrases, worldview, outlook on life, and values of virtual characters.
[0067] Optionally, character description information can be in one or more formats, including text, images, and audio. Of course, character description information can also be in other formats, which are not specifically limited here.
[0068] In one possible implementation, the role description information can be immutable role description information that is fixed to the system during the system development phase, or it can be role description information configured by the user, or a combination of both.
[0069] In a real system, there can be one or more virtual characters. The character description information of the virtual characters obtained in this embodiment refers to the character description information of the virtual characters that need to talk to the target user. For example, optionally, when the target user wants to talk to a certain virtual character, he / she can open a chat dialog box with the virtual character and enter the current dialogue information. At this time, this embodiment can obtain the character description information and the current dialogue information of the virtual character corresponding to the chat dialog box.
[0070] Step S102: Retrieve target memories related to the current dialogue information from the memory storage unit. The target memories include the emotional change information of the virtual character. The changed emotional state in the emotional change information is aligned with the target user's real emotional state.
[0071] Here, memory storage unit refers to a storage unit capable of storing memories, such as relational databases like MySQL and PostgreSQL.
[0072] In this embodiment, the virtual character's emotions can follow the target user's real emotional state changes, and the virtual character's emotional changes throughout can be summarized into emotional change information and then stored in the memory storage unit in the form of memory.
[0073] For example, the initial emotional state of a virtual character is a default emotional state, such as "calm and relaxed". When an emotional state needs to be refreshed, the changed emotional state is generated by combining the target user's real emotional state. For example, if the target user's real emotional state is sadness after losing a beloved pet, then the changed emotional state can be "sympathy and concern". At this time, the emotional change information is "Ouzai (the name of a virtual character)'s emotion has changed from calm to sympathy and concern because the user expressed sadness". This emotional change information can be stored as a memory in the memory storage unit.
[0074] Step S103: Generate dialogue response prompts based on the current dialogue information, character description information, and target memory.
[0075] In this embodiment, a dialogue response prompt instruction template can be generated in advance. The dialogue response prompt instruction template includes at least a dialogue content slot, a character description slot, and a memory slot, so as to guide the large dialogue model to generate reasonable dialogue response information through the information in each slot.
[0076] For example, a possible dialogue response prompt instruction template is:
[0077] "Let's assume you are {agent_name}, {core_characters}, and have the {personality} personality."
[0078] The current time is {current_time}, and you are having a conversation with {subject_involved}. The conversation content is {conversation}.
[0079] Background memory: {latest_memory}.
[0080] Task: Please respond in the first person, in a manner consistent with the character's established personality.
[0081] Among them, agent_name is used to fill in the name of the virtual character, such as "Ou Zai"; core_characters and personality are used to fill in the character description information. core_characters is used to fill in the core characteristics, such as "a very excellent detective", and personality is used to fill in personality-related parameters, such as "calm and meticulous, composed and confident, speaks to the point, and has an adventurous spirit"; current_time is used to fill in the current time; subject_involved is used to store the identifier of the target user, such as "Xiao Ming"; latest_memory is used to fill in the target memory retrieved from the previous text; and conversation is used to fill in the dialogue content, such as the current dialogue information.
[0082] Optionally, the conversation can also be populated with at least one round of dialogue history information preceding the current dialogue information to provide more information for the large dialogue model to perform deep reasoning.
[0083] Step S104: Input the dialogue response prompt command into the configured dialogue model to obtain the dialogue response information output by the model, which serves as the virtual character's dialogue response information to the current dialogue information.
[0084] Specifically, dialogue response prompts can be input into the large dialogue model, enabling the model to perform deep reasoning based on the prompts, generate dialogue response information that better matches the information in each slot of the prompts, and output the response information as a virtual character to the target user.
[0085] Understandably, the time in the target memory may be a time noun (such as next year). For accurate reasoning, the dialogue model can optionally convert the time noun into an actual date before proceeding with subsequent reasoning.
[0086] Optionally, the large dialogue model can be an existing large language model, such as the iFlytek Spark Large Model, the GPT series Large Model, the Doubao Large Model, etc. This application does not impose specific limitations.
[0087] The emotional companionship interaction method provided in this application obtains the character description information of a virtual character and the current dialogue information input by the target user. It retrieves target memories related to the current dialogue information from a memory storage unit. These target memories include information about the virtual character's emotional changes. Based on the current dialogue information, character description information, and target memories, it generates dialogue response prompts. These prompts are then input into a configured large-scale dialogue model to obtain the dialogue response information output by the model, which serves as the virtual character's response to the current dialogue information. Therefore, this application enables the virtual character's emotions to be variable during the dialogue between the virtual character and the target user. Furthermore, the information about the virtual character's emotional changes is stored in the memory storage unit in memory form, achieving continuous recording of the virtual character's emotions. Subsequent memory retrieval allows the virtual character's past emotional changes to be input into the large-scale dialogue model in memory form. This allows the large-scale dialogue model to generate dialogue response information that is more suitable for the dialogue scenario based on the dynamic evolution of the virtual character's emotions, improving the naturalness and fluency of the dialogue response information and thus enhancing the realism of the virtual character.
[0088] Furthermore, this application aligns the changed emotional state of the virtual character with the real emotional state of the target user, which can guide the large dialogue model to generate dialogue response information that is aligned with the target user's emotions. This deepens the emotional exchange between the virtual character and the target user, and the variable emotional state of the virtual character makes the emotional expression of the dialogue response information richer and more diverse, further enhancing the realism of the virtual character.
[0089] In some embodiments of this application, the process of determining the changed emotional state in the emotional change information in step S102 above is described in detail.
[0090] This embodiment allows for preset emotion refresh trigger conditions, which trigger emotion refresh when the conditions are met. For example, the emotion refresh trigger condition can be to refresh the emotion once every target number of dialogue rounds, optionally with a target number of rounds being 3; the emotion refresh trigger condition can also be manually triggered.
[0091] In order to determine the emotional state of a virtual character after a change, the latest user dialogue information and / or at least one round of dialogue history information before the current dialogue information can be obtained each time the emotion is refreshed. Here, user dialogue information refers to the current dialogue information entered by the user each time the emotion is refreshed. Then, personality quantification is performed based on the character description information to obtain the personality characteristic parameters of the virtual character.
[0092] Optionally, a personality quantification model containing 17 dimensions can be constructed based on the "Big Five personality theory" in psychology. Then, based on the individual quantification model and the character description information, the score values of each of the 17 dimensions can be obtained. The personality characteristic parameters of the virtual character can be obtained from the score values of each of the 17 dimensions.
[0093] The Big Five personality traits include: extraversion (assessing social tendencies and energy), neuroticism (assessing emotional stability), openness (assessing cognitive and aesthetic tendencies), agreeableness (assessing interpersonal affinity), and conscientiousness (assessing goal-orientation and sense of responsibility).
[0094] Based on the Big Five personality traits mentioned above, the 17 dimensions include the following: Extraversion includes three dimensions: social preference, energy level, and risk-taking willingness; Neuroticism includes three dimensions: emotion regulation, stress tolerance, and self-awareness; Openness includes four dimensions: curiosity, creativity, aesthetic interest, and inclusiveness; Agreeableness includes three dimensions: cooperativeness, friendliness, and compassion; and Conscientiousness includes four dimensions: responsibility, self-discipline, goal orientation, and commitment.
[0095] It should be noted that the above 17 dimensions are merely examples and are not intended to limit this application.
[0096] One possible scoring method is as follows: combine the role description information with 17 dimensions to form a personality quantification prompt instruction, input the personality quantification prompt instruction into the large dialogue model, obtain the score value and the reason for the score for each dimension output by the model, and store the score value and the reason for the score for each dimension output by the model in the Personality field in a preset format. Optionally, the preset format can be JSON (JavaScript Object Notation) format.
[0097] Optionally, each dimension can be rated on a scale of 1 to 5, with 3 being a neutral value.
[0098] Considering that the output format may be incorrect, in this embodiment, a retry mechanism can be initiated when the format is incorrect. For example, a maximum of 3 retries can be performed. If the correct format cannot be output after 3 retries, an error will be reported and no further dialogue will be conducted.
[0099] Furthermore, after obtaining the personality characteristic parameters of the virtual character, this embodiment can generate an emotion recognition prompt instruction based on the first dialogue context and the personality characteristic parameters, and input it into the large dialogue model to obtain the changed emotional state output by the model.
[0100] In one possible implementation, emotions can be pre-categorized, optionally into 34 emotions, such as "positive emotions: happiness, pride, calmness, ease, interest; negative emotions: anger, anxiety, sadness, disgust, fear; complex emotions: admiration, nostalgia, embarrassment, jealousy; etc."
[0101] One possible prompt template for emotion recognition prompts could be:
[0102] "Assuming you are {agent_name}, your personality dimensions are as follows:"
[0103] <Social Preferences: 4, Emotion Regulation: 3, Curiosity: 5...>
[0104] Please select 1-3 emotions from the following list that best describe your current mood:
[0105] [Happiness, anxiety, surprise, anger...]
[0106] Dialogue content: <User says "xx">
[0107] Output format: {"mood":["xx","xx"],"note":"xx"}.
[0108] Here, "mood" represents emotion, and "note" represents the reason for judgment. For example, {"mood":["frustrated","anxious"],"note":"feeling stressed due to project failure"}.
[0109] Optionally, after inputting the above emotion recognition prompts into the large dialogue model, the user's true emotional state can be determined based on the first dialogue context within the large dialogue model. Then, based on the user's true emotional state and personality characteristic parameters, the 1 to 3 most suitable emotions can be selected from 34 emotions as the changed emotional state.
[0110] It should be noted that the process of obtaining the changed emotional state through internal reasoning within the aforementioned large-scale dialogue model is merely an example and should not be construed as limiting this application. Furthermore, the process of generating emotion recognition prompts based on the first dialogue context and personality characteristic parameters is only an example. In addition, emotion recognition prompts can be generated based on other information, such as the first dialogue context, personality characteristic parameters, and worldview information from the character description.
[0111] This embodiment constructs a unique 17-dimensional personality trait parameter system, enabling a quantitative description of the virtual character's personality traits from multiple dimensions. These 17 dimensions cover various aspects of human character, emotions, and behavioral patterns. Through the analysis and modeling of a large amount of data, each personality trait is assigned a specific quantitative value. This personality quantification model provides a solid foundation for accurately understanding and simulating human emotions, and can more meticulously depict the emotional tendencies and behavioral characteristics of different individuals. Compared to traditional, simple personality classification methods, it has higher resolution and accuracy.
[0112] The large-scale dialogue model can utilize advanced sensor technology and data analysis algorithms to capture and analyze users' true emotional state in real time, accurately determining whether the user's current emotion is positive, negative, or neutral. This provides an emotional basis for subsequent interactions, enabling virtual characters to respond more appropriately based on the user's true emotional state and improving the user experience.
[0113] Furthermore, in this embodiment, the emotional state of the virtual character is updated according to the above process each time the emotion refresh trigger condition is met, and the emotional change information is recorded as a memory, thus forming a traceable emotional trajectory. The large dialogue model can infer based on this traceable emotional trajectory and generate dialogue response information that resonates with the user, thereby improving the accuracy of dialogue response information and the richness and naturalness of emotional expression, and enhancing the user's emotional companionship interaction experience.
[0114] Furthermore, in order to make the dialogue response information more compatible with the virtual character, that is, to make it look more like the dialogue response information actually output by a person with the same character characteristics as the virtual character, this embodiment can also determine the expression state information of the virtual character based on the personality characteristic parameters of the virtual character, and then inject the expression state information into the dialogue response generation process of the large dialogue model to guide the large dialogue model to generate dialogue response information that conforms to the expression state information.
[0115] Optionally, "injecting expressive state information into the dialogue response generation process of the large dialogue model" may include: concatenating the expressive state information with the current dialogue information, role description information and target memory in the preceding text to form a dialogue response prompt instruction; or, in the case that the large dialogue model is a conditional injection model, injecting the expressive state information as a condition into the dialogue response generation process.
[0116] Of course, there are other ways besides these, but we will not specify any particular method here.
[0117] Optionally, the expression state information includes the latest emotional state of the virtual character and / or the target language style expression paradigm. The latest emotional state is the emotional state after the most recent emotional refresh, and the target language style expression paradigm is used to represent a language style that is adapted to the virtual character.
[0118] Optionally, the process of determining the target language style expression paradigm may include: generating a scene description and a language expression paradigm under the scene description for each dimension included in the personality feature parameters; forming a grammatical information set by composing a grammatical information set by combining each dimension included in the personality feature parameters, the scene description corresponding to the dimension, and the language expression paradigm under the scene description; using the current dialogue information and / or at least one round of dialogue history information before the current dialogue information as the second dialogue context, filtering at least one grammatical information set from the grammatical information set based on the second dialogue context and / or target memory, and obtaining the target language style expression paradigm based on the at least one grammatical information set.
[0119] Optionally, a large dialogue model can be used to generate a corresponding scenario description and a language expression paradigm under the scenario description for each dimension of the personality trait parameters, based on the score value of each dimension and the role description information. Here, the language expression paradigm under the scenario description refers to the most likely words spoken in the application scenario described by the scenario description.
[0120] Optionally, the language expression paradigm can be a more colloquial one to better reflect real-life scenarios.
[0121] To save storage space, the word count of the language expression paradigm can be less than or equal to the target word count (e.g., 20 words).
[0122] Optionally, the grammatical information corresponding to any dimension may include: the dimension's identification (ID), the scene description, and the language expression paradigm under the scene description.
[0123] Taking "social preferences" as an example, assuming its identity ID is 1, the grammatical information can be:
[0124] “ID":"1",
[0125] "scene": "Ou Zai is at a detective gathering, exchanging ideas with other detectives."
[0126] "example": "I think this clue might be related to the case."
[0127] To make the system more flexible, it allows developers to add, delete, modify, and query the above scenario descriptions and language expression paradigms, enabling fine-tuning of the style.
[0128] Optionally, the second dialogue context and / or target memory, along with the grammatical information set, can be concatenated into style generation prompts and fed into the large dialogue model. Within the large dialogue model, the relevance of the scene description in the grammatical information set to the second dialogue context and target memory can be evaluated, and several (e.g., two) grammatical information items with the highest relevance can be selected to obtain the aforementioned target language style expression paradigm.
[0129] For example, the language expression paradigm contained in the selected grammatical information can be used as the target language style expression paradigm, or the language expression paradigm contained in the selected grammatical information can be adjusted according to the latest emotional state of the virtual character determined above, so as to obtain the target language style expression paradigm.
[0130] This embodiment can dynamically generate a matching language style based on personality trait parameters, and this language style can also be combined with the user's real-time emotional state. When the user is in a positive mood, the language style generated by the system may be more enthusiastic and lively; when the user is in a negative mood, the language style will be more gentle and comforting. This dynamic language style generation mechanism makes the interaction between the system and the user more natural and smooth, and can better evoke the user's emotional resonance, greatly enhancing the system's emotional interaction capabilities.
[0131] In other embodiments of this application, the memory in the aforementioned memory storage unit, as well as the memory generation and retrieval process, are described in detail.
[0132] Optionally, the memories in the memory storage unit may include short-term memory and long-term memory. The short-term memory includes information on the virtual character's emotional changes and behavioral event information extracted from the dialogue history. The long-term memory includes personality-related parameters of the virtual character and information summarized and extracted from one or more short-term memories.
[0133] More specifically, the memory storage unit can store multiple short-term memories and multiple long-term memories. Each short-term memory includes one or more pieces of information, such as information on the virtual character's emotional changes and information on behavioral events extracted from dialogue history. Each long-term memory includes one or more pieces of information, such as personality-related parameters of the virtual character and information summarized and extracted from more than one short-term memory.
[0134] For example, when configuring character description information, personality-related parameters such as hobbies, catchphrases, worldview, outlook on life, values, and personality descriptions can all be converted into one or more long-term memories of the virtual character.
[0135] Each time a round of dialogue is generated (i.e., a round of dialogue history information), behavioral event information can be extracted from the dialogue history information and recorded as a dialogue-type short-term memory.
[0136] When a virtual character's emotions are refreshed, the information about the changes in the virtual character's emotions is automatically recorded as an observational short-term memory.
[0137] When the preset conditions for generating long-term memory are met, a large model (such as OpenAI or Spark large model) is used to summarize and refine one or more short-term memories that meet the conditions for generating long-term memory, resulting in a long-term memory (optionally, an error retry mechanism can be set up; if it still fails after 3 retries, manual intervention is required to summarize and refine). For example, possible conditions for generating long-term memory are: every 10 short-term memories generated, or every 24 hours.
[0138] For example, long-term memories derived from several short-term memories could be: "User Xiaoming reported on July 19, 2024, that his pet cat had gone missing in the community garden and was feeling anxious. Detective Ou Zai has intervened in the investigation and collected preliminary clues such as time and location."
[0139] To enable those skilled in the art to better understand the Long Short-Term Memory (LSTM) in this application, the data structure of LSTM is described below, along with a short-term memory generation process based on the LSTM data structure.
[0140] Optionally, the data structure of short-term memory may include: a unique identifier for the short-term memory (interaction_ID), an identity identifier for the virtual role associated with the short-term memory (agent_ID), a memory type (including observation and conversation), a natural language description of the memory content (description), an ISO 8601 timestamp (create_time), an update timestamp (latest_access_time, which defaults to the same as create_time), and a list of participants in the observation or conversation (subject_involved). ISO 8601 is an international standard that specifies how to represent dates and times using numbers, dates, and times in a clear, unambiguous, and machine-processable manner.
[0141] Considering that the same entity can be described in different natural language terms in different dialogues (e.g., a cat can be described as a white cat, kitten, or tomcat), this embodiment can pre-configure an entity table to unify these identical entities described in different natural language terms in short-term memory. The entity table includes a generic name for the entity (e.g., cat) and its corresponding entity ID. This allows the entity ID to replace the specific entity in the natural language description of the remembered content and in the list of people participating in the observation or dialogue. Of course, the entity table can also contain other fields, such as type and attributes (e.g., JSON format), which are not limited in this application.
[0142] For example, a natural language description of the memory content could be: [u001] asks [agent001] for help, saying that [obj001] is missing at [time].
[0143] Based on this, optionally, the generation process of any short-term memory may include: acquiring memory content, which includes at least one of behavioral event information, dialogue history information, and emotional change information; identifying a first entity in the memory content, for example, using the spaCy library to identify first entities such as names of people, place names, behavioral events, virtual characters, and objects in the memory content, then determining the entity identifier of the second entity most similar to the first entity from a predefined entity table, replacing the first entity in the memory content with the entity identifier of the second entity, and generating a short-term memory that conforms to the above data structure based on the memory content after replacing the entity identifier.
[0144] Optionally, various matching algorithms, such as fuzzy matching or word vector similarity matching, can be used to determine the second entity most similar to the first entity from the entity table.
[0145] It is understandable that a certain first entity may not exist in the entity table. In this case, the matching algorithm described above will fail to match the entity. In such cases, the first entity can be registered in the entity table.
[0146] Furthermore, the entity table in this embodiment can also support developers to register, modify, and delete entities individually or in batches through the interface.
[0147] Optionally, the data structure of long-term memory may include: a unique identifier for the long-term memory (memory_ID), an identity identifier for the virtual character associated with the long-term memory (agent_ID), a memory type (including short-term memory summary, knowledge base, and character setting), an array of entity IDs involved (subject_involved), an array of associated short-term memory IDs (interaction_involved), a natural language description of the memory content (description), a recency value (ranging from 0 to 1, calculated using a forgetting function), an importance value (ranging from 0 to 1), an ISO 8601 timestamp (create_time), and an update timestamp (latest_access_time, which defaults to the same as create_time).
[0148] To obtain the timeliness and importance of long-term memory, alternatively, a rule engine or a large model can be used.
[0149] For example, in the character settings of virtual characters (such as character description information), attribute information (such as nationality) is automatically stored in long-term memory, and its validity value and importance are set to 1 through the rule engine.
[0150] For example, for long-term memories summarized and refined from several short-term memories, a large model (such as the Spark Large Model) can be used to score them, and the score value can be used as the importance.
[0151] For example, the forgetting function can be set through a rule engine, and a scheduled task can be used to periodically update the validity period of long-term memory. Here, the forgetting function can be: ,in, This represents the time-sensitivity value at time t. Indicates the moment when long-term memory is formed. and This indicates the preset adjustment coefficient.
[0152] Optionally, the rule engine in this embodiment may also provide functions such as enabling or disabling privacy protection, returning only the target user's long and short term memory during retrieval, configuring thresholds for each step of this application, and ensuring that configuration changes take effect promptly without restarting the service.
[0153] Next, the process of "retrieving the target memory related to the current dialogue information from the memory storage unit" in step S102 above will be introduced.
[0154] To enable accurate and efficient memory retrieval, this embodiment employs a vector retrieval method. Based on this, all memories within the memory storage unit can be converted into memory vectors, forming a set of memory vectors.
[0155] Alternatively, Sentence-BERT (Sentence Bidirectional Encoder Representations from Transformers) or CLIP (Contrastive Language-Image Pre-training) models can be used to process the natural language description of the memory content in the memory storage unit into a vector representation, which serves as the corresponding memory vector.
[0156] Optionally, the memory vector set can be stored in a vector database, such as FAISS, Pinecone, etc.
[0157] Next, a query vector corresponding to the current dialogue information can be generated. Optionally, a third entity can be extracted from the current dialogue information and processed into a vector representation, which can then be used as the query vector corresponding to the current dialogue information.
[0158] The third entity is the core content of the current dialogue information. Generating query vectors based solely on the third entity allows for more accurate retrieval of relevant memories.
[0159] Optionally, during memory retrieval, memories of non-target users can be filtered based on privacy settings to improve retrieval efficiency.
[0160] Furthermore, this embodiment can calculate the similarity between the memory vectors in the memory vector set and the query vectors, and use this similarity as the similarity of the memory corresponding to the memory vector, so as to obtain the similarity of each memory.
[0161] Optionally, the similarity can be cosine similarity or Euclidean distance, etc.
[0162] Finally, based on the similarity, importance, and timeliness values of all memories, the target memory related to the current dialogue information is retrieved from the memory storage unit.
[0163] Optionally, the recall score for each memory can be calculated using the following formula, and then several memories with high recall scores can be selected as target memories.
[0164] ;
[0165] in, This represents the recall score corresponding to the m-th memory. This represents the expiration value corresponding to the m-th memory. This represents the importance of the m-th memory. This represents the relevance of the m-th memory. , and This indicates the preset weighting coefficient.
[0166] Optionally, the top N memories with the highest recall scores can be selected as target memories, or the memories that meet the threshold condition can be selected. The memory of [the target memory] is used as the target memory, among which, This indicates the preset memory retrieval threshold.
[0167] As mentioned earlier, the data structure of short-term memory can be without time-to-sense and importance parameters. In this case, the time-to-sense and importance of short-term memory can both be 0.
[0168] Optionally, a vector database indexing acceleration strategy (such as IVF-PQ) or a strategy of caching high-frequency search results can be used to accelerate memory retrieval and improve memory retrieval efficiency.
[0169] In summary, this embodiment employs a long short-term memory (LSTM) storage architecture. STM is used for rapid processing and response to behavioral events during the current interaction, efficiently capturing and retaining relevant data to ensure continuity and real-time performance during dialogues or task execution. Simultaneously, STM stores information on the virtual character's emotional changes, achieving continuous memorization of the character's dynamic emotional shifts. Long-term memory stores more persistent and important information. This separate storage approach significantly improves the efficiency of information retrieval and use, avoiding slow retrieval and errors caused by a large amount of mixed information.
[0170] When storing memories in the form of entity identifiers, different natural language descriptions representing the same entity can be unified. This structured processing method can more clearly link various memories, which helps to understand and utilize information more comprehensively and deeply. When facing complex problems and situations, it can quickly integrate relevant memories and provide more accurate and comprehensive answers and decision-making basis.
[0171] When retrieving memories, a multi-dimensional retrieval strategy based on similarity, timeliness, and importance is employed. This strategy enables the retrieval of more important recent memories that are more relevant to the current conversation context. Responding to conversations based on these memories results in higher accuracy and more empathetic responses to the target user, leading to a better user experience.
[0172] It is understandable that in a dialogue scenario, the target user may exit the dialogue with the virtual character and then re-enter. To provide the target user with a better dialogue experience and maintain the continuity of the dialogue, optionally, this embodiment can retrieve one or more recently generated memories from the memory storage unit when the active dialogue trigger condition is met. These memories can be long-term memories and / or short-term memories. Then, based on these memories, an active dialogue with the target user is initiated. In this case, the virtual character has the same dialogue identifier for both active and passive dialogues with the same user.
[0173] Optionally, an active dialogue can be initiated based on one or more memories, combined with the personality trait parameters and / or grammatical information set determined above. For example, several (e.g., two) of the most relevant scene descriptions can be selected from the grammatical information set, and active dialogue content can be generated based on the language expression paradigms corresponding to these scene descriptions and one or more memories, thereby initiating an active dialogue.
[0174] For example, on January 1st, Xiaoming mentioned to Ouzai that his cat was missing. On January 2nd, when Xiaoming re-entered the chat interface with Ouzai, based on the latest long-term memory that "user Xiaoming reported on July 19, 2024, that his pet cat had gone missing in the community garden and was feeling anxious, and detective Ouzai had intervened in the investigation and collected preliminary clues such as time and location" and Ouzai's "adventurous spirit" personality, he could initiate an active conversation with user Xiaoming: "Xiaoming, regarding the matter of finding the cat, I have thought of some new directions for investigation. We can continue to discuss it when it is convenient for you."
[0175] For example, one possible implementation is as follows: When a user enters a dialogue interface with a virtual character and initiates a greeting, a dialogue identifier (chatid1) is generated, and the content of the greeting is stored. When the user asks a follow-up question, the system identifies the same dialogue identifier (chatid1), automatically inherits the content of the greeting, and constructs a complete context including the greeting content and the user's follow-up question. Based on this complete context and the virtual character's latest emotional state, a coherent response is generated.
[0176] In this embodiment, by using the same dialogue identifier for both active and passive dialogues between virtual characters and the same user, memory-driven active interaction can be achieved, maintaining the continuity of the dialogue and improving the user interaction experience.
[0177] Considering that users may input different dialogue information in dialogue scenarios, but the large dialogue model generates the same dialogue response information, this application provides a duplicate generation detection mechanism to address the problem of duplicate generation in the large model.
[0178] Optionally, if the large dialogue model generates the same dialogue response information for different user dialogue information under the same dialogue identifier, then the target parameters of the large dialogue model should be adjusted.
[0179] For example, after each generation of dialogue response information, precise text matching is performed from all dialogue history information under the stored current dialogue identifier to determine whether the large dialogue model generates the same dialogue response information for different user dialogue information. If so, the temperature parameter is increased to 0.8 during the first repetition, and the topN in the previous text is adjusted to 6. During two consecutive repetitions, the loop is forcibly exited to avoid infinite retries. After exiting, the adjusted target parameters are automatically restored to the default parameters.
[0180] This duplicate generation detection mechanism avoids the problem of large models generating identical responses, which can lead to a looping dialogue and negatively impact user experience, thus making the dialogue more efficient.
[0181] Understandably, there may be situations where users input the same dialogue information. Since the dialogue model consumes tokens to generate responses, and in some scenarios, tokens need to be purchased by users, to avoid wasting the tokens purchased by users, optionally, if the same user dialogue information is obtained in multiple consecutive rounds of dialogue under the same dialogue identifier, the dialogue response information output by the dialogue model in the first round of dialogue in the multiple consecutive rounds of dialogue will be used as the dialogue response information in subsequent rounds.
[0182] It is also understandable that, in order to improve the accuracy of subsequent dialogue responses, the generated dialogue history information from each round will be stored as a reference for subsequent dialogue responses.
[0183] However, there is a token limit for storing dialogue history information. To avoid clearing all stored dialogue history information when the token limit is reached, resulting in a large loss of context information and affecting the continuity of the dialogue, this embodiment can optionally store the dialogue history information of each round in the order of its generation time. Before each user inputs new dialogue information, the total number of tokens contained in the currently stored dialogue history information is calculated. If the total number of tokens is close to the token limit (for example, the difference between the two is less than a preset difference threshold), the dialogue history information is discarded in order of its generation time from oldest to newest. This ensures context relevance by prioritizing the retention of the most recent dialogue rounds.
[0184] It's also understandable that users might input sensitive words (such as prohibited words). To ensure a relatively coherent dialogue even when users input sensitive words without answering their questions, optionally, pre-defined conditions for generating prohibited text can be set. When these conditions are met, based on the preceding second dialogue context, a pre-defined response template is selected from a library to maintain semantic coherence, and a dialogue response is generated based on the selected template. Here, the response template used to maintain semantic coherence can be the one most semantically similar to the second dialogue context.
[0185] Optionally, the conditions for generating the blocking text include the current dialogue information containing preset sensitive words, the dialogue model not responding after a timeout, network anomalies, etc.
[0186] Optionally, "generating dialogue response information based on the selected response template" may include: determining the target tone to match based on the latest emotional state in the preceding text, and generating dialogue response information based on the selected response template and the target tone.
[0187] This embodiment, through a fallback blocking mechanism, can automatically and promptly generate a neutral response that conforms to the context, such as "Let's talk about this from a different angle?", when the target user's input contains sensitive words that cause the large dialogue model to refuse to answer. This avoids returning an error or causing the dialogue structure to become chaotic, thus maintaining the normal operation of the system and preserving the continuity of the dialogue.
[0188] This application also provides an emotional companionship interaction system, which includes a memory component, a contextual coherence component, and an emotion component. These three components work together to achieve better dialogue. Specifically, the memory component provides relevant memories through memory retrieval, the emotion component provides emotional guidance and language style based on real-time user emotion analysis, and the contextual coherence component combines relevant memories and emotional guidance to make a comprehensive judgment, determining the focus and approach of the response.
[0189] Corresponding to the previous embodiments, the memory component is used to generate, store, and retrieve memories; the emotion component is used to determine the emotional change information of the virtual character and the target language style expression paradigm; and the context coherence component is used to combine the target memories retrieved by the memory component, the emotional change information determined by the emotion component, and the target language style expression paradigm to call the dialogue big model to generate dialogue response information for the current dialogue information, as well as to perform the aforementioned proactive dialogue, repeated generation management, and blocking fallback management.
[0190] To help those skilled in the art better understand the three components described above, several possible component interaction processes are shown below.
[0191] First, when analyzing a user's true emotional state, the emotion component can refer to the relevant memories of the user's past behaviors and preferences in the memory component, thereby more accurately determining the user's current true emotional state.
[0192] Second, the emotion component feeds back information about the virtual character's emotional changes to the memory and contextual coherence components. The memory component stores this information in memory form, enriching its understanding of the user. The contextual coherence component then adjusts its dialogue strategy and responses based on this information, making the dialogue more aligned with the user's emotional needs. For example, when the emotion component detects that the user is angry, the contextual coherence component can adjust its language style to respond in a gentler, more soothing manner.
[0193] Third, during dialogue processing, the contextual coherence component promptly transmits newly generated historical dialogue information to the memory component for storage and management. Simultaneously, its proactive dialogue and reactive response strategies are also influenced by the sentiment analysis results. For example, when the sentiment component determines that a user is interested in a particular topic, the contextual coherence component can proactively initiate a deeper dialogue around that topic.
[0194] Fourth, the three components can also provide feedback to each other to achieve self-optimization and improvement. For example, if the memory component finds that some memories do not match the needs of the emotion component or the contextual coherence component during the storage and retrieval of memories, it can adjust the storage and retrieval strategies in a timely manner; the emotion component can continuously optimize the algorithms for real-time sentiment analysis and dynamic language style generation based on feedback from the contextual coherence component to improve the accuracy of sentiment analysis and the rationality of language style generation; if the user is not satisfied with the dialogue response, it can analyze which component is causing the problem and make targeted improvements.
[0195] The above describes an emotional companionship interaction method provided by the embodiments of this application. The following will describe the apparatus for performing the above-described emotional companionship interaction method.
[0196] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of an emotional companionship interaction device provided in an embodiment of this application. Figure 3 As shown, the emotional companionship interaction device may include:
[0197] The data input unit 301 is used to obtain the character description information of the virtual character and the current dialogue information input by the target user. The current dialogue information is used to have a dialogue with the virtual character.
[0198] The data query unit 302 is used to retrieve target memories related to the current dialogue information from the memory storage unit. The target memories include the emotional change information of the virtual character, and the changed emotional state in the emotional change information is aligned with the real emotional state of the target user.
[0199] The data processing unit 303 is used to generate dialogue response prompts based on the current dialogue information, character description information and target memory, input the dialogue response prompts into the configured large dialogue model, and obtain the dialogue response information output by the model, which serves as the virtual character's dialogue response information to the current dialogue information.
[0200] Each module in the aforementioned emotional companionship interaction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0201] This application also provides an electronic device, which may include at least one processor and a memory connected to the processor, wherein:
[0202] Memory is used to store computer programs;
[0203] The processor is used to execute computer programs to enable electronic devices to implement any of the emotional companionship interaction methods provided in the embodiments of this application.
[0204] refer to Figure 4 The diagram illustrates a structural schematic suitable for implementing the electronic device in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 4 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0205] like Figure 4 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. When the electronic device is powered on, the RAM 603 also stores various programs and data required for the operation of the electronic device. The processing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0206] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, memory cards, hard drives, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0207] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the emotional companionship interaction methods provided in this application.
[0208] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the emotional companionship interaction methods provided in this application.
[0209] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0210] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0211] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0212] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. An emotional companion interaction method, characterized by, include: Obtain the character description information of the virtual character and the current dialogue information input by the target user, wherein the current dialogue information is used to communicate with the virtual character; Retrieve target memories related to the current dialogue information from the memory storage unit. The target memories include the emotional change information of the virtual character, and the changed emotional state in the emotional change information is aligned with the actual emotional state of the target user. Generate dialogue response prompts based on the current dialogue information, the character description information, and the target memory; The dialogue response prompt instruction is input into the configured dialogue model to obtain the dialogue response information output by the model, which serves as the virtual character's dialogue response information to the current dialogue information. 2.The emotional companion interaction method of claim 1, wherein, The process of determining the changed emotional state in the emotional change information includes: Obtain a first dialogue context, which includes user dialogue information at the time of emotion refresh and / or at least one round of dialogue history information prior to the user dialogue information; Personality quantification is performed based on the character description information to obtain the personality characteristic parameters of the virtual character; An emotion recognition prompt instruction is generated based on the first dialogue context and the personality feature parameters, and then input into the large dialogue model to obtain the changed emotional state output by the model. 3.The emotional companion interaction method of claim 2, wherein, Also includes: Based on the personality trait parameters, the expression state information of the virtual character is determined. The expression state information includes the latest emotional state and / or the target language style expression paradigm. The latest emotional state is the emotional state after the most recent emotional refresh. The expressed state information is injected into the dialogue response generation process of the large dialogue model to guide the large dialogue model to generate dialogue response information that conforms to the expressed state information. 4.The emotional-companion interaction method of claim 3, wherein, The process of determining the target language style expression paradigm includes: For each dimension included in the personality trait parameters, generate a scene description corresponding to that dimension and a language expression paradigm under that scene description; Each dimension of the personality trait parameters, the scene description corresponding to the dimension, and the language expression paradigm under the scene description constitute a grammatical information set. Using the current dialogue information and / or at least one round of dialogue history information preceding the current dialogue information as the second dialogue context, at least one piece of grammatical information is filtered from the grammatical information set based on the second dialogue context and / or the target memory. The target language style expression paradigm is obtained based on at least one piece of grammatical information. 5.The emotional-companion interaction method of claim 1, wherein, The memories in the memory storage unit include short-term memory and long-term memory; The short-term memory includes information on the virtual character's emotional changes and information on behavioral events extracted from dialogue history. The long-term memory includes personality-related parameters of the virtual character and information summarized and extracted from one or more short-term memories. 6.The emotional-companion interaction method of claim 5, wherein, The process of generating any one of the aforementioned short-term memories includes: Acquire memory content, wherein the memory content includes at least one of the behavioral event information, the dialogue history information, and the emotion change information; Identify the first entity in the memory content, and determine the entity identifier of the second entity that is most similar to the first entity from a predefined entity table; The first entity in the memory content is replaced with the entity identifier of the second entity, and the short-term memory is generated based on the memory content after the entity identifier is replaced.
7. The emotional-companion interaction method according to any one of claims 1-6, characterized in that, The step of retrieving the target memory related to the current dialogue information from the memory storage unit includes: Generate a query vector corresponding to the current dialogue information; Obtain a set of memory vectors consisting of memory vectors corresponding to all memories in the memory storage unit, as well as the importance and timeliness values corresponding to all memories; Calculate the similarity between the memory vector in the memory vector set and the query vector, and use it as the similarity of the memory corresponding to the memory vector, so as to obtain the similarity of each memory. Based on the similarity, importance, and timeliness values corresponding to all the memories, the target memory related to the current dialogue information is retrieved from the memory storage unit. 8.The emotional-companion interaction method of claim 1, wherein, Also includes: When the conditions for triggering an active dialogue are met, one or more of the most recently generated memories are retrieved from the memory storage unit; Based on one or more memories, an active dialogue is initiated with the target user, wherein the virtual character has the same dialogue identifier for both active and passive dialogues with the same user. 9.The emotional-companion interaction method of claim 1, wherein, Also includes: If the large dialogue model generates the same dialogue response information for different user dialogue information under the same dialogue identifier, then the target parameters of the large dialogue model are adjusted. And / or, If the same user dialogue information is obtained in multiple consecutive rounds of dialogue under the same dialogue identifier, then the dialogue response information output by the large dialogue model in the first round of dialogue in the multiple consecutive rounds of dialogue will be used as the dialogue response information in subsequent rounds. 10.The emotional-companion interaction method of claim 1, wherein, Also includes: When the conditions for generating a blocking message are met, a reply template is selected from a preset reply template library to maintain the semantic coherence of the dialogue. The dialogue reply information is generated based on the selected reply template. The conditions for generating the blocking message include that the current dialogue information contains preset sensitive words.
11. A computer program product, characterised in that, It includes computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the emotional companionship interaction method as described in any one of claims 1 to 10.
12. An electronic device, comprising: It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to enable the electronic device to implement the emotional companionship interaction method as described in any one of claims 1 to 10.
13. A computer storage medium, characterized in that The storage medium carries one or more computer programs, which, when executed by an electronic device, enable the electronic device to implement the emotional companionship interaction method as described in any one of claims 1 to 10.