Session interaction method and device, electronic equipment and storage medium
By using natural language models to generate and evaluate candidate replies in group chat scenarios, the problem of limited interaction depth and efficiency of virtual characters is solved, and the rapid development of virtual plots and the improvement of user experience is achieved.
Patent Information
- Application Number
- CN202510327270.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-19
AI Technical Summary
The existing technology has limited depth and efficiency of virtual characters' interaction in group chat scenarios, resulting in slow development of virtual plots.
By detecting the interactive content to be replied in the target group chat, the natural language model is called to generate candidate replies, and based on evaluation indicators such as virtual character matching, scene matching, and plot matching, select the most matching reply to send.
It improves the interaction depth and efficiency of virtual characters in group chat scenarios, and promotes the development of virtual plots and immersiveness of user experience.
Smart Images

Figure CN120162418A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language models, and in particular, to a conversation interaction method, device, electronic device, and storage medium. Background Art
[0002] With the development of artificial intelligence technology, content interaction based on natural language models has been increasingly widely used in fields such as emotional companionship, psychological counseling, and medical Q&A. Among them, natural language models play an important role in the training of virtual characters.
[0003] For content interaction with virtual characters, the existing technology usually focuses on content interaction with a single virtual character. Even in the scenario of establishing a group chat with multiple virtual characters, the existing technology usually only allows the virtual characters specified by the user to reply to the conversation, with limited depth and low efficiency of content interaction, resulting in slow development of virtual plots. Summary of the Invention
[0004] The present invention provides a conversation interaction method, device, electronic device, and storage medium to improve the interaction depth and efficiency of virtual characters in a group chat scenario.
[0005] In a first aspect, an embodiment of the present invention provides a conversation interaction method, which includes:
[0006] If a target group chat detects content to be replied to and interacted with, call a natural language model, where the natural language model generates at least one candidate reply based on at least the content to be replied to and interacted with. The target group chat includes at least two virtual characters, and the candidate reply matches at least one virtual character in the group chat;
[0007] The natural language model determines the evaluation results of each candidate reply based on at least one reply evaluation index;
[0008] Wherein, the reply evaluation index includes at least one of virtual character matching degree, scene matching degree, plot matching degree, virtual character reply rate, and virtual character popularity;
[0009] Based on the evaluation results of each candidate reply, determine the candidate reply with the highest matching degree with the content to be replied to and interacted with as the target reply among each candidate reply, and send the target reply to the target group chat.
[0010] In a second aspect, an embodiment of the present invention further provides a conversation interaction device, which includes:
[0011] A candidate response generation module, configured to, if the target group chat detects the interaction content to be replied, call a natural language model, where the natural language model generates at least one candidate response based on at least the interaction content to be replied, at least two virtual characters are included in the target group chat, and the candidate response matches at least one virtual character in the group chat;
[0012] A candidate response evaluation module, configured to determine the evaluation result of each candidate response based on at least one response evaluation index by the natural language model;
[0013] Wherein, the response evaluation index includes at least one of virtual character matching degree, scene matching degree, plot matching degree, virtual character response rate, and virtual character popularity;
[0014] A target response determination module, configured to determine, according to the evaluation results of each candidate response, the candidate response with the highest matching degree with the interaction content to be replied as the target response in each candidate response, and send the target response to the target group chat.
[0015] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the session interaction method as described in any one of the embodiments of the present invention.
[0016] In a fourth aspect, an embodiment of the present invention further provides a storage medium storing computer-executable instructions, where the computer-executable instructions are used to execute the session interaction method as described in any one of the embodiments of the present invention when executed by a computer processor.
[0017] The technical solution of the embodiment of the present invention, when detecting the interaction content to be replied in the target group chat, calls a natural language model to generate at least one candidate response based on at least the interaction content to be replied, evaluates each candidate response by the natural language model in at least one response evaluation index dimension to obtain the evaluation results of each candidate response, wherein the response evaluation index includes at least one of virtual character matching degree, scene matching degree, plot matching degree, virtual character response rate, and virtual character popularity, and selects the target response from each candidate response according to the evaluation results of each candidate response and sends it to the target group chat. The technical solution of the embodiment of the present invention solves the problems in the prior art that focus on content interaction with a single virtual character, and the depth of content interaction is limited and the efficiency is low, resulting in the slow development of virtual plots. The technical solution of the embodiment of the present invention can improve the interaction depth and efficiency of virtual characters in the group chat scenario.
[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become readily understood from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 is a flowchart of a session interaction method provided in Embodiment 1 of the present invention;
[0021] Figure 2 is a flowchart of a session interaction method provided in Embodiment 2 of the present invention;
[0022] Figure 3 is a schematic structural diagram of a session interaction device provided in Embodiment 3 of the present invention;
[0023] Figure 4 is a schematic structural diagram of an electronic device provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0025] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In the embodiments of the present application, certain industry-existing solutions such as certain software, components, models, etc. may be mentioned, and they should be considered exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution.
[0026] In the technical solution of the present application, the acquisition, transmission, storage, use, processing, etc. of data all comply with the relevant regulations of national laws and regulations.
[0027] Embodiment 1
[0028] Figure 1 The flowchart of a session interaction method is provided for Embodiment 1 of the present invention. This embodiment is applicable to the situation where multiple virtual characters conduct session interaction. This method can be executed by a session interaction device, which can be implemented in the form of hardware and / or software. The session interaction device can be configured in a server or an electronic device to implement session interaction based on a natural language model.
[0029] As Figure 1 shown, the method includes:
[0030] S110. If the target group chat detects the interaction content to be replied, call the natural language model, and the natural language model generates at least one candidate reply based on at least the interaction content to be replied.
[0031] Among them, the target group chat includes at least two virtual characters, and the candidate reply matches at least one virtual character in the group chat.
[0032] The target group chat can be established in response to a group chat establishment instruction of the user interface, or can be automatically established according to the development of the plot. In addition to the user, the target group chat also needs to include at least two virtual characters. The virtual characters can be created in response to a role creation instruction of the user interface, or can be automatically created according to the development of the plot; at the same time, the virtual characters can be fictional characters created by the user or developer, or can be known characters in published works or plots, etc.
[0033] The interactive content to be replied is sent by the user through the user interface or by clicking a button. A natural language model is a model based on artificial intelligence technology that can understand, generate, and process natural languages (such as languages used by humans like English, Chinese, etc.).
[0034] Furthermore, the interactive content to be replied may include conversation content and / or other content (such as pictures, videos, or emoticons, etc.). The number of conversation content and / or other content can be one or at least two.
[0035] Specifically, if the time interval between consecutive conversation content and / or other content is less than or equal to a preset time interval (such as 10s), and the correlation degree between consecutive conversation content and / or other content is greater than or equal to a preset correlation degree threshold (such as 90%), then the consecutive conversation content and / or other content can be jointly regarded as a group of interactive content to be replied, and subsequently the natural language model generates candidate replies for this group of interactive content to be replied. Among them, the correlation degree can be judged by the natural language model. Exemplarily, taking the conversation content as "The weather is really nice today!" and the picture as blue sky and white clouds, and the time interval between the conversation content and the picture is 1s as an example, at this time the correlation degree between the conversation content and the picture is relatively high, and it can be regarded as a group of interactive content to be replied.
[0036] If the time interval between consecutive conversation content and / or other content is greater than the preset time interval, and / or the correlation degree between consecutive conversation content and / or other content is less than the preset correlation degree threshold, at this time, the different conversation content and / or other content are respectively regarded as interactive content to be replied, and then the natural language model is used to generate candidate replies respectively.
[0037] Candidate replies are replies generated by the natural language model for at least one virtual character, simulating the virtual character's reply to the interactive content to be replied. Similarly, candidate replies may include conversation content and / or other content (such as pictures, videos, or emoticons, etc.). The number of conversation content and / or other content can be one or at least two. At the same time, for the same virtual character, when the natural language model generates candidate replies for it based on the interactive content to be replied, the number of candidate replies can be one or at least two.
[0038] In this embodiment, when new interactive content to be replied is detected in the target group chat, candidate replies are generated for other virtual characters except the user in the target group chat through the natural language model. In an optional embodiment, if one or at least two virtual characters are specified in the interactive content to be replied, the natural language model generates corresponding candidate replies based on the interactive content to be replied for the specified virtual characters.
[0039] In another alternative embodiment, if no virtual character is specified in the interaction content to be replied, the natural language model may generate corresponding candidate replies for all other virtual characters in the target group chat except the user, at least based on the interaction content to be replied.
[0040] In yet another alternative embodiment, a preset number of virtual characters may also be selected from among the virtual characters according to the relevance between each virtual character and the interaction content to be replied, the user's preference for each virtual character, etc. The natural language model may generate corresponding candidate replies for the selected virtual characters at least based on the interaction content to be replied. Among them, the relevance between the virtual character and the interaction content to be replied may be judged by the natural language model. The user's preference for each virtual character may be determined according to the user's historical evaluations of each virtual character and historical interaction situations, etc.
[0041] The technical solution of this embodiment supports the session interaction between multiple virtual characters and the user in the group chat scenario. After the user sends the interaction content to be replied in the target group chat, the natural language model may generate candidate replies for one or at least two virtual characters at least based on the interaction content to be replied. Such a setting can achieve complex interactions between the user and multiple virtual characters.
[0042] Furthermore, the natural language model generating at least one candidate reply at least based on the interaction content to be replied may include: the natural language model generating at least one candidate reply at least based on the interaction content to be replied and the role data; where the role data includes at least one of the following: historical interaction content, virtual character settings, and virtual character relationship networks.
[0043] Among them, the historical interaction content may include the historical interaction content when the user has a separate session interaction with the virtual character, the historical interaction content when the user has a session interaction with the virtual character in the target group chat, and the historical interaction content when the user has a session interaction with the virtual character in other group chats. The virtual character settings may include, but are not limited to, the character, family background, personal background, appearance characteristics, etc. of the virtual character, which are used to describe the virtual character. The virtual character relationship network represents the relationship between the virtual character and the user and / or other virtual characters.
[0044] Among them, the virtual character settings and the virtual character relationship network may be generated in response to the input content or click of a button on the user interface when the virtual character is created, and / or when the virtual character is a known character in a published work or plot, the virtual character settings and / or the virtual character relationship network of the virtual character may be directly determined according to the role data of the known character.
[0045] Further, when the virtual character is created in response to a character creation instruction in the user interface, the natural language model can expand the virtual character settings and / or the virtual character relationship network of the virtual character based on data such as the virtual character name and the determined character settings. Specifically, the natural language model can retrieve known characters with relatively high similarity in the published works or plots based on data such as the virtual character name and the determined character settings; provide the character data of the retrieved known characters to the user interface for display; and in response to the selection operation of the character data of the known character on the user interface, supplement the character data of the known character to the character data of the virtual character.
[0046] Further, when the virtual character is a known character in a published work or plot, or has a relatively high similarity to a known character, in this embodiment, it is simultaneously supported to determine the virtual character settings and / or the virtual character relationship network of the virtual character based on the character data of the known character, and to support the user to generate the virtual character settings and / or the virtual character relationship network of the virtual character by inputting content or clicking a button on the user interface. Therefore, in the above situation, in this embodiment, the natural language model can be used to determine whether the virtual character settings and / or the virtual character relationship network generated based on the input content or click of the button on the user interface match the virtual character settings and / or the virtual character relationship network of the known character. If so, both the virtual character settings and / or the virtual character relationship network from the above two different sources can be used as the character data of the virtual character; otherwise, a prompt for non-conformity of the character settings and / or the relationship network can be given through the user interface.
[0047] In the technical solution of this embodiment, the natural language model generates candidate responses for the virtual character based on the interaction content to be replied and the character data of the virtual character. The advantage of this setting is that the generated candidate responses can fit the character settings of the virtual character, and the content of the candidate responses can be adjusted in real time according to the interaction content to be replied to adapt to the development of the plot. According to the interaction content to be replied, the historical interaction content, and the virtual character relationship network, the logical coherence and emotional consistency of the plot can be ensured.
[0048] S120. The natural language model determines the evaluation results of the candidate responses based on at least one response evaluation index.
[0049] In this embodiment, a response evaluation module can be deployed in the natural language model to evaluate each candidate response based on this module, or a response evaluation model can be trained based on the natural language model, and each candidate response can be evaluated by this response evaluation model. This embodiment does not limit this.
[0050] Among them, the reply evaluation index is used to evaluate the fitness between each candidate reply and the interaction content to be replied. The reply evaluation index includes at least one of the virtual character matching degree, the scene matching degree, the plot matching degree, the virtual character reply rate, and the virtual character popularity.
[0051] Among them, the virtual character matching degree represents the matching degree between the candidate reply and the virtual character's character setting. The scene matching degree represents the matching degree between the candidate reply and the current conversation scene. Among them, if the target group chat is established in response to the group chat establishment instruction of the user interface, the current conversation scene can be determined in response to the input content or the scene selection button of the user interface; if the target group chat is automatically created according to the plot development, the current conversation scene can be determined according to the known content of the published work. The plot matching degree represents the matching degree between the candidate reply and the plot development of the virtual plot. The higher the plot matching degree, the closer the candidate reply is to the user's preference, and / or the character setting, and / or can promote the development of the story. The virtual character reply rate represents the probability that the historical conversation of each virtual character is replied by the user during the historical interaction process of multiple virtual characters involved in the user and the group chat. Among them, the historical interaction process of multiple virtual characters can refer to the historical conversation interaction situation in other group chats except the target group chat, and / or the historical conversation interaction situation of the target group chat before the current moment. The virtual character popularity refers to the degree of being liked or the popularity of the virtual character by the user. The virtual character popularity can be determined according to the historical evaluation of the virtual character, the historical interaction situation (for example, it can be represented by the amount of data of the interaction with each virtual character in the user's historical conversation), and the historical recharge situation matching the virtual character, etc.
[0052] The evaluation result can be represented by a score (such as 1 - 100 points), or by text (such as unqualified, good, excellent, etc.), or by the number of stars or other graphics. Taking the representation by score as an example, the higher the score, the more the candidate reply can fit the character setting, the story development, and the plot logic, meet the user's preference, and is more suitable as a reply to the interaction content to be replied.
[0053] Furthermore, taking the evaluation result represented by score as an example, when the natural language model determines the evaluation result of each candidate reply based on the reply evaluation index, different weights can be set for different reply evaluation indexes, and the weights can be the same or different. For a candidate reply, determine its index value under each reply evaluation index respectively, and then perform weighted summation on the index values under each reply evaluation index, and the obtained value is used as the evaluation result of the candidate reply.
[0054] In the technical solution of this embodiment, the natural language model determines the evaluation results of each candidate response based on multi-dimensional response evaluation metrics, which can achieve a comprehensive, accurate, and objective quantitative evaluation of each candidate response, facilitating the subsequent selection of the optimal response to the interactive content to be replied to from each candidate response, thereby improving the intelligence of the virtual character's response, promoting the authenticity and coherence of the virtual plot deduction, and providing a more immersive group chat experience.
[0055] It should be noted that if one or at least two virtual characters are specified in the interactive content to be replied to, the candidate responses generated by the natural language model for the specified virtual characters are not evaluated for response in this embodiment, and the candidate responses corresponding to the specified virtual characters can be directly sent to the target group chat.
[0056] S130. According to the evaluation results of each candidate response, determine the candidate response with the highest matching degree with the interactive content to be replied to among each candidate response as the target response, and send the target response to the target group chat.
[0057] In this embodiment, taking the evaluation result being represented in the form of a score as an example, the highest matching degree can be the highest score; taking the evaluation result being represented in the form of text as an example, the highest matching degree can be that the evaluation result is "excellent"; taking the evaluation result being represented in the form of stars as an example, the highest matching degree can be represented as "★★★★★". Similarly, it can also be represented by the number of other symbols, for example
[0058] It should be noted that if the evaluation results of at least two candidate responses are all the highest in matching degree with the interactive content to be replied to, one candidate response can be randomly selected from the at least two candidate responses as the target response; or the candidate response corresponding to the virtual character with the highest virtual character response rate and / or virtual character popularity can be preferentially selected as the target response; or at least two candidate responses can be used as the target responses. This embodiment does not limit this.
[0059] Taking the evaluation result being represented in the form of a score as an example, in an optional embodiment, according to the evaluation results of each candidate response, determining the candidate response with the highest matching degree with the interactive content to be replied to among each candidate response as the target response can be that the natural language model determines the evaluation results of each candidate response based on each response evaluation metric, and directly takes the candidate response with the highest score of the evaluation result as the target response.
[0060] In another alternative embodiment, according to the evaluation results of each candidate response, the candidate response with the highest matching degree with the interaction content to be replied is determined as the target response among the candidate responses. The natural language model can also determine the evaluation results of each candidate response based on at least one of the virtual character matching degree, the scenario matching degree, and the plot matching degree, and according to the evaluation results of each candidate response corresponding to the same virtual character, select the candidate response with the highest score as the candidate response with the highest matching degree with the virtual character; according to the virtual character response rate and / or the virtual character popularity, among the candidate responses with the highest matching degree with each virtual character, select the candidate response with the highest score as the candidate response with the highest matching degree with the interaction content to be replied, and use it as the target response.
[0061] In the technical solution of this embodiment, the natural language model generates candidate responses for each virtual character based on the interaction content to be replied, and based on multi-dimensional response evaluation indicators, realizes a comprehensive, objective and accurate evaluation of the candidate responses, and finally selects the candidate response with the highest matching degree with the interaction content to be replied as the target response and sends it to the group chat. It ensures the consistency between the virtual character's response and the virtual character's character setting and relationship network, improves the intelligence of the virtual character's response, improves the logical rationality, plot deduction coherence and emotional consistency of the virtual plot, and provides an immersive group chat experience in the scenario of multi-virtual character conversation interaction.
[0062] Further, after S130, it further includes: if the target response self-learning condition is satisfied, the natural language model performs self-learning based on the interaction content in the target group chat; where the satisfaction of the target response self-learning condition includes at least one of the following: detecting a target response replacement instruction, detecting an input response matching the interaction content to be replied, and detecting a response evaluation result matching the target response.
[0063] This embodiment also provides a solution for the natural language model to perform self-learning based on the feedback of the target response.
[0064] Among them, detecting a target response replacement instruction means detecting a target response replacement instruction in the user interaction interface. The target response replacement instruction can be used to instruct the large model to regenerate the target response, or to instruct to switch the target response to another candidate response. Specifically, after the above embodiment sends the target response to the target group chat, if the user interaction interface detects a target response replacement instruction, a preset number (such as 2 or 3) of candidate responses with a relatively high matching degree with the interaction content to be replied can be provided to the user interaction interface for display, and in response to the user interaction interface's selection operation on the candidate response, the selected candidate response is sent to the target group chat as the new target response.
[0065] Detecting an input reply that matches the reply interaction content to be replied means that the user customizes an input through the user interaction interface that matches the reply interaction content to be replied.
[0066] Detecting a reply evaluation result that matches the target reply means that a reply evaluation result for the target reply is detected in the user interaction interface. Among them, the reply evaluation result can be expressed in the form of expressions (such as a smiling expression and an angry expression, a like expression and a dislike expression, √ or ×, etc.), letters (such as Y and N, etc.), and text to indicate the user's like or dislike evaluation of the target reply. For the reply evaluation of the target reply of the virtual character, it can be that every time the target reply is sent to the group chat, a reply evaluation button is attached and sent together. The user can click the reply evaluation button on the user interaction interface to achieve the reply evaluation of the target reply. It can also be at preset time intervals, or when the target group chat interface on the user interaction interface is closed, through the user interaction interface, provide a reply evaluation interface or button for the reply of the virtual character, etc. According to the feedback of the reply evaluation interface or reply evaluation button, etc. in the user interaction interface, the reply evaluation of the reply of the virtual character is achieved. It should be noted that at this time, the reply evaluation of the virtual character can take all the replies corresponding to each virtual character in this evaluation period as the evaluation object, or take a specific reply as the evaluation object. This embodiment does not limit this.
[0067] In this embodiment, when a target reply replacement instruction is detected, or an input reply that is more matched with the reply interaction content to be replied is detected, or a reply evaluation result that matches the target reply is detected, it indicates that the target reply does not meet the user's expectations. At this time, if there is a new target reply or input reply, the natural language model can be self-learned based on the new target reply or input reply; if there is no new target reply or input reply, the natural language model can also be self-learned based on the reply evaluation result of the target reply. The advantage of such a setting is that it can realize the iterative optimization of the natural language model based on the feedback of the target reply, continuously learn the user's preferences, and improve the generation effect of the reply of the virtual character.
[0068] Furthermore, this embodiment further includes: if the virtual character interaction condition is met, the natural language model generates a chat session that matches at least two virtual characters based on at least one of the reply interaction content to be replied, the historical interaction content, and the scene type, and sends the chat session to the target group chat; where the meeting of the virtual character interaction condition includes: the time difference between the current moment and the historical moment when the reply interaction content to be replied was last detected in the target group chat is greater than or equal to the first time interval, or a virtual character interaction trigger instruction is detected.
[0069] This embodiment also provides a solution for session interaction between virtual characters.
[0070] Specifically, if the difference between the current moment and the historical moment when the to-be-replied interactive content was last detected in the target group chat is greater than or equal to the first time interval, that is, after the user last sent the to-be-replied interactive content in the group chat, a long time has passed without sending new to-be-replied interactive content. At this time, the virtual characters can continue the conversation interaction. Or, when the user interaction interface detects an interaction instruction between virtual characters, the virtual characters can carry out the conversation interaction.
[0071] Specifically, when the virtual characters carry out the conversation interaction, they can continue the conversation interaction based on at least one of the to-be-replied interactive content, historical interactive content, and scene type. Among them, the historical interactive content includes the historical conversations in the target group chat that match at least one virtual character, that is, the historical conversations sent by the users in the target group chat and the historical conversations in which at least one virtual character replies to the historical conversations sent by the users. The scene type can include: the published plot type, the fictional plot type based on the characters in the published work, the newly created virtual character plot type, and the regular interaction type, etc. The published plot type means that the current scene is consistent with the known plot in the published work. The fictional plot type based on the characters in the published work means that while the virtual characters use the known characters in the published work, the current scene is a fictional plot. The newly created virtual character plot type means that the virtual characters are virtual characters created by the user, and the current scene is a fictional plot. The conversation of the regular interaction type means that the current scene is a daily conversation, current affairs discussion, etc., and the conversation that does not play a promoting role or plays a minor promoting role in the plot development.
[0072] Generate a chat session that matches at least two virtual characters based on at least one of the to-be-replied interactive content, historical interactive content, and scene type. Specifically, if the scene type is the published plot type, then based on the to-be-replied interactive content and historical interactive content, the known plot in the published work can be followed to continue the conversation interaction between the virtual characters. If the scene type is the fictional plot type based on the characters in the published work or the newly created virtual character plot type, then based on the to-be-replied interactive content, historical interactive content, and the character data of the virtual characters, a chat session that matches at least two virtual characters can be generated to continue the conversation interaction between the virtual characters. If the scene type is the regular interaction type, then based on the to-be-replied interactive content, historical interactive content, and the character data of the virtual characters, regular interaction content that matches at least two virtual characters can be generated to realize the conversation interaction between the virtual characters.
[0073] In the technical solution of this embodiment, when the chat stagnates, that is, when no new interaction content to be replied has been detected in the target group chat for a long time, or when a virtual character interaction command is detected, session interaction between virtual characters is implemented based on at least one of the interaction content to be replied, historical interaction content, and scene type. The advantage of this setting is that complex interactions between multiple virtual characters are realized, and at the same time, the session interaction between virtual characters can conform to the character settings of the virtual characters and the logical coherence of the plot development, enhancing the interaction depth of the target group chat and providing a more realistic and natural immersive group chat experience.
[0074] Furthermore, this embodiment further includes: if the virtual character relationship network update condition is met, the relationship network of the virtual character is updated according to at least one piece of historical interaction content in the target group chat; where the satisfaction of the virtual character relationship network update condition includes at least one of the following: the difference between the current time and the historical time when the virtual character relationship network was last updated is greater than or equal to the second time interval, after the virtual character relationship network was last updated, the new interaction content in the target group chat is greater than or equal to the first threshold, and a relationship network change trigger word is detected.
[0075] This embodiment also provides a technical solution for updating the virtual character relationship network.
[0076] Specifically, the virtual character relationship network can be updated every second time interval (for example, the virtual character relationship network is updated every 5 minutes), so when the difference between the current time and the historical time when the virtual character relationship network was last updated is greater than or equal to the second time interval, the virtual character relationship network is updated. It is also possible to monitor the new interaction content in the target group chat in real time. If the new interaction content in the current target group chat is greater than or equal to the preset first threshold compared to the last time the virtual character relationship network was updated, then the virtual character relationship network is updated once. At this time, the judgment can be based on either the number of new interaction contents or the data volume of the new interaction contents. For example, the virtual character relationship network can be updated every time ten new interaction contents are added. This embodiment does not limit this. It is also possible to update the virtual character relationship network once when a relationship network change trigger word is detected in the user interaction interface, where the relationship network change trigger word can include confession, proposal, quarrel, etc.
[0077] The relationship network of the virtual character is updated according to at least one piece of historical interaction content in the target group chat. Specifically, the keywords in the historical interaction content can be extracted, where the keywords are words referring to the relationship between people, such as brother, teacher, etc. The relationship network is updated according to the extracted keywords and the virtual characters matching the keywords.
[0078] Exemplarily, if the target group chat includes the historical conversation sent by the user "I went to the park with my brother @A's role name", the keyword "brother" is detected, and the virtual character A is mentioned in the historical conversation, then the relationship between the user and the virtual character A can be established.
[0079] Furthermore, when updating the virtual character relationship network, the virtual character relationship network can also be updated based on the historical interaction content during the separate session interaction with at least one virtual character and the historical interaction content within the target group chat.
[0080] Exemplarily, if in the separate session interaction with the virtual character A, there is a historical conversation sent by the virtual character A "Yesterday I went to the park with my brother", and according to the historical interaction content within the target group chat, the role of the virtual character B is the brother of the virtual character A, then it can be considered that the virtual character A and the virtual character B went to the park together yesterday, and the virtual character relationship network can be updated.
[0081] Furthermore, if the relationship to be updated is a known relationship in the relationship network, the relationship network is not updated. If the relationship to be updated is not an existing relationship, then it is further determined whether the relationship to be updated matches the existing relationship network. If it matches, the relationship network is updated; otherwise, a prompt for relationship confirmation is given through the user interaction interface. For example, if the relationship network already knows that the relationship between the virtual characters A and C is: A is the father and C is the son, and if the relationship to be updated is that A and C are brothers, then the relationship to be updated does not match the existing relationship network, and a relationship confirmation prompt is required.
[0082] The technical solution of this embodiment learns the virtual character relationships based on the natural language model and updates the virtual character relationship network. The advantage of this setting is that it can enrich the relationship network between virtual characters and users and between virtual characters in a timely manner, thereby providing rich data support for subsequent session interactions and improving the intelligence and accuracy of subsequent session interactions.
[0083] Furthermore, this embodiment further includes: if the data update condition is met, the historical interaction content in the target group chat is respectively stored in the role data of each virtual character, and the memory of the target group chat is released.
[0084] The advantage of this setting is that it can reduce memory occupancy, update the role data of virtual characters in a timely manner, improve the richness of role data, and thus make the generation of subsequent virtual character responses more intelligent. At the same time, due to the timely preservation of historical interaction content, it can make the subsequent virtual character session interactions have the coherence of story deduction and plot development, and improve the authenticity and naturalness of virtual character interactions.
[0085] In the technical solution of the embodiment of the present invention, when detecting the to-be-replied interaction content in the target group chat, a natural language model is called to generate at least one candidate reply based on at least the to-be-replied interaction content. The natural language model evaluates each candidate reply in at least one reply evaluation index dimension to obtain the evaluation results of each candidate reply. Among them, the reply evaluation index includes at least one of virtual character matching degree, scenario matching degree, plot matching degree, virtual character reply rate, and virtual character popularity. According to the evaluation results of each candidate reply, a target reply is selected from each candidate reply and sent to the target group chat. The technical solution of the embodiment of the present invention solves the problems in the prior art that focus on content interaction with a single virtual character, as well as the problems of limited depth and low efficiency of content interaction, resulting in slow development of virtual plots. The technical solution of the embodiment of the present invention can improve the interaction depth and efficiency of virtual characters in the group chat scenario.
[0086] Embodiment 2
[0087] Figure 2 FIG. is a flowchart of a session interaction method provided by Embodiment 2 of the present invention. On the basis of the above embodiment, the present invention further specifies the specific process of the natural language model generating at least one candidate reply based on at least the to-be-replied interaction content.
[0088] As Figure 2 shown, the method includes:
[0089] S210. If the target group chat detects the to-be-replied interaction content, call the natural language model.
[0090] S220. The natural language model generates at least one candidate reply based on at least the to-be-replied interaction content.
[0091] This embodiment provides different implementation manners when the natural language model generates candidate replies for virtual characters. The above embodiment has specifically described the manner in which the natural language model generates candidate replies based on the to-be-replied interaction content and character data, and this embodiment will not be elaborated herein.
[0092] Further, S220 may further include:
[0093] The natural language model generates at least one candidate reply based on at least the to-be-replied interaction content and historical interaction content; the historical interaction content includes historical conversations in the target group chat that match at least one virtual character.
[0094] Among them, the historical conversations in the target group chat that match at least one virtual character may further include historical conversations sent by the user in the target group chat, and historical conversations in which at least one virtual character replies to the historical conversations.
[0095] In this embodiment, a candidate response is generated for the virtual character based on the interaction content to be replied and the historical interaction content in the target group chat through the natural language model. The advantage of this setting is that the generated candidate response can be more in line with the plot and scene of the target group chat, promoting the coherence of the conversation interaction in the target group chat.
[0096] Furthermore, S220 can further include:
[0097] A1. Determine the scene type corresponding to the target group chat, and determine the data retrieval range matching the scene type;
[0098] Among them, the scene type includes at least one of the following: published plot type, fictional plot type based on the characters in the published works, newly created virtual character plot type, and regular interaction type;
[0099] A2. The natural language model generates at least one candidate response based on at least the interaction content to be replied and the data retrieval range.
[0100] Among them, when the natural language model generates a candidate response for the virtual character, in addition to based on the interaction content to be replied and the knowledge and patterns that the natural language model has learned, it can also perform data retrieval in the knowledge graph or knowledge base to assist in generating the candidate response.
[0101] The data retrieval range is used to represent the range when the natural language model performs data retrieval. Different scene types have different data retrieval ranges. In this embodiment, determining the scene type corresponding to the target group chat and determining the data retrieval range matching the scene type can be achieved through the scene classification module deployed in the natural language model, or through the scene classification model obtained by pre-training the natural language model.
[0102] Specifically, when the scene type is the published plot type, the data retrieval range can be restricted to the known plots of the published works and their related contexts. At this time, the natural language model generates a candidate response for the virtual character based on the interaction content to be replied and the known plots of the published works to continue to deduce the known plot.
[0103] When the scene type is the fictional plot type based on the characters in the published works, the data retrieval range can include all the known plots related to the character in the published works. At this time, the natural language model performs reasoning based on the interaction content to be replied, the known plots of the published works, and the character data of the virtual character, etc., to generate a candidate response for the virtual character, so that the generated candidate response can fit the character setting of the virtual character and conform to the development logic of the fictional plot of the characters in the published works.
[0104] When the scenario type is the new virtual character plot type, the virtual character is a fictional character that does not exist in the published works or plots. The data retrieval scope can include similar plots to the currently developed plot. At this time, the developed plot in the current target group chat can be determined, and similar plots with a relatively high similarity to the developed plot can be determined. The natural language model reasons based on the content of the interaction to be replied, similar plots, and the character data of the virtual character, etc., to generate candidate replies for the virtual character, so that the generated candidate replies can fit the character setting of the virtual character and conform to the development logic of the fictional plot.
[0105] When the scenario type is the regular interaction type, the data retrieval scope can be further determined according to the content of the interaction to be replied. For example, if the content of the interaction to be replied contains current affairs, weather, etc., the data retrieval scope can include current affairs news, recent weather data of the currently located city or a specified city, etc.; otherwise, if the content of the interaction to be replied is a daily conversation, the data retrieval scope can be unrestricted. The natural language model reasons based on the content of the interaction to be replied and the above data retrieval scope to infer candidate replies that conform to the character setting of the virtual character and can provide targeted intelligent feedback on the content of the interaction to be replied.
[0106] The technical solution of this embodiment can determine different data retrieval scopes according to different current scenario types, and reason about candidate replies based on different data retrieval scopes and the content of the interaction to be replied. The advantage of this setting is that it can reduce the amount of data retrieval, improve the generation efficiency and accuracy of candidate replies. At the same time, by assisting the generation of candidate replies through data retrieval, the intelligence of candidate replies can be improved, thereby improving the fitting degree of the virtual character's character setting and promoting the logical coherence of the plot development.
[0107] Furthermore, S220 can further include:
[0108] B1. Determine the sentiment analysis result that matches the content of the interaction to be replied;
[0109] B2. The natural language model generates at least one candidate reply based on at least the content of the interaction to be replied and the sentiment analysis result.
[0110] In this embodiment, determining the sentiment analysis result corresponding to the target group chat can be achieved through the sentiment analysis module deployed in the natural language model, or by pre-training the natural language model with text data labeled with sentiment data to implement the relevant functions of sentiment analysis.
[0111] The sentiment analysis result represents the user's emotion when sending the interactive content to be replied, and is used to assist in determining the emotion of the candidate reply. In this embodiment, if the interactive content to be replied includes an emoji, the text characters corresponding to the emoji are extracted. At the same time, punctuation marks, keywords, etc. in the interactive content to be replied are extracted, and the sentiment analysis result is determined according to at least one of the text characters corresponding to the emoji, punctuation marks, and keywords. The sentiment analysis result can be represented in the form of a sentiment tendency, such as a positive sentiment tendency or a negative sentiment tendency, etc., or can be represented by an exact sentiment judgment result (such as happy, sad, etc.).
[0112] In this embodiment, the natural language model generates at least one candidate reply based on at least the interactive content to be replied and the sentiment analysis result. Specifically, the natural language model infers the candidate reply based on the interactive content to be replied, the sentiment analysis result, and the relationship between the user and the virtual character. Further, it can also be determined according to the user's historical conversation habit whether the candidate reply can include an emoji. If so, the emoji matching the candidate reply can be determined and sent together according to the sentiment analysis result, the relationship between the user and the virtual character, and the content of the candidate reply.
[0113] In the technical solution of this embodiment, the natural language model generates a candidate reply based on the interactive content to be replied and the sentiment analysis result of the user. The advantage of this setting is that the generated candidate reply can be more in line with the user's current emotion, more intelligent, and improve the conversation interaction experience.
[0114] Further, S220 can further include:
[0115] C1. The natural language model generates at least one initial candidate reply based on at least the interactive content to be replied;
[0116] C2. Determine the auxiliary reply factors according to the initial candidate reply and the role data of the virtual character;
[0117] Wherein, the auxiliary reply factors include at least one of the following: emoji, action, and description of mental activity;
[0118] C3. Integrate each initial candidate reply and its corresponding auxiliary reply factor to generate a target candidate reply matching the virtual character.
[0119] In this embodiment, to determine the auxiliary reply factors, it can be implemented through an auxiliary reply module deployed in the natural language model, or can be implemented through an auxiliary reply model obtained by pre-training the natural language model.
[0120] The initial candidate reply refers to the reply content generated by the natural language model for the virtual character based on the interaction content to be replied. Subsequently, the initial candidate reply needs to be integrated with the auxiliary reply factors to generate the final target candidate reply for the virtual character.
[0121] The auxiliary reply factors are used to represent the expressions, actions, and descriptions of the mental activities, etc. of the virtual character when replying to the interaction content to be replied. Each initial candidate reply and its corresponding auxiliary reply factors are integrated to generate a target candidate reply that matches the virtual character. Specifically, if the initial candidate reply is in text form, the auxiliary reply factors can be displayed in the target candidate reply in forms such as parentheses, reducing the font size, changing the font color, etc. For example, when the virtual character A is scared, the target candidate reply generated by the natural language model can be: (with red eyes, said carefully) I'm fine. If the initial candidate reply is in voice form, the volume, intonation, pauses, and speech rate, etc. of the voice can be adjusted based on the auxiliary reply factors to obtain the final target candidate reply. If the initial candidate reply is in video form, the expressions, actions, postures, etc. of the virtual character image in the video can be adjusted based on the auxiliary reply factors to obtain the final target candidate reply.
[0122] The technical solution of this embodiment first generates an initial candidate reply based on the interaction content to be replied, then determines the auxiliary reply factors based on the initial candidate reply and the role data of the virtual character, and finally integrates the initial candidate reply with the auxiliary reply factors to obtain the final target candidate reply. The advantage of such a setting is that it can make the reply of the virtual character more vivid and natural, more in line with the character setting of the virtual character and the current group chat scenario, and provide a more immersive group chat experience.
[0123] It should be noted that different candidate reply generation methods are provided in this embodiment: the natural language model generates candidate replies based on the interaction content to be replied and the historical interaction content; determines the data retrieval range through the current scene type, and generates candidate replies based on different data retrieval ranges and the interaction content to be replied; generates candidate replies based on the interaction content to be replied and the result of the emotional analysis of the user; first generates an initial candidate reply based on the interaction content to be replied, then determines the auxiliary reply factors based on the initial candidate reply and the role data of the virtual character, and finally integrates the initial candidate reply with the auxiliary reply factors to obtain the final target candidate reply. One of the above several candidate reply generation methods can be selected, or at least two of them can be combined with each other to jointly play a role in generating the final candidate reply of the virtual character.
[0124] S230. The natural language model determines the evaluation results of each candidate reply based on at least one reply evaluation index.
[0125] S240. According to the evaluation results of each candidate response, determine the candidate response with the highest matching degree with the to-be-replied interaction content among the candidate responses as the target response, and send the target response to the target group chat.
[0126] The technical solution of this embodiment generates candidate responses for the virtual character based on the to-be-replied interaction content and historical interaction content in the target group chat, making the generated candidate responses more in line with the plot and scene of the target group chat, and promoting the coherence of the conversation interaction within the target group chat. Generating candidate responses based on different data retrieval scopes and the to-be-replied interaction content can reduce the data retrieval volume, improve the generation efficiency of candidate responses, enhance the intelligence of candidate responses, thereby improving the fitting degree of the virtual character's persona setting, and promoting the logical coherence of the plot development. Generating candidate responses based on the to-be-replied interaction content and the result of the emotional analysis of the user enables the generated candidate responses to be more in line with the user's current mood, more intelligent, and improves the conversation interaction experience. First, generate initial candidate responses based on the to-be-replied interaction content, then determine the auxiliary response factors based on the initial candidate responses and the role data of the virtual character, and finally fuse the initial candidate responses with the auxiliary response factors to obtain the final target candidate responses, which can make the responses of the virtual character more vivid and natural, more in line with the virtual character's persona setting and the current group chat scene, and provide a more immersive group chat experience. Based on multi-dimensional response evaluation indicators, evaluate each candidate response, realizing a comprehensive, objective, and accurate evaluation of the candidate responses. Selecting the candidate response with the highest matching degree with the to-be-replied interaction content as the target response ensures the fitting degree between the virtual character's response and the virtual character's persona setting and relationship network, improves the intelligence of the virtual character's response, improves the logical rationality, plot deduction coherence, and emotional consistency of the virtual plot, and provides an immersive group chat experience in the scenario of multi-virtual character conversation interaction.
[0127] Embodiment III
[0128] Figure 3 is a schematic structural diagram of a conversation interaction device provided in Embodiment III of the present invention. As Figure 3 shown, the device includes:
[0129] A candidate response generation module 310, configured to, if a to-be-replied interaction content is detected in the target group chat, call a natural language model, and the natural language model generates at least one candidate response based on at least the to-be-replied interaction content. The target group chat includes at least two virtual characters, and the candidate response matches at least one virtual character in the group chat;
[0130] A candidate response evaluation module 320, configured to determine the evaluation results of each candidate response based on at least one response evaluation indicator by the natural language model;
[0131] Among them, the reply evaluation metrics include at least one of virtual character matching degree, scene matching degree, plot matching degree, virtual character reply rate, and virtual character popularity.
[0132] A target reply determination module 330, configured to determine, according to the evaluation results of each candidate reply, the candidate reply with the highest matching degree with the to-be-replied interaction content among each candidate reply as the target reply, and send the target reply to the target group chat.
[0133] The technical solution of the embodiment of the present invention, when detecting the to-be-replied interaction content in the target group chat, calls a natural language model to generate at least one candidate reply based on at least the to-be-replied interaction content, evaluates each candidate reply by the natural language model in at least one reply evaluation metric dimension to obtain the evaluation results of each candidate reply. Among them, the reply evaluation metrics include at least one of virtual character matching degree, scene matching degree, plot matching degree, virtual character reply rate, and virtual character popularity. According to the evaluation results of each candidate reply, select the target reply among each candidate reply and send it to the target group chat. The technical solution of the embodiment of the present invention solves the problems in the prior art that focus on content interaction with a single virtual character, as well as the problems of limited depth and low efficiency of content interaction, resulting in slow development of virtual plots. The technical solution of the embodiment of the present invention can improve the interaction depth and efficiency of virtual characters in the group chat scenario.
[0134] On the basis of the above embodiment, optionally, the candidate reply generation module 310 includes:
[0135] A first candidate reply generation unit, configured to generate at least one candidate reply by the natural language model based on at least the to-be-replied interaction content and historical interaction content;
[0136] The historical interaction content includes historical conversations in the group chat that match at least one virtual character.
[0137] On the basis of the above embodiment, optionally, the candidate reply generation module 310 includes:
[0138] A data retrieval range determination unit, configured to determine the scene type corresponding to the target group chat and determine the data retrieval range that matches the scene type;
[0139] Among them, the scene type includes at least one of the following: published plot type, fictional plot type based on characters in published works, newly created virtual character plot type, and regular interaction type;
[0140] A second candidate reply generation unit, configured to generate at least one candidate reply by the natural language model based on at least the to-be-replied interaction content and the data retrieval range.
[0141] Based on the above embodiments, optionally, the candidate response generation module 310 includes:
[0142] An emotional analysis result determination unit, configured to determine an emotional analysis result that matches the interaction content to be replied;
[0143] A third candidate response generation unit, configured to use the natural language model to generate at least one candidate response based on at least the interaction content to be replied and the emotional analysis result.
[0144] Based on the above embodiments, optionally, the candidate response generation module 310 includes:
[0145] An initial candidate response generation unit, configured to use the natural language model to generate at least one initial candidate response based on at least the interaction content to be replied;
[0146] An auxiliary response factor determination unit, configured to determine auxiliary response factors according to the initial candidate response and the role data of the virtual character;
[0147] Wherein, the auxiliary response factors include at least one of the following: expression, action, and description of mental activity;
[0148] A target candidate response generation unit, configured to fuse each of the initial candidate responses and their corresponding auxiliary response factors to generate a target candidate response that matches the virtual character.
[0149] Based on the above embodiments, optionally, the device further includes:
[0150] A natural language model self-learning module, configured to, if a target response self-learning condition is satisfied, the natural language model performs self-learning based on the interaction content in the target group chat;
[0151] Wherein, the satisfaction of the target response self-learning condition includes at least one of the following: detecting a target response replacement instruction, detecting an input response that matches the interaction content to be replied, and detecting a response evaluation result that matches the target response.
[0152] Based on the above embodiments, optionally, the device further includes:
[0153] A virtual character interaction module, configured to, if a virtual character interaction condition is satisfied, the natural language model generates a chat session that matches at least two virtual characters based on at least one of the interaction content to be replied, historical interaction content, and scene type, and sends the chat session to the target group chat;
[0154] Among them, the satisfaction of the interaction conditions between virtual characters includes: the difference between the current moment and the historical moment when the to-be-replied interaction content was last detected in the target group chat is greater than or equal to the first time interval, or a virtual character interaction trigger instruction is detected.
[0155] Based on the above embodiments, optionally, the device further includes:
[0156] A virtual character relationship network update module, configured to update the relationship network of virtual characters according to at least one piece of historical interaction content in the target group chat if the virtual character relationship network update conditions are satisfied;
[0157] Among them, the satisfaction of the virtual character relationship network update conditions includes at least one of the following: the difference between the current moment and the historical moment when the virtual character relationship network was last updated is greater than or equal to the second time interval, after the last virtual character relationship network update, the new interaction content in the target group chat is greater than or equal to the first threshold, and a relationship network change trigger word is detected.
[0158] Based on the above embodiments, optionally, the candidate reply generation module 310 includes:
[0159] A fourth candidate reply generation unit, configured to generate at least one candidate reply by the natural language model based on at least the to-be-replied interaction content and character data;
[0160] Among them, the character data includes at least one of the following: historical interaction content, virtual character settings, and virtual character relationship networks.
[0161] The session interaction device provided by the embodiments of the present invention can execute the session interaction method provided by any embodiment of the present invention, and has corresponding function modules and beneficial effects for executing the method.
[0162] Embodiment 4
[0163] Figure 4 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described herein and / or claimed.
[0164] Such as Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0165] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0166] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the session interaction method.
[0167] In some embodiments, the session interaction method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the session interaction method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the session interaction method in any other appropriate manner (e.g., by means of firmware).
[0168] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0169] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0170] In the context of the present invention, a computer-readable storage medium may be a tangible medium that can contain, or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0171] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0172] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0173] A computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0174] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0175] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A conversation interaction method, characterized in that: include: If the target group chat detects the interaction content to be replied, calling the natural language model, the natural language model generates at least one candidate reply based on at least the interaction content to be replied, the target group chat includes at least two virtual characters, and the candidate reply matches at least one virtual character in the group chat; The natural language model determines an evaluation result of each candidate response based on at least one response evaluation indicator; The response evaluation index includes at least one of the following: virtual character matching, scene matching, plot matching, virtual character response rate, and virtual character likeability; According to the evaluation results of each candidate reply, the candidate reply with the highest matching degree with the interactive content to be replied is determined as the target reply among the candidate replies, and the target reply is sent to the target group chat.
2. The method according to claim 1, characterized in that The natural language model generates at least one candidate reply based on at least the interaction content to be replied, including: The natural language model generates at least one candidate reply based at least on the interaction content to be replied and the historical interaction content; The historical interaction content includes historical conversations matching at least one virtual character in the target group chat.
3. The method according to claim 1, characterized in that The natural language model generates at least one candidate reply based on at least the interaction content to be replied, and further includes: Determine the scenario type corresponding to the target group chat, and determine the data retrieval scope matching the scenario type; The scenario type includes at least one of the following: a published plot type, a fictional plot type based on a published work character, a newly created virtual character plot type, and a conventional interaction type; The natural language model generates at least one candidate reply based at least on the interaction content to be replied and the data retrieval scope.
4. The method according to claim 1, characterized in that The natural language model generates at least one candidate reply based on at least the interaction content to be replied, and further includes: Determine a sentiment analysis result that matches the interaction content to be replied; The natural language model generates at least one candidate reply based at least on the interaction content to be replied and the sentiment analysis result.
5. The method according to claim 1, characterized in that The natural language model generates at least one candidate reply based on at least the interaction content to be replied, including: The natural language model generates at least one initial candidate reply based at least on the interaction content to be replied; Determine auxiliary response factors based on the initial candidate responses and the role data of the virtual character; The auxiliary reply factor includes at least one of the following: facial expression, action and description of psychological activity; Each of the initial candidate responses and its corresponding auxiliary response factors are fused to generate a target candidate response that matches the virtual character.
6. The method according to claim 1, characterized in that After sending the target reply to the target group chat, the method further includes: If the target reply self-learning condition is met, the natural language model performs self-learning based on the interactive content in the target group chat; Among them, the conditions for satisfying the target reply self-learning include at least one of the following: detecting a target reply change instruction, detecting an input reply matching the interactive content to be replied to, and detecting a reply evaluation result matching the target reply.
7. The method according to claim 3, characterized in that The method further comprises: If the interaction condition between the virtual characters is met, the natural language model generates a chat session matching at least two virtual characters based on at least one of the interaction content to be replied, the historical interaction content, and the scene type, and sends the chat session to the target group chat; Among them, the conditions for interaction between virtual characters are met, including: the difference between the current moment and the historical moment when the interactive content to be replied to was last detected in the target group chat is greater than or equal to the first time interval, or a triggering instruction for interaction between virtual characters is detected.
8. The method according to claim 1, characterized in that The method further comprises: If the virtual character relationship network update condition is met, the virtual character relationship network is updated according to at least one historical interaction content in the target group chat; Among them, the conditions for updating the virtual character network relationship include at least one of the following: the difference between the current moment and the historical moment when the virtual character network relationship was last updated is greater than or equal to the second time interval, after the last virtual character network relationship was updated, the new interactive content in the target group chat is greater than or equal to the first threshold, and a network change trigger word is detected.
9. The method according to any one of claims 1 to 8, characterized in that: The natural language model generates at least one candidate reply based on at least the interaction content to be replied, and further includes: The natural language model generates at least one candidate reply based at least on the interaction content to be replied and the role data; The character data includes at least one of the following: historical interaction content, virtual character settings, and virtual character relationship network.
10. A conversation replying device, characterized in that: include: A candidate reply generation module, configured to call a natural language model if a target group chat detects interactive content to be replied, wherein the natural language model generates at least one candidate reply based on at least the interactive content to be replied, wherein the target group chat includes at least two virtual characters, and the candidate reply matches at least one virtual character in the group chat; A candidate response evaluation module, configured to determine an evaluation result of each candidate response based on at least one response evaluation indicator by the natural language model; The response evaluation index includes at least one of the following: virtual character matching, scene matching, plot matching, virtual character response rate, and virtual character likeability; The target reply determination module is used to determine the candidate reply with the highest matching degree with the interactive content to be replied as the target reply among the candidate replies according to the evaluation results of each candidate reply, and send the target reply to the target group chat.
Citation Information
Patent Citations
Conversation method and device, equipment and medium
CN115577081A
Archiving robot group chat method based on AI training
CN116402088A
Content interaction method and device and computer readable storage medium
CN117560337A
Page interaction method and device, equipment and storage medium
CN117707370A
Conversation interaction method based on artificial intelligence (AI) virtual character and electronic equipment
CN118551005A