Session-based interaction method and device, computer equipment and readable storage medium
By using virtual avatars to output conversation messages and input response text in the conversation interface, the problem of inaccurate expression of intent in traditional conversation interfaces is solved, and resource conservation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-17
AI Technical Summary
In traditional chat interfaces, conversation messages are in text format, which cannot accurately express the user's intent, resulting in wasted resources.
By displaying a virtual avatar in the chat interface, and using the virtual avatar to output chat messages and respond to text input operations, the intent can be accurately expressed.
This reduces the number of sessions and saves communication server resources.
Smart Images

Figure CN121879872A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a session-based interaction method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Technology
[0002] With the development of computer technology, more and more users are beginning to communicate through communication applications, such as text, voice, or video. Users can usually establish conversations with other users to achieve real-time dialogue.
[0003] In traditional technology, after a user selects another user to establish a session, the two parties in the session can usually communicate by sending session messages through the session interface.
[0004] However, conversation messages sent through the chat interface are typically in text format, which cannot accurately express the user's intent. Therefore, both parties in the conversation need to send multiple conversation messages to accurately convey their intentions. This leads to continuous consumption of communication server resources, resulting in resource waste. Summary of the Invention
[0005] Therefore, it is necessary to provide a resource-saving session-based interaction method, apparatus, computer device, computer-readable storage medium, and computer program product to address the aforementioned technical problems.
[0006] Firstly, this application provides a session-based interaction method, including:
[0007] Displays the conversation interface between the first user and the second user;
[0008] In the conversation interface, a first virtual avatar representing the first user is displayed, and the first virtual avatar outputs the conversation messages generated by the first user in the conversation;
[0009] In response to the text input operation triggered by the second user, the entered text information is displayed;
[0010] In the conversation interface, a second virtual avatar representing the second user is displayed, and the second virtual avatar interacts with the first virtual avatar according to the text information.
[0011] Secondly, this application also provides a session-based interactive device, comprising:
[0012] The conversation interface display module is used to display the conversation interface between the first user and the second user.
[0013] A conversation message display module is used to display a first virtual avatar representing the first user in the conversation interface, and to output conversation messages generated by the first user in the conversation from the first virtual avatar.
[0014] The text information input module is used to respond to the text input operation triggered by the second user and display the input text information;
[0015] An interaction module is used to display a second virtual avatar representing the second user in the session interface, and for the second virtual avatar to interact with the first virtual avatar according to the text information.
[0016] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0017] Displays the conversation interface between the first user and the second user;
[0018] In the conversation interface, a first virtual avatar representing the first user is displayed, and the first virtual avatar outputs the conversation messages generated by the first user in the conversation;
[0019] In response to the text input operation triggered by the second user, the entered text information is displayed;
[0020] In the conversation interface, a second virtual avatar representing the second user is displayed, and the second virtual avatar interacts with the first virtual avatar according to the text information.
[0021] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0022] Displays the conversation interface between the first user and the second user;
[0023] In the conversation interface, a first virtual avatar representing the first user is displayed, and the first virtual avatar outputs the conversation messages generated by the first user in the conversation;
[0024] In response to the text input operation triggered by the second user, the entered text information is displayed;
[0025] In the conversation interface, a second virtual avatar representing the second user is displayed, and the second virtual avatar interacts with the first virtual avatar according to the text information.
[0026] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0027] Displays the conversation interface between the first user and the second user;
[0028] In the conversation interface, a first virtual avatar representing the first user is displayed, and the first virtual avatar outputs the conversation messages generated by the first user in the conversation;
[0029] In response to the text input operation triggered by the second user, the entered text information is displayed;
[0030] In the conversation interface, a second virtual avatar representing the second user is displayed, and the second virtual avatar interacts with the first virtual avatar according to the text information.
[0031] The aforementioned session-based interaction method, apparatus, computer device, computer-readable storage medium, and computer program product display a session interface for a first user and a second user to engage in a session. In the session interface, a first virtual avatar representing the first user is displayed, and the first virtual avatar outputs session messages generated by the first user during the session. This method of outputting session messages through the first virtual avatar enables accurate expression of the first user's intent. Responding to a text input operation triggered by the second user, the input text information is displayed. In the session interface, a second virtual avatar representing the second user is displayed, and the second virtual avatar interacts with the first virtual avatar according to the text information. This interaction between the second and first virtual avatars enables accurate expression of the second user's intent within the text information. In this way, the intent of both parties in the session can be accurately conveyed to each other each time a session message is sent, effectively reducing the number of sessions and thus reducing the continuous occupation of communication server resources, achieving resource conservation. Attached Figure Description
[0032] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0033] Figure 1 This is a diagram illustrating an application scenario of a session-based interaction method in one embodiment.
[0034] Figure 2 This is a diagram illustrating an application scenario of a session-based interaction method in another embodiment.
[0035] Figure 3 This is a diagram illustrating the application environment of a session-based interaction method in one embodiment.
[0036] Figure 4 This is a flowchart illustrating a session-based interaction method in one embodiment;
[0037] Figure 5 This is a schematic diagram of a text input area in one embodiment;
[0038] Figure 6 This is a schematic diagram of a text input control in one embodiment;
[0039] Figure 7 This is a schematic diagram illustrating the interaction between the second virtual avatar and the first virtual avatar in one embodiment;
[0040] Figure 8 This is a schematic diagram illustrating the interaction between the second virtual avatar and the first virtual avatar in another embodiment;
[0041] Figure 9 This is a schematic diagram of a session user selection interface in one embodiment;
[0042] Figure 10 A schematic diagram of a session user selection interface in another embodiment;
[0043] Figure 11 This is a flowchart illustrating a conversation interface for interacting with a first user in one embodiment.
[0044] Figure 12 This is a schematic diagram illustrating the generation of a confirmation interface in one embodiment;
[0045] Figure 13 This is a schematic diagram of interface view switching in one embodiment;
[0046] Figure 14 This is a schematic diagram illustrating the viewing of historical session fragments in one embodiment;
[0047] Figure 15 This is a schematic diagram of the pre-session phase in one embodiment;
[0048] Figure 16 This is a schematic diagram of the stages in a session in one embodiment;
[0049] Figure 17 This is an interaction sequence diagram of a session-based interaction method in one embodiment;
[0050] Figure 18 This is a structural block diagram of a session-based interactive device in one embodiment;
[0051] Figure 19This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0053] The session-based interaction method provided in this application can be applied to, for example... Figure 1 In the application scenario shown, a client for session-based interaction is installed on the terminal. This client can be a communication application. The terminal displays a session interface provided by the communication application for a first user and a second user to interact. In the session interface, a first virtual avatar representing the first user is displayed, and the first virtual avatar outputs the session messages generated by the first user in the session. In response to a text input operation triggered by the second user, the entered text information is displayed. A second virtual avatar representing the second user is displayed in the session interface, and the second virtual avatar interacts with the first virtual avatar according to the text information. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc.
[0054] In one exemplary embodiment, the session-based interaction method provided in this application can be applied to, for example... Figure 2In the application scenario shown, a client for implementing session-based interaction is installed on the terminal. This client can be a communication application. The terminal displays a session interface provided by the communication application for a first user and a second user to communicate. In response to a text input operation triggered by the second user, the entered text information is displayed. A second virtual avatar representing the second user is displayed in the session interface, and the second virtual avatar interacts with the first virtual avatar according to the text information. A first virtual avatar representing the first user is displayed in the session interface, and the first virtual avatar outputs the session messages generated by the first user in the session. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc.
[0055] In one exemplary embodiment, the session-based interaction method provided in this application can be applied to, for example... Figure 3 In the application environment shown, terminal 302 communicates with server 304 via a network. A data storage system can store the data that server 304 needs to process. The data storage system can be integrated onto server 304 or placed on a cloud or other network server. Terminal 302 displays a conversation interface for a first user and a second user. When it receives conversation messages generated by the first user in the conversation from server 304, it displays a first virtual avatar representing the first user in the conversation interface, and the first virtual avatar outputs the conversation messages generated by the first user. In response to a text input operation triggered by the second user, it displays the entered text information and outputs the text information to server 304. When it receives a video clip generated based on the text information from server 304, it loads the video clip and plays it in the conversation interface. The content of the video clip is a second virtual avatar representing the second user, interacting with the first virtual avatar according to the text information.
[0056] Terminal 302 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. Server 304 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0057] In one exemplary embodiment, such as Figure 4 As shown, a session-based interaction method is provided. This method can be executed by the terminal or the server alone, or by the terminal and the server collaboratively. This method can be applied to... Figure 1 Taking the terminal in the example, the explanation includes the following steps 402 to 408. Wherein:
[0058] Step 402: Display the conversation interface between the first user and the second user.
[0059] In this context, "second user" refers to the user using the terminal. "First user" refers to the user selected by the second user for the conversation. For example, the first user could be a user added by the second user within the communication application, meaning a user with whom the second user has already established a social relationship within the application. Alternatively, the first user could be a user generated using artificial intelligence (AI) technology provided by the communication application, capable of engaging in conversation with the second user based on AI technology. For instance, the first user could be a character from a movie or novel generated using AI technology. The conversation interface refers to the interface provided by the communication application for the first and second users to communicate, allowing the second user to interact with the first user by entering text messages, etc.
[0060] For example, when a second user selects a first user for a conversation, the terminal displays a conversation interface between the first and second users. In a specific application, when logging into a communication application as a second user, the terminal displays a conversation user selection interface provided by the communication application. In response to the second user's selection of the first user, a conversation interface between the first and second users is displayed.
[0061] Step 404: In the conversation interface, display the first virtual avatar representing the first user, and have the first virtual avatar output the conversation messages generated by the first user in the conversation.
[0062] The first virtual avatar, also known as the first virtual character, refers to an image or model created and displayed in a virtual environment that represents the first user. It can be two-dimensional or three-dimensional. For example, if the first user is a user added by a second user in a communication application, the first virtual avatar could specifically refer to an image or model representing the real user. Or, if the first user is a user generated using artificial intelligence technology, the first virtual avatar could specifically refer to an image or model representing an AI agent. For instance, if the first user is a character from a movie or novel, then the first virtual avatar would be an image representing that character.
[0063] In this context, "conversational messages" refers to information sent and received on communication applications, social media, email, or other communication platforms. Conversational messages can take various forms, such as text and images, and are used in communication between users. In this embodiment, conversational messages specifically refer to information sent and received on a communication application, used in a conversation between a first user and a second user.
[0064] For example, when the first user generates a session message in the session, the session message generated by the first user in the session can be transmitted to the terminal by the server. After receiving the session message generated by the first user in the session, the terminal will display a first virtual avatar representing the first user in the session interface, and the first virtual avatar will output the session message generated by the first user in the session.
[0065] In practical applications, when the first user is a user added by the second user in the communication application, the first virtual avatar representing the first user can be set by the first user themselves or generated based on the first user's appearance in the communication application. When the first user is a user generated using artificial intelligence technology, the first virtual avatar can be generated by the model based on the user description information of the first user. This user description information refers to information used to describe the first user, such as their name and age. For example, if the first user is a character from a movie, TV show, or novel, the user description information could specifically refer to the character's name in the movie, TV show, or novel.
[0066] In practical applications, depending on the session mode, the terminal will use different methods to output the session messages generated by the first user through the first virtual avatar. For example, when the session's audio mode is enabled, the terminal will display the corresponding voice messages of the session messages generated by the first user through the first virtual avatar. When the session's mute mode is enabled, the terminal will display the corresponding lip movements of the session messages generated by the first user through the first virtual avatar, and also display the session messages generated by the first user.
[0067] Step 406: In response to the text input operation triggered by the second user, display the input text information.
[0068] The text input operation refers to the action of entering text information triggered by the second user. The second user can engage in conversation or other interactions with the first user by entering text information. The text information refers to the text-based information entered by the second user, which is used to instruct the second virtual avatar representing the second user to interact with the first virtual avatar representing the first user in the conversation interface.
[0069] For example, when a second user needs to have a conversation with a first user, they can trigger a text input operation in the conversation interface. In response to the text input operation triggered by the second user, the terminal will display the entered text information. In a specific application, a text input area is displayed in the conversation interface. The text input operation can specifically be a trigger operation targeting the text input area. In response to the second user's trigger operation on the text input area, the terminal will display a text input control. The second user can enter text information through the text input control, and the terminal will display the text information entered by the second user through the text input control.
[0070] In a specific application, such as Figure 5 As shown, the text input area can be displayed at the bottom of the conversation interface. The text input area includes a first area indicating voice input and a second area indicating text input. In response to a trigger operation at any location within the text input area, the terminal displays text input controls. Furthermore, in response to a trigger operation on an area within the text input area indicating different input methods, the terminal can display different text input controls.
[0071] For example, in response to a trigger operation targeting the first area of the text input area that indicates voice input, the terminal displays a text input control for voice input. When the text input control is displayed, the second user can initiate a trigger operation targeting the text input control to start voice input, and when ending voice input, initiate a second trigger operation targeting the text input control. In response to the second trigger operation targeting the text input control, the terminal converts the voice input by the second user into text and displays the text information entered by the second user through the text input control, i.e., the converted text.
[0072] For example, in response to a trigger operation on the second area of the text input area indicating text input, the terminal will display a text input control for entering text. This text input control can be as follows: Figure 6 As shown, this is a virtual keyboard. When the virtual keyboard is displayed, a second user can use it to input text information. During this input process, the terminal will display the content entered via the virtual keyboard in real time, i.e., display the entered text information (e.g., ...). Figure 6 The image shown is a scene of people eating, in the distance.
[0073] Step 408: In the conversation interface, a second virtual avatar representing the second user is displayed, and the second virtual avatar interacts with the first virtual avatar according to the text information.
[0074] The second virtual avatar, also known as the second virtual character, refers to an image or model created and displayed in a virtual environment that represents a second user. This image or model can be two-dimensional or three-dimensional. It is understood that since the second user is the user of the terminal, the second virtual avatar refers to an image or model representing the real user.
[0075] For example, when the terminal displays the entered text information, in response to the second user's sending operation of the text information, the terminal will display a second virtual avatar representing the second user in the conversation interface, and the second virtual avatar will interact with the first virtual avatar according to the text information.
[0076] In practical applications, before displaying the second virtual avatar representing the second user in the conversation interface, the second virtual avatar needs to be generated first. During the generation of the second virtual avatar, the second user needs to confirm its creation. For example, when generating the second virtual avatar, the terminal will display an avatar generation confirmation interface to instruct the second user to confirm the source image used to generate the second virtual avatar. For instance, this source image could be the second user's avatar from a communication application. The second user can choose to generate the second virtual avatar based on their avatar from the communication application, or they can choose to upload a local image to generate the second virtual avatar based on the uploaded image.
[0077] In specific applications, the text information may include at least one of the following: conversation content with the first user or screen description information describing the virtual scene. Then, the terminal can display a second virtual avatar representing the second user in the conversation interface according to the specific content of the text information, and the second virtual avatar can interact with the first virtual avatar according to the text information.
[0078] The aforementioned session-based interaction method displays a session interface for a first user and a second user to converse. Within this interface, a first virtual avatar representing the first user is displayed, and this avatar outputs session messages generated by the first user during the session. This method allows for the accurate expression of the first user's intent by outputting session messages through the first virtual avatar. In response to a text input operation triggered by the second user, the entered text information is displayed. A second virtual avatar representing the second user is then displayed in the session interface, and this second virtual avatar interacts with the first virtual avatar based on the text information. This interaction between the second and first virtual avatars allows for the accurate expression of the second user's intent within the text information. This approach ensures that the intent of both parties is accurately conveyed each time they send session messages, effectively reducing the number of sessions and thus reducing the continuous occupation of communication server resources, thereby conserving resources.
[0079] In one exemplary embodiment, in the session interface, displaying a second virtual avatar representing a second user, and having the second virtual avatar interact with the first virtual avatar according to text information, includes:
[0080] When the text message contains conversation content with the first user, a second virtual avatar representing the second user is displayed in the conversation interface, and the second virtual avatar outputs the conversation content and interacts with the first virtual avatar.
[0081] The "conversation content" refers to the conversation messages entered by the second user that require interaction with the first user. For example, if the first virtual avatar outputs the conversation messages generated by the first user, the conversation content could specifically be the second user's response to those messages. Alternatively, the conversation content could also be other content entered by the second user that requires communication with the first user.
[0082] For example, the text information may include at least one of the following: conversation content with the first user or screen description information describing the virtual scene. After the terminal obtains the entered text information, it will process the text information to determine the type of information contained in the text information. When the text information contains conversation content with the first user, the terminal will display a second virtual avatar representing the second user in the conversation interface, and the second virtual avatar will output the conversation content and interact with the first virtual avatar.
[0083] In practical applications, the terminal processes text information and converts it into a film / television script format for output, thereby determining the information type contained within the text. Furthermore, a second user can directly input text information using the film / television script format, further enhancing their experience and interactivity. When using the film / television script format to input text information, to simplify the input, the second user can use film / television script symbols to represent common descriptive information such as scene transitions in the script. That is, when the terminal recognizes a film / television script symbol, it can determine the descriptive information such as the scene transition method mapped by that symbol.
[0084] For example, the mapping relationship between script symbols and descriptive information such as shot transitions in film and television dramas (script symbols - descriptive information) can be: INT - Interior, EXT - Exterior, DAY - Daytime, NIGHT - Nighttime, FADE IN - Fade in (start), FADE OUT - Fade out (end), CUT TO - Switch to, DISSOLVETO - Fade to, PAN TO - Pan to, TILT UP / DOWN - High-angle shot / High-angle shot, ZOOM IN / OUT - Zoom in / Zoom out, CLOSE UP (CU) - Close-up, MEDIUM SHOT (MS) - Medium shot, LONG SHOT (LS) - Long shot, WIDE SHOT (WS) - Wide shot, OVERTHE SHOULDER (OTS) - Overhead shot, TWO SHOT - Two-person shot, INSERT - Insert shot, FLASHBACK - Flashback, FLASH FORWARD - Flash forward, DREAM SEQUENCE - Dream sequence, SFX - Sound effect Effects, VFX - Visual Effects, ON CAMERA - On camera, OFF CAMERA - Off camera, BEAT - BEAT, (VO) - Voiceover, (OS) - Off screen.
[0085] For example, a second user can use simple and easy-to-understand symbols to identify different content elements, making the text information more expressive and richer. To further illustrate, a second user can use quotation marks ("") to identify conversation content; when a second user uses quotation marks in text information, the terminal recognizes the content within the quotation marks as conversation content. Similarly, a second user can use parentheses (()) to identify screen description information; when a second user uses parentheses in text information, the terminal recognizes the content within the parentheses as screen description information.
[0086] To illustrate further, since text messages typically contain only two information types, the second user only needs to use symbols to identify one type of content element; the content not identified by a symbol represents the other type of content element. For example, the second user can use parentheses (()) to identify screen description information. When the second user uses parentheses in text messages, the terminal recognizes the content within the parentheses as screen description information and the content outside the parentheses as session content.
[0087] In practical applications, during the interaction between the second virtual avatar and the first virtual avatar, the terminal displays the conversation content on the conversation interface according to the progress of the second virtual avatar's output. Furthermore, when the conversation's audio mode is enabled, the second virtual avatar will output corresponding voice messages. When the conversation's mute mode is enabled, the second virtual avatar will output corresponding lip movements.
[0088] In this embodiment, when the text information contains conversation content, the intention of the second user in the conversation content can be accurately expressed by displaying a second virtual avatar representing the second user in the conversation interface and having the second virtual avatar output the conversation content and interact with the first virtual avatar. This allows the second user's intention to be accurately conveyed to the first user, thereby effectively reducing the number of conversations, reducing the continuous occupation of communication server resources, and saving resources.
[0089] In an exemplary embodiment, in the conversation interface, displaying a second virtual avatar representing a second user, and having the second virtual avatar output conversation content and interact with the first virtual avatar, includes:
[0090] In the conversation interface, a second virtual avatar representing the second user is displayed. When the conversation's sound mode is enabled, the second virtual avatar outputs the corresponding voice of the conversation content to interact with the first virtual avatar.
[0091] When the sound mode is enabled, it means that the sound-related functions in the communication application have been turned on and can be used normally, and the conversation content can be output in the form of voice.
[0092] For example, when the second virtual avatar needs to output the conversation content, the terminal will display the second virtual avatar representing the second user in the conversation interface, and when the conversation's sound mode is enabled, it will display the voice corresponding to the conversation content output by the second virtual avatar and interact with the first virtual avatar.
[0093] In practical applications, the session mode can be controlled by a second user using the terminal. That is, the second user can control whether the session's sound mode is enabled or silent mode is enabled, and switch between the two modes. When the session's sound mode is enabled, a primary icon indicating that the sound mode is enabled can be displayed in the session interface. The form of this primary icon can be preset; for example, the primary icon can be in the shape of a small speaker.
[0094] In practical applications, the voice content of the conversation output by the second virtual avatar is generated by the terminal based on the second user's voice data. This second user's voice data can specifically refer to voice data generated by the second user during conversations with other users using communication applications, or it can refer to voice data uploaded by the second user.
[0095] In a specific application, if there is sufficient voice data generated by the second user during conversations with other users using a communication application, the terminal can directly generate the corresponding voice for the conversation content output by the second virtual avatar based on this voice data. In this case, the terminal first generates a voice embedding vector for the second user based on the voice data generated during conversations. This voice embedding vector represents the second user's timbre characteristics or attributes, and then the corresponding voice for the conversation content is generated using this vector. The voice embedding vector can be generated using a pre-trained voice analysis model; by inputting the second user's voice data into the voice analysis model, the second user's voice embedding vector can be generated.
[0096] In a specific application, when generating the second user's voice embedding vector, the terminal first filters the voice data generated by the second user during conversations with other users using a communication application, storing the qualified voice segments and their corresponding text content. For example, the qualified voice segments and their corresponding text content can be stored in a content distribution server so that they can be directly retrieved via the content distribution link at the time of storage when needed.
[0097] In a specific application, voice data can be filtered from the following dimensions: First, voice quality, based on parameters such as sampling rate and bit depth, the voice content needs to be clear, without obvious background noise, with moderate volume, and without excessive distortion or clipping. Second, voice content, which needs to consider the length and vocabulary richness of the voice. There is usually a minimum duration requirement for the voice, which can be configured according to the actual application scenario. Third, text recognition accuracy, which involves verifying and scoring the translation of the text content corresponding to the existing and stored voice segments. If this score is too low, an update mechanism will be triggered to update the text content corresponding to the stored voice segments. Fourth, data diversity, given existing voice segments, if new voice data is added, it needs to be determined whether the new voice data covers more vocabulary, phrases, pronunciations, second-user characteristics, etc. If the new voice data is too similar to existing voice segments, it should be discarded.
[0098] In a specific application, the terminal can use a pre-trained speech scoring model to determine whether the speech data meets the requirements. The input parameters of the speech scoring model include the speech data to be scored. If there are stored speech segments, these segments will also be used as input parameters to the speech scoring model. The output is the score of the speech data to be scored. If the score of the speech data to be scored is greater than a score threshold, then the speech data to be scored can be considered a speech segment that meets the requirements.
[0099] In a specific application, if the second user has no voice recordings or the number of valid voice segments is less than the threshold (configurable according to the actual application scenario), the terminal will display a voice data upload prompt, instructing the second user to upload voice data to generate their own voice embedding vector. At this time, the second user can upload voice data through real-time recording or, if voice data exists locally, directly select and upload it from their local storage. It is understandable that if the second user chooses to upload voice data directly from their local storage, the uploaded voice data can be their own voice data or the voice data of other users that the second virtual avatar expects to be output during the conversation, such as the voice data of a character from a movie or TV show.
[0100] In a specific application, when a second user chooses to upload voice data directly from their local device, the terminal verifies the size and duration of the uploaded voice data. If it exceeds the preset voice data size or duration, the terminal prompts the second user to replace it. The preset voice data size and duration can be configured according to the actual application scenario. Upon receiving the uploaded voice data, the terminal performs format conversion and preprocessing. Preprocessing includes segmenting the uploaded voice data by sentence dimension, scoring each segment, and finally counting the number of segments meeting the score threshold and the average score of the segments. If the number of segments meeting the threshold is too small or the average score is too low, the terminal prompts the second user to re-upload the voice data until the uploaded data meets the requirements. Based on the uploaded voice data, the terminal generates the second user's voice embedding vector.
[0101] In a specific application, generating the second user's speech embedding vector can be a timed task. The terminal will periodically determine whether to regenerate the second user's speech embedding vector. Specifically, when generating the second user's speech embedding vector, if there is new speech data for the second user, the terminal will first score the new speech data using a speech data scoring model. If the score of the new speech data is greater than a score threshold, it will check whether the number and average score of the stored speech segments meet the requirements for regenerating the second user's speech embedding vector. If the number of stored speech segments is less than the number threshold and the average score is lower than the average score threshold, the second user's speech embedding vector will be regenerated based on the new speech data. Otherwise, it will wait until the proportion of the new speech data in the stored speech segments exceeds a proportion threshold before regenerating the second user's speech embedding vector. The score threshold, number threshold, average score threshold, and proportion threshold can all be configured according to the actual application scenario.
[0102] Furthermore, the generated second user's speech embedding vector can be stored in the content distribution server. The terminal only stores the link to the content distribution server. When it is necessary to generate the speech corresponding to the second user's conversation content, it will continue to download and cache it in memory for a certain period of time to reduce storage pressure, and further reduce the time spent generating the second user's speech and improve the quality and effect of speech generation.
[0103] In a specific application, the speech corresponding to the conversation content can be generated based on the conversation content and the speech embedding vector of the second user. The terminal can obtain the speech corresponding to the conversation content by inputting the conversation content and the speech embedding vector of the second user into a pre-trained speech synthesis model. It should be noted that the speech synthesis model, while outputting the speech corresponding to the conversation content, also outputs segment information of the speech corresponding to the conversation content, such as duration, voice quality, and emotion.
[0104] Furthermore, before the second virtual avatar outputs the corresponding voice for the conversation content, the terminal displays a first preview area for the voice, instructing the second user to evaluate the generated voice. In response to the second user's triggering action in the first preview area, the terminal plays the voice. If the second user is not satisfied with the generated voice, they can trigger a regeneration action in the first preview area. The terminal will then re-generate the voice until the second user is satisfied. Finally, the terminal can trigger a sending action in the first preview area, and in response, it will display the voice output by the second virtual avatar interacting with the first virtual avatar. By separating the generation of the voice from the steps of processing text information and generating the second user's voice embedding vector, the time required to generate new voice can be significantly reduced.
[0105] In this embodiment, when the voice mode of the conversation is enabled, the second virtual avatar directly outputs the corresponding voice of the conversation content to interact with the first virtual avatar. This enriches the form of conversation content output, accurately expresses the intention of the second user in the conversation content, and enables the second user's intention to be accurately conveyed to the first user. This effectively reduces the number of conversations, reduces the continuous occupation of communication server resources, and saves resources.
[0106] In one exemplary embodiment, the session-based interaction method further includes:
[0107] When the mute mode of the conversation is enabled, the second virtual avatar outputs the corresponding lip movements of the conversation content to interact with the first virtual avatar and display the conversation content.
[0108] When silent mode is enabled, it means that the sound output function of the communication application is turned off or turned to the lowest level.
[0109] For example, when the second virtual avatar needs to output conversation content, the terminal will display a second virtual avatar representing the second user in the conversation interface. If the conversation's mute mode is enabled, the second virtual avatar will interact with the first virtual avatar by lip-syncing the conversation content, and then display the conversation content. In specific applications, when the conversation's mute mode is enabled, a second icon indicating that the mute mode is enabled can be displayed in the conversation interface. The form of this second icon can be preset; for example, the second icon can be a small speaker icon with a slash.
[0110] In this embodiment, when the mute mode of the conversation is enabled, the second virtual avatar outputs the corresponding lip movements of the conversation content to interact with the first virtual avatar and display the conversation content. This enriches the form of conversation content output, accurately expresses the intention of the second user in the conversation content, and enables the second user's intention to be accurately conveyed to the first user, thereby effectively reducing the number of conversations, reducing the continuous occupation of communication server resources, and saving resources.
[0111] In one exemplary embodiment, in the session interface, displaying a second virtual avatar representing a second user, and having the second virtual avatar interact with the first virtual avatar according to text information, includes:
[0112] When the text information contains a scene description that describes a virtual scene, the scene that matches the scene description will be displayed.
[0113] In the matched scene, a second virtual avatar representing the second user is displayed. If the scene description information contains first information indicating the interaction method, the second virtual avatar interacts with the first virtual avatar according to the interaction method indicated by the first information.
[0114] Virtual scene visuals refer to the visual effects of a virtual world generated through computer graphics technology. Scene description information refers to information that describes the virtual scene visuals; in other words, the virtual scene visuals can be reconstructed through the scene description information. For example, the scene description information may specifically include primary information indicating the interaction method, where the interaction method refers to the way the second virtual avatar interacts with the first virtual avatar.
[0115] For example, when the text information includes screen description information describing a virtual scene, the terminal displays a scene screen that matches the screen description information. In the matched scene screen, a second virtual avatar representing the second user is displayed. If the screen description information includes first information indicating the interaction method, the second virtual avatar interacts with the first virtual avatar according to the interaction method indicated by the first information. In a specific application, if the screen description information does not contain first information indicating the interaction method, the second virtual avatar will maintain the current interaction method and interact with the first virtual avatar.
[0116] In practical applications, the displayed scene that matches the screen description information can be a dynamic scene, which can be understood as a video clip generated by the terminal. It's understood that when the text information does not contain conversation content, the video clip is a segment without sound. In a specific application, the terminal can generate a video clip that matches the screen description information based on the screen description information, the first virtual avatar, the second virtual avatar, and the configuration information of the video clip, and display this video clip as the matched scene. The configuration information refers to the information obtained after configuring the parameters of the video clip, specifically including the video resolution, whether the video has sound, the images corresponding to the video character models (i.e., the first virtual avatar and the second virtual avatar), and the voice files (i.e., speech embedding vectors) of the video character roles (i.e., the second user and the first user).
[0117] In a specific application, by inputting screen description information, a first virtual avatar, a second virtual avatar, and video clip configuration information into a pre-trained video generation model, a video clip matching the screen description information can be obtained. Furthermore, in addition to screen description information, the first virtual avatar, the second virtual avatar, and video clip configuration information, the terminal can also obtain information about a second user and other additional information as input to enrich the video clip. Specifically, the second user's information can be information provided by the second user in communication applications, such as the user's appearance (specifically, elements of the user's appearance such as traditional Chinese style, trendy style, or academic style), registered anniversaries, and age. Other additional information can include the date the text information was generated, the corresponding holiday, the current weather, and the second user's geographical location. In other words, the input to the video generation model can include: screen description information, the first virtual avatar, the second virtual avatar, video clip configuration information, the second user's information, and other additional information.
[0118] Furthermore, the terminal can also use the previous session message corresponding to the screen description information and a video segment of the previous session message as input to generate a coherent video segment that conforms to physical laws. It should be noted that the video segment in this embodiment can also be a video segment corresponding to at least two session messages generated during the conversation between the first user and the second user. In this case, the video segment duration needs to be specified when generating the video segment, and the specified video segment duration can be configured according to the actual application scenario. For example, the specified video segment duration can be 10 seconds.
[0119] In this embodiment, when the text information includes screen description information, a scene screen matching the screen description information is displayed. In the matching scene screen, a second virtual avatar representing the second user is displayed. When the screen description information includes first information indicating the interaction method, the second virtual avatar interacts with the first virtual avatar according to the interaction method indicated by the first information. This can accurately express the second user's intention in the screen description information, so that the second user's intention can be accurately conveyed to the first user, thereby effectively reducing the number of sessions, reducing the continuous occupation of communication server resources, and saving resources.
[0120] In an exemplary embodiment, the scene description information further includes second information indicating the camera view; when the text information contains scene description information describing a virtual scene, displaying a scene scene matching the scene description information includes:
[0121] When the text information contains a description of the virtual scene and the virtual scene is displayed in the session interface, the scene shown in the camera lens indicated by the second information is displayed.
[0122] In film and video production, a "view shot" typically refers to a single view or scene captured by a camera—that is, a single image seen by the viewer on the screen. In this embodiment, it refers to a single image seen by the second user on the session interface, which can be understood as the virtual scene image seen by the second user on the session interface. It is understood that the second information indicating the view shot allows switching of the perspective and other parameters of the virtual scene image displayed on the session interface. This second information can describe the visual range of the shot (e.g., long shot, medium shot, close-up, feature shot, etc.), fade-in / fade-out effects (e.g., fade-in, fade-out), and relationship shots (e.g., shoulder shot, two-person shot, etc.).
[0123] For example, when the text information contains screen description information describing the virtual scene screen, and the virtual scene screen is displayed in the session interface, the terminal will display the scene screen under the screen lens indicated by the second information. In the scene screen under the screen lens indicated by the second information, a second virtual image representing the second user is displayed. If the screen description information contains first information indicating the interaction method, the second virtual image interacts with the first virtual image according to the interaction method indicated by the first information.
[0124] In a specific application, such as Figure 7 As shown, when a virtual scene is displayed in the conversation interface, if the scene description information contains the first piece of information indicating the interaction method (such as...) Figure 7 The image shown is a "dining scene") and the second information indicating the camera angle (such as...). Figure 7(As shown in the background), the terminal will display a dining scene. In the dining scene, a second virtual avatar representing the second user will be displayed. The second virtual avatar will interact with the first virtual avatar according to the interaction method (i.e., dining) indicated by the first information.
[0125] In this embodiment, by using the second information indicating the camera lens, the virtual scene screen displayed in the conversation interface can be switched, and the intention of the second user in the second information can be accurately expressed.
[0126] In one exemplary embodiment, in the session interface, displaying a second virtual avatar representing a second user, and having the second virtual avatar interact with the first virtual avatar according to text information, includes:
[0127] When the text information contains screen description information describing the virtual scene and conversation content of the conversation with the first user, the scene screen that matches the screen description information is displayed;
[0128] In the matched scene, a second virtual avatar representing the second user is displayed, and the second virtual avatar outputs conversation content and interacts with the first virtual avatar.
[0129] For example, when the text information includes a scene description describing a virtual scene and conversation content with the first user, the terminal will display a scene matching the scene description. In the matched scene, a second virtual avatar representing the second user will be displayed, and the second virtual avatar will output conversation content and interact with the first virtual avatar. In specific applications, if the conversation's audio mode is enabled, the terminal will display the voice corresponding to the conversation content output by the second virtual avatar interacting with the first virtual avatar in the matched scene. If the conversation's mute mode is enabled, the terminal will display the lip movements corresponding to the conversation content output by the second virtual avatar interacting with the first virtual avatar in the matched scene, and will also display the conversation content. It should be noted that when displaying the voice corresponding to the conversation content output by the second virtual avatar interacting with the first virtual avatar, the terminal can also display the conversation content.
[0130] In specific applications, such as Figure 8 As shown, the text information includes screen description information describing the virtual scene (such as...). Figure 8 The image shown is a "close-up" and includes the conversation content with the first user (e.g., ...). Figure 8 In the case shown as "Want to go out and play after dinner?", the terminal will display a scene that matches the screen description information (close-up). In the matching scene, a second virtual avatar representing the second user will be displayed, and the second virtual avatar will output the conversation content ("Want to go out and play after dinner?") to interact with the first virtual avatar.
[0131] In practical applications, the displayed scene matching the screen description information can be a dynamic scene, which can be understood as a video clip generated by the terminal. It is understood that since the text information contains conversation content, and the video clip is a segment with sound, the sound is the corresponding audio of the conversation content. Whether the audio is played depends on the conversation's mode. If the conversation's audio mode is enabled, the audio will play; if the conversation's mute mode is enabled, the audio will not play. What the user sees is the second virtual avatar outputting the lip movements corresponding to the conversation content and interacting with the first virtual avatar, displaying the conversation content.
[0132] In a specific application, if the voice corresponding to the conversation content is generated solely based on the voice data of the second user, the terminal can generate a video segment that matches the screen description information based on the screen description information, the first virtual avatar, the second virtual avatar, the configuration information of the video segment, and the voice corresponding to the conversation content, and display the video segment as the matched scene screen.
[0133] In a specific application, by inputting screen description information, a first virtual avatar, a second virtual avatar, video segment configuration information, and corresponding audio from the conversation into a pre-trained video generation model, a video segment matching the screen description information can be obtained. Furthermore, in addition to screen description information, the first virtual avatar, the second virtual avatar, video segment configuration information, and corresponding audio from the conversation, the terminal can also obtain information about a second user and other additional information as input to enrich the video segment. That is, the input to the video generation model can include: screen description information, the first virtual avatar, the second virtual avatar, video segment configuration information, the second user's information, and other additional information.
[0134] Furthermore, the terminal can also use the previous session message corresponding to the screen description information and a video segment of the previous session message as input to generate a coherent video segment that conforms to physical laws. It should be noted that the video segment in this embodiment can also be a video segment corresponding to at least two session messages generated during the conversation between the first user and the second user. In this case, the video segment duration needs to be specified when generating the video segment, and the specified video segment duration can be configured according to the actual application scenario. For example, the specified video segment duration can be 10 seconds.
[0135] In a specific application, generating video clips using a pre-trained video generation model can be divided into the following stages:
[0136] The first step is to analyze the text information and generate the content required for the plot:
[0137] Here, a natural language model is used to understand the input text information, and combined with the second user's information, additional information and other parameters other than audio and video segments, to generate the prompts required by the video generation model, and the prompts are also in the format of a movie script.
[0138] In practical applications, the terminal can process input text information using a pre-trained conversational message formatting model. This model can be invoked after a second user inputs text information. When invoking the model, the second user's information, video clip configuration information, and other additional information can be passed in simultaneously. This facilitates the model's ability to generate the information needed for the video clip, inputting it according to the format of a film or television script. This provides rich material sources for the corresponding speech and video generation models. In use, the conversational message formatting model can generate interesting and coherent text descriptions related to the conversational scenario based on the text information and additional information, thereby enhancing the user's conversational experience.
[0139] In a specific application, the session message formatting model runs on the terminal. To ensure operational efficiency and model scalability on the terminal, the lightweight TinyBERT model, a variant of the Transformer architecture, can be selected. Furthermore, knowledge distillation techniques can be used for model compression and optimization, and the model can be deployed using the mobile deep learning framework TensorFlow Lite, thereby reducing the model's consumption of terminal hardware resources.
[0140] Second, the construction of the training dataset:
[0141] The first step is to prepare the relevant training sets: First, a training set of film script texts. Specifically, a large amount of film script content needs to be prepared for unsupervised training, and labeled film script texts need to be prepared for subsequent supervised training. Second, a training set of static images of 2D characters and corresponding video clips (movie clips) needs to be prepared. Specifically, a large amount of video film script content of the characters needs to be prepared for unsupervised training, and labeled film script texts need to be prepared for subsequent supervised training. Finally, a video clip training set needs to be prepared manually according to dimensions such as subject matter, cinematography style, establishing shots, and character shots, and then labeled accordingly to create a high-quality training set that meets the requirements.
[0142] The second step is dataset preprocessing: First, the videos are cropped and scaled to achieve the same resolution. Next, the text is broken down into smaller text blocks, and the video is broken down into smaller video segments to be fed into the video generation model. Then, the text blocks and video segments are normalized to ensure their values fall within a suitable range.
[0143] Thirdly, the training of the video generation model:
[0144] Based on functionality, the video generation model is divided into several modules, such as a text analysis and story generation module, a character image reading module, a video generation module, and a video sound synthesis module. The text analysis and story generation module, the character image reading module, and the video sound synthesis module will initially undergo unsupervised training to generate corresponding base models. Then, supervised learning will be performed using the training set to achieve initial model fine-tuning. The fine-tuned model will then be trained, and reinforcement learning will be applied to complete the model training process.
[0145] Fourth, the output of video clips:
[0146] When the input parameter type is the first time a video clip is generated, the relevant parameters will be generated and assembled and input into the video generation model. After the video generation model outputs the video clip, the video compression algorithm is used to compress the video. After compression, the video clip is automatically uploaded to the content distribution server and the link of the content distribution server is stored in the terminal. The terminal will then render the video clip in a specified area of the screen.
[0147] It's important to note that before rendering the video clip, the terminal displays a second preview area for the clip, allowing a second user to evaluate the generated clip. In response to the second user's actions in the preview area, the terminal plays the clip. If the second user is dissatisfied with the result, they can trigger a regeneration operation in the preview area. The terminal will then regenerate the clip until the user is satisfied. Finally, the user can trigger a send operation in the preview area, and the terminal will display the clip in response. This method eliminates the need for a separate parameter generation and assembly process when a user chooses to regenerate the video; instead, the parameters are directly input into the video generation model to generate a new clip. This approach reuses reusable resources throughout the process and significantly reduces the time required for subsequent generation processes.
[0148] In this embodiment, when the text information includes screen description information and session content, by displaying a scene screen that matches the screen description information, a second virtual avatar representing the second user is displayed in the matched scene screen, and the second virtual avatar outputs session content to interact with the first virtual avatar. This can achieve an accurate expression of the second user's intention in the text information, so that the second user's intention can be accurately conveyed to the first user, thereby effectively reducing the number of sessions, reducing the continuous occupation of communication server resources, and achieving resource conservation.
[0149] In one exemplary embodiment, the session interface displaying the conversation between the first user and the second user includes:
[0150] When logging into the communication application as a second user, the session user selection interface provided by the communication application is displayed.
[0151] In the session user selection interface, at least one candidate user item is displayed;
[0152] In response to a first selection operation on any candidate user item, the user indicated by the first selection operation is designated as the first user, and a session interface for engaging in conversation with the first user is displayed.
[0153] For example, when logging into a communication application as a second user, the terminal displays a session user selection interface provided by the communication application. This interface displays at least one candidate user item for the second user to choose from. In response to a first selection operation on any candidate user item, the terminal designates the user indicated by the first selection operation as the first user and displays a session interface for engaging in conversation with the first user. In specific applications, each candidate user item indicates a candidate user, who can be a user added by the second user in the communication application or a user generated using artificial intelligence technology provided by the communication application.
[0154] In this embodiment, by displaying a session user selection interface, the second user can be instructed to select the user for the session. In response to the first selection operation of any candidate user item, the first user can be determined and a session interface for engaging in a session with the first user can be displayed.
[0155] In one exemplary embodiment, displaying at least one candidate user item on the session user selection interface includes:
[0156] In the session user selection interface, at least one user category item is displayed; the selected user category item is highlighted among the at least one user category item.
[0157] Display at least one candidate user item under the selected user category; at each candidate user item, display a user profile indicating the candidate user.
[0158] For example, in the user selection interface, the terminal displays at least one user category item. The selected user category item is highlighted. If a selected user category item exists, the terminal displays at least one candidate user item under that selected category, and at each candidate user item, displays a user profile indicating the candidate user, allowing a second user to select a candidate user based on the profile. The user profile may include the user name and user tags. Furthermore, user tags may include personality tags, MBTI (Myers-Briggs Type Indicator) type, and, if the candidate user is a user generated using artificial intelligence technology provided by the communication application, the user profile may also include user dialogue. User tags may also include era / clothing tags (e.g., ancient costume, modern costume, etc.).
[0159] In practical applications, the user category indicated by the user category field can be one of the preset categories, which can be configured according to the actual application scenario. For example, the preset categories can specifically be real users and virtual users. Real users refer to users added by a second user in the communication application, while virtual users refer to users generated using artificial intelligence technology provided by the communication application. The displayed session user selection interface can then be as follows: Figure 9 As shown, there are two user category items: real users and virtual users. When the selected user category item is real users, at least one candidate user item under the real users category item is displayed, indicating the user added by the second user in the communication application (e.g., real users). Figure 9 The image shows User 1, User 2, and User 3, and displays user profiles indicating the candidate users in the candidate user section. It is understood that a second user can switch between user categories by selecting the "Virtual User" category, and then select from at least one candidate user under the displayed user category.
[0160] In a specific application, candidate users can be virtual users (i.e., users generated using artificial intelligence technology provided by the communication application). The preset categories can further subdivide virtual users. When virtual users are further divided into multiple categories, at least one user category item displayed on the session user selection interface can be as follows: Figure 10 As shown, where, as Figure 10 As shown, the preset categories include Recommendation, Companionship, MBTI, Stories, Challenges, Celebrities, and Others. The selected user category "Recommended" is highlighted, and the terminal displays at least one candidate user under the "Recommended" user type. Each candidate user displays a brief introduction of the candidate user (e.g., ...). Figure 10 (Including tags and dialogue), so that a second user can select candidate users based on the user profile.
[0161] In this embodiment, the second user can be instructed to refer to the displayed user category items and the user profiles of the candidate users under the user category items to achieve accurate selection of the first user.
[0162] In an exemplary embodiment, in response to a first selection operation on any candidate user item, the user indicated by the first selection operation is designated as the first user, and a session interface for engaging in a conversation with the first user is displayed, including:
[0163] In response to the first selection operation on any candidate user item, the user indicated by the first selection operation is selected as the first user, and the session scenario selection interface is displayed.
[0164] In response to a scene selection operation triggered on the scene selection interface, the conversation interface for engaging in conversation with the first user is displayed according to the virtual scene indicated by the scene selection operation.
[0165] For example, in response to the second user's first selection operation on any candidate user item, the terminal will take the user indicated by the first selection operation as the first user, display the conversation scenario selection interface, instruct the second user to select a conversation scenario, and after the second user initiates the scenario selection operation, in response to the scenario selection operation triggered by the second user on the conversation scenario selection interface, display the conversation interface for having a conversation with the first user according to the virtual scenario indicated by the scenario selection operation.
[0166] In a specific application, the conversation scene selection interface displays at least one candidate scene item. In response to the second user's scene selection operation on any candidate scene item, the terminal will display the conversation interface for the conversation with the first user according to the virtual scene indicated by the scene selection operation.
[0167] In this embodiment, after the second user selects the first user, a conversation scenario selection interface is displayed, providing a virtual scenario for the second user to choose from. Thus, when the second user selects a virtual scenario, the conversation interface for engaging with the first user is displayed according to the selected virtual scenario. In this way, the initial setting of the conversation interface for engaging with the first user can be achieved, allowing the second user to engage in conversation with the first user within the set virtual scenario, thereby enabling the full expression of the second user's intentions during the conversation.
[0168] In an exemplary embodiment, the session scene selection interface includes a rendering style selection interface and a session location selection interface; in response to a scene selection operation triggered on the session scene selection interface, the session interface for engaging in a session with the first user is displayed according to the virtual scene indicated by the scene selection operation, including:
[0169] When at least one candidate style item is displayed in the rendering style selection interface, the session location selection interface is displayed in response to a second selection operation on any candidate style item.
[0170] On the session location selection screen, at least one candidate location item is displayed;
[0171] In response to a third selection operation on any candidate location item, a session interface for engaging in conversation with the first user is displayed, according to the rendering style indicated by the second selection operation and the virtual location indicated by the third selection operation.
[0172] The conversation scene selection interface includes a rendering style selection interface and a conversation location selection interface. The rendering style selection interface instructs the second user to choose a rendering style for the conversation interface. For example, the rendering style can be one of the following: anime, realistic, or Miyazaki style. The conversation location selection interface instructs the second user to choose a conversation location from the conversation interface. For example, the conversation location can be one of the following: rooftop, classroom, or restaurant.
[0173] For example, when the session scene selection interface includes a rendering style selection interface and a session location selection interface, the terminal may first display the rendering style selection interface. If at least one candidate style item is displayed in the rendering style selection interface, the second user can select the rendering style indicated by the candidate style item. In response to the second selection operation of any candidate style item, the terminal will further display the session location selection interface. In the session location selection interface, at least one candidate location item is displayed to instruct the second user to further select a meeting place. In response to the third selection operation of any candidate location item, the terminal will display the session interface for having a session with the first user according to the rendering style indicated by the second selection operation and the virtual location indicated by the third selection operation.
[0174] In specific applications, the terminal can also first display a session location selection interface. If at least one candidate location item is displayed in the session location selection interface, the second user can select the virtual location indicated by the candidate location item. In response to a third selection operation on any candidate location item, the terminal will further display a rendering style selection interface to instruct the second user to further select a rendering style. In response to a second selection operation on any candidate style item, the terminal will display a session interface for conversing with the first user according to the rendering style indicated by the second selection operation and the virtual location indicated by the third selection operation.
[0175] In a specific application, taking the display of the rendering style selection interface first, followed by the session location selection interface, as an example, the process of displaying the session interface for the first user can be as follows: Figure 11 As shown, at least one candidate style item is displayed in the rendering style selection interface (such as...). Figure 11 In the case of candidate style 1, candidate style 2, candidate style 3, etc., the response is a second selection operation on any candidate style item (assuming it is the candidate style item indicating candidate style 2) (select candidate style item 2 (indicated by a bold black box) and click). Figure 11 (Next step) Displays the session location selection interface. In the session location selection interface, at least one candidate location item is displayed (e.g., Figure 11 The example shown includes rooftops, classrooms, and restaurants. In response to a third selection action on any candidate location (e.g., the candidate location indicating a restaurant), the third selection action is performed (selecting the candidate location indicating a restaurant and clicking...). Figure 11 (In the selection process), according to the rendering style indicated by the second selection operation (candidate style 2) and the virtual location (restaurant) indicated by the third selection operation, a conversation interface for communicating with the first user is displayed.
[0176] In this embodiment, by instructing the second user to select a rendering style and a virtual location, and displaying a conversation interface for the first user to engage in dialogue according to the second user's selected rendering style and virtual location, the initial setting of the conversation interface for the first user to engage in dialogue can be achieved, so that the second user can engage in dialogue with the first user in the set virtual scene, and the second user's intention can be fully expressed in the conversation.
[0177] In an exemplary embodiment, in response to a scene selection operation triggered on the scene selection interface, displaying a session interface for conversing with the first user according to the virtual scene indicated by the scene selection operation includes:
[0178] In response to a scene selection operation triggered on the conversation scene selection interface, a confirmation interface for the generation of the second virtual avatar is displayed;
[0179] The avatar generation confirmation screen displays the user's outfit used to generate the second virtual avatar;
[0180] In response to the image generation confirmation operation triggered on the image generation confirmation interface, a conversation interface for engaging with the first user is displayed; the image generation confirmation operation is used to instruct the generation of a second virtual image according to the rendering style indicated by the user's avatar and scene selection operation.
[0181] Among them, the user avatar refers to the virtual image that a second user creates by dressing up themselves in a communication application, representing the second user.
[0182] For example, in response to a scene selection operation triggered on the scene selection interface, the terminal displays a second virtual avatar image generation confirmation interface. On the image generation confirmation interface, the user's avatar image used to generate the second virtual avatar is displayed to instruct the second user to confirm that the user's avatar image is the source of the second virtual avatar. If the second user confirms that the user's avatar image is the source of the generation, the terminal can trigger an image generation confirmation operation on the image generation interface. In response to the image generation confirmation operation triggered on the image generation confirmation interface, the terminal displays a conversation interface with the first user and generates the second virtual avatar according to the rendering style indicated by the user's avatar image and the scene selection operation.
[0183] In this embodiment, by displaying an image generation confirmation interface, which shows the user's avatar used to generate the second virtual avatar, the second user can be instructed to confirm that the user's avatar is the source of the second virtual avatar. Therefore, when the second user triggers the image generation confirmation operation, the second virtual avatar can be generated according to the rendering style indicated by the user's avatar and scene selection operation. This approach enhances the second user's interactive experience.
[0184] In one exemplary embodiment, the session-based interaction method further includes:
[0185] In response to an image generation adjustment operation triggered on the image generation confirmation screen, display at least one local image;
[0186] In response to a fourth selection operation on any local image, a conversation interface for engaging with the first user is displayed; the fourth selection operation is used to instruct the generation of a second virtual avatar according to the rendering style indicated by the selected local image and the scene selection operation.
[0187] For example, if the second user believes that other images need to be used as the generation source, he / she can trigger an image generation adjustment operation on the image generation confirmation interface. In response to the image generation adjustment operation triggered by the second user on the image generation confirmation interface, the terminal will display at least one local image to instruct the second user to select a local image as the generation source. In response to the fourth selection operation of any local image, the terminal will display a conversation interface for communicating with the first user and generate a second virtual image according to the selected local image and the rendering style indicated by the scene selection operation.
[0188] In specific applications, such as Figure 12 As shown, the avatar generation confirmation interface displays the user's outfit image used to generate the second virtual avatar, and a first control (such as...) that triggers the avatar generation confirmation operation. Figure 12 The second control shown is "Start Chat" and triggers the image generation and adjustment operation (such as...). Figure 12 (As shown in the image "Think about it again"), the second user can trigger the image generation confirmation operation by clicking the first control, or trigger the image generation adjustment operation by clicking the second control.
[0189] In this embodiment, in response to the image generation and adjustment operation, at least one local image is displayed, which can instruct the second user to select a local image as the generation source. Therefore, if the second user selects any local image, the second virtual image can be generated according to the selected local image and the rendering style indicated by the scene selection operation. This approach enhances the second user's interactive experience.
[0190] In one exemplary embodiment, the session-based interaction method further includes:
[0191] During the process of the first virtual avatar outputting the conversation messages generated by the first user in the conversation, the conversation interface displays the conversation messages generated by the first user in the conversation according to the progress of the first virtual avatar outputting the conversation messages.
[0192] For example, during the process of the first virtual avatar outputting the session messages generated by the first user in the session, the terminal displays the session messages generated by the first user in the session interface according to the progress of the first virtual avatar's output of session messages. It is understood that if the session messages generated by the first user cannot all be displayed at once, the terminal can display the session messages generated by the first user in the session according to the progress of the first virtual avatar's output of session messages. This allows the second user to better understand the first user's intent by combining the displayed session messages, effectively reducing the number of sessions and thus reducing the continuous occupation of communication server resources, achieving resource conservation.
[0193] In one exemplary embodiment, the session-based interaction method further includes:
[0194] The interface view switching entry is displayed in the chat interface; the interface view switching entry is used to indicate how to switch the view displayed in the chat interface.
[0195] In response to the triggering operation of the interface view switching entry, at least one session message generated by the first user and the second user in the session and the corresponding identity icon of the session message are displayed in the session interface.
[0196] The identity icon refers to the icon of the user who published the conversation message. For example, the identity icon can specifically refer to the avatar of the user who published the conversation message. When the user who published the message is a real user (i.e., a second user or a first user added by a second user in the communication application), the avatar of the user who published the message can be pre-configured according to the user's preferences.
[0197] For example, the session interface displays a view switching entry, which is used to indicate the switching of the view displayed in the session interface. If the second user needs to switch the view displayed in the session interface (including the view of the first virtual avatar or the second virtual avatar), he / she can initiate a trigger operation on the view switching entry. In response to the trigger operation on the view switching entry, the terminal will display at least one session message generated by the first user and the second user in the session and the corresponding identity icon of the session message in the session interface, so that the second user can quickly browse the at least one session message generated, so as to understand the intention of the first user by combining the at least one session message, thereby reducing the number of sessions and reducing the continuous occupation of the communication server's resources, thus saving resources.
[0198] In specific applications, such as Figure 13 As shown, the conversation interface displays a view switching entry, and the displayed view includes a second virtual avatar representing the second user. If the second user needs to quickly browse at least one conversation message generated during the conversation, they can initiate a trigger operation on the view switching entry. In response to this trigger operation, the terminal displays at least one conversation message generated by the first and second users during the conversation, along with the corresponding identity icon for that message. Furthermore, in... Figure 13 In the middle, the conversation messages generated by the first user and the corresponding identity icons of the conversation messages are displayed from left to right, while the conversation messages generated by the second user and the corresponding identity icons of the conversation messages are displayed from right to left.
[0199] In one exemplary embodiment, the session-based interaction method further includes:
[0200] The chat interface displays an entry point for viewing historical chats;
[0201] In response to a trigger operation on the historical session viewing entry, at least one historical session fragment item corresponding to the session message generated in the session is displayed; at the historical session fragment item, at least one virtual image of the session message output in the indicated historical session fragment is displayed;
[0202] In response to the selection of any historical session segment item, play the historical session segment indicated by the selected historical session segment item.
[0203] For example, the conversation interface displays a historical conversation viewing entry, which is used to indicate the viewing and management of historical conversation segments. If a second user needs to view historical conversation segments, they can initiate a trigger operation on the historical conversation viewing entry. In response to the trigger operation, the terminal displays at least one historical conversation segment item corresponding to the conversation message generated in the conversation, and displays at least one virtual image of the conversation message output in the indicated historical conversation segment at the historical conversation segment item. The second user can view the historical conversation segment indicated by any historical conversation segment item by selecting it. In response to the selection operation of any historical conversation segment item, the terminal plays the historical conversation segment indicated by the selected historical conversation segment item.
[0204] In practical applications, each historical session fragment item can correspond to a single session message generated in a session, or it can correspond to at least two session messages generated in a session. For example, a historical session fragment can correspond to two adjacent session messages generated in a session, and the session messages indicated by different historical session fragments are not repeated.
[0205] In practical applications, the historical session segment item displays a segment playback entry. A second user can trigger the selection of the historical session segment item containing any segment playback entry by initiating a trigger operation on any segment playback entry.
[0206] In a specific application, taking an example where the number of at least four historical session fragment items corresponding to the session messages generated in the session, such as... Figure 14As shown, the session interface displays a historical session viewing entry. In response to a trigger operation on the historical session viewing entry, the terminal displays four historical session fragment items corresponding to the session messages generated in the session. At each session fragment item, at least one virtual object that outputs session messages in the indicated historical session fragment is displayed, along with a fragment playback entry. In response to a trigger operation on any fragment playback entry, a second user can trigger a selection operation on a historical session fragment item containing any fragment playback entry. Then, the terminal can respond to the selection operation on any historical session fragment item by playing the historical session fragment indicated by the selected historical session fragment item.
[0207] In this embodiment, the second user can be instructed to quickly view and manage at least one historical session fragment in order to understand the intent of the first user by combining at least one historical session fragment, thereby effectively reducing the number of sessions, reducing the continuous occupation of communication server resources, and achieving resource conservation.
[0208] In an exemplary embodiment, the session-based interaction method of this application is used as an example to illustrate the interaction between a user (i.e., a first user) generated using artificial intelligence technology and a second user provided by a communication application.
[0209] The inventors argue that in current communication applications, users typically engage in asynchronous conversations via text or voice. In many scenarios, such as conversations with a user generated using artificial intelligence (hereinafter referred to as the first user), the second user can only visualize the conversation from the text, making it difficult to achieve an immersive experience. Furthermore, conversing solely through text cannot accurately convey the user's intent; therefore, both parties need to send multiple messages to accurately express their intentions. This leads to continuous resource consumption on the communication server, resulting in resource waste.
[0210] Based on this, this application provides a session-based interaction method that combines session messages generated during a session with information from the first and second users to generate coherent session fragments that fit the current session atmosphere. This visualizes the session, providing a richer, more immersive, and more realistic session experience, significantly enhancing the immersive experience for both parties in the current session. Simultaneously, the generated session fragments can be used to accurately express the user's intent. In this way, the intent of both parties can be accurately conveyed to each other each time they send a session message, effectively reducing the number of sessions and thus reducing the continuous consumption of communication server resources, achieving resource conservation.
[0211] In practical applications, the session-based interaction method provided in this application is mainly divided into two parts: the pre-session stage and the in-session stage. The two parts are described below.
[0212] First, in the pre-session phase, the second user needs to set necessary information to provide the basic basis for AI generation. Specifically, such as... Figure 15 As shown, the second user first needs to select the first user for the conversation on the conversation user selection interface, then select the rendering style on the rendering style selection interface, then select the virtual location on the conversation location selection interface, and finally confirm the source of the second virtual avatar on the avatar generation confirmation interface. After completing the above operations, the terminal will display the conversation interface with the first user according to the rendering style and virtual location selected by the second user, and generate the second virtual avatar according to the operation triggered by the second user on the avatar generation confirmation interface.
[0213] In specific applications, such as Figure 15 As shown, in the session user selection interface, at least one user category item is displayed, and the selected user category item is highlighted (e.g., ...). Figure 15 The terminal displays multiple candidate user items under the "Recommended" user category for the second user to choose from. Each candidate user item includes a brief description of the user. If the second user selects a first user under the "Recommended" category, they simply trigger a selection operation on any candidate user item. The terminal, in response to this selection, will designate the user indicated by the first selection operation as the first user. Alternatively, the second user can choose to switch to another user category before making a selection. Understandably, if the second user switches to another user category, the terminal will display at least one candidate user item under the new user category for the second user to choose from.
[0214] In specific applications, such as Figure 15 As shown, in the rendering style selection interface, at least one candidate style item is displayed for the second user to select, and in the session location selection interface, at least one candidate location item is displayed for the second user to select. After the second user completes the selection of the candidate style item and the candidate location item, the terminal will display the session interface for the first user to have a session according to the rendering style indicated by the second selection operation and the virtual location indicated by the third selection operation.
[0215] In specific applications, such as Figure 15As shown, after the second user completes the selection of candidate style and candidate location, the terminal will display a confirmation interface for the generation of the second virtual avatar. This interface will show the user's avatar outfit used to generate the second virtual avatar. If the second user selects the user's avatar outfit as the source for generating the second virtual avatar, the confirmation operation will be triggered (i.e., clicking the confirmation button). Figure 15 In the "Start Chat" section, the terminal responds to the avatar generation confirmation operation by displaying a conversation interface with the first user according to the rendering style indicated by the second selected operation and the virtual location indicated by the third selected operation, and instructs the generation of a second virtual avatar according to the rendering style indicated by the user's avatar and scene selection operation.
[0216] In practical applications, if a second user wishes to select a different image as the source for generating the second virtual avatar, they can trigger an avatar generation adjustment operation (i.e., click...). Figure 15 In response to the image generation adjustment operation ("Think again"), the terminal will display at least one local image for the second user to select. If the second user selects any local image, in response to the fourth selection operation on any local image, the terminal will display a conversation interface for the first user to communicate with the second user according to the rendering style indicated by the second selection operation and the virtual location indicated by the third selection operation, and generate a second virtual image according to the selected local image and the rendering style indicated by the second selection operation.
[0217] Secondly, during the chat phase, the default setting is an AI-generated immersive conversation interface to begin the conversation.
[0218] In specific applications, in a conversation, such as Figure 16 As shown, if the second user selects to generate a second virtual avatar according to the rendering style indicated by the user's appearance and scene selection operation, then the generated second virtual avatar (such as...) Figure 16 The image shown is of a woman wearing a dress and is based on the user's avatar (e.g., ...). Figure 15 The image shown is of a woman wearing a dress. AI will generate a second virtual avatar based on these features, combined with the rendering style selected by the second user. The first virtual avatar, representing the first user, is generated based on the rendering style selected by the second user. Figure 16 In the game, the first virtual character is a man wearing glasses.
[0219] In specific applications, such as Figure 16 As shown, if a session can be initiated by the first user, then the session interface will display a first virtual avatar representing the first user, and the first virtual avatar will output the session messages generated by the first user in the session (such as...). Figure 16(The image shown is "The call is over"). Based on this, a second user can trigger a text input operation to enter text information, such as... Figure 16 As shown, in response to a text input operation triggered by a second user, the terminal will display the input text information (such as...). Figure 16 The text shown is titled "Finished (looking at him and asking)"). The text entered by the second user includes the conversation content with the first user (e.g., ...). Figure 16 The text shows "finished" and descriptions of the virtual scene (specifically, the first piece of information indicating the interaction method, such as...). Figure 16 The image shown is titled "Looking at him gently, he said." Therefore, the terminal will display a scene screen that matches the screen description information. In the matched scene screen, a second virtual avatar representing the second user will be displayed. If the screen description information contains first information indicating the interaction method, the second virtual avatar will output the session content (such as...) according to the interaction method indicated by the first information. Figure 16 The text shows an interaction between the user and the first virtual avatar (indicated by "finished"). Further, the first user can reply to the conversation content output by the second virtual avatar representing the second user, such as... Figure 16 As shown, when the reply is "Come over for dinner," the terminal responds to the first user's reply in the session by displaying a virtual avatar representing the first user, and the first virtual avatar outputs the reply content generated by the first user (e.g., ...). Figure 16 The example shown is "Come over and have some food." It's understandable that during the conversation between the first and second users, whichever user outputs the conversation message will have their virtual avatar displayed on the conversation interface.
[0220] In practical applications, if the text information includes a scene description that describes the virtual scene, the terminal can render the scene based on the scene description, that is, display the scene that matches the scene description. In a specific application, such as... Figure 7 As shown, if the text information is "(eating scene, distant view)", then the terminal displays a scene that matches the scene description information. In this scene, a second virtual avatar representing the second user is displayed. This second virtual avatar interacts with the first virtual avatar according to the interaction method ("eating") indicated by the first information ("eating scene") in the scene description information. Furthermore, since the scene description information also includes a second information ("distant view") indicating the camera angle, the displayed scene that matches the scene description information is actually the scene under the "distant view" indicated by the second information. In a specific application, such as... Figure 8As shown, if the text information is "Want to go out and play after dinner (close-up)", which includes the conversation content ("Want to go out and play after dinner") and the second information ("close-up"), the terminal will display the scene shown in the shot ("close-up") indicated by the second information. In the scene, a second virtual avatar representing the second user is displayed, and the second virtual avatar outputs the conversation content ("Want to go out and play after dinner") to interact with the first virtual avatar. It should be noted that in this embodiment, parentheses are used to identify the scene description information. The content within the parentheses in the text information is the scene description information, and the content outside the parentheses is the conversation content.
[0221] In specific applications, the session interface displays an interface view switching entry. This entry is used to indicate the switching of the view displayed in the session interface so as to quickly browse previous session messages. In response to the triggering operation of the interface view switching entry, the terminal displays at least one session message generated by the first user and the second user in the session, along with the corresponding identity icon of the session message.
[0222] In specific applications, the generated video clips are saved simultaneously, and a second user can access and manage them in the history list. In one specific application, the session interface displays a historical session viewing entry. In response to a trigger operation targeting this entry, the terminal displays at least one historical session clip item corresponding to the session message generated in the session. At each historical session clip item, at least one virtual image of the session message output in the indicated historical session clip is displayed. In response to selecting any historical session clip item, the terminal plays the historical session clip indicated by the selected item.
[0223] In specific applications, the main technologies involved in the session-based interaction method of this application can be simply summarized as follows: based on the specified session message, combined with the user's information and other necessary information, a video segment with text and speech of a specified duration and resolution is generated. The video segment can record the content of the previous video segment and output the same style and main character image (i.e., the first virtual image and the second virtual image). Based on the commonalities of the above two modes, the entire process can be divided into three stages: the first step is text analysis and generation, the second step is the synthesis of audio files, and the third step is video segment generation.
[0224] The first step, text analysis and generation, mainly includes preprocessing the conversation messages and processing the input text information using a pre-trained conversation message formatting model to generate the information needed for the video clip. Specifically, if the text information is input via voice, the voice needs to be converted into text first, and user information, additional information, and video clip configuration information need to be obtained. Finally, the text information is converted into a film / television script format for output.
[0225] In the second step of synthesizing the audio file, the main steps include generating the second user's speech embedding vector based on the second user's speech data, and then synthesizing speech segments based on the second user's speech embedding vector and text information.
[0226] In the third step of video clip generation, the main steps include inputting text information, the first virtual avatar, the second virtual avatar, the configuration information of the video clip, the information of the second user and other additional information, the previous conversation message corresponding to the screen description information, and the video clip of the previous conversation message into a pre-trained video generation model to generate a video clip that matches the text information.
[0227] In an exemplary embodiment, taking the session-based interaction method of this application as being implemented through the interaction between a terminal and a server, and the first user being a user (i.e., an AI agent) generated using artificial intelligence technology provided by a communication application as an example, the interaction sequence diagram of the session-based interaction method of this application can be as follows: Figure 17 As shown.
[0228] Specifically, when the second user clicks the "Start Chat" button on the terminal, triggering the entry into the chat interface, the terminal transmits the AI agent's identifier, rendering style identifier, and chat location identifier to the business backend server. The business backend server then transmits the AI agent's identifier to the server running the AI agent's AI model, enabling the server running the AI agent's AI model to transmit the AI agent's opening message to the business backend server. The business backend server then transmits the opening message to the terminal, which renders the chat interface with the AI agent and displays the opening message. If the second user enters text information and clicks send, the terminal displays the entered text information in the chat interface and sends it to the business backend server. The business backend server then sends the text information to the server running the AI agent's AI model, enabling the server running the AI agent's AI model to transmit the AI's response information to the business backend server. The AI response information and the second user's voice embedding vector are transmitted to the server running the sound segment AI model. This allows the server to generate a sound segment, upload it to the content distribution server, and transmit the generated sound segment file via a first retrieval link from the content distribution server to the business backend server. The business backend server then transmits this first retrieval link, along with other necessary information (including text information, the second user's information, additional information, and video segment configuration information), to the server running the video segment AI model. This allows the server running the sound segment AI model to generate a video segment, upload it to the content distribution server, and transmit the generated video segment file via a second retrieval link from the content distribution server to the business backend server. The business backend server then transmits this second retrieval link to the terminal, allowing the terminal to load the video segment and play it on the session interface. It should be noted that... Figure 17 The interaction with the content delivery server is not shown.
[0229] It is understood that the session-based interaction method of this application has the following advantages:
[0230] 1. It can visualize chat, providing a richer, more immersive, and more realistic chat experience.
[0231] 2. By supporting movie script formats, the accuracy of model-generated videos and the user's control over them are improved.
[0232] 3. In the text analysis and generation step, a lightweight model is selected to run locally on the terminal, thereby reducing the server load and shortening the time required for a single generation.
[0233] 4. Storing the generated files on the content distribution server reduces the performance loss to the content and AI model servers, facilitates resource reuse, and makes it easier to clean up when users delete them.
[0234] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0235] Based on the same inventive concept, this application also provides a session-based interaction device for implementing the session-based interaction method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more session-based interaction device embodiments provided below can be found in the limitations of the session-based interaction method described above, and will not be repeated here.
[0236] In one exemplary embodiment, such as Figure 18 As shown, a session-based interactive device is provided, including: a session interface display module 1802, a session message display module 1804, a text information input module 1806, and an interaction module 1808, wherein:
[0237] The conversation interface display module 1802 is used to display the conversation interface between the first user and the second user.
[0238] The conversation message display module 1804 is used to display a first virtual avatar representing the first user in the conversation interface, and to output the conversation messages generated by the first user in the conversation from the first virtual avatar.
[0239] The text information input module 1806 is used to respond to the text input operation triggered by the second user and display the input text information;
[0240] The interaction module 1808 is used to display a second virtual avatar representing the second user in the conversation interface, and the second virtual avatar interacts with the first virtual avatar according to text information.
[0241] The aforementioned session-based interactive device displays a session interface for a first user and a second user to converse. The session interface displays a first virtual avatar representing the first user, which outputs session messages generated by the first user during the session. This method of outputting session messages through the first virtual avatar accurately expresses the first user's intent. In response to a text input operation triggered by the second user, the device displays the input text information. The session interface also displays a second virtual avatar representing the second user, which interacts with the first virtual avatar based on the text information. This interaction between the second and first virtual avatars accurately expresses the second user's intent within the text information. This approach ensures that the intent of both parties is accurately conveyed each time they send session messages, effectively reducing the number of sessions and thus reducing the continuous occupation of communication server resources, thereby saving resources.
[0242] In an exemplary embodiment, the interaction module is further configured to, when the text information contains conversation content with the first user, display a second virtual avatar representing the second user in the conversation interface, and have the second virtual avatar output the conversation content to interact with the first virtual avatar.
[0243] In an exemplary embodiment, the interaction module is further configured to display a second virtual avatar representing the second user in the conversation interface, and, when the conversation's sound mode is enabled, have the second virtual avatar output voice corresponding to the conversation content to interact with the first virtual avatar.
[0244] In an exemplary embodiment, the interaction module is further configured to, when the mute mode of the conversation is enabled, have the second virtual avatar output lip movements corresponding to the conversation content to interact with the first virtual avatar and display the conversation content.
[0245] In an exemplary embodiment, the interaction module is further configured to, when the text information contains screen description information describing a virtual scene screen, display a scene screen that matches the screen description information, display a second virtual image representing the second user in the matched scene screen, and, if the screen description information contains first information indicating the interaction method, have the second virtual image interact with the first virtual image according to the interaction method indicated by the first information.
[0246] In an exemplary embodiment, the screen description information further includes second information indicating the screen lens, and the interaction module is further configured to display the scene screen under the screen lens indicated by the second information when the text information contains screen description information describing the virtual scene screen and the virtual scene screen is displayed in the session interface.
[0247] In an exemplary embodiment, the interaction module is further configured to, when the text information contains screen description information describing a virtual scene and conversation content for a conversation with a first user, display a scene screen that matches the screen description information, display a second virtual avatar representing a second user in the matched scene screen, and have the second virtual avatar output conversation content to interact with the first virtual avatar.
[0248] In an exemplary embodiment, the session interface display module is further configured to, when logging into the communication application as a second user, display a session user selection interface provided by the communication application, display at least one candidate user item in the session user selection interface, and, in response to a first selection operation on any candidate user item, designate the user indicated by the first selection operation as the first user and display a session interface for having a session with the first user.
[0249] In an exemplary embodiment, the session interface display module is further configured to display at least one user category item in the session user selection interface; the selected user category item is highlighted among the at least one user category item, and at least one candidate user item is displayed under the selected user category item; at the candidate user item, a user profile indicating the candidate user is displayed.
[0250] In an exemplary embodiment, the conversation interface display module is further configured to, in response to a first selection operation on any candidate user item, designate the user indicated by the first selection operation as the first user, display a conversation scene selection interface, and, in response to a scene selection operation triggered on the conversation scene selection interface, display a conversation interface for conversing with the first user according to the virtual scene indicated by the scene selection operation.
[0251] In an exemplary embodiment, the session scene selection interface includes a rendering style selection interface and a session location selection interface. The session interface display module is further configured to, in response to a second selection operation on any candidate style item when at least one candidate style item is displayed in the rendering style selection interface, display the session location selection interface, display at least one candidate location item in the session location selection interface, and, in response to a third selection operation on any candidate location item, display the session interface for having a session with the first user according to the rendering style indicated by the second selection operation and the virtual location indicated by the third selection operation.
[0252] In an exemplary embodiment, the conversation interface display module is further configured to display a second virtual avatar image generation confirmation interface in response to a scene selection operation triggered on the conversation scene selection interface. On the image generation confirmation interface, a user avatar for generating the second virtual avatar is displayed. In response to an image generation confirmation operation triggered on the image generation confirmation interface, a conversation interface for conversing with the first user is displayed. The image generation confirmation operation is used to instruct the generation of the second virtual avatar according to the rendering style indicated by the user avatar and the scene selection operation.
[0253] In an exemplary embodiment, the conversation interface display module is further configured to display at least one local image in response to an image generation adjustment operation triggered on the image generation confirmation interface, and to display a conversation interface for conversing with the first user in response to a fourth selection operation on any local image; the fourth selection operation is configured to instruct the generation of a second virtual image according to the selected local image and the rendering style indicated by the scene selection operation.
[0254] In an exemplary embodiment, the conversation message display module is further configured to, during the process of the first virtual avatar outputting conversation messages generated by the first user in the conversation, display the conversation messages generated by the first user in the conversation interface according to the progress of the first virtual avatar outputting conversation messages.
[0255] In an exemplary embodiment, the session-based interaction device further includes a view switching module, which is used to display an interface view switching entry in the session interface. The interface view switching entry is used to indicate the switching of the view displayed in the session interface. In response to the triggering operation of the interface view switching entry, at least one session message generated by the first user and the second user in the session and the corresponding identity icon of the session message are displayed in the session interface.
[0256] In an exemplary embodiment, the session-based interactive device further includes a session viewing module, which is configured to display a historical session viewing entry in the session interface, and in response to a trigger operation on the historical session viewing entry, display at least one historical session fragment item corresponding to the session message generated in the session, and at the historical session fragment item, display at least one virtual image of the session message output in the indicated historical session fragment, and in response to a selection operation on any historical session fragment item, play the historical session fragment indicated by the selected historical session fragment item.
[0257] The modules in the aforementioned session-based interactive device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can invoke and execute the operations corresponding to each module.
[0258] In one exemplary embodiment, a computer device is provided. This computer device can be a terminal or a server. Taking the computer device as a terminal as an example, its internal structure diagram can be as follows: Figure 19 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a session-based interactive method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0259] Those skilled in the art will understand that Figure 19 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0260] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0261] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0262] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0263] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0264] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0265] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0266] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A session-based interaction method, characterized in that, The method includes: Displays the conversation interface between the first user and the second user; In the conversation interface, a first virtual avatar representing the first user is displayed, and the first virtual avatar outputs the conversation messages generated by the first user in the conversation; In response to the text input operation triggered by the second user, the entered text information is displayed; In the conversation interface, a second virtual avatar representing the second user is displayed, and the second virtual avatar interacts with the first virtual avatar according to the text information.
2. The method according to claim 1, characterized in that, The step of displaying a second virtual avatar representing the second user in the conversation interface, and having the second virtual avatar interact with the first virtual avatar according to the text information, includes: When the text information contains conversation content with the first user, a second virtual avatar representing the second user is displayed in the conversation interface, and the second virtual avatar outputs the conversation content to interact with the first virtual avatar.
3. The method according to claim 2, characterized in that, The step of displaying a second virtual avatar representing the second user in the conversation interface, and having the second virtual avatar output the conversation content and interact with the first virtual avatar, includes: In the conversation interface, a second virtual avatar representing the second user is displayed. When the voice mode of the conversation is enabled, the second virtual avatar outputs the corresponding voice of the conversation content to interact with the first virtual avatar.
4. The method according to claim 3, characterized in that, The method further includes: When the mute mode of the conversation is enabled, the second virtual avatar outputs the lip movements corresponding to the conversation content to interact with the first virtual avatar and display the conversation content.
5. The method according to claim 1, characterized in that, The step of displaying a second virtual avatar representing the second user in the conversation interface, and having the second virtual avatar interact with the first virtual avatar according to the text information, includes: When the text information contains screen description information describing a virtual scene, the scene screen that matches the screen description information is displayed; In the matched scene, a second virtual avatar representing the second user is displayed. If the scene description information contains first information indicating the interaction method, the second virtual avatar interacts with the first virtual avatar according to the interaction method indicated by the first information.
6. The method according to claim 5, characterized in that, The image description information also includes second information indicating the camera angle; when the text information contains image description information describing a virtual scene, displaying a scene image matching the image description information includes: When the text information contains screen description information describing the virtual scene, and the virtual scene is displayed in the session interface, the scene shown in the camera lens indicated by the second information is displayed.
7. The method according to claim 1, characterized in that, The step of displaying a second virtual avatar representing the second user in the conversation interface, and having the second virtual avatar interact with the first virtual avatar according to the text information, includes: When the text information includes screen description information describing a virtual scene and conversation content of a conversation with the first user, a scene screen matching the screen description information is displayed; In the matched scene, a second virtual avatar representing the second user is displayed, and the second virtual avatar outputs the conversation content and interacts with the first virtual avatar.
8. The method according to claim 1, characterized in that, The session interface displaying the conversation between the first user and the second user includes: When logging into the communication application as the second user, the session user selection interface provided by the communication application is displayed. The session user selection interface displays at least one candidate user item; In response to a first selection operation on any candidate user item, the user indicated by the first selection operation is designated as the first user, and a session interface for engaging in conversation with the first user is displayed.
9. The method according to claim 8, characterized in that, Displaying at least one candidate user item on the session user selection interface includes: In the session user selection interface, at least one user category item is displayed; the selected user category item is highlighted among the at least one user category items. Display at least one candidate user item under the selected user category; at the candidate user item, display a user profile indicating the candidate user.
10. The method according to claim 8, characterized in that, The step of responding to a first selection operation on any candidate user item, and displaying a conversation interface for engaging in conversation with the first user as the first user, includes: In response to a first selection operation on any candidate user item, the user indicated by the first selection operation is selected as the first user, and the session scenario selection interface is displayed. In response to a scenario selection operation triggered on the scenario selection interface, a scenario interface for engaging in a conversation with the first user is displayed according to the virtual scenario indicated by the scenario selection operation.
11. The method according to claim 10, characterized in that, The conversation scene selection interface includes a rendering style selection interface and a conversation location selection interface; the process of displaying the conversation interface for conversing with the first user according to the virtual scene indicated by the scene selection operation, in response to a scene selection operation triggered on the conversation scene selection interface, includes: When at least one candidate style item is displayed in the rendering style selection interface, the session location selection interface is displayed in response to a second selection operation on any candidate style item. The session location selection interface displays at least one candidate location item; In response to a third selection operation on any candidate location item, a session interface for engaging in conversation with the first user is displayed, according to the rendering style indicated by the second selection operation and the virtual location indicated by the third selection operation.
12. The method according to claim 10, characterized in that, The step of displaying a conversation interface for engaging in conversation with the first user in response to a scene selection operation triggered on the conversation scene selection interface, according to the virtual scene indicated by the scene selection operation, includes: In response to a scene selection operation triggered on the session scene selection interface, a confirmation interface for the generation of the second virtual avatar is displayed; The avatar generation confirmation interface displays the user's outfit used to generate the second virtual avatar; In response to the image generation confirmation operation triggered on the image generation confirmation interface, a conversation interface for engaging with the first user is displayed; the image generation confirmation operation is used to instruct the generation of the second virtual image according to the rendering style indicated by the user's avatar and the scene selection operation.
13. The method according to claim 12, characterized in that, The method further includes: In response to an image generation adjustment operation triggered on the image generation confirmation interface, at least one local image is displayed; In response to a fourth selection operation on any local image, a session interface for engaging with the first user is displayed; the fourth selection operation is used to instruct the generation of the second virtual avatar according to the selected local image and the rendering style indicated by the scene selection operation.
14. The method according to any one of claims 1 to 13, characterized in that, The method further includes: During the process of the first virtual avatar outputting the conversation messages generated by the first user in the conversation, the conversation interface displays the conversation messages generated by the first user in the conversation according to the progress of the first virtual avatar outputting the conversation messages.
15. The method according to any one of claims 1 to 13, characterized in that, The method further includes: The session interface displays an interface view switching entry; the interface view switching entry is used to indicate the switching of the view displayed in the session interface. In response to the triggering operation of the interface view switching entry, at least one session message generated by the first user and the second user in the session and the corresponding identity icon of the session message are displayed in the session interface.
16. The method according to any one of claims 1 to 13, characterized in that, The method further includes: The session interface displays an entry point for viewing historical sessions; In response to a trigger operation on the historical session viewing entry, at least one historical session fragment item corresponding to the session message generated in the session is displayed; at the historical session fragment item, at least one virtual image of the session message output in the indicated historical session fragment is displayed; In response to the selection of any historical session segment item, play the historical session segment indicated by the selected historical session segment item.
17. A conversation-based interactive device, characterized in that, The device includes: The conversation interface display module is used to display the conversation interface between the first user and the second user. A conversation message display module is used to display a first virtual avatar representing the first user in the conversation interface, and to output conversation messages generated by the first user in the conversation from the first virtual avatar. The text information input module is used to respond to the text input operation triggered by the second user and display the input text information; An interaction module is used to display a second virtual avatar representing the second user in the conversation interface, and for the second virtual avatar to interact with the first virtual avatar according to the text information.
18. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 16.
19. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 16.
20. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 16.