Voice interaction method and device, computer equipment, readable storage medium and program product

By receiving user voice in the voice interactive page and generating virtual object voice, and updating the model with feedback mechanism, the problem of single interaction mode of virtual object is solved, improving the richness of interaction and user cultivation.

CN120340490APending Publication Date: 2025-07-18XIAOHONGSHU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510702011.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the interaction between users and virtual objects is single, and the sense of cultivation is insufficient.

Method used

Provide a speech interaction method, which is to interact with the target virtual object by displaying a speech interaction page, receive user voice and generate virtual object voice, update the speech generation model based on feedback operations, and support multi-theme interaction and feedback mechanism.

Benefits of technology

It enhances the richness and fun of the interaction between users and virtual objects, improves the user's sense of cultivation of virtual objects, and optimizes the voice generation model through feedback operations to better meet user preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340490A_ABST
    Figure CN120340490A_ABST
Patent Text Reader

Abstract

The invention relates to a voice interaction method and device, computer equipment, a readable storage medium and a computer program product. The method comprises the steps of displaying a voice interaction page, wherein the voice interaction page is used for performing voice interaction with a target virtual object based on a first theme; after a first voice is received on the voice interaction page, generating a second voice of the target virtual object based on the first voice, and playing the second voice on the voice interaction page; and receiving a feedback operation on the second voice, and based on the feedback operation, generating a third voice of the target virtual object based on the first voice or performing voice interaction based on a second theme, the feedback operation being used for updating a voice generation model of the target virtual object. And the interaction mode is enriched.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technologies, and in particular, to a voice interaction method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art

[0002] Currently, a user can interact with a virtual object on a conversation page of the virtual object. For example, the user can input text, voice messages, etc. on the conversation page. However, this way of dialogue interaction is single and lacks a sense of cultivation. Summary of the Invention

[0003] Based on this, in view of the above technical problems, it is necessary to provide a voice interaction method, apparatus, computer device, computer-readable storage medium, and computer program product that can enrich the interaction method.

[0004] In a first aspect, the present application provides a voice interaction method, including:

[0005] Display a voice interaction page for voice interaction with a target virtual object based on a first theme;

[0006] After receiving a first voice on the voice interaction page, generate a second voice of the target virtual object based on the first voice, and play the second voice on the voice interaction page;

[0007] Receive a feedback operation on the second voice, and based on the feedback operation, generate a third voice of the target virtual object based on the first voice or perform voice interaction based on a second theme, where the feedback operation is used to update a voice generation model of the target virtual object.

[0008] In some embodiments, generating the second voice of the target virtual object based on the first voice includes: using the voice generation model to generate a voice with the same text content as the first voice as the second voice.

[0009] In some embodiments, based on the feedback operation, generating the third voice of the target virtual object based on the first voice or performing voice interaction based on a second theme includes: when the feedback operation is an operation satisfied with the second voice, performing voice interaction based on a second theme; when the feedback operation is an operation dissatisfied with the second voice, using the voice generation model again to generate a voice with the same text content as the first voice as the third voice; and playing the third voice on the voice interaction page.

[0010] In some embodiments, the voice interaction page displays interaction guidance information regarding a first topic. Before generating the second voice of the target virtual object based on the first voice, the method further includes: determining whether the first voice is related to the first topic; if so, generating the second voice of the target virtual object based on the first voice; if not, prompting the user to input a voice related to the first topic.

[0011] In some embodiments, the method further includes: displaying a conversation page with the target virtual object; after the conversation page first receives a conversation voice input by the user, displaying a voice reply message matching the type of the target virtual object on the conversation page; in response to a trigger operation on the voice reply message, playing the voice reply message and displaying a voice interaction prompt; in response to a trigger operation on the voice interaction prompt, displaying the voice interaction page; in response to an operation of returning from the voice interaction page to the conversation page, displaying a prompt at the location where the voice interaction page entry is located on the conversation page.

[0012] In some embodiments, the first topic and the second topic are sub-topics, and the first topic and the second topic belong to the same upper-level topic; the method further includes: after voice interactions for all sub-topics under the upper-level topic are completed, displaying an interaction summary page, which displays an interaction summary block, a content publishing control, and a content sharing control; the interaction summary block is used to display a summary of the voice interactions corresponding to all sub-topics; the content publishing control is used to publish the interaction summary block, and the content sharing control is used to share the interaction summary block.

[0013] In some embodiments, the method further includes: after voice interactions for all sub-topics under the upper-level topic are completed, if an operation to open the conversation page with the target virtual object is received, displaying an interaction record block in the form of a conversation record on the conversation page; in response to a trigger operation on the interaction record block, displaying an interaction record pop-up window, which is used to display the records of the voice interactions corresponding to all sub-topics, and the interaction record pop-up window is also used to display the interaction summary block.

[0014] In some embodiments, the display of the voice interaction page, which is used for voice interaction with a target virtual object based on a first theme, includes: in response to a trigger operation on the entrance of the voice interaction page, displaying an interaction progress page, which shows the current language level of the target virtual object, and the language level is determined based on the number of themes completed by the target virtual object; the interaction progress page also shows theme progress information, which includes statistical information on completed themes and statistical information on uncompleted themes; the interaction progress page also includes an interaction trigger control; in response to a trigger operation on the interaction trigger control, displaying a voice interaction page, which is used for voice interaction with the target virtual object based on the uncompleted theme.

[0015] In a second aspect, the present application also provides a voice interaction device, including:

[0016] A display module, configured to display a voice interaction page, which is used for voice interaction with a target virtual object based on a first theme;

[0017] A generation module, configured to generate a second voice of the target virtual object based on the first voice after the first voice is received on the voice interaction page, and play the second voice on the voice interaction page;

[0018] A feedback module, configured to receive a feedback operation on the second voice, and based on the feedback operation, generate a third voice of the target virtual object based on the first voice or perform voice interaction based on a second theme, and the feedback operation is used to update the voice generation model of the target virtual object.

[0019] In a third aspect, the present application also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0020] Display a voice interaction page, which is used for voice interaction with a target virtual object based on a first theme;

[0021] After the first voice is received on the voice interaction page, generate a second voice of the target virtual object based on the first voice, and play the second voice on the voice interaction page;

[0022] Receive a feedback operation on the second voice, and based on the feedback operation, generate a third voice of the target virtual object based on the first voice or perform voice interaction based on a second theme, and the feedback operation is used to update the voice generation model of the target virtual object.

[0023] Fourthly, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0024] Display a voice interaction page, which is used for voice interaction with a target virtual object based on a first theme;

[0025] After receiving a first voice on the voice interaction page, generate a second voice of the target virtual object based on the first voice, and play the second voice on the voice interaction page;

[0026] Receive a feedback operation on the second voice. Based on the feedback operation, generate a third voice of the target virtual object based on the first voice or perform voice interaction based on a second theme. The feedback operation is used to update the voice generation model of the target virtual object.

[0027] Fifthly, the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0028] Display a voice interaction page, which is used for voice interaction with a target virtual object based on a first theme;

[0029] After receiving a first voice on the voice interaction page, generate a second voice of the target virtual object based on the first voice, and play the second voice on the voice interaction page;

[0030] Receive a feedback operation on the second voice. Based on the feedback operation, generate a third voice of the target virtual object based on the first voice or perform voice interaction based on a second theme. The feedback operation is used to update the voice generation model of the target virtual object.

[0031] The above-mentioned voice interaction method, device, computer equipment, computer-readable storage medium and computer program product display a voice interaction page, the voice interaction page is used to perform voice interaction with the target virtual object based on the first theme; after the voice interaction page receives the first voice, the second voice of the target virtual object is generated based on the first voice, and the second voice is played on the voice interaction page; a feedback operation on the second voice is received, and based on the feedback operation, a third voice of the target virtual object is generated based on the first voice or voice interaction is performed based on the second theme, and the feedback operation is used to update the voice generation model of the target virtual object. The user can perform voice interaction with the target virtual object according to the preset theme, which increases the richness and fun of the interaction. Moreover, the user's feedback operation on the voice of the target virtual object can be used to update the voice generation model of the target virtual object, so that the voice output by the voice generation model of the target virtual object can increasingly meet the user's preferences and enhance the user's sense of cultivation of the target virtual object. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0033] Figure 1 A diagram of an application environment of a voice interaction method in an embodiment;

[0034] Figure 2 A flowchart of a voice interaction method in one embodiment;

[0035] Figure 3 A schematic diagram of a voice interaction page in one embodiment;

[0036] Figure 4 A schematic diagram of an operation of a user inputting voice in one embodiment;

[0037] Figure 5 A schematic diagram of displaying a voice interaction page during the process of generating a second voice in one embodiment;

[0038] Figure 6 is a schematic diagram of an interactive voice message in one embodiment;

[0039] Figure 7 A schematic diagram of a user performing a satisfactory operation in one embodiment;

[0040] Figure 8 is a schematic diagram of a voice interaction prompt 11 in an embodiment;

[0041] Figure 9 Schematic diagram for prompting the location of the voice interaction page entry in an embodiment;

[0042] Figure 10 Schematic diagram of the interaction summary page in an embodiment;

[0043] Figure 11 Schematic diagram of the interaction record pop-up window in an embodiment;

[0044] Figure 12 Schematic diagram of the interaction progress page in an embodiment;

[0045] Figure 13 Schematic diagram for real-time saving of interaction progress in an embodiment;

[0046] Figure 14 Schematic diagram of the interaction between the terminal and the server in an embodiment;

[0047] Figure 15 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0048] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0049] The voice interaction method provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or other network servers. The voice interaction method provided by the embodiments of the present application can be implemented by the cooperation of the terminal 102 and the server 104; it can also be implemented by the terminal 102 alone. The embodiments of the present application do not make any limitations in this regard.

[0050] Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0051] In an exemplary embodiment, as Figure 2 shown, a voice interaction method is provided. Taking the method applied to the Figure 1 terminal in it as an example for illustration, it includes the following steps 202 to 206. Among them:

[0052] Step 202, display a voice interaction page, where the voice interaction page is used to perform voice interaction with a target virtual object based on a first theme.

[0053] Optionally, an application (APP) is installed on the terminal. The voice interaction method provided in the embodiments of the present application can be executed by the APP. Exemplarily, the APP can be of the content sharing type, social type, shopping type, video type, etc. The embodiments of the present application do not make any limitations in this regard.

[0054] Optionally, the voice interaction page is a page for performing voice interaction with a target virtual object. The target virtual object can be a virtual character, a virtual pet, etc. The embodiments of the present application do not make any limitations in this regard.

[0055] Optionally, the terminal provides an entrance to the voice interaction page. When the user performs a trigger operation on the entrance to the voice interaction page, the terminal will display the voice interaction page.

[0056] Optionally, the entrance to the voice interaction page can be set in any page provided by the terminal. Exemplarily, the entrance to the voice interaction page can be set on the home page; or, as Figure 3 shown, the entrance to the voice interaction page can be set in the conversation page with the target virtual object. After the user performs a trigger operation on the entrance to the voice interaction page, the terminal displays the voice interaction page.

[0057] Optionally, the conversation page with the target virtual object is used to have a conversation with the target virtual object. The conversation page with the target virtual object supports the user and the target virtual object to input conversation messages in any form such as text, voice, image, video, file, etc. The embodiments of the present application do not make any limitations in this regard.

[0058] Optionally, a first theme is displayed on the voice interaction page, and the user can have a voice interaction with the target virtual object on the voice interaction page based on the first theme. Exemplarily, referring to Figure 3 as shown, the first theme can be, for example, "Greeting".

[0059] Optionally, interaction guiding information for the first theme is also displayed on the voice interaction page. The interaction guiding information can be an instruction or a question related to the first theme. Subsequently, the user can input corresponding voice according to the instruction or the question on the voice interaction page.

[0060] Exemplarily, referring to Figure 3 as shown, the first theme can be, for example, "Greeting", and the interaction guiding information for the first theme can be: I want to learn how to greet. Can you teach me? For example, "Hello".

[0061] Optionally, the first theme includes a plurality of voice interaction links with a preset specific order. The voice interaction page is used to have a voice interaction with the target virtual object for each voice interaction link in sequence according to this order. Each voice interaction link contains corresponding interaction guiding information. When having a voice interaction for a certain voice interaction link, the interaction guiding information corresponding to this voice interaction link and the first theme are displayed on the voice interaction page. Exemplarily, the first theme can be "Greeting", and the first theme includes two voice interaction links. The interaction guiding information for the first voice interaction link can be: I want to learn how to greet. Can you teach me? For example, "Hello". The interaction guiding information for the second voice interaction link can be: How do you usually greet your friends? I also want to learn.

[0062] Step 204, after receiving the first voice on the voice interaction page, generate a second voice of the target virtual object based on the first voice, and play the second voice on the voice interaction page.

[0063] Optionally, continuing to refer to Figure 3 as shown, the voice interaction page includes a voice input control 10. The user can input voice through the voice input control 10, and the terminal takes the received voice as the first voice. As described above, the first theme is displayed on the voice interaction page, and the user can input voice related to the first theme through the above-mentioned voice input control 10.

[0064] Further, when the interactive guidance information of the first theme is displayed on the voice interaction page, the user can input the voice guided by the interactive guidance information through the above-mentioned voice input control 10. For example, if the interactive guidance information is an instruction to say specific words (related to the first theme), the user can input the voice corresponding to these words through the above-mentioned voice input control 10. Another example is that if the interactive guidance information is a question (related to the first theme), the user can input the voice of the reply corresponding to the question through the above-mentioned voice input control 10.

[0065] Exemplarily, the user can complete voice input by speaking while holding down the voice input control 10. Specifically, the user holds down the voice input control 10 and simultaneously speaks something related to the first theme, and then releases the voice input control 10. The terminal takes the voice received during the period from when the voice input control 10 is held down to when it is released as the first voice. Among them, when the user holds down the voice input control 10, the display of the voice interaction page is as Figure 4 shown.

[0066] Optionally, after receiving the first voice, the terminal can generate a second voice of the target virtual object based on the first voice by using the voice generation model of the target virtual object.

[0067] Exemplarily, during the process of the terminal generating the second voice, the display of the voice interaction page is as Figure 5 shown, and at this time, the voice input control 10 is in a disabled state.

[0068] Among them, after obtaining the second voice, the terminal can directly play the second voice on the voice interaction page, or display an unplayed interactive voice message on the voice interaction page. The user can perform a trigger operation on the interactive voice message, and the terminal responds to the operation and then plays the second voice.

[0069] Optionally, as shown in Figure 6 shown, the above-mentioned interactive voice message includes a voice identifier and an unplayed status identifier. The unplayed status identifier is used to remind the user that this interactive voice message has not been played yet. The voice identifier includes the voice duration, and the voice duration is the duration of the second voice. The user can perform a trigger operation on the voice identifier, and the terminal responds to the operation, plays the second voice, and closes the above-mentioned unplayed status identifier.

[0070] Step 206, receive a feedback operation on the second voice, and based on the feedback operation, generate a third voice of the target virtual object based on the first voice or perform voice interaction based on the second theme. The feedback operation is used to update the voice generation model of the target virtual object.

[0071] Among them, the user can judge the second voice played by the terminal. If satisfied, the user can perform an operation indicating satisfaction with the second voice on the voice interaction page; if not satisfied, the user can perform an operation indicating dissatisfaction with the second voice on the voice interaction page. Both of these operations belong to the feedback operations on the second voice.

[0072] In a possible implementation manner, after obtaining the second voice, the terminal can directly play the second voice on the voice interaction page, and display a first control and a second control on the voice interaction page. Relevant text and / or patterns indicating dissatisfaction can be displayed on the first control, and relevant text and / or patterns indicating satisfaction can be displayed on the second control. When the user is not satisfied with the second voice, the user can perform a trigger operation on the first control; when the user is satisfied with the second voice, the user can perform a trigger operation on the second control.

[0073] In another possible implementation manner, continue to refer to Figure 6 As shown, after obtaining the second voice, the terminal first displays an unplayed interactive voice message on the voice interaction page, and at the same time displays a first control and a second control on the voice interaction page. However, at this time, the first control and the second control are in a disabled state. After the user performs a trigger operation on the interactive voice message, the terminal responds to this operation, plays the second voice, and switches the first control and the second control to an operable state. When the user is not satisfied with the second voice, the user can perform a trigger operation on the first control; when the user is satisfied with the second voice, the user can perform a trigger operation on the second control. Figure 6 In

[0074] Among them, after receiving the feedback operation, the terminal judges the feedback operation. If the feedback operation is an operation indicating satisfaction with the second voice, continue the voice interaction based on the second topic; multiple topics with a specific order can be preset, and the voice interaction for each topic can be carried out with the target virtual object in sequence according to this order. If the feedback operation is an operation indicating dissatisfaction with the second voice, regenerate the third voice of the target virtual object based on the first voice.

[0075] Among them, the feedback operation received by the terminal can be used to update the voice generation model of the target virtual object. Specifically, when the feedback operation is an operation indicating satisfaction with the second voice, the second voice can be used as a training positive sample to update the voice generation model of the target virtual object; when the feedback operation is an operation indicating dissatisfaction with the second voice, the second voice can be used as a training negative sample to update the voice generation model of the target virtual object.

[0076] Among them, the voice generation model of the target virtual object is used to convert text content into voice. Voice has characteristics such as timbre, intonation, rhythm, and speech rate. The performance of the voice generation model of the target virtual object in these aspects is related to the training samples. As described above, the user's feedback operation will generate a training sample, and the user's feedback operation will also determine the type of the training sample. Specifically, when the user performs an operation that is satisfied with the second voice, the second voice will be used as a positive training sample, and when the user performs an operation that is not satisfied with the second voice, the second voice will be used as a negative training sample. In this way, as the voice interaction process continues, the training samples increase accordingly. By continuously training the voice generation model of the target virtual object based on the training samples, the voice output by the voice generation model of the target virtual object can increasingly meet the user's preferences.

[0077] Optionally, after the terminal obtains the second voice, it can play the second voice. The user can judge whether the second voice played by the terminal is similar to their own speech in terms of timbre, intonation, rhythm, speech rate, etc. If so, the user performs an operation that is satisfied with the second voice, and the second voice will subsequently be used as a positive training sample to update the voice generation model of the target virtual object; if not, the user performs an operation that is not satisfied with the second voice, and the second voice will subsequently be used as a negative training sample to update the voice generation model of the target virtual object. Through continuous updates, the voice output by the voice generation model of the target virtual object will become more and more similar to the user's own speech in terms of timbre, intonation, rhythm, speech rate, etc.

[0078] In the above embodiment, a voice interaction page is displayed. The voice interaction page is used for voice interaction with the target virtual object based on the first theme; after receiving the first voice on the voice interaction page, the second voice of the target virtual object is generated based on the first voice, and the second voice is played on the voice interaction page; a feedback operation on the second voice is received, and based on the feedback operation, the third voice of the target virtual object is generated based on the first voice or voice interaction is performed based on the second theme. The feedback operation is used to update the voice generation model of the target virtual object. The user can perform voice interaction with the target virtual object according to the preset theme, which increases the interaction interest. Moreover, the user's feedback operation on the voice of the target virtual object can be used to update the voice generation model of the target virtual object, so that the voice output by the voice generation model of the target virtual object can increasingly meet the user's preferences, and the user's sense of cultivation of the target virtual object is enhanced.

[0079] In some embodiments, generating the second voice of the target virtual object based on the first voice includes the following steps: using the voice generation model to generate a voice with the same text content as the first voice as the second voice.

[0080] As described above, the voice generation model of the target virtual object can be used to convert text content into voice. After the terminal receives the first voice, it can perform text conversion on the first voice to obtain the text content of the first voice, and then input the text content into the voice generation model of the target virtual object. The voice generation model will convert the text content into voice output, and the voice output by the voice generation model can be used as the second voice of the target virtual object. The second voice is consistent with the text content of the first voice.

[0081] In the above embodiment, the voice generation model is used to generate a voice that is consistent with the text content of the first voice as the second voice. In this way, a voice interaction of "I say and you learn" is formed between the user and the target virtual object, further enhancing the sense of interaction.

[0082] In some embodiments, based on the feedback operation, generating the third voice of the target virtual object based on the first voice or performing voice interaction based on the second theme includes the following steps: when the feedback operation is an operation that is satisfied with the second voice, performing voice interaction based on the second theme; when the feedback operation is an operation that is not satisfied with the second voice, using the voice generation model again to generate a voice that is consistent with the text content of the first voice as the third voice; and playing the third voice on the voice interaction page.

[0083] Among them, after the terminal receives the feedback operation, it judges the feedback operation. If the feedback operation is an operation that is satisfied with the second voice, continue to perform voice interaction based on the second theme. That is, the second theme is displayed on the voice interaction page, and the user can perform voice interaction with the target virtual object on the voice interaction page based on the second theme. Similarly, the voice interaction page can display the interaction guiding information of the second theme, and the interaction guiding information can be an instruction or a question related to the second theme. Subsequently, the user can input the corresponding voice on the voice interaction page according to the instruction or the question. The processing process after the terminal receives the voice is similar to the processing process after receiving the first voice in the above embodiment, and reference can be made to the introduction in the foregoing embodiment, which will not be elaborated here.

[0084] Exemplarily, as shown in Figure 7 When the terminal plays the second voice and the user performs an operation that is satisfied with the second voice on the voice interaction page, the second theme and the interaction guiding information of the second theme are displayed on the voice interaction page. The second theme can be, for example, "talking about food", and the interaction guiding information of the second theme can be: Besides fish, what other delicious foods do you know. The user can input the corresponding voice through the voice input control 10 according to the interaction guiding information of the second theme.

[0085] Among them, after the terminal receives the feedback operation, it judges the feedback operation. If the feedback operation is an operation dissatisfied with the second voice, the voice generation model is used again to generate a voice that is the same as the content of the first voice text as the third voice; and the third voice is played on the voice interaction page. Similar to generating the second voice, during the process of the terminal generating the third voice, the display of the voice interaction page is as Figure 5 shown, and at this time the voice input control 10 is in a disabled state.

[0086] Optionally, the terminal can input the text content obtained by converting the first voice into the voice generation model of the target virtual object again. The voice generation model will convert the text content into the voice of the target virtual object again, and the voice can be used as the third voice. Since the internal processing process of the model is very complex, even if the text content input twice is the same, there will still be differences in the output voices, that is, the second voice and the third voice are usually different. The processing steps of the terminal after obtaining the third voice are similar to those after obtaining the second voice, and are not described in detail in this embodiment of the present application.

[0087] In the above embodiment, when the feedback operation is an operation satisfied with the second voice, voice interaction is performed based on the second theme; when the feedback operation is an operation dissatisfied with the second voice, the voice generation model is used again to generate a voice that is the same as the content of the first voice text as the third voice; and the third voice is played on the voice interaction page. Allowing users to evaluate the voice of the target virtual object improves the user's sense of control over the voice interaction process.

[0088] In some embodiments, the voice interaction page displays interactive guidance information related to the first theme. Before generating the second voice of the target virtual object based on the first voice, the method further includes the following steps: judging whether the first voice is related to the first theme; if so, generating the second voice of the target virtual object based on the first voice; if not, prompting the user to input a voice related to the first theme.

[0089] Among them, the voice interaction page displays interactive guidance information related to the first theme. The interactive guidance information related to the first theme can be an instruction or a question related to the first theme. For example, the first theme can be "greeting", and the interactive guidance information of the first theme can be: I want to learn how to greet~Can you teach me~For example, "Hello", the interactive guidance information of the first theme can also be: How do you usually greet your friends~I also want to learn.

[0090] After receiving the first voice, the terminal can convert the first voice into text to obtain the text content of the first voice, and then determine whether the text content is relevant to the first theme.

[0091] In a possible implementation, each theme corresponds to a related word library. The similarity between the text content of the first voice and each word in the related word library corresponding to the first theme can be calculated. If the similarity is greater than or equal to a preset threshold, it is determined that the first voice is relevant to the first theme. If it is less than the preset threshold, it is determined that the first voice is not relevant to the first theme.

[0092] In a possible implementation, the first theme, the text content of the first voice, and the question of whether they are relevant can be input into a large model, and the large model is used to determine whether the first voice is relevant to the first theme.

[0093] If the judgment result is that the first voice is relevant to the first theme, the second voice of the target virtual object is generated based on the first voice. The specific generation method and the processing process after obtaining the second voice are as described above. If the judgment result is that the first voice is not relevant to the first theme, the user is prompted to input a voice relevant to the first theme.

[0094] Optionally, the terminal can display a prompt pop-up window on the voice interaction page, and display a prompt message in the prompt pop-up window to prompt the user to input a voice relevant to the first theme.

[0095] In the above embodiment, after receiving the first voice, it can be first determined whether the first voice is relevant to the first theme; if so, the second voice of the target virtual object is generated based on the first voice; if not, the user is prompted to input a voice relevant to the first theme. This ensures the relevance between the user input voice and the first theme.

[0096] In some embodiments, the voice interaction method provided by the embodiments of the present application further includes: displaying a conversation page of the target virtual object; after the conversation voice input by the user is first received on the conversation page, displaying a voice reply message matching the type of the target virtual object on the conversation page; in response to a trigger operation on the voice reply message, playing the voice reply message and displaying a voice interaction prompt 11; in response to a trigger operation on the voice interaction prompt 11, displaying the voice interaction page; in response to an operation of returning from the voice interaction page to the conversation page, displaying a prompt at the position where the voice interaction page entry is located on the conversation page.

[0097] Among them, the terminal provides a session page entry. When the user performs a triggering operation on the session page entry, the terminal will display the session page of the target virtual object. As described above, the session page of the target virtual object is used to have a conversation with the target virtual object. The session page of the target virtual object supports the user and the target virtual object to input conversation messages in any form such as text, voice, image, video, file, etc.

[0098] Among them, as shown in Figure 8 After the session page first receives the user's input session voice, that is, after the session page first receives the user's input conversation message in voice form, a voice reply message matching the type of the target virtual object is displayed on the session page. The type of the target virtual object can be a person, a dog, a cat, etc., which can be freely selected by the user. When the type of the target virtual object is an animal, the voice reply message is the original sound of the animal. Exemplarily, when the type of the target virtual object is a cat, the voice reply message can be: Meow~. After the user performs a triggering operation on the voice reply message, the terminal responds to this operation, plays the voice reply message and displays the voice interaction prompt 11. After the user performs a triggering operation on the voice interaction prompt 11, the terminal responds to this operation and will display the voice interaction page.

[0099] Among them, when the user enters the voice interaction page by triggering Figure 8 the voice interaction prompt 11 in, as shown in Figure 9 if the user performs an operation to return to the session page on the voice interaction page, the terminal will display the session page and display a prompt at the location of the voice interaction page entry on the session page. This prompt can be displayed through a floating layer. When the user clicks on any position of the floating layer, the prompt will close.

[0100] In the above embodiments, displaying the session page of the target virtual object; after the session page first receives the user's input session voice, displaying a voice reply message matching the type of the target virtual object on the session page; in response to the triggering operation on the voice reply message, playing the voice reply message and displaying the voice interaction prompt 11, realizing the popularization of the voice interaction function; in response to the triggering operation on the voice interaction prompt 11, displaying the voice interaction page; in response to the operation of returning from the voice interaction page to the session page, displaying a prompt at the location of the voice interaction page entry on the session page, which is convenient for the user to know where the entry is when they want to use the voice interaction function next time, and improves the user operation efficiency.

[0101] In some embodiments, the first topic and the second topic are sub-topics, and the first topic and the second topic belong to the same upper-level topic; the method further includes the following steps: after voice interactions for all sub-topics under the upper-level topic are completed, an interaction summary page is displayed, and the interaction summary page displays an interaction summary block 12, a content publishing control 13, and a content sharing control 14; the interaction summary block 12 is used to display summary content of the voice interactions corresponding to all sub-topics; the content publishing control 13 is used to publish the interaction summary block 12, and the content sharing control 14 is used to share the interaction summary block 12.

[0102] Among them, the first topic and the second topic are sub-topics, and the first topic and the second topic belong to the same upper-level topic. There are N pre-set sub-topics with a specific order under this upper-level topic. The voice interaction page is used to perform voice interactions for each topic with the target virtual object in sequence according to this order. The voice interaction process for each topic can be referred to the above description. After voice interactions for all sub-topics under the upper-level topic are completed, the terminal can display an interaction summary page, and the interaction summary page displays an interaction summary block 12, a content publishing control 13, and a content sharing control 14.

[0103] Among them, the interaction summary block 12 is used to display summary content of the voice interactions corresponding to all sub-topics. This summary content can, for example, include the number of times the user performs a satisfactory operation and the number of times the user performs an unsatisfactory operation during the voice interactions corresponding to all sub-topics, and can also include the common points of the voices corresponding to all satisfactory operations, the common points of the voices corresponding to all unsatisfactory operations, etc. The embodiments of the present application do not limit the summary content.

[0104] Among them, the user can perform a trigger operation on the content publishing control 13. In response to this trigger operation, the terminal can display a content editing page, and the content editing page displays the above-mentioned interaction summary block 12. The user can add a title, text, etc. to the interaction summary block 12 on the content editing page, and then click the publish button on the content editing page. The terminal will generate publish content based on the information on the content editing page for publishing processing.

[0105] Among them, the user can perform a trigger operation on the content sharing control 14. In response to this trigger operation, the terminal can display a list of objects related to the currently logged-in object. The user can select a target object from this list of objects, and the terminal will send the above-mentioned interaction summary block 12 to the target object. The target object can view the interaction summary block 12 in the conversation page with the currently logged-in object. The above-mentioned related relationship can be, for example, a following relationship, a being-followed relationship, etc. The embodiments of the present application do not limit this.

[0106] Exemplarily, the upper-level theme is, for example, "Greetings". Under this upper-level theme, there are 3 pre-set sub-themes in a specific order, namely the first theme, the second theme, and the third theme. The first theme can be, for example, greeting a pet, the second theme can be, for example, greeting a friend, and the third theme can be, for example, greeting a teacher. Refer to Figure 10 As shown, after all sub-themes under the upper-level theme have completed voice interaction, the terminal can display an interaction summary page, which shows an interaction summary block 12, a content publishing control 13, and a content sharing control 14. The interaction summary page also includes a return control and a continue interaction control. After the user triggers the return control, the terminal can jump to the page where the voice interaction page entry is located. After the user triggers the continue interaction control, the terminal can display the voice interaction page and continue the voice interaction of the next upper-level theme on the voice interaction page.

[0107] In the above embodiment, after all sub-themes under the upper-level theme have completed voice interaction, an interaction summary page is displayed, which shows an interaction summary block 12, a content publishing control 13, and a content sharing control 14; the user can view the summary content of all sub-themes under the current upper-level theme in the interaction summary block 12, and can also publish the interaction summary block 12 through the content publishing control 13 and share the interaction summary block 12 with friends through the content sharing control 14, improving the user's interaction experience.

[0108] In some embodiments, the voice interaction method provided by the embodiments of the present application further includes the following steps: after all sub-themes under the upper-level theme have completed voice interaction, if an operation to open a conversation page with the target virtual object is received, an interaction record block 15 is displayed in the conversation page in the form of a conversation record; in response to a trigger operation on the interaction record block 15, an interaction record pop-up window is displayed, and the interaction record pop-up window is used to display the records of the voice interactions corresponding to all sub-themes, and the interaction record pop-up window is also used to display the interaction summary block 12.

[0109] Among them, refer to Figure 11 As shown, after all sub-themes under the upper-level theme have completed voice interaction, if the user performs an operation to open a conversation page with the target virtual object, in response to this operation, the terminal will display the conversation page with the target virtual object, and an interaction record block 15 will be displayed in the conversation page in the form of a conversation record. After the user performs a trigger operation on the interaction record block 15, the terminal will display an interaction record pop-up window, and the interaction record pop-up window displays the records of the voice interactions corresponding to all sub-themes under the upper-level theme and the interaction summary block 12.

[0110] In the above embodiments, after all sub - topics under the upper - level topic have completed voice interaction, the interaction record block 15 will be presented in the conversation page in the form of a conversation record, enabling the user to view the records of voice interactions corresponding to each sub - topic on the conversation page, further enhancing the interaction experience.

[0111] In some embodiments, the display of the voice interaction page, which is used for voice interaction with a target virtual object based on a first topic, includes: in response to a trigger operation on the voice interaction page entry, displaying an interaction progress page, which shows the current language level of the target virtual object, and the language level is determined based on the number of topics completed by the target virtual object; the interaction progress page also shows topic progress information, and the topic progress information includes statistical information on completed topics and statistical information on uncompleted topics; the interaction progress page also includes an interaction trigger control 16; in response to a trigger operation on the interaction trigger control 16, displaying a voice interaction page, which is used for voice interaction with the target virtual object based on the uncompleted topics.

[0112] As described above, the terminal provides a voice interaction page entry, and the voice interaction page entry can be set on any page provided by the terminal, such as the home page or the conversation page with the target virtual object. After the user performs a trigger operation on the voice interaction page entry, the terminal responds to the trigger operation and first displays an interaction progress page, which shows the current language level of the target virtual object, and the language level is determined based on the number of topics completed by the target virtual object.

[0113] Optionally, when the user triggers Figure 10 the return control in, the terminal can jump to this interaction progress page.

[0114] Optionally, assuming that the total number of all topics is N and there are M language levels in total, a linear mapping relationship between the number of topics and the language levels can be established according to N and M. Whenever the number of topics completed by the target virtual object increases, the latest language level is calculated based on the increased number and the above - mentioned linear mapping relationship, and the language level on the interaction progress page is updated to this latest language level.

[0115] Optionally, the interaction progress page also shows topic progress information. Specifically, the topic progress information includes statistical information on completed topics and statistical information on uncompleted topics.

[0116] Optionally, the statistical information on completed topics can include the number of completed topics and the percentage of the total number of all topics, and the statistical information on uncompleted topics can include the number of uncompleted topics and the percentage of the total number of all topics.

[0117] Optionally, an interaction trigger control 16 is displayed on the interaction progress page. Since multiple preset themes have a specific order, after the user performs a trigger operation on the interaction trigger control 16, the terminal can display a voice interaction page. At this time, the theme displayed on the voice interaction page is the first unfinished theme determined based on the above order.

[0118] In a possible implementation, multiple upper-level themes with a specific order are preset, and each upper-level theme includes multiple sub-themes with a specific order. The multiple upper-level themes can be displayed on the interaction progress page. For each upper-level theme, the number of sub-themes that have completed voice interaction under this upper-level theme can be obtained, and by dividing this number by the total number of all sub-themes under this upper-level theme, the completed percentage corresponding to this upper-level theme can be obtained, and this percentage can be displayed under this upper-level theme on the interaction progress page. After the user performs a trigger operation on the interaction trigger control 16, in response to this operation, the terminal will display a voice interaction page. At this time, the theme displayed on the voice interaction page is the first unfinished sub-theme among the upper-level themes that have not been fully completed in voice interaction.

[0119] Exemplarily, see Figure 12 As shown, after the user performs a trigger operation on the voice interaction page entry, the terminal responds to this trigger operation and displays an interaction progress page. The interaction progress page displays the current language level of the target virtual object. It also displays three upper-level themes: "Say Hello", "Talk about Food", and "Talk about Interests". Among them, the completed percentage corresponding to "Say Hello" is 30%, and the completed percentages corresponding to "Talk about Food" and "Talk about Interests" are 0%. After the user performs a trigger operation on the interaction trigger control 16, in response to this operation, the terminal will display a voice interaction page. At this time, the theme displayed on the voice interaction page is the first unfinished sub-theme in "Say Hello".

[0120] Optionally, see Figure 13 As shown, each time the user exits the voice interaction page, the terminal can pop up a dialog box asking whether to confirm the exit. When the user triggers the exit in this dialog box, the terminal saves the voice interaction progress of the target virtual object and displays the latest theme completion status on the interaction progress page. In this way, when the user clicks to enter the voice interaction page next time, the theme displayed by the terminal is the theme at the time of exit.

[0121] In the above embodiments, an interaction progress page is provided. The user can view the current language level of the target virtual object and the theme completion status on this interaction progress page, and through the interaction trigger control 16, continue the voice interaction following the unfinished theme, which can avoid repeating the interaction based on a certain theme and further improve the user's interaction experience.

[0122] In some embodiments, an interaction method between a terminal and a server is provided. See Figure 14As shown, the method includes: S1. The terminal displays a voice interaction page. S2. The user inputs a first voice on the voice interaction page. S3. The terminal sends the first voice to the server. S4. The server generates a second voice of the target virtual object based on the first voice. S5. The server sends the second voice to the terminal. S6. The terminal plays the second voice. S7. The user inputs a feedback operation for the second voice. S8. The terminal, based on the feedback operation, generates a third voice of the target virtual object based on the first voice or conducts a voice interaction based on a second theme.

[0123] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the indication of the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages, and these steps or stages do not necessarily need to be executed at the same moment, but can be executed at different moments, and the execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0124] Based on the same inventive concept, an embodiment of the present application also provides a voice interaction device for implementing the above-mentioned voice interaction method. The implementation solutions provided by this device for solving problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the following voice interaction devices can refer to the limitations on the voice interaction method in the above text, and will not be elaborated here.

[0125] In an exemplary embodiment, a voice interaction device is provided, including:

[0126] A display module, configured to display a voice interaction page, and the voice interaction page is used to conduct a voice interaction with a target virtual object based on a first theme;

[0127] A generation module, configured to, after receiving a first voice on the voice interaction page, generate a second voice of the target virtual object based on the first voice, and play the second voice on the voice interaction page;

[0128] A feedback module, configured to receive a feedback operation for the second voice, and based on the feedback operation, generate a third voice of the target virtual object based on the first voice or conduct a voice interaction based on a second theme, and the feedback operation is used to update the voice generation model of the target virtual object.

[0129] In some embodiments, a generation module is configured to: generate speech that is consistent with the first speech text content by using the speech generation model as the second speech.

[0130] In some embodiments, a feedback module is configured to: when the feedback operation is an operation that is satisfied with the second speech, perform voice interaction based on the second theme; when the feedback operation is an operation that is not satisfied with the second speech, generate speech that is consistent with the first speech text content again by using the speech generation model as the third speech; and play the third speech on the voice interaction page.

[0131] In some embodiments, the voice interaction page displays interaction guidance information related to the first theme, and a generation module is configured to: determine whether the first speech is related to the first theme; if so, generate the second speech of the target virtual object based on the first speech; if not, prompt the user to input speech related to the first theme.

[0132] In some embodiments, the feedback module is further configured to: display a conversation page with the target virtual object; after receiving the conversation speech input by the user for the first time on the conversation page, display a voice reply message that matches the type of the target virtual object on the conversation page; in response to a trigger operation on the voice reply message, play the voice reply message and display a voice interaction prompt 11; in response to a trigger operation on the voice interaction prompt 11, display the voice interaction page; in response to an operation of returning from the voice interaction page to the conversation page, display a prompt at the location where the voice interaction page entry is located on the conversation page.

[0133] In some embodiments, the first theme and the second theme are sub-themes, and the first theme and the second theme belong to the same upper-level theme; a generation module is configured to: after voice interaction for all sub-themes under the upper-level theme is completed, display an interaction summary page, and the interaction summary page displays an interaction summary block 12, a content publishing control 13, and a content sharing control 14; the interaction summary block 12 is used to display the summary content of the voice interaction corresponding to all sub-themes; the content publishing control 13 is used to publish the interaction summary block 12, and the content sharing control 14 is used to share the interaction summary block 12.

[0134] In some embodiments, a generation module is configured to: after voice interaction for all sub-themes under the upper-level theme is completed, if an operation of opening a conversation page with the target virtual object is received, display an interaction record block 15 in the form of a conversation record on the conversation page; in response to a trigger operation on the interaction record block 15, display an interaction record pop-up window, and the interaction record pop-up window is used to display the records of the voice interaction corresponding to all sub-themes, and the interaction record pop-up window is further used to display the interaction summary block 12.

[0135] In some embodiments, a display module is configured to: in response to a trigger operation on an entry of a voice interaction page, display an interaction progress page, where the interaction progress page displays the current language level of the target virtual object, and the language level is determined based on the number of topics completed by the target virtual object; the interaction progress page further displays topic progress information, and the topic progress information includes statistical information on completed topics and statistical information on uncompleted topics; the interaction progress page further includes an interaction trigger control 16; in response to a trigger operation on the interaction trigger control 16, display a voice interaction page, and the voice interaction page is used to perform voice interaction with the target virtual object based on the uncompleted topics.

[0136] Each module in the above voice interaction device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0137] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 15 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program stored in the non-volatile storage medium to run. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. The computer program, when executed by the processor, implements a voice interaction method. The display unit of the computer device is configured to form a visually visible picture, and may be a display screen, a projection device, or a virtual reality imaging device. The display screen may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0138] Those skilled in the art can understand thatFigure 15 The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0139] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0140] Display a voice interaction page for voice interaction with a target virtual object based on a first theme;

[0141] After receiving a first voice on the voice interaction page, generate a second voice of the target virtual object based on the first voice, and play the second voice on the voice interaction page;

[0142] Receive a feedback operation on the second voice. Based on the feedback operation, generate a third voice of the target virtual object based on the first voice or perform voice interaction based on a second theme. The feedback operation is used to update the voice generation model of the target virtual object.

[0143] In one embodiment, when the processor executes the computer program, the following steps are further implemented: Generate a voice consistent with the text content of the first voice using the voice generation model as the second voice.

[0144] In one embodiment, when the processor executes the computer program, the following steps are further implemented: When the feedback operation is an operation satisfied with the second voice, perform voice interaction based on a second theme; when the feedback operation is an operation not satisfied with the second voice, generate a voice consistent with the text content of the first voice again using the voice generation model as the third voice; and play the third voice on the voice interaction page.

[0145] In one embodiment, the voice interaction page displays interactive guidance information about the first theme. When the processor executes the computer program, the following steps are further implemented: Determine whether the first voice is related to the first theme; if so, generate a second voice of the target virtual object based on the first voice; if not, prompt the user to input a voice related to the first theme.

[0146] In one embodiment, when the processor executes the computer program, the following steps are further implemented: displaying a session page of the target virtual object; after receiving the session voice input by the user on the session page for the first time, displaying a voice reply message matching the type of the target virtual object on the session page; in response to a trigger operation on the voice reply message, playing the voice reply message and displaying a voice interaction prompt; in response to a trigger operation on the voice interaction prompt, displaying the voice interaction page; in response to an operation of returning from the voice interaction page to the session page, displaying a prompt for the location where the voice interaction page entry is located on the session page.

[0147] In one embodiment, the first theme and the second theme are sub-themes, and the first theme and the second theme belong to the same upper-level theme; when the processor executes the computer program, the following steps are further implemented: after all sub-themes under the upper-level theme have completed voice interaction, displaying an interaction summary page, where the interaction summary page displays an interaction summary block, a content publishing control, and a content sharing control; the interaction summary block is used to display the summary content of the voice interaction corresponding to all sub-themes; the content publishing control is used to publish the interaction summary block, and the content sharing control is used to share the interaction summary block.

[0148] In one embodiment, when the processor executes the computer program, the following steps are further implemented: after all sub-themes under the upper-level theme have completed voice interaction, if an operation of opening the session page of the target virtual object is received, displaying an interaction record block in the form of a session record on the session page; in response to a trigger operation on the interaction record block, displaying an interaction record pop-up window, where the interaction record pop-up window is used to display the records of the voice interaction corresponding to all sub-themes, and the interaction record pop-up window is also used to display the interaction summary block.

[0149] In one embodiment, when the processor executes the computer program, the following steps are further implemented: in response to a trigger operation on the voice interaction page entry, displaying an interaction progress page, where the interaction progress page displays the current language level of the target virtual object, and the language level is determined based on the number of themes completed by the target virtual object; the interaction progress page also displays theme progress information, where the theme progress information includes statistical information on the completed themes and statistical information on the uncompleted themes; the interaction progress page also includes an interaction trigger control; in response to a trigger operation on the interaction trigger control, displaying a voice interaction page, where the voice interaction page is used to perform voice interaction with the target virtual object based on the uncompleted themes.

[0150] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0151] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the steps in the above-mentioned method embodiments.

[0152] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0153] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., and are not limited thereto.

[0154] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope recorded in this application.

[0155] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.

Claims

1. A voice interaction method, characterized in that, The method includes: Displaying a voice interaction page for voice interaction with a target virtual object based on a first theme; After receiving a first voice on the voice interaction page, generating a second voice of the target virtual object based on the first voice and playing the second voice on the voice interaction page; Receiving a feedback operation on the second voice, and based on the feedback operation, generating a third voice of the target virtual object based on the first voice or performing voice interaction based on a second theme, where the feedback operation is used to update the voice generation model of the target virtual object.

2. The method according to claim 1, wherein The generating the second voice of the target virtual object based on the first voice includes: Using the voice generation model to generate a voice with the same text content as the first voice as the second voice.

3. The method according to claim 2, wherein The generating the third voice of the target virtual object based on the first voice or performing voice interaction based on the second theme based on the feedback operation includes: When the feedback operation is an operation of being satisfied with the second voice, performing voice interaction based on the second theme; When the feedback operation is an operation of being dissatisfied with the second voice, using the voice generation model again to generate a voice with the same text content as the first voice as the third voice; and playing the third voice on the voice interaction page.

4. The method according to claim 1, wherein The voice interaction page displays interaction guidance information about the first theme. Before generating the second voice of the target virtual object based on the first voice, the method further includes: Determining whether the first voice is related to the first theme; If so, generating the second voice of the target virtual object based on the first voice; If not, prompting the user to input a voice related to the first theme.

5. The method according to claim 1, wherein The method further includes: Displaying a conversation page with the target virtual object; After receiving the conversation voice input by the user for the first time on the conversation page, displaying a voice reply message matching the type of the target virtual object on the conversation page; In response to a trigger operation on the voice reply message, playing the voice reply message and displaying a voice interaction prompt; In response to a trigger operation on the voice interaction prompt, displaying the voice interaction page; In response to an operation of returning from the voice interaction page to the conversation page, displaying a prompt for the location where the voice interaction page entrance is located on the conversation page.

6. The method according to claim 1, wherein The first theme and the second theme are sub-themes, and the first theme and the second theme belong to the same upper theme; the method further includes: After voice interaction for all sub-themes under the upper theme is completed, displaying an interaction summary page, where the interaction summary page displays an interaction summary block, a content publishing control, and a content sharing control; the interaction summary block is used to display the summary content of the voice interaction corresponding to all sub-themes; the content publishing control is used to publish the interaction summary block, and the content sharing control is used to share the interaction summary block.

7. The method according to claim 6, wherein The method further includes: After all sub - topics under the upper - layer topic have completed voice interaction, if an operation to open the conversation page with the target virtual object is received, an interaction record block is displayed in the conversation page in the form of a conversation record. In response to a trigger operation on the interaction record block, an interaction record pop - up window is displayed. The interaction record pop - up window is used to display the records of voice interactions corresponding to all sub - topics, and the interaction record pop - up window is also used to display the interaction summary block.

8. The method according to claim 1, wherein The display of the voice interaction page, which is used to perform voice interaction with the target virtual object based on the first topic, includes: In response to a trigger operation on the voice interaction page entry, an interaction progress page is displayed. The interaction progress page shows the current language level of the target virtual object, which is determined based on the number of topics completed by the target virtual object; the interaction progress page also shows topic progress information, which includes statistical information on completed topics and statistical information on uncompleted topics; the interaction progress page also includes an interaction trigger control. In response to a trigger operation on the interaction trigger control, a voice interaction page is displayed, which is used to perform voice interaction with the target virtual object based on uncompleted topics.

9. A voice interaction device, characterized by The device includes: A display module, which is used to display a voice interaction page for performing voice interaction with the target virtual object based on the first topic. A generation module, which is used to generate a second voice of the target virtual object based on the first voice after the first voice is received on the voice interaction page, and play the second voice on the voice interaction page. A feedback module, which is used to receive a feedback operation on the second voice, and based on the feedback operation, generate a third voice of the target virtual object based on the first voice or perform voice interaction based on the second topic. The feedback operation is used to update the voice generation model of the target virtual object.

10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 8.