Information interaction method and apparatus, and computer device and storage medium

By directly displaying text information after voice input and generating response results in real time, the problem of low efficiency in voice interaction is solved, and more efficient information exchange is achieved.

WO2025031338A9PCT designated stage expired Publication Date: 2026-04-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-06
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Current methods of information interaction via voice input are inefficient, involving multiple steps such as voice input, conversion to text, user confirmation, and intelligent agent generating responses, which are cumbersome and time-consuming.

Method used

After obtaining the user's voice input, the system directly displays the corresponding text information without requiring secondary confirmation from the user, and uses a content generation model to generate real-time response results, simplifying the operation process.

Benefits of technology

It improves the efficiency of information exchange, reduces user waiting time, and provides faster interactive responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024109991_02042026_PF_FP_ABST
    Figure CN2024109991_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present disclosure are an information interaction method and apparatus, and a computer device and a storage medium. The method comprises: presenting an information interaction page; in response to a first trigger operation for a target control in the information interaction page, receiving voice information, which is inputted by a user; in response to the end of the input of the voice information, presenting, on the information interaction page, an information sending result after information is sent to a target object, wherein the information sending result comprises first text information corresponding to the voice information; and presenting, on the information interaction page, a first real-time reply result, which is fed back by the target object, wherein the first real-time reply result is generated in real time by using a content generation model and on the basis of the voice information. In the present disclosure, after voice information which is inputted by a user is acquired, an information sending result which comprises corresponding text information is directly presented to the user without secondary confirmation by the user, and during the process of the user inputting the voice information, a content generation model is used to generate a real-time reply result in real time on the basis of the voice information, thereby improving the interaction efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

An information interaction method and device, computer device, and storage medium

[0001] The present application claims priority to Chinese Patent Application No. 202310988669.5, filed on August 7, 2023, the disclosure of which is incorporated herein in its entirety as part of the present application. TECHNICAL FIELD

[0002] The present disclosure relates to an information interaction method and device, computer device, and storage medium. BACKGROUND

[0003] As an important branch of artificial intelligence technology, intelligent question answering is widely applied to various information interaction scenarios, such as information search, intelligent chat, intelligent customer service, etc. An artificial intelligence (AI) body can automatically generate corresponding reply content according to the input information of a user. Voice input is a convenient input method adopted by many users. However, the current information interaction method through voice input has the problem of low interaction efficiency.

[0004] SUMMARY

[0005] The present disclosure provides at least an information interaction method and device, computer device, and storage medium.

[0006] In a first aspect, the present disclosure provides an information interaction method, comprising:

[0007] displaying an information interaction page; the information interaction page is used for information interaction with a target object;

[0008] in response to a first trigger operation on a target control in the information interaction page, receiving voice information input by a user;

[0009] in response to the end of the voice information input, displaying an information sending result after information sending to the target object on the information interaction page; the information sending result includes first text information corresponding to the voice information;

[0010] displaying a first real-time reply result fed back by the target object on the information interaction page; wherein the first real-time reply result is generated in real time based on the voice information by using a content generation model.

[0011] In this way, after obtaining the voice information input by the user, without the user confirming again, the information sending result including the text information corresponding to the voice information is directly displayed to the user, and in the process of inputting the voice information by the user, the real-time reply result is generated based on the voice information in real time by using the content generation model; and then, the real-time reply result can be displayed to the user without consuming much time, and the interaction efficiency is improved.

[0012] In a possible implementation, the method further includes:

[0013] In response to a second trigger operation on the information sending result, a modification control of the information sending result is displayed, and the first text information is displayed on the modification control.

[0014] In response to a modification operation on the first text information displayed on the modification control, second text information is generated.

[0015] The first text information displayed on the information interaction page is replaced with the second text information.

[0016] In this way, the modification control can be displayed by directly triggering the information sending result, the second text information can be obtained by directly modifying the first text information by using the modification control, and then the first text information displayed on the information interaction page is replaced with the second text information, so that the user can modify the first text information without performing a withdrawal operation on the first text information, and the user does not need to input new text information, that is, the user can conveniently modify the first text information, and the operation of the user is facilitated.

[0017] In a possible implementation, the method further includes:

[0018] Recommended modification information corresponding to the first text information is displayed; the recommended modification information includes at least one of the following: indication information of a recommended modification word in the first text information, a candidate word for replacing at least part of the content in the first text information, and text information obtained by content extraction on the first text information.

[0019] The modification operation includes a third trigger operation on the recommended modification information.

[0020] In this way, by displaying the recommended modification information, the user can be further facilitated to modify the first text information, the operation of the user is facilitated, and the modification efficiency is improved.

[0021] In a possible implementation, after the second text information is generated, the method further includes:

[0022] training a target model based on the voice information and the second text information; the target model is used to convert the voice information into corresponding text information in real time.

[0023] In this way, the target model can be trained for each user according to different pronunciation of each user, so as to improve the accuracy of converting voice information into text information.

[0024] In a possible implementation, the first real-time reply result belongs to at least part of the first reply result corresponding to the first text information.

[0025] The method further includes:

[0026] In response to receiving the second trigger operation on the information sending result in a case where the first real-time reply result does not contain all results of the first reply result, suspending updating the first real-time reply result.

[0027] In this way, in a case where the user gives up modifying the first text content, the first real-time reply result can continue to be updated according to the previous update progress of the first real-time reply result. In a case where the user confirms modifying the first text content into the second text content, the first real-time reply result is no longer updated, and the second real-time reply result fed back by the target object based on the second text content is directly displayed to the user, so as to improve the user experience in the reply result updating process.

[0028] In a possible implementation, after the first text information displayed on the information interaction page is replaced with the second text information, the method further includes:

[0029] displaying, on the information interaction page, a second real-time reply result fed back by the target object based on the second text information; the second real-time reply result belongs to at least part of a second reply result corresponding to the second text information.

[0030] In this way, the second real-time reply result generated after the reply result is modified along with the modification of the first text information can be displayed to the user.

[0031] In a possible implementation, after the first text information displayed on the information interaction page is replaced with the second text information, the method further includes:

[0032] deleting the first real-time reply result displayed on the information interaction page.

[0033] In this way, after the first text information is modified, the first real-time reply result corresponding to the first text information can be deleted from the information interaction page, only information useful to the user is retained in the information interaction page, and unnecessary information is avoided from interfering with the user.

[0034] In a possible implementation, the displaying, on the information interaction page, of the second real-time reply result fed back by the target object based on the second text information comprises:

[0035] In response to the second real-time reply result being the same as the first real-time reply result, the first real-time reply result displayed on the information interaction page is continued to be displayed, and the other reply results in the second reply result except the first real-time reply result are continued to be displayed.

[0036] In this way, the continuity of the reply result update is ensured, and the user experience is improved.

[0037] In a possible implementation, the method further comprises:

[0038] In response to the voice playing control displayed on the information interaction page being in a use state, the first real-time reply result is played based on the tone corresponding to the target object.

[0039] In response to the second trigger operation on the information sending result, the playing of the first real-time reply result is paused.

[0040] In this way, the first real-time reply result can be played based on the tone corresponding to the target object by using the voice playing control, and more interaction modes are provided for the user.

[0041] In a second aspect, the present disclosure also provides an information interaction apparatus, comprising:

[0042] a display module configured to display an information interaction page; the information interaction page is configured to perform information interaction with a target object;

[0043] a receiving module configured to, in response to a first trigger operation on a target control in the information interaction page, receive voice information input by a user;

[0044] the display module is further configured to, in response to the voice information input being ended, display, on the information interaction page, an information sending result after information is sent to the target object; the information sending result comprises the first text information; and the display module is further configured to display, on the information interaction page, a first real-time reply result fed back by the target object; the first real-time reply result is generated by a content generation model based on the voice information in real time.

[0045] In a possible implementation, the apparatus further comprises a modification module configured to:

[0046] in response to a second triggering operation on the information sending result, displaying a modification control on the information sending result, and displaying the first text information in the modification control;

[0047] in response to a modification operation on the first text information displayed in the modification control, generating second text information;

[0048] The display module is further configured to replace the first text information displayed on the information interaction page with the second text information.

[0049] In a possible implementation, the display module is further configured to:

[0050] display recommended modification information corresponding to the first text information; the recommended modification information includes at least one of the following: indication information of a recommended modification word in the first text information, a candidate word for replacing at least part of the content in the first text information, and text information obtained by content extraction on the first text information;

[0051] The modification operation includes a third triggering operation on the recommended modification information.

[0052] In a possible implementation, the method further includes: a training module configured to, after the second text information is generated, train a target model based on the voice information and the second text information; the target model is used to convert the voice information into corresponding text information in real time.

[0053] In a possible implementation, the first real-time reply result belongs to at least part of the first reply result corresponding to the first text information.

[0054] The display module is further configured to:

[0055] in response to receiving the second triggering operation on the information sending result in a case where the first real-time reply result does not contain all results of the first reply result, pausing updating the first real-time reply result.

[0056] In a possible implementation, after the display module replaces the first text information displayed on the information interaction page with the second text information, the display module is further configured to:

[0057] display a second real-time reply result of the target object based on the second text information on the information interaction page; the second real-time reply result belongs to at least part of a second reply result corresponding to the second text information.

[0058] In a possible implementation, the display module, after replacing the first text information displayed on the information interaction page with the second text information, is further configured to:

[0059] delete the first real-time reply result displayed on the information interaction page.

[0060] In a possible implementation, the display module, when the information interaction page displays a second real-time reply result fed back by the target object based on the second text information, is configured to:

[0061] in response to the second real-time reply result being the same as the first real-time reply result, continue to display the other reply results in the second reply result except the first real-time reply result, by taking the first real-time reply result displayed on the information interaction page.

[0062] In a possible implementation, the apparatus further includes a voice playing module, configured to:

[0063] in response to the voice playing control displayed on the information interaction page being in a use state, play the first real-time reply result based on the tone corresponding to the target object.

[0064] in response to the second trigger operation on the information sending result, pause playing the first real-time reply result.

[0065] In a third aspect, an optional implementation of the present disclosure provides a computer device, a processor, and a memory. The memory stores machine readable instructions executable by the processor. The processor is configured to execute the machine readable instructions stored in the memory. When the machine readable instructions are executed by the processor, the machine readable instructions perform the steps of the first aspect or any possible implementation of the first aspect.

[0066] In a fourth aspect, an optional implementation of the present disclosure provides a computer readable storage medium, which stores a computer program. When the computer program is run, the computer program performs the steps of the first aspect or any possible implementation of the first aspect.

[0067] For the effects of the information interaction apparatus, the computer device, and the computer readable storage medium, refer to the description of the information interaction method, which will not be repeated here.

[0068] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, rather than limiting the technical solutions of the present disclosure.

[0069] In order to make the above objectives, characteristics and advantages of the present disclosure more obvious and easy to understand, the following preferred embodiments are specifically described below, and the accompanying drawings are described in detail as follows. BRIEF DESCRIPTION OF DRAWINGS

[0070] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following will briefly introduce the drawings needed to be used in the embodiments. The drawings herein are incorporated into the specification and form a part of the specification, which show the embodiments consistent with the present disclosure, and are used to explain the technical solutions of the present disclosure together with the specification. It should be understood that the following drawings only show some embodiments of the present disclosure, and therefore should not be regarded as a limitation on the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor.

[0071] FIG. 1 shows a flowchart of an information interaction method according to some embodiments of the present disclosure;

[0072] FIG. 2 shows one of the examples of an information interaction page according to some embodiments of the present disclosure;

[0073] FIG. 3 shows another example of an information interaction page according to some embodiments of the present disclosure;

[0074] FIG. 4 shows a schematic diagram of an information interaction device according to some embodiments of the present disclosure;

[0075] FIG. 5 shows a schematic diagram of a computer device according to some embodiments of the present disclosure. DETAILED DESCRIPTION

[0076] In order to make the objectives, technical solutions and advantages of the embodiments of the present disclosure more clear, the following will combine the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, but not all the embodiments. The components of the embodiments of the present disclosure described and shown herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the claimed present disclosure, but only represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present disclosure.

[0077] It is found through research that, when interacting with an intelligent agent, a user can input voice through a voice input control provided in an information interaction page; after voice input is completed, the input voice information is converted into text information and displayed to the user, so that the user can confirm whether to send the displayed text information to the intelligent agent; after the user confirms to send the displayed text information to the intelligent agent, an information sending operation is performed; after the intelligent agent receives the information sent by the user, the intelligent agent generates a reply content according to the information and displays the reply content to the user. In this process, multiple links such as voice input, conversion of voice information into text and display to the user, user confirmation, sending, and generation of a reply by the intelligent agent according to the text are involved, the operation process is complicated, a lot of time is consumed, and the process of generating a reply by the intelligent agent according to the text also needs a certain amount of time, thereby causing the problem of low interaction efficiency.

[0078] Based on the above research, the present disclosure provides an information interaction method, after obtaining voice information input by a user, without the user's second confirmation, an information sending result including text information corresponding to the voice information is directly displayed to the user, and in the process of inputting the voice information by the user, a real-time reply result is generated in real time based on the voice information by using a content generation model; thereby, without consuming a lot of time, the real-time reply result can be displayed to the user, and the interaction efficiency is improved.

[0079] The defects of the above solutions are the results of the inventors after practice and careful research, therefore, the discovery process of the above problems and the solutions proposed by the present disclosure to solve the above problems should be the contributions of the inventors to the present disclosure.

[0080] It should be noted that: similar reference numerals and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0081] It can be understood that, before using the technical solutions disclosed by the embodiments of the present disclosure, the type, use range, use scenario, etc. of personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.

[0082] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will need to obtain and use personal information of the user. Therefore, the user can voluntarily choose whether to provide personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium, etc. that performs the operation of the technical solutions of the present disclosure according to the prompt information.

[0083] As an optional but non-limiting implementation, in response to receiving the active request of the user, the manner of sending the prompt information to the user may, for example, be a pop-up window manner, in which the prompt information may be presented in the form of text. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0084] It can be understood that the above notification and user authorization obtaining process is only illustrative and does not limit the implementation of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0085] For the convenience of understanding the present embodiment, first, a kind of information interaction method disclosed by the present embodiment is introduced in detail, the execution subject of the information interaction method provided by the present embodiment is generally a computer device with certain computing power, which includes, for example: terminal device or server or other processing device, the terminal device can be user equipment (User Equipment, UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (Personal Digital Assistant, PDA), handheld device, computing device, vehicle-mounted device, wearable device, etc. In some possible implementation manners, the information interaction method can be realized by the processor calling the computer readable instructions stored in the memory.

[0086] The information interaction method provided by the present embodiment is described below taking the terminal device as an example.

[0087] Referring to FIG. 1, a flowchart of the information interaction method provided by the present embodiment is shown, and the method includes steps S101-S104, wherein:

[0088] S101: display an information interaction page; the information interaction page is used for information interaction with a target object.

[0089] In a specific implementation, the target object may, for example, include an intelligent agent; the intelligent agent is a computing entity based on artificial intelligence technology; the intelligent agent can carry an artificial intelligence model and can generate a corresponding reply result according to received text information or voice information by using the artificial intelligence model. The intelligent agent may, for example, have different identities in different application scenarios, such as an intelligent customer service in a customer service scenario, an intelligent chat robot in a chat scenario, etc.

[0090] In addition, in the chat scenario, the identity of the agent can also be set by the user; for example, the user wants to practice a certain language, and can set the identity of the agent as a virtual "person" familiar with the language. Based on the set identity of the user, the agent can generate a corresponding reply result based on the text information or voice information of the user. The reply result includes, for example, a reply to a question, integration of information, a reply to chat content, and the like, which are set according to actual application needs.

[0091] The information interaction page can also be different in different application scenarios; for example, in the chat scenario, the information interaction interface includes, for example, a chat interface; in the information query scenario, the chat interface can include, for example, an information query interface.

[0092] S102: In response to a first trigger operation on the target control in the information interaction page, receiving voice information input by the user.

[0093] In specific implementation, the target control is, for example, a control for voice input control of the user set in the information interaction page. The first trigger operation on the target control includes, for example, the following long press, click, slide, double click, and the like. After the user performs the first trigger operation on the target control, the target control enters a voice input valid state, at which time the terminal device can capture external sound to obtain voice information input by the user.

[0094] As shown in the example of FIG. 2a, a specific example of an information interaction page is shown; in the information interaction page, an information display area s1 for showing information interaction content to the user is included; in the information display area s1, text information converted from voice input by the user can be displayed; text content sent by a target object can also be displayed, which can include, for example, text content obtained by the target object replying based on the text information or voice information of the user. In this example, the user has previously input another text content, and the reply result corresponding to the other text content is displayed in the information display area s1. In addition, the target control s2 can also be included in the information interaction page; at the position of the target control s2, "press to talk" instruction information is also displayed to prompt the user to trigger input of voice information by long pressing the target control s2.

[0095] As shown in the example of FIG. 2b, the state of the target control s2 corresponding to the target control after the user performs a first trigger operation on the target control by long pressing the target control s2 is shown. In this state, the terminal device can receive voice information input by the user.

[0096] The above embodiment only shows a specific example of the first trigger operation, and other first trigger operations can be set according to actual application requirements, which are not limited herein.

[0097] S103: in response to the end of the voice information input, displaying, on the information interaction page, an information sending result after information is sent to the target object; the information sending result includes the first text information.

[0098] In a specific implementation, the trigger operation of the end of the voice information input can be different for different first trigger operations. When the trigger operation of the end of the voice information input is received, the voice of the outside world is stopped from being received.

[0099] For example, when the first trigger operation includes a long-press operation on the target button, the voice information input is considered to be ended when the user releases the pressed target button.

[0100] For example, when the first trigger operation includes a click operation or a double-click operation on the target button, the terminal device enters a voice acquisition state, captures the voice of the outside world, and receives the voice information input by the user after the user clicks or double-clicks the target button. In this case, the user can trigger the end of the voice input by clicking or double-clicking the target button again or directly triggering the end of the voice input by using a set voice instruction. For example, the voice instruction includes "end input" with a specific semantic or a specific short sentence, such as a call for the target object, and the short sentence can be set as "OK" or "how do you think". The embodiments of the present disclosure are not limited.

[0101] For example, when the first trigger operation includes a sliding operation on the target button, the target button can be slid to a preset area of the information interaction page by using the sliding operation. At this time, the terminal device enters a voice acquisition state, captures the voice of the outside world, and receives the voice information input by the user. In this case, the user can trigger the end of the voice input by dragging the target button out of the preset area. In addition, the sliding operation can also be a sliding operation in a first direction. When the user performs the sliding operation in the first direction, the terminal device enters a voice acquisition state, captures the voice of the outside world, and receives the voice information input by the user. In this case, the user can trigger the end of the voice input by performing the sliding operation in a second direction again.

[0102] In addition, other trigger operations for ending the voice input can also be set, and the embodiments of the present disclosure are not limited.

[0103] After the voice information input is completed, the information sending result after the information is sent to the target object is displayed on the information interaction page, for example, the first text information obtained by converting the voice information input by the user.

[0104] Specifically, during the process of receiving the voice information input by the user, the terminal device can convert the received voice into text information in real time. A short time after the voice information input is completed, the first text information corresponding to the voice information input by the user can be obtained. Then, the first text information is sent to the target object, and the information sending result after the information is sent to the target object by the user is displayed on the information interaction page.

[0105] In addition, the information sending result can include an indication of sending success or sending failure in addition to the first text information.

[0106] As shown in FIG. 2c, a specific example of displaying the information sending result after the voice information input is completed on the information interaction page is shown. In this example, the first voice information in the information sending result includes “What is quantum?”, and the first text is displayed in the information display area s1.

[0107] Based on the above S103, the information interaction method provided by the embodiment of the present disclosure further includes:

[0108] S104: Displaying the first real-time reply result fed back by the target object on the information interaction page; wherein the first real-time reply result is generated in real time by the content generation model based on the voice information.

[0109] In a specific implementation, the first text information corresponds to the first reply result. In a possible implementation, the process of generating the first reply result by the content generation model is a process lasting for a certain length of time. Therefore, the process of feeding back the reply result by the target object can also be a process lasting for a certain length of time; and the content that has been fed back is referred to as the first real-time reply content. The first real-time reply result belongs to at least part of the content of the first reply result corresponding to the first text information.

[0110] For example, assuming that the first reply result corresponding to the first text information includes three parts of content a1, a2, and a3, the first real-time reply result can be at least part of the content a1, a2, and a3.

[0111] The target object is an intelligent agent carrying the content generation model; the target object can be deployed on another device different from the terminal device, or can be deployed on the terminal device.

[0112] In the case where the target object is deployed on the terminal device:

[0113] The first text information corresponding to the voice information is obtained by converting the real-time voice input by the user in the process of inputting the voice information. Therefore, in the case where the content generation model is deployed on the terminal device, the terminal device can input the real-time generated text into the content generation model, and generate the first real-time reply content by using the content generation model in the process where the voice input by the user has not ended. In this process, with the continuous generation of the first text information, the content generation model can be used to continuously generate new first real-time reply results or update the already generated first real-time reply results, so as to quickly obtain the first real-time reply results after the user ends the voice input, and display the first real-time reply results in the information interaction page.

[0114] In the case where the target object is deployed on the terminal device:

[0115] The first text information corresponding to the voice information is obtained by converting the real-time voice input by the user in the process of inputting the voice information. With the continuous generation of the words in the first text information, the search keywords can be extracted according to the already generated words; if the search keywords are extracted, the extracted search keywords are sent to the other device in real time; and the target object deployed in the other device inputs the search keywords into the content generation model, and starts to generate the first real-time reply results corresponding to the search keywords while the user inputs the voice information.

[0116] When the user ends the voice input, since part of the first real-time reply results have been generated in the process of voice input, the generated first real-time reply results can be displayed in the information interaction interface; at the same time, if the first real-time reply results belong to part of the first reply results corresponding to the first text information, the remaining reply results can be generated while the already generated first real-time reply results are displayed.

[0117] Taking the above example as an example, assuming that the first reply results corresponding to the first text information include k1, k2, and k3, a total of three parts of content. Assuming that k1 and k2 have been generated according to the voice information after the voice input ends; k1 and k2 can be continuously displayed in the information interaction page, while k3 is generated; and after k3 is generated, k3 is displayed in the information interaction page by connecting k1 and k2.

[0118] As shown in FIG. 2d, a specific example of displaying a first real-time reply result in an information interaction page is provided; in this example, the first text information includes "what is quantum"; the first reply content corresponding to the first text information includes "a physical quantity is quantized if it has a minimum unit and cannot be continuously divided, and the minimum unit is called a quantum. Its basic concept is that all tangible properties may be "quantizable". "Quantization" means that the value of the physical quantity will be some specific value, rather than any value. For example, in (rest state) atoms, the energy of the electron is quantized, which can determine the stability and general problem of the atom." The first real-time result displayed in the information display area s1 belongs to part of the above first reply result, including "a physical quantity is quantized if it has a minimum".

[0119] The information interaction method provided by the embodiments of the present disclosure can receive voice information input by a user in response to a first triggering operation on a target control in an information interaction page; display an information sending result after information is sent to the target object in the information interaction page in response to the end of voice information input; the information sending result includes first text information corresponding to the voice information; and display a first real-time reply result fed back by the target object in the information interaction page; wherein the first real-time reply result is generated by a content generation model based on the voice information in real time. In this way, after obtaining the voice information input by the user, the information sending result including the text information corresponding to the voice information is directly displayed to the user without the need for the user to confirm again, and the real-time reply result is generated by the content generation model based on the voice information in real time during the process of inputting the voice information by the user; and then the real-time reply result can be displayed to the user without consuming much time, thereby improving the interaction efficiency.

[0120] In the information interaction method provided by another embodiment of the present disclosure, the method further includes: displaying a modification control of the information sending result in response to a second triggering operation on the information sending result, and displaying the first text information in the modification control; generating second text information in response to a modification operation on the first text information displayed in the modification control; and replacing the first text information displayed in the information interaction page with the second text information.

[0121] In specific implementation, after the information sending result after information is sent to the target object is displayed in the information interaction page, the second triggering operation on the information sending result can be, for example, a click operation, a double-click operation, a sliding operation, etc.

[0122] Exemplary:

[0123] a1: In the case that the second trigger operation includes a click operation or a double-click operation on the information sending result, the user can trigger the modification of the first text information by clicking or double-clicking the first text information displayed on the information interaction page.

[0124] At this time, the corresponding modification control can be displayed in the preset area of the information interaction page, and the first text information can be displayed in the modification control.

[0125] a2: In the case that the second trigger operation includes a sliding operation on the information sending result, for example, a specific trigger area can be set in the information interaction page; the first text information can be dragged into the trigger area by using the sliding operation.

[0126] After the first text information is dragged into the trigger area, the corresponding modification control is triggered to be displayed in the interaction control page, and the first text information is displayed in the modification control.

[0127] Here, the modification control can be separately set and be used for modifying the first text information. In another possible implementation, the target control s1 can switch states, i.e., the target control s1 can be switched between a voice input state and a text input state; when the target control is in the voice input state, the user can trigger the input of voice information by the first trigger operation on the target control; when the target control is in the text input state, a virtual keyboard can be displayed to the user; the user can directly input text by the virtual keyboard in the target control. At this time, the target control in the text input state can be used as the modification control.

[0128] As shown in FIG. 3a, an example of displaying the modification control in the information interaction page and displaying the first text information in the modification control after the second trigger operation is shown. In this example, the target control can be a modification control used for modifying the first text information. The modification control can be displayed in the information interaction page, for example.

[0129] The modification control includes, for example, an information input box s3 used for displaying the first text information, the information input box s3 displays the first text information "what is quantum?", and a virtual keyboard s4 used for editing the first text information displayed in the information input box s3. The virtual keyboard s4 can be used to implement the modification operation of the first text information displayed in the information input box s3, such as deleting at least part of the words, adding words, format adjustment, etc.

[0130] In addition, the first button s5 used for triggering the abandonment of the modification of the first text information and the second button s6 used for confirming the result of the modification of the first text information can also be included in the modification control.

[0131] When the user triggers the first button s5, i.e. the user gives up the modification of the first text information, the modification control is hidden, the information interaction page is switched to the state shown in Fig. 2d, and the subsequent part of the first reply content is continuously updated. When the user modifies the first text information in the information input box s3 and triggers the second button s6, the content in the information input box s3 is taken as the second text content, and the first text information displayed by the information interaction page is replaced by the second text information.

[0132] In another possible implementation, when the target control in the text input state is taken as the modification control, the style of the target control may, for example, be as shown in the example in Fig. 3a. For example, if the target control in the text input state is triggered by a second trigger operation, the first button s5 and the second button s6 may be displayed in the target control, for giving up the modification of the first text information or accepting the modification of the first text information.

[0133] In order to distinguish the normal input of text content from the target control, when the target control is triggered by switching the state of the target control (i.e. the state of the target control is switched from the voice input state to the text input state), the state of the first button s5 and the second button s6 may be set to an invalid state, in which the first button s5 and the second button s6 do not respond to any operation of the user, i.e. do not affect the normal information input function of the target control in the text input state.

[0134] Alternatively, an association between the first text content in the text input box in the target control and the first text content displayed in the information display area s1 may be established, and the first text content displayed in the information display area s1 may be specially marked. When the first text content in the text input box is modified, if the user further triggers the confirmation of the modification, the first text content displayed in the information display area s1 is replaced by the second text information displayed in the text input box according to the association between the first text content in the text input box in the target control and the first text content displayed in the information display area s1, so as to realize the modification of the first text information by the target control without adding another modification control.

[0135] In this way, by using the above method, the user can trigger the modification of the first text information without having to perform a withdrawal operation on the sent first text information, and does not need to input new text information, i.e. the modification of the sent first text information can be conveniently realized, and the operation of the user is facilitated.

[0136] In another embodiment of the present disclosure, in response to a second triggering operation on the information sending result, recommended modification information corresponding to the first text information can also be displayed.

[0137] The recommended modification information includes at least one of the following, for example: indication information of a word in the first text information that needs to be modified, candidate words for replacing at least part of the content in the first text information, and text information obtained by content summarization of the first text information.

[0138] b1: In the case where the recommended modification information includes indication information of a word in the first text information that needs to be modified:

[0139] For example, there can be a short sentence or a word in the first text information that cannot be adapted to the content expressed by the first text information as a whole or that is not logically coherent; at this time, the word or the short sentence with the problem can be marked by using the indication information, for example, to prompt the user that there can be an inaccurate expression in the first text information, so as to remind the user to modify the word or the short sentence at the corresponding position.

[0140] Here, the indication information can include at least one of the following, for example: special color marking of the word or phrase that needs to be indicated, underlining or an indication arrow near the word or phrase that needs to be indicated, etc. The specific marking can be determined according to actual needs.

[0141] In this case, the user can directly trigger the modification of the word or phrase indicated by the indication information by a third triggering operation on the indication information, such as a click, double-click, long press, etc.

[0142] b2: In the case where the recommended modification information includes candidate words for replacing at least part of the content in the first text information:

[0143] After determining the word in the first text information that can have a problem, candidate words for replacing the word with the problem can be predicted according to the content expressed by the first text information as a whole or according to language logic.

[0144] After the corresponding candidate words are predicted, the candidate words are displayed to the user; the user can modify the word or phrase with the problem by a third triggering operation on the candidate words, such as a selection operation (such as clicking the candidate words) or a sliding operation (such as dragging the candidate words to the position of the word with the problem in the first text information by using the sliding operation).

[0145] b3: In the case where the recommended modification information includes text information obtained by content summarization of the first text information:

[0146] Exemplarily, when a user inputs voice information, in order to express the intended meaning clearly, the user may describe a certain problem in a rather cumbersome language, resulting in the first text information being not refined enough; or the problem the user wants to express is not particularly clear. At this time, the first text information can be further refined, and the refined text information can be more refined and can more accurately express the meaning the user wants to express. Then, the refined text information is displayed to the user. At this time, the user can perform a third trigger operation on the refined text information (such as a click operation or a double-click operation on the refined text information) to modify the first text information.

[0147] As shown in b of FIG. 3, a specific example of displaying recommended modification information corresponding to the first text information is shown; in this example, the first text information includes "What is quantum"; at the same time, there may be other words for the pronunciation "liangzi", such as "beef between two parties", "handsome boy"; and there are also other words related to "quantum", such as "quantum entanglement"; at this time, using bolding and underlining the word as the indication information, the word "quantum" is marked out from it. At the same time, in the information interaction page, three other optional candidate words "beef between two parties", "handsome boy" and "quantum entanglement" are also displayed.

[0148] As shown in c of FIG. 3, after the user selects the candidate word "quantum entanglement" through the third trigger operation, the first text content "What is quantum" displayed in the information display area s1 is modified to "What is quantum entanglement".

[0149] As shown in d of FIG. 3, following c of FIG. 3 above, if the user triggers the confirmation of modification, then the "What is quantum" displayed in the information display area s1 is also replaced with "What is quantum entanglement". The modification control is hidden, and at the same time, the second reply result determined for the second text content includes: "Quantum entanglement, or quantum entanglement, is a quantum mechanical phenomenon. Its definition describes a special quantum state of a composite system (with more than two member systems). This quantum state cannot be decomposed into the tensor product of the quantum states of the member systems. Quantum entanglement technology is an encryption technology for secure information transmission and has nothing to do with faster-than-light information transmission. Although we know that the speed of "communication" between these particles is very fast, we cannot use this connection to control and transmit information at such a fast speed. Therefore, the rule proposed by Einstein, that is, the speed of any information transmission cannot exceed the speed of light, still holds. In fact, the entanglement effect is not very far, and once one of the parties is interfered with, the entangled state will automatically disappear."

[0150] The second real-time reply result belongs to a part of the second reply result, which includes: "Quantum entanglement, or quantum entanglement, is a quantum mechanical phenomenon that defines the complex", so as to display the second real-time reply result below the second text information.

[0151] In another embodiment of the present disclosure, the conversion of semantic information into text information is achieved by using a target model; in order to enable the target model to be more accurate, the target model can be retrained by using voice information and second text information obtained by modifying the first text information, so as to continuously enable the target model to be more accurate for the pronunciation characteristics of a specific user.

[0152] In another embodiment of the present disclosure, the first real-time reply result belongs to at least part of the first reply result corresponding to the first text information. The information interaction method provided by the embodiment of the present disclosure further includes:

[0153] In response to receiving the second trigger operation of the information sending result in the case that the first real-time reply result does not contain all results of the first reply result, the updating of the first real-time reply result is paused.

[0154] In a specific implementation, the process of the target object feeding back the reply result is, for example, a process that lasts for a certain duration; and the content that has been fed back is referred to as first real-time reply content. The first real-time reply content belongs to at least part of the first reply result corresponding to the first text information. In order to facilitate the user to read, the received first real-time reply content can be displayed to the user while receiving the first real-time reply content fed back by the target object. In addition, if the target object feeds back the complete first reply result at one time, the first reply content can also be displayed to the user at a certain updating rate (at this time, the content updated in real time is the first implementation reply result). In this way, the display of the first reply content is a continuous process, rather than a one-time full display, which facilitates the user to read.

[0155] As shown in the example of FIG. 2d, after the first text information is sent to the target object, a specific example of displaying the first real-time reply result to the user, in which example, the first real-time reply result only contains part of the reply result in the first reply result.

[0156] Further, in a case where the first real-time reply result only belongs to a part of the first reply result, if a second trigger operation on the information sending result is received at this time, updating the first real-time reply result is suspended, that is, the part of the first reply result that is not displayed to the user is suspended from being displayed to the user. In this way, it can be ensured that after the user gives up modifying the first text content, the user can continue to update other content in the first reply result according to the previous update progress of the first real-time reply result. In addition, after the user confirms that the first text content is modified to the second text content, the second real-time reply result fed back by the target object based on the second text content can be directly displayed to the user, thereby improving the user experience in the reply result updating process.

[0157] In another embodiment of the present disclosure, after the first text information displayed on the information interaction page is replaced by the second text information, the following further includes:

[0158] displaying a second real-time reply result fed back by the target object based on the second text information on the information interaction page; the second real-time reply result belongs to at least part of a second reply result corresponding to the second text information.

[0159] Here, the generation manner and the display manner of the second real-time reply result are similar to those of the first real-time reply result, and will not be described here again.

[0160] In addition, in another embodiment, after the first text information displayed on the information interaction page is replaced by the second text information, the following further includes:

[0161] deleting the first real-time reply result displayed on the information interaction page.

[0162] As shown in the example in FIG. 3d, after replacing “what is quantum” displayed in the information display area s1 with “what is quantum entanglement”, the previously displayed first real-time reply result “one physical quantity if there is minimum” is deleted from the information display area s1, and the second real-time reply result fed back by the target object is displayed.

[0163] In addition, when the information interaction page displays the second real-time reply result fed back by the target object based on the second text information, the following can be used, for example:

[0164] in response to the second real-time reply result being the same as the first real-time reply result, the first real-time reply result displayed on the information interaction page is continued, and other reply results in the second reply result except the first real-time reply result are continued to be displayed.

[0165] For example, assume that the first reply result includes a1, a2, and a3, and the second reply result includes a1, a2, a4, and a5. The first real-time reply result displayed to the user includes a1. Since a1 in the second reply result belongs to the same reply result as a1 in the first reply result, when the second reply result is displayed, a1 that has already been displayed can be continued to be displayed, and a2, a4, and a5 can be displayed without deleting a1 that has been originally displayed.

[0166] In this way, the continuity of the reply result update is ensured.

[0167] In addition, in order to enable the user to perceive that the determined reply result is updated after the first text information is updated to the second text information, the second reply result can also be specially marked to prompt the user that the determined reply result is also adjusted to a certain extent.

[0168] In another embodiment of the present disclosure, a voice playing control is also displayed in the information interaction page. In response to the voice playing control displayed in the information interaction page being in a use state, the first real-time reply result is played based on the timbre corresponding to the target object.

[0169] In response to the second trigger operation on the information sending result, the playing of the first real-time reply result is paused.

[0170] Here, the timbre corresponding to the target object can be any one of a plurality of selectable timbres, for example, and can include adult male voice, adult female voice, sweet female voice, urban female voice, magnetic male voice, child female voice, child male voice, etc. The selectable timbres of the target object can be set according to actual needs. The user can trigger a timbre changing control based on the information interaction page to change the timbre corresponding to the target object.

[0171] The voice playing control being in a use state can exist in the following several cases, for example.

[0172] d1: The voice playing control is in a default open state. In this state, as long as the real-time reply result is displayed, the displayed real-time reply result is played.

[0173] d2: The voice playing control is in a default closed state. After the real-time reply result is received, in response to a trigger operation on the voice playing control in the default closed state, the voice playing control is switched from the closed state to the open state, and at this time, the displayed real-time reply result is played.

[0174] After the playing of the reply result is completed, the voice playing control in the open state can be switched to the default closed state again.

[0175] In the example shown in FIG. 3, a voice playing control s7 is displayed. When the voice playing control s7 is in use, the second real-time reply result is played by using the timbre corresponding to the target object.

[0176] Those skilled in the art can understand that, in the above method of the specific implementation, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0177] Based on the same inventive concept, the information interaction method in the embodiments of the present disclosure is also provided. Since the principle of solving problems of the device in the embodiments of the present disclosure is similar to the above-mentioned information interaction method, the implementation of the device can be referred to the implementation of the method, and the repeated parts will not be described here.

[0178] Referring to FIG. 4, a schematic diagram of an information interaction device provided by the embodiments of the present disclosure is shown. The device comprises:

[0179] The display module 41 is configured to display an information interaction page; the information interaction page is configured to perform information interaction with a target object;

[0180] The acquisition module 42 is configured to, in response to a first triggering operation on a target control in the information interaction page, receive voice information input by a user;

[0181] The display module 41 is further configured to, in response to the end of the voice information input, display, on the information interaction page, an information sending result after information is sent to the target object; the information sending result comprises the first text information; and the display module 41 is further configured to display, on the information interaction page, a first real-time reply result fed back by the target object; wherein the first real-time reply result is generated by a content generation model in real time based on the voice information.

[0182] In a possible implementation, the device further comprises a modification module 43; the modification module 43 is configured to:

[0183] In response to a second triggering operation on the information sending result, display a modification control of the information sending result, and display the first text information on the modification control;

[0184] In response to a modification operation on the first text information displayed on the modification control, generate second text information;

[0185] The display module 41 is further configured to replace the first text information displayed on the information interaction page with the second text information.

[0186] In a possible implementation, the display module 41 is further configured to:

[0187] displaying recommended modification information corresponding to the first text information; the recommended modification information comprises at least one of the following: indication information of a recommended modification vocabulary in the first text information, a candidate vocabulary for replacing at least part of the content in the first text information, and text information obtained by content summarization of the first text information;

[0188] The modification operation comprises a third trigger operation on the recommended modification information.

[0189] In a possible implementation, the method further comprises: training, by a training module 44, a target model based on the voice information and the second text information after the second text information is generated; and the target model is used to convert the voice information into corresponding text information in real time.

[0190] In a possible implementation, the first real-time reply result belongs to at least part of the first reply result corresponding to the first text information.

[0191] The display module 41 is further configured to:

[0192] In response to receiving the second trigger operation on the information sending result in a case where the first real-time reply result does not contain all results of the first reply result, the display module 41 is further configured to:

[0193] In a possible implementation, after the display module 41 replaces the first text information displayed on the information interaction page with the second text information, the display module 41 is further configured to:

[0194] displaying a second real-time reply result of the target object on the information interaction page, the second real-time reply result being based on the second text information and belonging to at least part of a second reply result corresponding to the second text information.

[0195] In a possible implementation, after the display module 41 replaces the first text information displayed on the information interaction page with the second text information, the display module 41 is further configured to:

[0196] deleting the first real-time reply result displayed on the information interaction page.

[0197] In a possible implementation, when the display module 41 displays a second real-time reply result of the target object on the information interaction page, the second real-time reply result being based on the second text information, the display module 41 is configured to:

[0198] In response to the second real-time reply result being the same as the first real-time reply result, the first real-time reply result displayed by the information interaction page is accepted, and the other reply results in the second reply result except the first real-time reply result are continuously displayed.

[0199] In a possible implementation, the apparatus further includes a voice playing module 45 configured to:

[0200] In response to the voice playing control displayed by the information interaction page being in a use state, the first real-time reply result is played based on the tone of the target object.

[0201] In response to the second trigger operation on the information sending result, the playing of the first real-time reply result is paused.

[0202] The processing procedure of each module in the apparatus and the interaction procedure between the modules can refer to the related description in the method embodiments, and will not be described in detail here.

[0203] The embodiments of the present disclosure further provide a computer device, as shown in FIG. 5, which is a structural schematic diagram of a computer device provided by the embodiments of the present disclosure, and includes:

[0204] a processor 51 and a memory 52; the memory 52 stores machine readable instructions executable by the processor 51, and the processor 51 is configured to execute the machine readable instructions stored in the memory 52, and when the machine readable instructions are executed by the processor 51, the processor 51 performs the following steps:

[0205] an information interaction page is displayed; the information interaction page is used for information interaction with a target object;

[0206] In response to a first trigger operation on a target control in the information interaction page, voice information input by a user is received.

[0207] In response to the voice information input being ended, an information sending result after information sending to the target object is displayed on the information interaction page; the information sending result includes first text information corresponding to the voice information.

[0208] A first real-time reply result fed back by the target object is displayed on the information interaction page; the first real-time reply result is generated by a content generation model based on the voice information in real time. The memory 52 includes an internal memory 521 and an external memory 522; the internal memory 521 is also called an internal storage, and is used for temporarily storing operation data in the processor 51 and exchanging data with the external memory 522 such as a hard disk, and the processor 51 exchanges data with the external memory 522 through the internal memory 521.

[0209] The specific execution process of the above instructions can refer to the steps of the information interaction method described in the embodiments of the present disclosure, which will not be repeated here.

[0210] The embodiments of the present disclosure also provide a computer-readable storage medium, which stores a computer program. When the computer program is run by a processor, the steps of the information interaction method described in the above method embodiments are executed. The storage medium can be a volatile or non-volatile computer-readable storage medium.

[0211] The embodiments of the present disclosure also provide a computer program product, which carries a program code. The instructions included in the program code can be used to execute the steps of the information interaction method described in the above method embodiments. For details, refer to the above method embodiments, which will not be repeated here.

[0212] The computer program product can be specifically implemented by hardware, software or a combination thereof. In an optional embodiment, the computer program product is specifically embodied as a computer storage medium. In another optional embodiment, the computer program product is specifically embodied as a software product, such as a software development kit (SDK) and the like.

[0213] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system and device can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here. In several embodiments provided by the present disclosure, it should be understood that the disclosed system, device and method can be implemented by other ways. The above-described device embodiments are only schematic, for example, the division of the units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some communication interfaces, devices or units, and can be electrical, mechanical or other forms.

[0214] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.

[0215] In addition, each functional unit in the various embodiments of the present disclosure can be integrated into one processing unit, or each unit can exist physically, or two or more units can be integrated into one unit.

[0216] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a nonvolatile computer readable storage medium executable by a processor. Based on this understanding, the technical solutions of the present disclosure essentially or the part of the prior art that contributes to the present disclosure or the part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0217] Finally, it should be noted that: the above-described embodiments are merely specific embodiments of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limit them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can make modifications or easily think of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed by the present disclosure, or make equivalent replacements to some of the technical features. Such modifications, changes or replacements do not cause the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. An information interaction method, comprising: displaying an information interaction page; the information interaction page is used for information interaction with a target object; in response to a first triggering operation on a target control in the information interaction page, receiving voice information input by a user; in response to the end of the voice information input, displaying an information sending result after information sending to the target object on the information interaction page; the information sending result includes first text information corresponding to the voice information; displaying a first real-time reply result fed back by the target object on the information interaction page; wherein the first real-time reply result is generated in real time based on the voice information by using a content generation model.

2. The method of claim 1, further comprising: in response to a second triggering operation on the information sending result, displaying a modification control of the information sending result, and displaying the first text information on the modification control; in response to a modification operation on the first text information displayed on the modification control, generating second text information; replacing the first text information displayed on the information interaction page with the second text information.

3. The method of claim 2, further comprising: displaying recommended modification information corresponding to the first text information; the recommended modification information includes at least one of the following: indication information of a recommended modification vocabulary in the first text information, candidate vocabulary for replacing at least part of the content in the first text information, and text information obtained by content extraction of the first text information; the modification operation includes a third triggering operation on the recommended modification information.

4. The method of claim 2 or 3, wherein, After generating the second text information, further comprising: training a target model based on the voice information and the second text information; the target model is used to convert the voice information into corresponding text information in real time.

5. The method according to any one of claims 3-4, wherein, The first real-time reply result belongs to at least part of the first reply result corresponding to the first text information; the method further comprises: in response to receiving the second triggering operation on the information sending result in a case where the first real-time reply result does not contain all results of the first reply result, pausing updating the first real-time reply result.

6. The method of claim 5, wherein, After replacing the first text information displayed on the information interaction page with the second text information, further comprising: displaying a second real-time reply result fed back by the target object based on the second text information on the information interaction page; the second real-time reply result belongs to at least part of the second reply result corresponding to the second text information.

7. The method of claim 5, wherein, After replacing the first text information displayed on the information interaction page with the second text information, further comprising: deleting the first real-time reply result displayed on the information interaction page.

8. The method of claim 6, wherein, The displaying a second real-time reply result fed back by the target object based on the second text information on the information interaction page, comprises: In response to the second real-time reply result being the same as the first real-time reply result, the first real-time reply result displayed by the information interaction page is accepted, and the other reply results in the second reply result except the first real-time reply result are continuously displayed.

9. The method of any one of claims 5-8, further comprising: In response to the voice playing control displayed by the information interaction page being in a use state, the first real-time reply result is played based on the tone of the target object; In response to the second trigger operation on the information sending result, the playing of the first real-time reply result is paused.

10. An information interaction apparatus, comprising: a display module configured to display an information interaction page; the information interaction page is used for information interaction with a target object; an acquisition module configured to, in response to a first trigger operation on a target control in the information interaction page, receive voice information input by a user; the display module is further configured to, in response to the voice information input being ended, display an information sending result after information sending to the target object on the information interaction page; the information sending result includes first text information; and display a first real-time reply result fed back by the target object on the information interaction page; wherein the first real-time reply result is generated by a content generation model based on the voice information in real time. a processor and a memory; the memory stores machine readable instructions executable by the processor; the processor is configured to execute the machine readable instructions stored in the memory; when the machine readable instructions are executed by the processor, the processor executes the steps of the method in any one of claims 1-9.

11. A computer device comprising: a computer program is stored on the computer readable storage medium; when the computer program is run by a computer device, the computer device executes the steps of the method in any one of claims 1-9.

12. A computer readable storage medium, wherein, ​