Voice interaction method, device, electronic device and medium

By generating dialogue requests and outputting answer voices, combined with interaction scene detection and virtual interaction image output, the problem of poor voice interaction in existing technologies is solved, an interactive experience closer to real-life conversations is achieved, and the effect of oral practice is improved.

CN114582339BActive Publication Date: 2025-09-09BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210089127.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2025-09-09
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

The voice interaction mode of the existing technology is not effective, resulting in low effectiveness of oral practice.

Method used

By generating dialogue requests, sending answer texts and outputting answer voices, combined with real-time detection of interaction scenarios and output of virtual interaction images, the interaction experience is improved.

Benefits of technology

It achieves an interactive process that is closer to real-life conversation, improves the voice interaction effect and experience, and enhances the immersive conversation interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114582339B_ABST
    Figure CN114582339B_ABST
Patent Text Reader

Abstract

The present disclosure provides a voice interaction method, apparatus, device, medium, and product, relating to the field of artificial intelligence technology, specifically natural language processing, speech recognition, and other technical fields. The voice interaction method includes: generating a dialogue request in response to receiving a target voice; sending the dialogue request; and, in response to receiving a textual response to the dialogue request, outputting a voice response based on the textual response. Furthermore, during the voice interaction process, an interaction scenario at any point in time is determined, and a virtual interaction avatar is output based on an interaction mode associated with the interaction scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically natural language processing, speech recognition and other technical fields, and more specifically, to a voice interaction method, device, electronic device, medium and program product. Background Art

[0002] In the related art, users can practice spoken English by performing voice interaction with electronic devices, for example, by performing voice interaction with electronic devices. However, the voice interaction method in the related art is not effective, resulting in low oral practice results. Summary of the Invention

[0003] The present disclosure provides a voice interaction method, device, electronic device, storage medium, and program product.

[0004] According to one aspect of the present disclosure, a voice interaction method is provided, comprising: generating a dialogue request in response to receiving a target voice; sending the dialogue request; outputting a response voice based on the response text in response to receiving a response text for the dialogue request; wherein, during the voice interaction process, an interaction scene at any moment is determined, and a virtual interaction image is output based on an interaction mode associated with the interaction scene.

[0005] According to another aspect of the present disclosure, a voice interaction device is provided, comprising: a first generation module, a first sending module, and a first output module. The first generation module is configured to generate a dialogue request in response to receiving a target voice; the first sending module is configured to send the dialogue request; and the first output module is configured to output a response voice based on a response text received in response to the dialogue request. The voice interaction device further comprises a determination module and a second output module, wherein the determination module is configured to determine an interaction scenario at any moment during the voice interaction process; and the second output module is configured to output a virtual interaction avatar based on an interaction mode associated with the interaction scenario.

[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above-described voice interaction method.

[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the above-mentioned voice interaction method.

[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program / instruction, which implements the steps of the above-mentioned voice interaction method when executed by a processor.

[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0011] Figure 1 The following schematically illustrates a system architecture for voice interaction according to an embodiment of the present disclosure;

[0012] Figure 2 The following schematically shows a flow chart of a voice interaction method according to an embodiment of the present disclosure;

[0013] Figure 3 The following schematically shows a flow chart of a voice interaction method according to another embodiment of the present disclosure;

[0014] Figure 4 The following schematically shows a flow chart of a voice interaction method according to another embodiment of the present disclosure;

[0015] Figure 5 A block diagram schematically shows a voice interaction device according to an embodiment of the present disclosure; and

[0016] Figure 6 The block diagram is a block diagram of an electronic device for performing voice interaction for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION

[0017] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0018] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0019] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0020] When expressions such as "at least one of A, B and C, etc." are used, they should generally be interpreted in accordance with the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0021] Figure 1 The system architecture of voice interaction according to an embodiment of the present disclosure is schematically shown. It should be noted that, Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments or scenarios.

[0022] like Figure 1 As shown, the system architecture 100 according to this embodiment may include clients 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the clients 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0023] Users can use clients 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on clients 101, 102, and 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0024] The clients 101, 102, 103 may be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, etc. The clients 101, 102, 103 of the embodiment of the present disclosure may, for example, run applications.

[0025] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using clients 101, 102, and 103 (for example only). The backend management server can analyze and process received data such as user requests, and feed back the processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the client. Alternatively, server 105 can be a cloud server, i.e., server 105 has cloud computing capabilities.

[0026] It should be noted that the voice interaction method provided in the embodiment of the present disclosure can be executed by the client 101, 102, 103. Accordingly, the voice interaction device provided in the embodiment of the present disclosure can be set in the client 101, 102, 103.

[0027] In one example, the clients 101, 102, and 103 may generate a dialogue request based on the target speech and send the dialogue request to the server 105 via the network 104. The server 105 may receive the dialogue requests from the clients 101, 102, and 103 via the network 104, obtain a response text based on the dialogue request, and return the response text to the clients 101, 102, and 103 via the network 104.

[0028] It should be understood that Figure 1 The number of clients, networks, and servers in the embodiment is merely illustrative. Any number of clients, networks, and servers may be used depending on the implementation requirements.

[0029] The following combination Figure 1 The system architecture of Figures 2 to 4 To describe the voice interaction method according to the exemplary embodiment of the present disclosure. The voice interaction method of the embodiment of the present disclosure can be, for example, Figure 1 The client is shown to execute, Figure 1 The client shown is, for example, the same as or similar to the electronic device described below.

[0030] Figure 2 The flowchart of the voice interaction method according to an embodiment of the present disclosure is schematically shown.

[0031] like Figure 2 As shown, the voice interaction method 200 of the embodiment of the present disclosure may, for example, include operations S210 to S250.

[0032] In operation S210, in response to receiving the target voice, a dialogue request is generated.

[0033] In operation S220, a conversation request is sent.

[0034] In operation S230, in response to receiving the answer text to the dialogue request, an answer voice is output based on the answer text.

[0035] In operation S240 , an interaction scenario at any time during the voice interaction process is determined.

[0036] In operation S250 , a virtual interaction avatar is output based on the interaction mode associated with the interaction scenario.

[0037] For example, the electronic device can collect the target voice of the target object through the recording function, and the target object is, for example, a user. In one scenario, when the target user needs to practice spoken English, he or she can interact with the electronic device in an English conversation, and the target voice is, for example, English voice.

[0038] The electronic device generates a dialogue request based on the target voice, and then sends the dialogue request to the server, which processes the dialogue request and generates a response text.

[0039] The server then sends the reply text to the electronic device. After receiving the reply text, the electronic device converts the reply text into a reply voice and outputs the reply voice, which is, for example, a British accent. After hearing the reply voice, the user continues to generate the target voice based on the reply voice and continues to perform the above operations, thereby achieving voice interaction.

[0040] Exemplarily, operations S240 to S250 may be performed during the entire voice interaction process, that is, before or after operation S210 , before or after operation S220 , or before or after operation S230 .

[0041] For example, the voice interaction process includes the time period between a first moment before receiving the target voice and a second moment after outputting the response voice, where any moment is any moment or moments within the time period. After determining the interaction scene at any moment, an interaction mode associated with the interaction scene can be determined, and a virtual interaction avatar can be output based on the interaction mode, enabling voice interaction between the user and the virtual interaction avatar, providing the user with an immersive conversational interaction experience.

[0042] According to the embodiments of the present disclosure, an electronic device interacts with a server to generate a text response based on a target voice. This allows for free and unrestricted conversation content, without being constrained by fixed content. This makes the interaction process more similar to a real-life conversation, improving the voice interaction effect and user experience. Furthermore, during the voice interaction process, the electronic device determines the interaction scenario in real time and outputs a virtual interaction avatar based on the interaction mode associated with that scenario. This ensures a close connection between the interaction avatar and the interaction scenario, enhancing the voice interaction experience.

[0043] In one example, the interaction scenario includes receiving a target voice scenario or waiting to receive a target voice scenario, and the interaction mode includes a listening mode. Outputting the virtual interaction avatar based on the interaction mode associated with the interaction scenario includes: controlling the virtual interaction avatar to generate a listening action based on the listening mode.

[0044] For example, when receiving the user's target voice or waiting for the user to make a target voice, the electronic device controls the virtual interactive image to make a listening action, making the user feel that the other party is listening to him or her, giving the user an immersive dialogue interaction experience.

[0045] In another example, the interaction scenario includes outputting an answering voice scenario, and the interaction mode includes an answering mode. Based on the interaction mode associated with the interaction scenario, outputting the virtual interaction image includes: controlling the virtual interaction image to generate an answering action based on the answering mode.

[0046] For example, after the electronic device generates an answer voice based on the answer text from the server, during the process of the electronic device outputting the answer voice, the electronic device controls the virtual interactive image to make an answer action, making the user feel that the other party is speaking, giving the user an immersive dialogue interaction experience.

[0047] In another example, the interaction scenario includes a waiting for outputting a response voice scenario, and the interaction mode includes a thinking mode. Based on the interaction mode associated with the interaction scenario, outputting the virtual interaction image includes: controlling the virtual interaction image to generate a thinking action based on the thinking mode.

[0048] For example, when the electronic device is waiting for the server's answer text or has received the answer text but has not output the answer voice, the electronic device controls the virtual interactive image to make thinking movements, making the user feel that the other party is thinking about how to answer, giving the user an immersive dialogue interaction experience.

[0049] According to the embodiments of the present disclosure, during the voice interaction process, the electronic device detects the current interaction scene in real time, and outputs a virtual interaction image based on different interaction modes according to different interaction scenes, thereby improving the intelligence level of the virtual interaction image and thus improving the dialogue interaction experience.

[0050] Figure 3 The following schematically shows a flow chart of a voice interaction method according to another embodiment of the present disclosure.

[0051] like Figure 3 As shown, the voice interaction method of the embodiment of the present disclosure may include operations S301A to S306A and operations S301B to S304B, wherein operations S301A to S306A are performed by an electronic device, and operations S301B to S304B are performed by a server.

[0052] In operation S301A, a target voice is received.

[0053] In operation S302A, speech recognition is performed on the target speech to obtain a target text.

[0054] For example, the electronic device recognizes the target voice through a voice recognition function to obtain the target text.

[0055] In operation S303A, a dialogue request is generated based on the target text.

[0056] Exemplarily, the dialogue request includes, for example, a target text.

[0057] In operation S304A, a conversation request is sent.

[0058] In operation S301B, a conversation request is received.

[0059] In operation S302B, a response text is generated based on the dialogue request.

[0060] For example, the server uses a natural language processing model to obtain a response text based on a target text in a conversation request.

[0061] In operation S303B, the answer text is verified.

[0062] For example, before the server obtains the answer text and sends the answer text to the electronic device, the server may verify the answer text, for example, verify whether the answer text is legal and compliant.

[0063] In operation S304B, the reply text is sent.

[0064] If the answer text passes the verification, it means that the answer text is legal and compliant, that is, the answer text does not contain sensitive information. The server can send the verified answer text to the electronic device.

[0065] In operation S305A, a reply text is received.

[0066] In operation S306A, an answer voice is output based on the answer text.

[0067] After the electronic device outputs the answer voice, when the user hears the answer voice, the user can continue to issue the next target voice based on the answer voice, and then continue to perform the above operations, thereby achieving dialogue interaction.

[0068] During the entire voice interaction process, the electronic device can detect the interaction scene in real time to output a virtual interaction image based on the interaction mode associated with the interaction scene.

[0069] According to an embodiment of the present disclosure, an electronic device interacts with a server, and the server obtains a response text based on a natural language processing model, so that the content of the conversation is free and unrestricted, without being restricted by fixed conversation content. The interaction process is closer to a real-life conversation, thereby improving the voice interaction effect.

[0070] Figure 4 The following schematically shows a flow chart of a voice interaction method according to another embodiment of the present disclosure.

[0071] like Figure 4 As shown, the voice interaction method of the embodiment of the present disclosure may include operations S401A to S404A and operations S401B to S404B, wherein operations S401A to S404A are performed by an electronic device, and operations S401B to S404B are performed by a server.

[0072] In operation S401A, under a preset condition, a prompt request for a target voice is generated.

[0073] For example, after the electronic device outputs the answer voice, if the electronic device does not receive the user's next target voice, it can generate a prompt request to request the server to provide a prompt text.

[0074] Exemplarily, the preset condition includes, for example, that the target voice is not received for a preset period of time.

[0075] Alternatively, the preset condition may also include not receiving the target voice but receiving the request voice, that is, when the user cannot make the target voice, the user may make the request voice to request the server to provide the prompt text.

[0076] Alternatively, the preset condition may further include receiving a click operation of the user on the help control. That is, when the user is unable to pronounce the target voice, the user may click the help control to request the server to provide a prompt text.

[0077] In operation S402A, a prompt request is sent.

[0078] In operation S401B, a prompt request is received.

[0079] In operation S402B, a prompt text is generated based on the prompt request.

[0080] For example, since the user cannot provide the target voice for the last answer voice, the server can generate a prompt text based on the answer text corresponding to the last answer voice. If the last answer text is a question, the prompt text can be the answer to the question; if the last answer text is the answer, the prompt text can be the next question.

[0081] In operation S403B, the prompt text is verified.

[0082] In operation S404B, the prompt text is sent.

[0083] For example, before the server obtains the prompt text and sends it to the electronic device, the server can verify the prompt text, for example, to check whether the prompt text is legal and compliant. If the prompt text passes the verification, it means that the prompt text is legal and compliant, and the server can send the verified prompt text to the electronic device.

[0084] In operation S403A, a prompt text is received.

[0085] In operation S404A, a prompt text is output.

[0086] Exemplarily, the prompt text is used to prompt the target object (user) to generate the target voice. For example, the user can utter the next target voice according to the prompt text output by the electronic device. For example, the electronic device can output the prompt text in the form of a pop-up window.

[0087] According to an embodiment of the present disclosure, a variety of preset conditions are configured to help users who have difficulty in conversation, so as to provide prompt information in time and improve the fluency of conversation interaction.

[0088] In another example of the present disclosure, the user can also configure the output mode of the answering voice according to their needs. For example, after the electronic device receives the user's configuration operation, the electronic device can select the output mode for the answering voice based on the configuration operation. The output mode includes, for example, a timbre mode and a pronunciation speed mode. The timbre mode includes, for example, an American pronunciation mode and a British pronunciation mode.

[0089] Figure 5 The block diagram schematically shows a voice interaction device according to an embodiment of the present disclosure.

[0090] like Figure 5 As shown, the voice interaction device 500 of the embodiment of the present disclosure includes, for example, a first generating module 510 , a first sending module 520 , a first output module 530 , a determining module 540 and a second output module 550 .

[0091] The first generating module 510 can be used to generate a dialogue request in response to receiving the target voice. According to the embodiment of the present disclosure, the first generating module 510 can, for example, execute the above reference Figure 2 The operation S210 described above will not be repeated here.

[0092] The first sending module 520 can be used to send a conversation request. According to the embodiment of the present disclosure, the first sending module 520 can, for example, execute the above reference Figure 2The operation S220 described above will not be repeated here.

[0093] The first output module 530 can be used to respond to the received answer text for the dialogue request and output the answer voice based on the answer text. According to the embodiment of the present disclosure, the first output module 530 can, for example, perform the above reference Figure 2 The operation S230 described above will not be repeated here.

[0094] The determination module 540 can be used to determine the interaction scene at any time during the voice interaction process. According to the embodiment of the present disclosure, the determination module 540 can, for example, execute the above reference Figure 2 Operation S240 described above will not be repeated here.

[0095] The second output module 550 can be used to output a virtual interactive image based on the interactive mode associated with the interactive scene. According to an embodiment of the present disclosure, the second output module 550 can, for example, perform the above reference Figure 2 Operation S250 described above will not be described again in detail.

[0096] According to an embodiment of the present disclosure, the interaction scenario includes receiving a target voice scenario or waiting to receive a target voice scenario; the interaction mode includes a listening mode; and the second output module 550 is further used to: based on the listening mode, control the virtual interaction image to generate a listening action.

[0097] According to an embodiment of the present disclosure, the interaction scenario includes an answer voice output scenario; the interaction mode includes an answer mode; and the second output module 550 is further used to: based on the answer mode, control the virtual interaction image to generate an answer action.

[0098] According to an embodiment of the present disclosure, the interaction scenario includes a waiting output answer voice scenario; the interaction mode includes a thinking mode; and the second output module 550 is further used to: based on the thinking mode, control the virtual interaction image to generate a thinking action.

[0099] According to an embodiment of the present disclosure, the apparatus 500 may further include: a second generation module, a second sending module, and a third output module. The second generation module is configured to generate a prompt request for the target speech under preset conditions; the second sending module is configured to send the prompt request; and the third output module is configured to output a prompt text in response to receiving a prompt text in response to the prompt request, wherein the prompt text is used to prompt the target subject to generate the target speech.

[0100] According to an embodiment of the present disclosure, the preset condition includes at least one of the following: failure to receive the target voice for a preset period of time; receipt of a request voice; and receipt of a click operation on the help control.

[0101] According to an embodiment of the present disclosure, the first generation module 510 includes: a recognition submodule for performing speech recognition on the target speech upon receiving the target speech to obtain the target text; and a generation submodule for generating a dialogue request based on the target text.

[0102] According to an embodiment of the present disclosure, the dialogue request includes a target text; and the answer text is obtained based on the target text in the dialogue request using a natural language processing model.

[0103] According to an embodiment of the present disclosure, the apparatus 500 may further include: a selection module for selecting an output mode for the answering speech based on a configuration operation, wherein the output mode includes at least one of a timbre mode and a pronunciation speed mode.

[0104] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0105] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0106] Figure 6 The block diagram is a block diagram of an electronic device for performing voice interaction for implementing an embodiment of the present disclosure.

[0107] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device 600 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0108] like Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0109] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0110] The computing unit 601 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as the voice interaction method. For example, in some embodiments, the voice interaction method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the voice interaction method described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the voice interaction method in any other appropriate manner (e.g., by means of firmware).

[0111] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0112] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable voice interaction device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0113] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0114] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0115] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0116] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0117] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0118] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A voice interaction method, comprising: generating a dialogue request in response to receiving a target speech, wherein the target speech is English speech; sending the conversation request; In response to receiving a response text to the dialogue request, outputting a response voice based on the response text, the response voice being a British voice; Under preset conditions, generating a prompt request for the target voice; sending the prompt request; as well as In response to receiving a prompt text for the prompt request, outputting the prompt text, wherein the prompt text is used to prompt the target object to generate the next target speech, and the prompt text is generated based on the answer text, In the process of voice interaction, the interaction scene at any moment is determined, and a virtual interaction image is output based on the interaction mode associated with the interaction scene. The interaction scenario includes a waiting for output answer voice scenario; the interaction mode includes a thinking mode; Determining the interaction scenario at any time during the voice interaction process includes: Outputting a virtual interactive image based on the interaction mode associated with the interaction scene includes: Based on the thinking pattern associated with the waiting-for-output answer voice scenario, the virtual interactive character is controlled to generate a thinking action.

2. The method according to claim 1, wherein The interaction scenario includes a receiving target voice scenario or a waiting to receive target voice scenario; The interaction mode includes a listening mode; and outputting a virtual interaction image based on the interaction mode associated with the interaction scenario includes: Based on the listening mode, the virtual interactive image is controlled to generate a listening action.

3. The method according to claim 1, wherein The interaction scenario includes an answer voice output scenario; the interaction mode includes an answer mode; and outputting a virtual interaction image based on the interaction mode associated with the interaction scenario includes: Based on the answer mode, the virtual interactive image is controlled to generate an answer action.

4. The method according to claim 1, wherein The preset conditions include at least one of the following: The target voice is not received for a preset period of time; Receive a request voice; Receives a click action on the help control.

5. The method according to claim 1, wherein In response to receiving the target voice, generating a dialogue request includes: In response to receiving the target speech, performing speech recognition on the target speech to obtain a target text; and The dialogue request is generated based on the target text.

6. The method according to claim 5, wherein: The dialogue request includes the target text; the answer text is obtained based on the target text in the dialogue request using a natural language processing model.

7. The method according to any one of claims 1 to 6, further comprising: Based on the configuration operation, an output mode for the answer voice is selected, The output mode includes at least one of a timbre mode and a pronunciation speed mode.

8. A voice interaction device, comprising: A first generating module is configured to generate a dialogue request in response to receiving a target speech, wherein the target speech is an English speech; A first sending module, configured to send the conversation request; a first output module, configured to, in response to receiving a response text to the dialogue request, output a response voice based on the response text, wherein the response voice is a British voice; A second generating module is used to generate a prompt request for the target voice under preset conditions; A second sending module, configured to send the prompt request; as well as a third output module, configured to output the prompt text in response to receiving a prompt text for the prompt request, wherein the prompt text is used to prompt the target object to generate the next target speech, and the prompt text is generated based on the answer text; The voice interaction device further includes a determination module and a second output module. The determination module is used to determine the interaction scene at any time during the voice interaction process. The second output module is used to output a virtual interaction image based on the interaction mode associated with the interaction scene. The interaction scenario includes a waiting for output answer voice scenario; the interaction mode includes a thinking mode; The determining module is further configured to: determine, during the voice interaction process, an interaction scenario of a moment of waiting for the answer text as the waiting-for-output-answer-voice scenario, or, during the voice interaction process, determine an interaction scenario of a moment of receiving the answer text but before outputting the answer voice as the waiting-for-output-answer-voice scenario; The second output module is further configured to control the virtual interactive image to generate a thinking action based on the thinking mode associated with the waiting output answer voice scenario.

9. The device according to claim 8, wherein The interaction scenario includes a receiving target voice scenario or a waiting to receive target voice scenario; the interaction mode includes a listening mode; and the second output module is further configured to: Based on the listening mode, the virtual interactive image is controlled to generate a listening action.

10. The device according to claim 8, wherein The interaction scenario includes an answer voice output scenario; the interaction mode includes an answer mode; and the second output module is further configured to: Based on the answer mode, the virtual interactive image is controlled to generate an answer action.

11. The device according to claim 8, wherein The preset conditions include at least one of the following: The target voice is not received for a preset period of time; Receive a request voice; Receives a click action on the help control.

12. The device according to claim 8, wherein The first generation module includes: a recognition submodule, configured to, in response to receiving a target speech, perform speech recognition on the target speech to obtain a target text; and A generating submodule is used to generate the dialogue request based on the target text.

13. The device according to claim 12, wherein The dialogue request includes the target text; the answer text is obtained based on the target text in the dialogue request using a natural language processing model.

14. The apparatus according to any one of claims 8 to 13, further comprising: A selection module is configured to select an output mode for the answer voice based on a configuration operation. The output mode includes at least one of a timbre mode and a pronunciation speed mode.

15. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.

17. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Voice interaction method, device and system, and voice processing method and device

    CN109346076A

  • Information interaction method and device, electronic equipment, medium and program product

    CN113392201A