Voice interaction method, device, equipment, medium and program product

By establishing a cache mechanism for session ID and TTS parameters between the voice terminal and the cloud, the problem of long wait time for human-computer voice interaction reply process in the prior art is solved, faster voice response is achieved, and user experience is improved.

CN120164464APending Publication Date: 2025-06-17HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510376327.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the prior art, the waiting time of human-computer voice interaction reply process is relatively long, which affects the user experience.

Method used

By establishing a cache mechanism for session ID and TTS parameters between the voice terminal and the cloud, the local TTS parameter processing of the terminal device is reduced, and the voice request and session ID are directly sent to the cloud for processing, and the generated audio stream is pushed from the cloud to the terminal device for playback.

Benefits of technology

It effectively reduces the waiting time for voice interactive replies, improves user experience, and avoids multiple interactions and extended replies synthesis time caused by TTS synthesis by terminal devices and clouds respectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164464A_ABST
    Figure CN120164464A_ABST
Patent Text Reader

Abstract

The invention provides a voice interaction method and device, equipment, a medium and a program product, and relates to the technical field of smart home / smart home. The method comprises the following steps: in response to a wake-up instruction of a user, generating a session ID, and sending the session ID and a TTS parameter corresponding to the session to a cloud end, so that the cloud end caches the TTS parameter and the session ID; in response to a voice request of a user, sending the voice request and the session ID to a cloud, so that the cloud determines a corresponding reply text according to the voice request, determines a corresponding TTS parameter according to the session ID, and generates a corresponding audio stream according to the reply text and the TTS parameter; and obtaining and playing an audio stream generated by the cloud. According to the method, the voice interaction waiting time is shortened, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of smart home / smart family, and particularly relates to a voice interaction method, device, equipment, medium and program product. Background Art

[0002] With the rapid development of artificial intelligence, voice interaction functions are becoming more and more common in the field of smart home, and human-machine voice interaction has gradually become the main interaction method for many products. Therefore, the intelligent requirements for intelligent voice interaction will gradually increase.

[0003] If the terminal device responds to the user's voice slowly during the voice interaction between the user and the terminal device, the user may be dissatisfied or bored with the waiting time, affecting the user experience. In the prior art, during the human-machine voice interaction process, the terminal device needs to send requests to the cloud multiple times to complete the voice response, which will cause the entire voice interaction response process to be long and affect the user experience.

[0004] Therefore, this application proposes a voice interaction method to solve the above problems. Summary of the Invention

[0005] This application provides a voice interaction method, device, equipment, medium and program product to solve the problem in the prior art that the waiting time of the human-machine voice interaction response process is long and affects the user experience.

[0006] In a first aspect, this application provides a voice interaction method, which is applied to a voice terminal and includes:

[0007] Responding to the user's wake-up instruction, generating a session ID, and sending the session ID and the TTS parameters corresponding to this session to the cloud, so that the cloud caches the TTS parameters and the session ID;

[0008] Responding to the user's voice request, sending the voice request and the session ID to the cloud, so that the cloud determines the corresponding reply text according to the voice request, determines the corresponding TTS parameters according to the session ID, and generates the corresponding audio stream according to the reply text and the TTS parameters;

[0009] Obtaining the audio stream generated by the cloud and playing it.

[0010] In a possible implementation manner, the voice terminal is used to respond to the user's voice request in at least two ways, and the at least two ways include a voice playback method and a non-voice playback method;

[0011] Obtaining the audio stream generated by the cloud and playing it includes:

[0012] Obtain the operation information sent by the cloud, and perform corresponding operations according to the operation information, where the operation information is determined by the cloud according to the voice request, and the operation belongs to a non-playing voice method;

[0013] After generating the corresponding audio stream in the cloud, obtain the corresponding audio stream and play it.

[0014] In a possible implementation manner, the operation includes displaying the screen corresponding to the reply text;

[0015] Correspondingly, the obtaining the operation information sent by the cloud and performing corresponding operations according to the operation information includes:

[0016] Obtain the information to be executed sent by the cloud, where the information to be executed includes the screen identifier corresponding to the reply text;

[0017] Display the corresponding screen according to the screen identifier.

[0018] In a possible implementation manner, the voice terminal includes an application layer and a terminal SDK, and the application layer is used to obtain the wake-up instruction input by the user and wake up the terminal SDK;

[0019] The session ID is created by the terminal SDK after being awakened and saved as a global shared variable, so that the application layer sends the TTS parameters and the voice request to the terminal SDK according to the session ID, and the terminal SDK sends them to the cloud.

[0020] In a possible implementation manner, the voice terminal includes an application layer and a terminal SDK;

[0021] The obtaining the operation information sent by the cloud and performing corresponding operations according to the operation information includes:

[0022] Obtain the operation information sent by the cloud through the terminal SDK, and send the operation information to the application layer, so that the application layer performs corresponding operations according to the operation information;

[0023] Correspondingly, the obtaining the corresponding audio stream and playing it after generating the corresponding audio stream in the cloud includes:

[0024] After generating the corresponding audio stream in the cloud, obtain the corresponding audio stream through the terminal SDK and store it;

[0025] Obtain the operation completion instruction sent by the application layer through the terminal SDK, and play the corresponding audio stream after obtaining the operation completion instruction; where the operation completion instruction is generated by the application layer after executing the corresponding operation.

[0026] Second aspect, the present application provides a voice interaction method, which is applied to the cloud. The method includes:

[0027] Receiving a session ID and TTS parameters corresponding to the current session from a voice terminal, and caching the TTS parameters and the session ID;

[0028] Receiving a voice request of a user and the session ID through the voice terminal, determining a corresponding reply text according to the voice request, determining corresponding TTS parameters according to the session ID, and generating a corresponding audio stream according to the reply text and the TTS parameters;

[0029] Sending the audio stream to the voice terminal.

[0030] In a possible implementation, the cloud includes a network connection and protocol adaptation layer, a scheduling service, and a speech synthesis service;

[0031] The network connection and protocol adaptation layer is used to receive the session ID and the TTS parameters corresponding to the current session, cache the TTS parameters and the session ID, receive the voice request of the user, and send the voice request and the session ID to the scheduling service;

[0032] The scheduling service is used to determine a corresponding reply text according to the voice request, determine corresponding TTS parameters according to the session ID, and send the corresponding reply text and the corresponding TTS parameters to the synthesis service;

[0033] The speech synthesis service is used to generate a corresponding audio stream according to the reply text and the TTS parameters, and send the audio stream to the voice terminal through the network connection and protocol adaptation layer.

[0034] Third aspect, the present application provides a voice interaction device, which is applied to a voice terminal and includes:

[0035] A first processing module, configured to generate a session ID in response to a wake-up instruction of a user, and send the session ID and TTS parameters corresponding to the current session to the cloud, so that the cloud caches the TTS parameters and the session ID;

[0036] A second processing module, configured to send the voice request and the session ID to the cloud in response to a voice request of the user, so that the cloud determines a corresponding reply text according to the voice request, determines corresponding TTS parameters according to the session ID, and generates a corresponding audio stream according to the reply text and the TTS parameters;

[0037] A playback module, configured to obtain the audio stream generated by the cloud and play it.

[0038] Fourthly, the present application provides a voice interaction device applied to the cloud, including:

[0039] A cache module, configured to receive a session ID and TTS parameters corresponding to the current session from a voice terminal, and cache the TTS parameters and the session ID;

[0040] A generation module, configured to receive a user's voice request and a session ID through the voice terminal, determine a corresponding reply text according to the voice request, determine corresponding TTS parameters according to the session ID, and generate a corresponding audio stream according to the reply text and the TTS parameters;

[0041] A transmission module, configured to send the audio stream to the voice terminal.

[0042] Fifthly, the present application provides an electronic device, including: at least one processor and a memory;

[0043] The memory stores computer-executable instructions;

[0044] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the voice interaction method as described above.

[0045] Sixthly, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the voice interaction method as described above are implemented.

[0046] Seventhly, the present application provides a computer program product, the computer program product includes instructions, and when the instructions are executed on an electronic device, the electronic device is enabled to implement the voice interaction method as described above.

[0047] A voice interaction method, device, equipment, medium and program product provided by the present application generate a session ID in response to a user's wake-up instruction, and send the session ID and TTS parameters corresponding to the current session to the cloud, so that the cloud caches the TTS parameters and the session ID; in response to a user's voice request, send the voice request and the session ID to the cloud, so that the cloud determines a corresponding reply text according to the voice request, determines corresponding TTS parameters according to the session ID, and generates a corresponding audio stream according to the reply text and the TTS parameters; obtain the audio stream generated by the cloud and play it.

[0048] In the above method, the voice terminal and the cloud cooperate with each other to perform voice processing and give a voice reply to the user; the user wakes up the voice terminal through a wake-up command, generates a session ID for this session, and sends the TTS parameters corresponding to this session to the cloud, allowing the cloud to take over the TTS parameters of the voice terminal; the user initiates a voice request and uses the same session ID; the cloud receives the voice request and session ID, generates a corresponding reply text according to the voice request, finds the corresponding TTS parameters in the cache according to the session ID, and completes the assembly of the reply text and TTS parameters to form an audio stream; the voice terminal obtains the audio stream and plays it; the reply is not synthesized based on the TTS parameters at the voice terminal, so as to avoid multiple interactions and prolonged reply synthesis time caused by the voice terminal and the cloud performing reply synthesis separately. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0050] Figure 1 A hardware environment diagram of voice interaction provided in an embodiment of the present application;

[0051] Figure 2 A flow chart of a voice interaction method provided in an embodiment of the present application Figure 1 ;

[0052] Figure 3 A flow chart of a voice interaction method provided in an embodiment of the present application Figure 2 ;

[0053] Figure 4 A flow chart of a voice interaction method provided in an embodiment of the present application Figure 3 ;

[0054] Figure 5 A flow chart of a voice interaction method provided in an embodiment of the present application Figure 4 ;

[0055] Figure 6 A voice interaction device provided by an embodiment of the present invention Figure 1 ;

[0056] Figure 7 A voice interaction device provided by an embodiment of the present invention Figure 2 ;

[0057] Figure 8 A hardware schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0059] It should be noted that the terms "first", "second" etc. of the present application are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable in appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, the process, method, system, product or equipment comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0060] The existing complete voice interaction full-link process is as follows: the user wakes up the voice terminal and initiates a voice request to the voice terminal, that is, an Automatic Speech Recognition (ASR) request, which corresponds to a session; the voice terminal receives the user's voice request and uploads the user's voice request to the cloud. The cloud begins to convert the voice of the user's voice request into text to understand and assemble a reply text. The cloud feeds the assembled reply text back to the voice terminal, and the first session corresponding to the ASR request is completed; after completing the first session, the voice terminal performs secondary assembly according to the voice terminal's requirements. The voice terminal uploads the secondary assembled content of the voice terminal to the cloud and initiates a Text To Speech (TTS) request to the cloud, which corresponds to another session; after receiving the TTS request, the cloud begins to convert the secondary assembled content into voice, generates an audio stream and pushes it to the voice terminal, completing the voice reply to the user.

[0061] The processing process of the existing technology is relatively complicated, which increases the speech synthesis feedback time; different voice terminals perform secondary assembly according to different logics and then request the cloud, which will also result in different interaction times due to different processing logics, which is not conducive to interaction efficiency.

[0062] Therefore, the present application proposes a voice interaction method to solve the above problems.

[0063] The following describes an implementation process of voice interaction proposed in the present application in conjunction with the accompanying drawings and specific embodiments.

[0064] Figure 1 A hardware environment diagram of a voice interaction provided in an embodiment of the present application; According to one aspect of an embodiment of the present application, a voice interaction method is provided. The voice interaction method is widely used in smart home (Smart Home), smart home, smart home device ecology, smart residential (Intelligence House) ecology and other whole-house intelligent digital control application scenarios. Optionally, in this embodiment, the above-mentioned voice interaction method can be applied to Figure 1 In the hardware environment composed of the terminal device 102 and the server 104 shown in FIG. Figure 1 As shown, the server 104 is connected to the terminal device 102 via a network, and can be used to provide services (such as application services, etc.) for the terminal or a client installed on the terminal. A database can be set on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data computing services for the server 104.

[0065] The network may include but is not limited to at least one of the following: wired network, wireless network. The wired network may include but is not limited to at least one of the following: wide area network, metropolitan area network, local area network, and the wireless network may include but is not limited to at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 may be but is not limited to a PC, a mobile phone, a tablet computer, a smart air conditioner, a smart range hood, a smart refrigerator, a smart oven, a smart stove, a smart washing machine, a smart water heater, a smart washing equipment, a smart dishwasher, a smart projection equipment, a smart TV, a smart clothes drying rack, a smart curtain, a smart audio and video, a smart socket, a smart speaker, a smart speaker, a smart fresh air equipment, a smart kitchen and bathroom equipment, a smart bathroom equipment, a smart sweeping robot, a smart window cleaning robot, a smart mopping robot, a smart air purification equipment, a smart steamer, a smart microwave oven, a smart kitchen treasure, a smart purifier, a smart water dispenser, a smart door lock, etc.

[0066] Figure 2 A flow chart of a voice interaction method provided in an embodiment of the present application Figure 1 .like Figure 2 As shown, the method is applied to a voice terminal, and the method includes:

[0067] S201. In response to a user's wake-up instruction, generate a session ID, and send the session ID and TTS parameters corresponding to this session to the cloud, so that the cloud caches the TTS parameters and the session ID.

[0068] A wake-up instruction is a command used to activate a device or system. Common ones include voice commands, button presses, and gestures. Its purpose is to bring the device from the standby state into the active state so that it can receive further instructions or requests. In voice interaction, a voice terminal often presets a wake-up word as the corresponding wake-up instruction, and the wake-up word is usually a specific phrase. When a user interacts with the voice terminal and gives a wake-up instruction, the voice terminal responds to the wake-up instruction to generate a session ID, and sends the session ID and the TTS parameters corresponding to this session to the cloud. The cloud will cache the TTS parameters and the session ID.

[0069] The voice terminal has its own set of processing logics for replying to the user's voice requests, including the TTS parameter configuration logic that affects the final voice quality, style, and characteristics, etc. The TTS parameters include speech rate, volume, pitch, timbre, language, pause, emotional expression, etc. After establishing a session connection with the cloud, the terminal can upload its TTS parameters to the cloud, allowing the cloud to take over the processing related to the terminal's TTS parameters, avoiding the terminal from initiating another request for the processing result of the TTS parameters.

[0070] S202. In response to the user's voice request, send the voice request and the session ID to the cloud, so that the cloud can determine the corresponding reply text according to the voice request, determine the corresponding TTS parameters according to the session ID, and generate a corresponding audio stream according to the reply text and the TTS parameters.

[0071] After the user wakes up the voice terminal, there will be a further voice request. This voice request establishes the same session as the wake-up instruction and uses the same session ID. The voice terminal sends the voice request and the session ID to the cloud, and the cloud can then track the same session according to the session ID, obtain the TTS parameters corresponding to this session ID, and the cloud generates a corresponding audio stream according to the reply text and the TTS parameters. During this process, the voice terminal does not initiate two sessions again.

[0072] The user's voice request includes the voice content spoken by the user during the voice interaction with the terminal. For example, the weather in Beijing. The voice terminal forwards the received user voice request to the cloud for voice recognition and processing. After receiving the user's voice, the cloud understands and analyzes the user's intention and needs based on the voice understanding function, selects a response method, and modifies the response method based on the TTS parameters to obtain a modified voice reply text. The voice reply text is, for example, the weather and temperature in Beijing, and the modification is, for example, a gentle tone.

[0073] S203. Obtain the audio stream generated by the cloud and play it.

[0074] After the voice terminal makes a voice request, other work related to the voice request can be processed while the cloud generates the audio stream. If the speed at which the cloud generates the audio stream is faster than the speed at which the voice terminal processes other work related to the voice request, the voice terminal can cache the audio stream and wait for the work related to the voice request to be completed before responding to the user. Otherwise, the voice terminal needs to wait for the audio stream to be generated by the cloud and respond to the user together with the work related to the completed voice request.

[0075] After the audio stream playback is completed, the next time the user wakes up the voice terminal, the session ID will be established again.

[0076] TTS parameters are data pre-arranged and stored in voice terminals. Each voice terminal has its own set of TTS parameters. The TTS parameters of voice terminals of the same type are generally the same, and the TTS parameters of voice terminals of different types may be the same or different. The TTS parameters of each voice terminal do not change in real time. Whether the TTS parameters change depends on the settings of the developer, including when the TTS parameters of the voice terminal are updated and changed when the upgrade is carried out according to the demand or the TTS parameter configuration is adjusted.

[0077] In order to allow the user-initiated session to obtain new TTS parameters, each time the user wakes up the voice terminal, the voice terminal will obtain the TTS parameters corresponding to this session. After the cloud synthesizes the audio stream, the TTS parameters can be released at the set time.

[0078] In an embodiment of the present application, the terminal and the cloud cooperate with each other to perform voice processing and provide a voice reply to the user. The user wakes up the voice terminal through a wake-up command, generates a session ID for this session, and sends the TTS parameters corresponding to this session to the cloud, allowing the cloud to take over the TTS parameters of the voice terminal; the user initiates a voice request and uses the same session ID; the cloud receives the voice request and session ID, generates a corresponding reply text according to the voice request, finds the corresponding TTS parameters in the cache according to the session ID, and completes the assembly of the reply text and TTS parameters to form an audio stream; the voice terminal obtains the audio stream and plays it; the reply is not synthesized based on the TTS parameters in the voice terminal, so as to avoid multiple interactions and prolonged reply synthesis time caused by separate reply synthesis by the voice terminal and the cloud.

[0079] Figure 3 A flow chart of a voice interaction method provided in an embodiment of the present application Figure 2 .like Figure 3 As shown, the voice terminal is used to respond to the user's voice request in at least two ways, wherein the at least two ways include a voice playback way and a non-voice playback way. The method includes:

[0080] S301. Obtain the operation information sent by the cloud, and perform the corresponding operation according to the operation information. The operation information is determined by the cloud according to the voice request, and the operation belongs to a non-voice-playing method.

[0081] Before the cloud sends the operation information to the voice terminal, the synthesis of the reply text has been completed. According to the content of the reply text, the operation information is obtained; the voice terminal can also perform corresponding operations according to the operation information; this operation is not in the form of playing voice. During the process of the voice terminal performing the logic processing related to the operation, the cloud synchronously synthesizes the audio stream according to the reply text and the TTS parameters; synchronizes the reply process, thereby saving the time to respond to the user and promptly responding to the user's voice request.

[0082] Specific content of the operation:

[0083] Exemplarily, the operation includes displaying the screen corresponding to the reply text;

[0084] Correspondingly, the obtaining of the operation information sent by the cloud and performing the corresponding operation according to the operation information includes:

[0085] Obtain the information to be executed sent by the cloud, and the information to be executed includes the screen identifier corresponding to the reply text;

[0086] Display the corresponding screen according to the screen identifier.

[0087] If it is necessary for the voice terminal to display the screen related to the content in the reply text, then the cloud can generate the information to be executed according to the reply text. The information to be executed is the screen related to the content in the reply text. This screen can be a kind of identifier, and the display interface of the voice terminal can display the corresponding screen at a preset position; for example, when the weather is sunny, the display interface of the voice terminal can display a sun symbol.

[0088] S302. After the corresponding audio stream is generated in the cloud, obtain the corresponding audio stream and play it.

[0089] The playing of the audio stream and the operation of the voice terminal can be set to be synchronously responsive; when the voice terminal starts to play the audio stream, it also starts to perform the operation; if any one of the two processes of generating the operation and generating the audio stream is completed later, the process that is completed first can wait for the process that is completed later.

[0090] In the embodiments of the present application, by displaying the screen related to the reply text, the user can not only hear the voice reply but also obtain visual feedback, making the information transmission more intuitive.

[0091] Figure 4 It is a schematic flowchart of a voice interaction method provided by the embodiments of the present application Figure 3 . Such asFigure 4 As shown, the voice terminal includes an application layer and a terminal SDK. The application layer is used to obtain a wake-up instruction input by a user and wake up the terminal SDK;

[0092] The session ID is created and saved as a global shared variable by the terminal SDK after being woken up, so that the application layer can send the TTS parameters and the voice request to the terminal SDK according to the session ID, and the terminal SDK sends them to the cloud.

[0093] The terminal SDK, that is, the Software Development Kit (SDK) of the terminal, is used to develop and operate terminal devices. These terminal devices can be various types of hardware devices, such as mobile devices, Internet of Things devices, smart home devices, etc.; the terminal SDK can provide interfaces and libraries; the application layer of the terminal can interact with users; the terminal SDK can establish connections with the application layer and the cloud;

[0094] The application layer receives the user's wake-up instruction and sends the wake-up instruction to the terminal SDK; the terminal SDK receives the wake-up instruction, generates a session ID; saves the session ID as a global shared variable, so that the user voice requests of the same user can be tracked;

[0095] The process of the terminal SDK establishing connections with the application layer and the cloud can be after establishing the session ID;

[0096] The terminal SDK sends feedback to the application layer to save the session ID as a global shared variable; the terminal SDK establishes a communication connection with the application layer and a communication connection with the cloud; after the terminal SDK saves the session ID as a global shared variable, it gives feedback to the application layer to complete the establishment of the session ID, and completes the process of establishing the session ID; after completing the process of establishing the session ID, a communication connection is established between the terminal SDK and the application layer, and between the terminal SDK and the cloud; the application layer and the cloud interact through the terminal SDK.

[0097] As a global shared variable, the session ID can ensure that all devices / services processing this session use the same identifier, avoiding errors caused by data inconsistency; for example, the session ID can be sent from the application layer to the terminal SDK and then to the cloud. Correspondingly, the TTS parameters and the voice request can be sent from the application layer to the terminal SDK and then to the cloud corresponding to this session ID. After the cloud caches the TTS parameters, it can feedback to the terminal.

[0098] The specific structure and processing of the voice terminal:

[0099] Exemplarily, the voice terminal includes an application layer and a terminal SDK;

[0100] Obtaining the operation information sent by the cloud and performing corresponding operations according to the operation information includes:

[0101] Obtaining the operation information sent by the cloud through the terminal SDK and sending the operation information to the application layer so that the application layer performs corresponding operations according to the operation information;

[0102] Correspondingly, after generating a corresponding audio stream in the cloud, obtaining and playing the corresponding audio stream includes:

[0103] After generating a corresponding audio stream in the cloud, obtaining and storing the corresponding audio stream through the terminal SDK;

[0104] Obtaining the operation completion instruction sent by the application layer through the terminal SDK, and after obtaining the operation completion instruction, playing the corresponding audio stream; wherein, the operation completion instruction is generated by the application layer after executing the corresponding operation.

[0105] When the cloud synthesizes the voice reply text, it will give two feedbacks to the terminal. One is the operation information. After receiving this operation information, the audio terminal does not need to send a new request, and only needs to process the operations related to the voice reply text within the terminal; the terminal SDK receives the operation information and forwards it to the application layer to let the application layer process the operations related to the voice reply text in the terminal. For example, for the display panel of the terminal, when it is confirmed that the weather in Beijing is sunny, a sun icon can be displayed.

[0106] Another feedback given to the terminal when the cloud synthesizes the voice reply text is the audio stream; after the cloud synthesizes the voice reply text and performs audio synthesis, after the cloud synthesizes the audio stream of the voice reply text, the feedback is pushed to the terminal SDK, and the terminal SDK stores the received audio stream; when the application layer of the terminal completes the relevant operation processing, it notifies the terminal SDK to perform voice broadcast, and the terminal SDK broadcasts the voice reply text synthesized by the cloud. For example, Hello, the current weather in Beijing is sunny, and the temperature is 30°C.

[0107] If the audio stream is synthesized first, it can be stored in the terminal SDK; after the application layer processes the logic of the relevant operations, it sends a feedback of completing the relevant operations to the terminal SDK. After receiving the feedback of completing the relevant operations, the terminal SDK finds the audio stream of the voice reply text corresponding to the user in the audio stream queue and directly plays it.

[0108] In the embodiments of the present application, by synthesizing the reply text and generating the operation information in the cloud, when the voice terminal processes the operation information, the cloud synthesizes the audio stream. The synchronous work can quickly respond to user requests, ensure the smoothness of the user experience, and reduce the waiting time.

[0109] By sending TTS parameters during the interaction between the edge side and the cloud side in the speech recognition stage, there is no need for the terminal side to request the cloud for TTS synthesis again. After obtaining the result of speech recognition, semantic understanding is directly performed to obtain the response result. On the one hand, the relevant logic for terminal processing is returned, and on the other hand, the speech script for broadcasting is assembled in parallel and TTS pre-synthesis is carried out in advance to push the audio stream to the terminal. In this way, the terminal side can perform speech broadcasting after processing the relevant behavior logic or in parallel. The terminal processing logic and TTS synthesis are carried out in parallel. The terminal provides an audio cache channel, and the audio synthesized by the cloud can be directly pushed to the terminal in a streaming manner for caching or directly for broadcasting.

[0110] Minimize the impact on the terminal application layer and do not add new processing modules; under the existing process of the application layer, parallelize the TTS processing lead time and cloud speech synthesis.

[0111] Figure 5 Schematic flow of a voice interaction method provided by an embodiment of the present application Figure 4 As Figure 5 shown, applied to the cloud, the method includes:

[0112] Receive the session ID and the TTS parameters corresponding to the current session from the voice terminal, and cache the TTS parameters and the session ID;

[0113] Receive the user's voice request and the session ID through the voice terminal, determine the corresponding reply text according to the voice request, determine the corresponding TTS parameters according to the session ID, and generate the corresponding audio stream according to the reply text and the TTS parameters;

[0114] Send the audio stream to the voice terminal.

[0115] The cooperation of the cloud with the voice terminal on the other side to generate the audio stream to complete the process of voice interaction has been described in detail in the above embodiments and will not be elaborated here.

[0116] Specific structure and processing of the cloud:

[0117] Exemplarily, the cloud includes a network connection and protocol adaptation layer, a scheduling service, and a speech synthesis service;

[0118] The network connection and protocol adaptation layer is used to receive the session ID and the TTS parameters corresponding to the current session, cache the TTS parameters and the session ID, receive the user's voice request, and send the voice request and the session ID to the scheduling service;

[0119] The scheduling service is used to determine the corresponding reply text according to the voice request, determine the corresponding TTS parameters according to the session ID, and send the corresponding reply text and the corresponding TTS parameters to the synthesis service;

[0120] The voice synthesis service is used to generate a corresponding audio stream according to the reply text and TTS parameters, and send the audio stream to the voice terminal through the network connection and protocol adaptation layer.

[0121] To accelerate the audio stream push, during the operation information response process on the terminal side, there is also corresponding cooperation on the cloud side; the cloud includes a network connection and protocol adaptation layer and a scheduling service. Among them, the network connection and protocol adaptation layer is used to establish a network connection and perform protocol adaptation with the terminal to ensure communication between the terminal and the cloud; the scheduling service is used to synthesize the voice reply text.

[0122] The network connection and protocol adaptation layer receives the user voice request from the terminal and sends it to the scheduling service. After receiving the request, the scheduling service performs speech recognition, semantic understanding, and information retrieval based on the speech understanding function, and combines the TTS parameters to synthesize the corresponding voice reply text.

[0123] The scheduling service is responsible for generating the voice reply text; the network connection and protocol adaptation layer is responsible for interacting with the terminal SDK of the terminal, and feedback the situation of the scheduling service synthesizing the voice reply text to the terminal: synchronize the voice reply text to the terminal SDK, and the terminal SDK then synchronizes it to the application layer; while synchronizing the text, the network connection and protocol adaptation layer also modifies the existing protocol and notifies the terminal to perform direct push streaming. The terminal knows that there is no need to request again and starts to complete the relevant services.

[0124] In the embodiment of the present application, the cloud receives the TTS parameters of the terminal. The scheduling service synthesizes the voice reply text for the user's reply according to the speech understanding function and the TTS parameters uploaded by the terminal, sends a feedback of the voice reply text to the terminal, and notifies the terminal to perform the subsequent relevant service processing of the voice reply text without requesting again, thereby increasing the processing efficiency.

[0125] The cloud also includes a voice synthesis service for synthesizing an audio stream according to the voice reply text; the scheduling service synthesizes the voice reply text and sends two feedbacks. The operation information is fed back to the terminal SDK through the network connection and protocol adaptation layer to notify the terminal to perform relevant service processing.

[0126] When the scheduling service sends out the operation information, it sends the voice reply text to the voice synthesis service. The voice synthesis service generates an audio stream according to the voice reply text and generates a second feedback instruction carrying the audio stream; the voice synthesis service sends the second feedback instruction to the network connection and protocol adaptation layer, and the network connection and protocol adaptation layer sends the second feedback instruction to the terminal SDK. The terminal SDK obtains the audio stream in the second feedback instruction and caches it; when the terminal application layer completes the relevant service, it can notify the terminal SDK to read the audio stream for broadcasting.

[0127] In the embodiments of the present application, the speech response text is synthesized in the cloud, and two feedbacks are performed. The terminal is allowed to perform related service processing, and the text-to-audio processing is also synchronized to improve the response efficiency.

[0128] Figure 6 A speech interaction device provided by an embodiment of the present invention Figure 1 , as Figure 6 shown, is applied to a speech terminal. The device includes: a first processing module 601, a second processing module 602, and a playback module 603;

[0129] The first processing module 601 is configured to generate a session ID in response to a user's wake-up instruction, and send the session ID and the TTS parameters corresponding to the current session to the cloud, so that the cloud caches the TTS parameters and the session ID;

[0130] The second processing module 602 is configured to send the voice request and the session ID to the cloud in response to a user's voice request, so that the cloud determines a corresponding response text according to the voice request, determines corresponding TTS parameters according to the session ID, and generates a corresponding audio stream according to the response text and the TTS parameters;

[0131] The playback module 603 is configured to obtain and play the audio stream generated by the cloud.

[0132] The playback module 603 is further configured to the speech terminal is used to respond to the user's voice request in at least two ways, and the at least two ways include a voice playback method and a non-voice playback method;

[0133] Obtaining and playing the audio stream generated by the cloud includes:

[0134] Obtaining the operation information sent by the cloud, and performing a corresponding operation according to the operation information, where the operation information is determined by the cloud according to the voice request, and the operation belongs to a non-playing voice method;

[0135] After the cloud generates a corresponding audio stream, obtaining and playing the corresponding audio stream.

[0136] The playback module 603 is further configured to the operation includes displaying a screen corresponding to the response text;

[0137] Correspondingly, the obtaining the operation information sent by the cloud and performing a corresponding operation according to the operation information includes:

[0138] Obtaining the to-be-executed information sent by the cloud, where the to-be-executed information includes a screen identifier corresponding to the response text;

[0139] Displaying a corresponding screen according to the screen identifier.

[0140] Figure 7 A voice interaction device provided by an embodiment of the present invention Figure 2 , as Figure 7 shown, is applied to the cloud. The device includes: a cache module 701, a generation module 702, and a transmission module 703;

[0141] The cache module 701 is configured to receive a session ID and TTS parameters corresponding to the current session from a voice terminal, and cache the TTS parameters and the session ID.

[0142] The generation module 702 is configured to receive a user's voice request and a session ID through the voice terminal, determine a corresponding reply text according to the voice request, determine corresponding TTS parameters according to the session ID, and generate a corresponding audio stream according to the reply text and the TTS parameters.

[0143] The transmission module 703 is configured to send the audio stream to the voice terminal.

[0144] The present application also provides an electronic device, including: at least one processor and a memory;

[0145] The memory stores computer-executable instructions;

[0146] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes a voice interaction method.

[0147] Figure 8 It is a hardware schematic diagram of the electronic device provided by an embodiment of the present invention. As Figure 8 shown, the electronic device 80 provided in this embodiment includes: at least one processor 801 and a memory 802. The device 80 further includes a communication component 803. Among them, the processor 801, the memory 802, and the communication component 803 are connected through a bus 804.

[0148] In a specific implementation process, at least one processor 801 executes the computer-executable instructions stored in the memory 802, so that at least one processor 801 executes the above voice interaction method.

[0149] For the specific implementation process of the processor 801, reference can be made to the above method embodiment, and its implementation principle and technical effects are similar, so they will not be elaborated here in this embodiment.

[0150] In the above Figure 8In the illustrated embodiments, it should be understood that the processor may be a central processing unit (CPU for short), or other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0151] The memory may include a random access memory (RAM), and may also include non-volatile memory (NVM), such as at least one disk memory.

[0152] The bus may be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.

[0153] This application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the method described above is implemented.

[0154] For the above-mentioned computer-readable storage medium, the above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a disk or an optical disc. The readable storage medium can be any available medium accessible by a general or special-purpose computer.

[0155] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0156] The division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the couplings or direct couplings or communication connections shown or discussed between each other can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0157] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0158] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: ROM, RAM, magnetic disk, optical disk, or other media that can store program code.

[0159] Finally, it should be noted that: After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other implementation schemes of the present invention. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. A voice interaction method, characterized in that: Applied to a voice terminal, the method comprises: In response to the user's wake-up instruction, a session ID is generated, and the session ID and the TTS parameters corresponding to the current session are sent to the cloud, so that the cloud caches the TTS parameters and the session ID; In response to the user's voice request, the voice request and the session ID are sent to the cloud, so that the cloud determines a corresponding reply text according to the voice request, determines a corresponding TTS parameter according to the session ID, and generates a corresponding audio stream according to the reply text and the TTS parameter; Get the audio stream generated by the cloud and play it.

2. The method according to claim 1, characterized in that The voice terminal is used to respond to the voice request of the user in at least two ways, and the at least two ways include a voice playing way and a non-voice playing way; Get the audio stream generated by the cloud and play it, including: Obtaining operation information sent by the cloud, and performing a corresponding operation according to the operation information, wherein the operation information is determined by the cloud according to the voice request, and the operation is a non-voice playback method; After the corresponding audio stream is generated in the cloud, the corresponding audio stream is obtained and played.

3. The method according to claim 2, characterized in that The operation includes displaying a screen corresponding to the reply text; Accordingly, the obtaining of the operation information sent by the cloud and performing corresponding operations according to the operation information includes: Obtaining the to-be-executed information sent by the cloud, wherein the to-be-executed information includes a screen identifier corresponding to the reply text; The corresponding picture is displayed according to the picture identifier.

4. The method according to any one of claims 1 to 3, characterized in that: The voice terminal includes an application layer and a terminal SDK, wherein the application layer is used to obtain a wake-up instruction input by a user and wake up the terminal SDK; The session ID is created by the terminal SDK after being awakened and saved as a global shared variable, so that the application layer sends the TTS parameters and the voice request to the terminal SDK according to the session ID, and the terminal SDK sends them to the cloud.

5. The method according to claim 2 or 3, characterized in that: The voice terminal includes an application layer and a terminal SDK; The obtaining of the operation information sent by the cloud and performing corresponding operations according to the operation information includes: Acquire operation information sent by the cloud through the terminal SDK, and send the operation information to the application layer, so that the application layer performs corresponding operations according to the operation information; Correspondingly, after the corresponding audio stream is generated in the cloud, the corresponding audio stream is obtained and played, including: After the corresponding audio stream is generated in the cloud, the corresponding audio stream is obtained and stored through the terminal SDK; The operation completion instruction sent by the application layer is obtained through the terminal SDK, and after obtaining the operation completion instruction, the corresponding audio stream is played; wherein the operation completion instruction is generated by the application layer after executing the corresponding operation.

6. A voice interaction method, characterized in that: Applied to the cloud, the method includes: Receive a session ID and TTS parameters corresponding to the current session from the voice terminal, and cache the TTS parameters and session ID; Receiving a user's voice request and a session ID through the voice terminal, determining a corresponding reply text according to the voice request, determining a corresponding TTS parameter according to the session ID, and generating a corresponding audio stream according to the reply text and the TTS parameter; The audio stream is sent to the voice terminal.

7. The method according to claim 6, characterized in that The cloud includes a network connection and protocol adaptation layer, a scheduling service, and a speech synthesis service; The network connection and protocol adaptation layer is used to receive the session ID and the TTS parameters corresponding to the current session, cache the TTS parameters and the session ID, receive the user's voice request and send the voice request and the session ID to the scheduling service; The scheduling service is used to determine the corresponding reply text according to the voice request, determine the corresponding TTS parameters according to the session ID, and send the corresponding reply text and the corresponding TTS parameters to the synthesis service; The speech synthesis service is used to generate a corresponding audio stream according to the reply text and TTS parameters, and send the audio stream to the voice terminal through the network connection and protocol adaptation layer.

8. An electronic device, characterized in that: include: at least one processor and memory; The memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the voice interaction method as described in any one of claims 1-7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the voice interaction method as described in any one of claims 1 to 7 are implemented.

10. A computer program product, characterized in that The computer program product includes instructions, which, when executed on an electronic device, enable the electronic device to implement the voice interaction method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Audio data playing method and device based on voice interaction, storage medium and computer program product

    CN120729660A