Voice processing method and system
By analyzing and updating the input data on the voice processing server, the problem of users being unable to interrupt machine output is solved, achieving natural and smooth human-computer interaction. The machine can quickly respond to user input, improving the interactive experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, users cannot interrupt the machine's output responses during human-computer interaction, resulting in unnatural interactions and an inability to achieve natural and fluent dialogue.
By analyzing the input data on the voice processing server to determine whether the voice processing conditions are met, a status update command is sent to change the voice output state of the device to the voice input state, and the input data is parsed to obtain the response data, so that the device can still receive input data and respond quickly even in the voice output state.
It enables machines to speak and receive feedback simultaneously with extremely low latency, achieving a natural human-computer interaction effect. Users can interrupt the machine's output and respond promptly when appropriate.
Smart Images

Figure CN121747597A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present specification relate to the technical field of speech processing, and in particular to a speech processing method and system. BACKGROUND
[0002] At present, large models have advantages in text-related understanding and generation tasks and are widely used. However, when the large model is applied, the user and the machine (the machine uses the large model to generate and output a reply to the user's input question) usually have a one-question-one-answer conversation, and the delay of the large model reply is longer.
[0003] With the improvement of user demand, users expect a natural and smooth conversation with the machine, and the machine can speak and receive feedback at the same time and give a quick response, like human-to-human communication. However, recent research has less explored how to achieve full-duplex interaction of multi-modal large models, and the existing multi-modal speech interaction model has the problem that the user cannot interrupt the AI (Artificial Intelligence), which leads to the user being unable to interrupt the machine's output reply when the human-machine full-duplex interaction is performed, and the natural human-machine interaction cannot be achieved. SUMMARY
[0004] Therefore, the embodiments of the present specification provide a speech processing method. One or more embodiments of the present specification also provide a speech processing system, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects in the prior art that the user cannot interrupt the machine's output reply when the human-machine full-duplex interaction is performed, and the natural human-machine interaction cannot be achieved.
[0005] According to a first aspect of the embodiments of the present specification, a speech processing method is provided, applied to a speech processing server, comprising: In a case where the device end is in a speech output state and input data sent by the device end is received, performing data analysis on the input data to obtain a data analysis result; In a case where the data analysis result represents that the speech data meets a first speech processing condition, sending a first state update instruction to the device end, wherein the first state update instruction is used to update the device end from the speech output state to a speech input state; Analyzing the speech data to obtain reply data corresponding to the speech data, and sending the reply data to the device end.
[0006] According to a second aspect of the embodiments of the present specification, a speech processing system is provided, comprising a speech processing server and a device end, wherein: The device end is configured to send received input data to the speech processing server. The voice processing server is configured to, in a case where the device end is in a voice output state and input data sent by the device end is received, perform data analysis on the voice data to obtain a data analysis result, and in a case where the data analysis result indicates that the input data satisfies a first voice processing condition, send a first state update instruction to the device end. The device end is further configured to update from the voice output state to a voice input state according to the first state update instruction. The voice processing server is further configured to parse the input data to obtain reply data corresponding to the input data, and send the reply data to the device end.
[0007] According to a third aspect of an embodiment of the present specification, a computing device is provided, including: a memory and a processor; The memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, which realize the steps of the voice processing method described above when executed by the processor.
[0008] According to a fourth aspect of an embodiment of the present specification, a computer readable storage medium is provided, which stores computer programs / instructions, which realize the steps of the voice processing method described above when executed by the processor.
[0009] According to a fifth aspect of an embodiment of the present specification, a computer program product is provided, including computer programs / instructions, which realize the steps of the voice processing method described above when executed by the processor.
[0010] The voice processing method provided by one embodiment of the present specification, in a case where the device end is in a voice output state and input data sent by the device end is received, judges whether the input data satisfies a first voice processing condition according to a data analysis result obtained by performing data analysis on the input data, and in a case where the input data satisfies the first voice processing condition, changes the voice output state of the device end to a voice input state by sending a first state update instruction to the device end, parses the input data to obtain reply data, and sends the reply data to the device end; that is, the device end can still receive input data and forward the input data to the voice processing server in a case where the device end is in a voice output state, and the voice processing server can update the state of the device end in a case where the input data satisfies the first voice processing condition by analyzing the input data, and by timely receiving voice data and updating the state of the device end, the machine can speak and receive feedback at the same time, and in a very low time delay, obtain reply data to give a quick response, so that the human-computer interaction can achieve a very natural effect. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 This is a schematic diagram of a scenario illustrating a speech processing method provided in one embodiment of this specification; Figure 2 This is a flowchart illustrating a speech processing method provided in one embodiment of this specification; Figure 3 This is a schematic diagram of the dialogue state control model in a speech processing method provided in one embodiment of this specification; Figure 4 This is a flowchart illustrating the processing procedure of a speech processing method provided in one embodiment of this specification. Figure 5 This is a flowchart illustrating the processing procedure of another speech processing method provided in one embodiment of this specification. Figure 6 This is a schematic diagram of the structure of a speech processing system provided in one embodiment of this specification; Figure 7 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0012] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0013] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0014] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0015] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0016] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0017] Full-duplex refers to a communication system in which both parties can send and receive data simultaneously, similar to a telephone conversation where both parties can speak and hear each other at the same time.
[0018] Multimodal Large Model: A multimodal large model is a large machine learning model that can process and associate information from multiple modalities (such as text, speech, and images), and is often used to improve the machine's understanding and interaction capabilities.
[0019] Speech encoding is the process of converting speech signals into a digital format that can be processed by a computer, typically involving sampling, quantization, and compression; in the embodiments of this specification, it is used to convert audio into feature inputs for a multimodal large model.
[0020] Turn-Taking is the process of alternating speaking in a dialogue. The key is to effectively identify when one participant finishes speaking and another begins speaking.
[0021] User interruption: This refers to the behavior of a user actively interrupting or speaking during the system's voice output, requiring the system to quickly recognize and process the user's new input.
[0022] AI: Artificial Intelligence.
[0023] This specification provides a speech processing method, and also relates to a speech processing system, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0024] See Figure 1 , Figure 1 A schematic diagram of a scenario for a speech processing method provided according to an embodiment of this specification is shown. Specifically, this voice processing method is implemented using device 102 and voice processing server 104. Device 102 is used to send the user's voice data to voice processing server 104; for example, the voice data is "Wait a minute, why...". In practical applications, if the user can input voice data into device 102 via text, then device 102 will also include corresponding text processing components, such as text parsing, text-to-speech, and speech synthesis modules, to convert the text input by the user into voice data. This specification does not impose any restrictions on this.
[0025] A multimodal large model is trained in the speech processing server 104. When the device 102 is in a speech output state and receives speech data sent by the device 102, the multimodal large model performs data analysis on the speech data to obtain data analysis results. If, based on the data analysis results, it is determined that the speech data meets the first speech processing condition, the state management unit of the speech processing server 104 sends a first state update instruction to the device 102, so that the device 102 updates from the speech output state to the speech input state according to the first state update instruction. The multimodal large model parses the speech data to obtain the corresponding response data (for example, the response data is: "Okay, because..."), and sends the response data to the device 102. The speech processing server 104 sends a second state update instruction to the device 102, and the device 102 updates from the speech input state to the speech output state according to the second state update instruction and outputs the response data.
[0026] Device 102 can include browsers, apps, web applications such as H5 (Hypertext Markup Language 5) applications, lightweight applications (also known as mini-programs), or cloud applications. The device can be developed based on the software development kit (SDK) provided by the server, such as a real-time communication (RTC) SDK. The device can be deployed on electronic devices and depends on the device or certain apps running on the device. Electronic devices can have displays and support information browsing, such as personal mobile terminals like smartphones, tablets, and personal computers. Various other types of applications can also be configured on electronic devices, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platforms.
[0027] The voice processing server 104 can be understood as a server providing various services, including physical servers and cloud servers. For example, it could be a server providing communication services to multiple clients, a server supporting the training of models used on clients, or a server processing data sent by clients. It should be noted that the voice processing server 104 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The voice processing server 104 can also be a server in a distributed system, or a server integrated with blockchain. Furthermore, the voice processing server 104 can be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0028] It is worth noting that the speech processing method provided in the embodiments of this specification can be executed by the speech processing server 104. In other embodiments of this specification, the multimodal large model can be deployed in the device 102, so that the device 102 can also have similar functions to the speech processing server 104, thereby executing the speech processing method provided in the embodiments of this specification. In other embodiments, the speech processing method provided in the embodiments of this specification can also be jointly executed by the device 102 and the speech processing server 104.
[0029] See Figure 2 , Figure 2 A flowchart of a speech processing method provided in one embodiment of this specification is shown, which specifically includes the following steps.
[0030] Step 202: When the device is in voice output mode, and input data is received from the device, perform data analysis on the input data to obtain data analysis results; The voice output state can be understood as the state in which the device is outputting data. For example, when the device is broadcasting voice, the voice output state is the "speaking" state.
[0031] Input data can be understood as the current voice data received by the device, that is, the audio stream received by the device from the user. Of course, input data can also be voice data that has been processed by the device on the user's audio stream, such as noise reduction processing, cropping processing, etc.
[0032] In practical applications, the voice processing server can store historical voice data associated with the current voice data and historical response data corresponding to the historical voice data. Therefore, when the voice processing server performs data analysis on the input data, it can perform more accurate data analysis on the current voice data based on information such as historical voice data, historical response data, and their correlation with the current voice data. The historical voice data associated with the current voice data can be understood as the audio stream of the user in the previous k rounds (k can be set according to the actual situation) before the current dialogue round. The dialogue data of the previous k rounds is determined by the historical voice data and historical response data.
[0033] Specifically, when the device is in voice output mode and receives the current voice data sent by the device, it is necessary to determine whether the user wants to interrupt the content being output. Therefore, the received current voice data is analyzed to obtain the data analysis results, and then the correct processing is made based on the data analysis results.
[0034] In one or more embodiments of this specification, the voice processing server includes a state control unit; the state control unit performs data interruption and data integrity analysis on the current voice data to obtain corresponding data interruption analysis results and data integrity analysis results, and updates the device's state using the different results. Specific implementation methods are as follows: When the device is in voice output mode and input data is received from the device, data analysis is performed on the input data to obtain data analysis results, including: When the state control unit receives input data sent by the device and determines that the device is in voice output state, it performs data interruption and data integrity analysis on the input data to obtain data interruption analysis results and data integrity analysis results.
[0035] The data interruption analysis can be understood as analyzing whether the input data meets the first speech processing condition. The data interruption analysis can be performed by analyzing information such as the correlation with the previous k rounds of dialogue data and the duration of the content. The first speech processing condition is that the data interruption analysis result indicates that the input data is used to interrupt the device output of the device. The data interruption analysis result includes the result that the input data is used to interrupt the device output of the device, or the result that the input data is not used to interrupt the device output of the device. That is, if the data interruption analysis result is that the input data is used to interrupt the device output of the device, it is determined that the input data meets the first speech processing condition. If the data interruption analysis result is that the input data is not used to interrupt the device output of the device, it is determined that the input data does not meet the first speech processing condition.
[0036] Data integrity analysis can be understood as analyzing whether the input data meets the second speech processing condition. Data integrity analysis can be performed using information such as the correlation with the previous k rounds of dialogue data and the duration of pauses. The second input data processing condition is that the data integrity analysis result indicates that the input data meets the preset data integrity. The data integrity analysis result includes either a result indicating that the input data is complete data or a result indicating that the input data is incomplete data. That is, if the data integrity analysis result indicates that the input data is complete data, it is determined that the input data meets the second speech processing condition; if the data integrity analysis result indicates that the input data is incomplete data, it is determined that the input data does not meet the second speech processing condition.
[0037] Specifically, the state control unit may include a large dialogue state control model. The large dialogue state control model is used to perform data interruption analysis on the received input data. Specifically, in the case of data interruption analysis, the semantics, content length (duration), and correlation with historical dialogue data of the current voice data are analyzed to obtain the data interruption analysis results. That is, the data interruption analysis results include the semantics, content length (duration), and correlation with historical dialogue data of the current voice data.
[0038] The large-scale dialogue state control model is used to perform data integrity analysis on the received input data. Specifically, the data integrity analysis is performed on the semantic integrity, pause duration, and correlation with historical dialogue data of the current speech data to obtain the data integrity analysis results. That is, the data integrity analysis results include the semantic integrity, pause duration, and correlation with historical dialogue data of the current speech data.
[0039] The speech processing method provided in the embodiments of this specification performs data interruption and data integrity analysis on speech data through the large dialogue state control model of the state control unit, thereby quickly and accurately obtaining the data interruption analysis results and data integrity analysis results, so as to provide the correct direction for subsequent processing based on the data interruption analysis results and data integrity analysis results.
[0040] Step 204: When the data analysis results indicate that the input data meets the first voice processing condition, a first state update instruction is sent to the device, wherein the first state update instruction is used to update the device from the voice output state to the voice input state.
[0041] In one or more embodiments of this specification, the data analysis result includes a data interruption analysis result. Based on the data interruption analysis result, it is determined whether the input data is input data used to interrupt the device output. If so, the state of the device is changed by sending a first state update command to the device. Specific implementation methods are as follows: When the data analysis results indicate that the input data meets the speech processing conditions, sending a first state update instruction to the device includes: If, based on the data interruption analysis results, it is determined that the input data is used to interrupt the device output, a first state update command is sent to the device.
[0042] Among them, device output can be understood as the response data output by the device in the previous round, in relation to the user's current voice data; voice input state can be understood as a state in which no data is output, but the received voice data is processed, such as the voice input state being "listen".
[0043] Specifically, when the device is in voice output mode and is outputting device data, the user inputs current voice data. The voice processing server performs a comprehensive analysis of the current voice data and historical dialogue data to obtain the data interruption analysis result. It then determines whether the user's current voice data is data used to interrupt the device's output, i.e., whether the user wants to interrupt the device's output. If so, it indicates that the user wants to interrupt the device's output, so it needs to send a first state update command to the device. This allows the device to update from voice output mode to voice input mode according to the first state update command, such as updating the "speak" mode to the "listen" mode. The device then terminates the device's output and receives and forwards the user's current voice data to the voice processing server.
[0044] In practical applications, when the device is in the "speaking" state, it will still listen to the user's audio input in real time. The large model will determine whether the user needs to insert into the conversation based on the dialogue state. If it is determined that the user needs to insert, the device state will be updated to the "listening" state.
[0045] The voice processing method provided in the embodiments of this specification can accurately determine whether the input data is data output by the device used to interrupt the device based on the data interruption analysis results. If so, the state of the device is changed by sending a first state update command to the device in a timely manner.
[0046] In one or more embodiments of this specification, the voice processing server includes a state control unit and a state management unit; the state control unit makes judgments based on data analysis results, and the state management unit interacts with the device to send a first state update command to the device. Specific implementation methods are as follows: When the data analysis results indicate that the voice data meets the first voice processing condition, sending a first state update instruction to the device includes: When the data analysis results indicate that the voice data meets the first voice processing conditions, the state control unit sends a first state update instruction to the state management unit. The status management unit sends the first status update instruction to the device.
[0047] The state management unit is responsible for actually performing the state update operation. It receives the first state update instruction from the state control unit and forwards the first state update instruction to the device, thereby changing the current state of the device.
[0048] Specifically, the state management unit receives a first state update instruction from the state control unit and sends the first state update instruction to the device, instructing the device to update its state according to the first state update instruction. That is, the device updates itself from the current voice output state (such as the "speaking" state) to the voice input state (such as the "listening" state) according to the first state update instruction, so that the device can stop the current output and start receiving and processing the user's new voice input.
[0049] See Figure 3 , Figure 3 This diagram illustrates the processing of a large-scale dialogue state control model in a speech processing method provided in one embodiment of this specification.
[0050] in, Figure 3The multimodal large model includes a dialogue state control large model. The model input is historical dialogue data from the last K rounds (including historical voice data and historical response data). The current voice data and historical voice data are encoded to obtain voice features. Through feature alignment, the dimension of the voice features is transformed to align the voice features with the text feature space, thus obtaining aligned voice features. The historical response data can be voice modal response data or text modal response data. In the embodiments of this specification, the historical response data is taken as text modal response data as an example. The historical response data is encoded to obtain response text features. The current state of the device (such as the voice output state) is input into the dialogue state control large model. By analyzing the aligned voice features and response text features, if it is determined that the current voice data is used to interrupt the output of the device, the state command for the next state is output, that is, the first state command is output.
[0051] The voice processing method provided in this specification includes a voice processing server comprising a state control unit and a state management unit. The state control unit is used to determine whether the state of the device needs to be updated based on the data analysis results of the input data. The state management unit is responsible for actually performing the state update operation. Through this design, the voice processing server can flexibly adjust the state of the device according to the user's voice input, achieving a more natural and smooth interactive experience.
[0052] Step 206: Parse the input data to obtain the response data corresponding to the input data, and send the response data to the device.
[0053] The response data can be in the form of speech, text, a combination of text and speech, or even video; there are no restrictions on this.
[0054] In one or more embodiments of this specification, input data is analyzed, and based on the analysis results, it is determined whether the input data meets the second speech processing condition. If so, the input data is parsed to obtain the corresponding response data. Specific implementation methods are described below: The step of parsing the voice data to obtain the corresponding response data includes: If the data analysis results indicate that the voice data meets the second voice processing conditions, the voice data is parsed to obtain the response data corresponding to the voice data.
[0055] In practical applications, the data analysis results include data integrity analysis results, and the second voice data processing condition is that the data integrity analysis results indicate that the input data meets the preset data integrity. When the data analysis results indicate that the speech data meets the second speech processing condition, parsing the input data to obtain the response data corresponding to the input data includes: If, based on the data integrity analysis results, it is determined that the voice data meets the preset data integrity requirements, the input data is parsed to obtain the response data corresponding to the input data.
[0056] Among them, preset data integrity can be understood as preset conditions that determine that the voice data is complete voice data, such as semantic integrity, pause duration, etc.
[0057] Specifically, based on the data integrity analysis results, if the voice data meets the preset data integrity, it is considered that the user's current input voice data is coherent and complete. That is, when the user finishes speaking and waits for the system to respond, the turn order needs to be switched, the user ends speaking, and the device needs to output response data.
[0058] In practical applications, when the device is in "listening" mode, the user's voice input is sent to the dialogue state control model in real time in segments until the model determines that the user has finished speaking. At this point, a text response is generated based on the user's voice input data. If the response data is in text modality, the generated text response can be output by the device. If the response data is in speech modality, the generated text response is sent to the streaming speech synthesis module for audio synthesis to obtain speech modality response data.
[0059] The voice processing method provided in the embodiments of this specification accurately determines whether the current voice data meets the preset data integrity by analyzing the data integrity results, and processes it correctly based on the judgment results. Only when the voice data meets the preset data integrity can the voice data be parsed to obtain accurate response data, thereby improving the user experience.
[0060] In one or more embodiments of this specification, the voice processing server includes a status control unit and a voice response unit; the status control unit is used to determine whether the input data meets the second voice processing condition, and the voice response unit is used to parse and process the input data to obtain response data. Specific implementation methods are as follows: The step of parsing the voice data to obtain the corresponding response data includes: The status control unit, upon determining, based on the data analysis results, that the voice data meets the second voice processing conditions, sends the voice data to the voice response unit. The voice response unit parses the voice data to obtain the response data corresponding to the voice data.
[0061] The voice response unit can be understood as the unit in the system responsible for generating and sending response information. The voice response unit receives voice data from the status control unit that meets the second voice processing conditions, and parses and generates a response based on the voice data.
[0062] Specifically, the voice response unit includes a large-scale voice semantic model. This model receives voice data that meets the second voice processing condition, parses the received voice data (including but not limited to speech recognition and further semantic analysis) to accurately understand the user's intent and needs, and generates corresponding response data based on the parsing results. This response data can be in text format for subsequent text responses, or it can be directly converted into voice format for voice responses. In practical applications, text responses and voice responses can be output together as response data.
[0063] The voice processing method provided in the embodiments of this specification, through close cooperation between the state control unit and the voice response unit, realizes the processing of voice data that meets the second voice processing conditions, obtains the corresponding response data, and ensures the accuracy of the response data when the voice data is complete.
[0064] In one or more embodiments of this specification, when the response data is in speech modality, the response text of the input data is first obtained, and then speech synthesis is performed on the response text to obtain the speech modality response data. Specific implementation methods are as follows: The step of parsing the input data to obtain the corresponding response data includes: Parse the input data to obtain the response text of the input data; The reply text is processed by speech synthesis to obtain the reply data corresponding to the input data.
[0065] Specifically, the voice response unit includes a text response module and a speech synthesis module. The text response module parses the voice data to generate corresponding response text, while the speech synthesis module is responsible for converting the response text generated by the text response module into a natural and fluent voice response. The speech synthesis module uses text-to-speech technology to convert text information into audible voice signals.
[0066] Specifically, the text response module includes a large speech semantic model. The large speech semantic model receives speech data that meets the second speech processing conditions from the state control unit, generates a text response in a streaming manner, and sends the generated response text to the speech synthesis module for subsequent speech synthesis processing.
[0067] The speech synthesis module uses text-to-speech technology to convert text into speech. This process includes, but is not limited to, multiple steps such as text analysis, speech encoding, and sound synthesis. The goal is to generate a natural and fluent speech signal that matches the original text content and output the synthesized speech response to the device. This can be achieved by playing it through a speaker, saving it to a file, or transmitting it over a network.
[0068] In one or more embodiments of this specification, the speech synthesis module uses a speech synthesis model to obtain the text features of the response text, uses a vocoder to generate a speech waveform, and synthesizes response data in a speech modality. Specific implementation methods are as follows: The step of performing speech synthesis on the reply text to obtain reply data corresponding to the input data includes: Using a speech synthesis model, the text features of the text response are determined, and the text features are input into a vocoder; Using the vocoder, speech waveforms are generated from the text features to obtain response data corresponding to the input data.
[0069] Specifically, the speech synthesis process employs a progressive synthesis strategy for streaming text-to-speech (StreamingTTS), meaning speech synthesis begins simultaneously with text input, utilizing prediction and parallel processing techniques to reduce latency in speech generation. Furthermore, with the aforementioned large-scale speech-semantic model generating text responses in a streaming manner, the large-scale model outputs predicted text tokens (the basic unit of model processing) one by one. Each output predicted text token is sent to the speech synthesis model (e.g., the NeuralTTS model). The NeuralTTS model uses deep learning language models (e.g., GPT, T5) to convert text into intermediate speech representations for use as input to the streaming vocoder. The streaming vocoder employs a streaming neural vocoder (e.g., WaveRNN, ParallelWaveGAN) to provide efficient real-time speech waveform generation, synthesizing response data corresponding to the input data.
[0070] The speech processing method provided in the embodiments of this specification generates text responses in a streaming manner through a large speech semantic model and performs streaming speech synthesis using a progressive synthesis strategy. This greatly reduces the latency of speech synthesis, achieves a fast response to user input speech, conforms to the characteristics of real interaction, and can respond with extremely low latency (theoretically, the lowest response time can be within 100 milliseconds, combined with the mainstream streaming TTS response time of within 500 milliseconds).
[0071] In one or more embodiments of this specification, upon receiving response data, a second state update command is sent to the device to update the device from a voice input state to a voice output state, so that the response data can be output in the voice output state. Specific implementation methods are described below: After parsing the voice data to obtain the corresponding response data, the process further includes: Send a second status update instruction and the response data to the device, wherein the second status update instruction is used to update the device from the voice input state to the voice output state and output the response data.
[0072] Specifically, after the voice processing server obtains the response data corresponding to the voice data, it needs to output the response data to the user through the device. Therefore, at this time, the state management unit needs to send a second state update instruction to the device. Upon receiving the second state update instruction, the device updates from the voice input state ("listen" state) to the voice output state ("speak" state) and outputs the response data.
[0073] In one or more embodiments of this specification, the voice processing server can receive feedback information from the device, such as a device status update command or a data output completion command sent by the device, so as to issue a corresponding status update command to the device according to the device's command and update the current status of the device. Specific implementation methods are described below: After sending the response data to the device, the process further includes: Upon receiving a device status update command from the device and determining that the device is in the voice output state, a first status update command is sent to the device, wherein the first status update command is used to update the device from the voice output state to the voice input state; or Upon receiving a response data output completion instruction from the device and determining that the device is in the voice output state, a first state update instruction is sent to the device.
[0074] Specifically, user interactions on the device can trigger the device to send a device status update command to the voice processing server. For example, if there are interrupt or exit controls on the user interface of the device, clicking these controls will trigger the device to send a device status update command to the voice processing server. When the voice processing server receives the device status update command from the device and determines that the device is in voice output mode, it will send a first status update command to the device through the status management unit.
[0075] Alternatively, if the user does not interact with the device and there is no interruption, the device can send a reply data output completion instruction to the voice processing server after completing the output of the reply data. This indicates that the device has finished outputting and needs to switch from the "speaking" state to the "listening" state. At this time, the voice processing server, after receiving the reply data output completion instruction from the device and confirming that the device is in the voice output state, sends a first state update instruction to the device.
[0076] The voice processing method provided in the embodiments of this specification allows the voice processing server to receive feedback information from the device, such as device status update instructions corresponding to "user clicks to interrupt" and "user exits", as well as reply data output completion instructions for "playback completed", and to accurately adjust the global state using these instructions.
[0077] In one or more embodiments of this specification, the voice processing method further includes: If it is determined that the device is in the voice input state, the initial voice data sent by the device is received; Perform data integrity analysis on the initial voice data to obtain the data integrity analysis results; When the data integrity analysis result indicates that the initial voice data meets the second voice processing condition, a second state update instruction is sent to the device, wherein the second state update instruction is used to update the device from the voice input state to the voice output state; The initial voice data is parsed to obtain the initial response data corresponding to the initial voice data, and the initial response data is sent to the device.
[0078] Specifically, when the device is in "listening" mode, it receives the user's initial voice data and performs data integrity analysis on the initial voice data. Based on the data integrity analysis results, the initial voice data is then processed. For details, please refer to the above embodiments, which will not be repeated here.
[0079] The voice processing method provided in the embodiments of this specification is based on a state control unit. It avoids the need to explicitly define a long silence time for judgment in existing methods when detecting "round switching". This naturally results in a fixed waiting time and an unnatural effect in human-computer interaction. The state control unit judges "round switching" and "user interruption" in different states. The judgment is made by combining the current voice data and historical dialogue data, which conforms to the characteristics of real interaction and can respond with extremely low waiting time based on the voice response unit.
[0080] See Figure 4 , Figure 4A flowchart illustrating the processing procedure of a speech processing method provided in one embodiment of this specification is shown.
[0081] This specification describes the voice processing method in detail using the example of the device being in voice input mode.
[0082] When the device is in "listening" state, the large dialogue state control model (for example, this model is modified and optimized based on the Audio-text Large Language Models audio response speech large model; note that this model can share the model structure and parameters with the subsequent speech semantic large model) receives the audio stream. By comprehensively analyzing the audio stream and dialogue history, and judging based on multiple factors such as semantic completeness, pause duration, and contextual relevance, it decides whether the AI side (i.e., the speech processing server in the above embodiment) should start responding to the user. 0 indicates that the user is in a speaking state or the user has not started speaking at all, in which case the AI side does not need to respond; 1 indicates that the user has completed a new round of input, in which case the AI side needs to respond accordingly.
[0083] Specifically, the user's audio stream is fed in real-time into the dialogue state control model of the multimodal big data model. The user's historical dialogue data (including the user's historical audio stream and the multimodal big data model's text responses), the current audio stream, and the device's current state are input into the dialogue state control model. Based on the device's "listening" state, the dialogue state control model identifies whether the semantics of the received current audio stream are complete. It makes judgments based on multiple factors such as semantic completeness, pause duration, and contextual relevance. If it determines that the user has finished speaking, it outputs "1". In practical applications, the output... <eos>, indicating that the user has finished speaking, and starting to report into "speak" state (here, an over-parameter can be set to dynamically control the tolerance level); otherwise, output "0", in practical applications, output <listen>, indicating that the user continues to speak.
[0084] In the case that the dialogue state control large model outputs "1", that is, the user has finished speaking, the historical dialogue data of the user and the current audio stream are input into the voice semantic large model, and the voice semantic large model can generate a more accurate text reply for the current audio stream according to the context of the historical dialogue data; the text reply output by the voice semantic large model is input into the stream speech synthesis module. Specifically, the stream speech synthesis module converts the text reply into speech features by using a deep learning language model, and inputs the speech features into a vocoder to generate a broadcast voice. In the process of generating the broadcast voice by the stream speech synthesis module, a progressive synthesis strategy is adopted, that is, instead of inputting the complete text reply generated by the voice reply module into the stream speech synthesis module, the stream speech synthesis module generates speech every time a text Token is predicted by the voice reply module, so as to reduce the delay of generating the voice.
[0085] In the case that the dialogue state control large model outputs "1", that is, the user has finished speaking, the output result of the dialogue state control large model is sent to the cloud dialogue state management module (that is, the state management unit in the above embodiment) to make the cloud dialogue state management module send a state switching instruction (that is, the second state update instruction in the above embodiment) to the device end, and the state switching instruction of the cloud dialogue state management module and the broadcast voice generated by the stream speech synthesis module are sent to the device end to make the device end switch to the "speak" state and start broadcasting the generated broadcast voice.
[0086] Referring to Figure 5 , Figure 5 A processing process flowchart of another voice processing method provided by an embodiment of the present specification is shown.
[0087] The embodiments of the present specification take the device end in a voice output state as an example to explain the voice processing method in detail.
[0088] When the device end state is the "speak" state, the dialogue state control large model receives the audio stream, analyzes the audio stream and the dialogue history, and comprehensively judges whether the user's speech at this time constitutes an interruption to the AI end according to the current input audio semantics, content length and other factors; 0 indicates that the user's input audio stream has a relatively complete semantic to try to insert into the dialogue, and the AI end needs to stop speaking at this time; 1 indicates that the user's input audio stream does not have complete semantics or interruption ability, and the AI end continues to broadcast.
[0089] When the device is in "speaking" mode, it also listens for the user's audio input. When the device detects the user's audio stream, it inputs the user's audio stream, historical dialogue data, and the device's current state into the dialogue state control model. The dialogue state control model comprehensively judges whether the user's speech constitutes an interruption to the AI based on factors such as the semantics and content length of the current input audio. If the judgment result is yes, that is, it is determined that the user wants to interrupt, it outputs "0". In practical applications, the output is... <stop>, indicating that the user is ready to intervene, interrupt the broadcast into the "listen" state (here allows to set a hyperparameter, for dynamic control tolerance); and in the case of determining that the user has no to break, output "1", the actual application, output <speak>, which means to continue the broadcast.
[0090] In the case that the dialogue state control large model outputs "0", that is, the user is ready to intervene, the cloud dialogue state management module sends a state switching instruction (which is the first state update instruction in the above embodiment) to the device end according to the output of the dialogue state control large model, so that the device end switches the state of the device end to the "listen" state according to the state switching instruction.
[0091] The user can click on the device end user interaction interface to interrupt, and the device end can send a user interruption instruction to the cloud dialogue state management module according to the user's click interruption action, so that the cloud dialogue state management module returns a state switching response based on the user interruption instruction, and switches the state of the device end to the "listen" state.
[0092] In actual application, in the case that the device end completes the broadcast of the broadcast voice, the device end can send a broadcast completion instruction to the cloud dialogue state management module, so that the cloud dialogue state management module returns a state switching response based on the broadcast completion instruction, thereby switching the state of the device end to the "listen" state.
[0093] The voice processing method provided by the embodiments of the present specification can, in the case that the device end is in a voice output state and the input data sent by the device end is received, determine whether the input data meets the first voice processing condition by analyzing the input data, and in the case of meeting, change the voice output state of the device end to a voice input state by sending a first state update instruction to the device end, analyze the voice data to obtain reply data, and send the reply data to the device end. That is, the device end can still receive voice data and forward it to the voice processing service end in the voice output state, and the voice processing service end can update the state of the device end in a timely manner in the case that the voice data meets the first voice processing condition through analysis of the voice data. By receiving voice data in a timely manner and updating the state of the device end, the machine can speak and receive feedback at the same time, and obtain reply data quickly to give a reaction under very low latency, so that the human-computer interaction can achieve a very natural effect.
[0094] Corresponding to the above method embodiments, the present specification also provides a voice processing system embodiment, Figure 6 The structure schematic diagram of a voice processing system provided by an embodiment of the present specification is shown. As shown in the figure, Figure 6 The system includes a voice processing service end 104 and a device end 102, wherein, The device end 102 is configured to send the received input data to the voice processing service end 104. The voice processing server 104 is configured to, in a case where the device end 102 is in a voice output state and the input data sent by the device end 102 is received, perform data analysis on the input data to obtain a data analysis result, and in a case where the data analysis result indicates that the voice data meets a voice processing condition, send a first state update instruction to the device end 102. The device end 102 is further configured to update the voice output state to a voice input state according to the first state update instruction. The voice processing server 104 is further configured to parse the input data to obtain reply data corresponding to the input data, and send the reply data to the device end 102.
[0095] Optionally, the voice processing server 104 is further configured to send a second state update instruction to the device end 102. Optionally, the device end 102 is further configured to update the voice input state to the voice output state according to the second state update instruction, and output the reply data.
[0096] Optionally, the voice processing server 104 is further configured to: In a case where the data analysis result indicates that the input data meets a second voice processing condition, parse the input data to obtain reply data corresponding to the input data.
[0097] Optionally, the voice processing server 104 is further configured to: In a case where it is determined according to the data integrity analysis result that the input data meets a preset data integrity, parse the input data to obtain reply data corresponding to the input data.
[0098] Optionally, the voice processing server 104 is further configured to: Parse the input data to obtain reply text of the input data. Perform voice synthesis on the reply text to obtain the reply data corresponding to the input data.
[0099] Optionally, the voice processing server 104 is further configured to: Determine a text feature of the text reply by using the voice synthesis model, and input the text feature into a vocoder. Perform voice waveform generation on the text feature by using the vocoder to obtain the reply data corresponding to the input data.
[0100] Optionally, the voice processing server 104 is further configured to send a first state update instruction to the device end 102 in a case that the voice processing server 104 receives a device state update instruction sent by the device end 102 and determines that the device end 102 is in the voice output state. Optionally, the voice processing server 104 is further configured to send a first state update instruction to the device end 102 in a case that the voice processing server 104 receives a reply data output completion instruction sent by the device end 102 and determines that the device end 102 is in the voice output state. Optionally, the voice processing server 104 is further configured to receive initial voice data sent by the device end 102 in a case that the voice processing server 104 determines that the device end 102 is in the voice input state, perform data integrity analysis on the initial voice data, and obtain a data integrity analysis result; in a case that the data integrity analysis result indicates that the initial voice data meets a second voice processing condition, send a second state update instruction to the device end 102, analyze the initial voice data, obtain initial reply data corresponding to the initial voice data, and send the initial reply data to the device end 102.
[0101] The above is a schematic scheme of the voice processing system of the embodiment. It should be noted that the technical scheme of the voice processing system and the technical scheme of the voice processing method described above belong to the same concept, and the details of the technical scheme of the voice processing system that are not described in detail can be referred to the description of the technical scheme of the voice processing method.
[0102] Figure 7 A structural block diagram of a computing device 700 is shown according to an embodiment of the present specification. The components of the computing device 700 include but are not limited to a memory 710 and a processor 720. The processor 720 is connected with the memory 710 through a bus 730, and a database 750 is used to save data.
[0103] The computing device 700 also includes an access device 740 that enables the computing device 700 to communicate via one or more networks 760. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or combinations of such networks, such as the Internet. The access device 740 can include one or more of any type of network interface (for example, a network interface card (NIC)) such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC).
[0104] In one embodiment of the present specification, the above-mentioned components of the computing device 700 and other components not shown in the Figure 7 may be connected to each other, such as through a bus. It should be understood that Figure 7 The computing device structure diagram shown is only for the purpose of example, and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.
[0105] The computing device 700 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (for example, a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, and the like), a mobile phone (for example, a smartphone), a wearable computing device (for example, a smart watch, smart glasses, and the like), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 700 can also be a mobile or stationary server.
[0106] The processor 720 is configured to execute computer program / instructions that implement the steps of the voice processing method described above when the computer program / instructions are executed by the processor.
[0107] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts among the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, the computing device embodiment is described simply because it is basically similar to the speech processing method embodiment, and the relevant part can be referred to the description of the speech processing method embodiment.
[0108] An embodiment of the specification further provides a computer readable storage medium storing computer programs / instructions, which are executed by a processor to implement the steps of the speech processing method.
[0109] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts among the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, the computer readable storage medium embodiment is described simply because it is basically similar to the speech processing method embodiment, and the relevant part can be referred to the description of the speech processing method embodiment.
[0110] An embodiment of the specification further provides a computer program product including computer programs / instructions, which are executed by a processor to implement the steps of the speech processing method.
[0111] The above is a schematic scheme of the computer program product of the embodiment. It should be noted that the technical scheme of the computer program product and the technical scheme of the speech processing method described above belong to the same concept, and the details of the technical scheme of the computer program product which are not described in detail can be referred to the description of the technical scheme of the speech processing method.
[0112] The above describes specific embodiments of the specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than the order in which they are recited and still accomplish desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.
[0113] The computer readable medium can include any entity or apparatus capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, software distribution medium, etc. It should be noted that the computer readable medium can include appropriate additions or subtractions according to the requirements of patent practice. For example, according to the patent practice in some regions, the computer readable medium does not include electrical carrier signals and telecommunication signals.
[0114] It should be noted that, for the foregoing method embodiments, in order to facilitate description, they are all described as a combination of a series of actions, but those skilled in the art should know that the embodiments of the present specification are not limited by the order of the described actions, because according to the embodiments of the present specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of the present specification.
[0115] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0116] The preferred embodiments of the present specification disclosed above are only used to help explain the present specification. The alternative embodiments do not describe all the details and limit the invention to the specific embodiments described. Obviously, according to the content of the embodiments of the present specification, many modifications and changes can be made. The present specification selects and describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present specification, so that those skilled in the art can well understand and use the present specification. The present specification is limited only by the claims and their entire scope and equivalents.< / speak> < / stop> < / listen> < / eos>
Claims
1. A speech processing method, applied to a speech processing server, comprising: When the device is in voice output mode, and input data is received from the device, the input data is analyzed to obtain the data analysis results. When the data analysis results indicate that the input data meets the first voice processing condition, a first state update instruction is sent to the device, wherein the first state update instruction is used to update the device from the voice output state to the voice input state. The input data is parsed to obtain the corresponding response data, and the response data is sent to the device.
2. The voice processing method according to claim 1, wherein sending the response data to the device comprises: Send a second status update instruction and the response data to the device, wherein the second status update instruction is used to update the device from the voice input state to the voice output state and output the response data.
3. The voice processing method according to claim 1, wherein the data analysis result includes a data interruption analysis result, and the first voice processing condition is that the data interruption analysis result indicates that the input data is used to interrupt the device output of the device.
4. The speech processing method according to claim 1, wherein parsing the input data to obtain the response data corresponding to the input data includes: If the data analysis results indicate that the input data meets the second speech processing condition, the input data is parsed to obtain the response data corresponding to the input data.
5. The speech processing method according to claim 4, wherein the data analysis result includes a data integrity analysis result, and the second speech processing condition is that the data integrity analysis result indicates that the input data meets a preset data integrity requirement; When the data analysis results indicate that the input data meets the second speech processing condition, the step of parsing the input data to obtain the response data corresponding to the input data includes: If, based on the data integrity analysis results, it is determined that the input data meets the preset data integrity requirements, the input data is parsed to obtain the response data corresponding to the input data.
6. The speech processing method according to claim 1, wherein parsing the input data to obtain the response data corresponding to the input data includes: Parse the input data to obtain the response text of the input data; The reply text is processed by speech synthesis to obtain the reply data corresponding to the input data.
7. The speech processing method according to claim 6, wherein the step of performing speech synthesis on the reply text to obtain reply data corresponding to the input data includes: Using a speech synthesis model, the text features of the text response are determined, and the text features are input into a vocoder; Using the vocoder, speech waveforms are generated from the text features to obtain response data corresponding to the input data.
8. The voice processing method according to claim 1, further comprising, after sending the response data to the device: Upon receiving a device status update command from the device and determining that the device is in the voice output state, a first status update command is sent to the device, wherein the first status update command is used to update the device from the voice output state to the voice input state; or Upon receiving a response data output completion instruction from the device and determining that the device is in the voice output state, a first state update instruction is sent to the device.
9. The speech processing method according to any one of claims 1-8, wherein the input data includes the current speech data received by the device.
10. The speech processing method according to claim 1, further comprising: If it is determined that the device is in the voice input state, the initial voice data sent by the device is received; Perform data integrity analysis on the initial voice data to obtain the data integrity analysis results; When the data integrity analysis result indicates that the initial voice data meets the second voice processing condition, a second state update instruction is sent to the device, wherein the second state update instruction is used to update the device from the voice input state to the voice output state; The initial voice data is parsed to obtain the initial response data corresponding to the initial voice data, and the initial response data is sent to the device.
11. A voice processing system, comprising a voice processing server and a device, wherein, The device is used to send the received input data to the voice processing server. The voice processing server is configured to perform data analysis on the voice data when the device is in voice output state and receives input data sent by the device, obtain data analysis results, and send a first state update command to the device when the data analysis results indicate that the input data meets the first voice processing conditions. The device is also configured to update from the voice output state to the voice input state according to the first state update instruction; The voice processing server is also used to parse the input data, obtain the response data corresponding to the input data, and send the response data to the device.
12. The speech processing system according to claim 11, wherein, The voice processing server is also used to send a second status update instruction to the device. The device is further configured to update from the voice input state to the voice output state according to the second state update instruction, and output the response data.
13. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the speech processing method according to any one of claims 1 to 10.
14. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the speech processing method according to any one of claims 1 to 10.
15. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the speech processing method according to any one of claims 1 to 10.