Voice processing method and system

By analyzing user input data and sending status update commands on the voice processing server to change the device's state, the problem of users being unable to interrupt machine output is solved, achieving natural human-computer interaction and rapid response.

WO2026066814A1PCT designated stage Publication Date: 2026-04-02ALIBABA (CHINA) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

In existing technologies, users cannot interrupt the machine's output responses during human-computer interaction, resulting in unnatural interactions and the inability to achieve full-duplex interaction.

Method used

By analyzing the input data on the voice processing server to determine whether the voice processing conditions are met, a status update command is sent to change the voice output state of the device to the input state, and the input data is parsed to obtain the response data, so that the device can receive and process user input in the voice output state.

Benefits of technology

It enables machines to speak and receive feedback simultaneously with extremely low latency, achieving a natural human-computer interaction effect. Users can interrupt the machine's output at any time and receive a timely response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025115501_02042026_PF_FP_ABST
    Figure CN2025115501_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a voice processing method and system. The method is applied to a voice processing server, and comprises: when a device is in a voice output state and input data sent by the device is received, performing data analysis on the input data to obtain a data analysis result; when the data analysis result indicates that voice data meets a first voice processing condition, sending a first state update instruction to the device, wherein the first state update instruction is used to update the device from the voice output state to a voice input state; and analyzing the voice data to obtain response data corresponding to the voice data, and sending the response data to the device. By promptly receiving the voice data and updating the state of the device, the machine can speak and receive feedback simultaneously and obtain the response data with extremely low latency to react quickly, thereby enabling human–machine interaction to achieve a very natural effect.
Need to check novelty before this filing date? Find Prior Art

Description

Speech processing method and system

[0001] The present disclosure claims priority to Chinese Patent Application No. 202411361141.6, filed on September 26, 2024, with the Chinese Patent Office, entitled "Speech processing method and system", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] Embodiments of the present specification relate to the technical field of speech processing, in particular to a speech processing method and system. BACKGROUND

[0003] At present, large models have advantages in text-related understanding and generation tasks and are widely used. However, when the large model is applied, the user and the machine (the machine uses the large model to generate and output a reply to the user's input question) usually have a one-question-one-answer conversation, and the delay of the large model reply is longer.

[0004] With the improvement of user demand, users expect natural and smooth conversation with machines, and machines can speak and receive feedback at the same time and give quick responses, like human-to-human communication. However, recent research on how to realize full-duplex interaction of multi-modal large models is less explored, and the existing multi-modal speech interaction model has the problem that users cannot interrupt AI (Artificial Intelligence), which leads to the fact that when humans and machines have full-duplex interaction, users cannot interrupt the machine's output reply, and natural human-machine interaction cannot be achieved. SUMMARY

[0005] Therefore, embodiments of the present specification provide a speech processing method. One or more embodiments of the present specification also relate to a speech processing system, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects in the prior art that when humans and machines have full-duplex interaction, users cannot interrupt the machine's output reply, and natural human-machine interaction cannot be achieved.

[0006] According to a first aspect of embodiments of the present specification, a speech processing method is provided, applied to a speech processing server, comprising:

[0007] In the case that the device end is in a speech output state and the input data sent by the device end is received, the input data is analyzed to obtain a data analysis result;

[0008] In a case where the data analysis result indicates that the voice data satisfies a first voice processing condition, a first state update instruction is sent to the device end, where the first state update instruction is used to update the device end from the voice output state to a voice input state.

[0009] The voice data is parsed to obtain reply data corresponding to the voice data, and the reply data is sent to the device end.

[0010] According to a second aspect of the embodiments of the present specification, a voice processing method system is provided, comprising a voice processing server and a device end, wherein,

[0011] The device end is configured to send received input data to the voice processing server.

[0012] The voice processing server is configured to, in a case where the device end is in a voice output state and the input data sent by the device end is received, perform data analysis on the voice data to obtain a data analysis result, and in a case where the data analysis result indicates that the input data satisfies a first voice processing condition, send a first state update instruction to the device end.

[0013] The device end is further configured to update from the voice output state to a voice input state according to the first state update instruction.

[0014] The voice processing server is further configured to parse the input data to obtain reply data corresponding to the input data, and send the reply data to the device end.

[0015] According to a third aspect of the embodiments of the present specification, a computing device is provided, comprising:

[0016] a memory and a processor;

[0017] The memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, which realize the steps of the voice processing method when executed by the processor.

[0018] According to a fourth aspect of the embodiments of the present specification, a computer readable storage medium is provided, which stores computer programs / instructions, which realize the steps of the voice processing method when executed by the processor.

[0019] According to a fifth aspect of the embodiments of the present specification, a computer program product is provided, comprising computer programs / instructions, which realize the steps of the voice processing method when executed by the processor.

[0020] The voice processing method provided by one embodiment of the present specification can change the voice output state of the device end to a voice input state when the input data satisfies the first voice processing condition, and can obtain the reply data by analyzing the input data, and send the reply data to the device end. That is, the device end can still receive the input data and forward it to the voice processing server when it is in the voice output state. The voice processing server can update the state of the device end in time when the input data satisfies the first voice processing condition by analyzing the input data. By receiving the voice data in time and updating the state of the device end, the machine can speak and receive feedback at the same time, and obtain the reply data quickly to give a reaction under very low latency, so that the human-computer interaction can achieve a very natural effect. BRIEF DESCRIPTION OF DRAWINGS

[0021] FIG. 1 is a scene schematic diagram of a voice processing method provided by one embodiment of the present specification;

[0022] FIG. 2 is a flowchart of a voice processing method provided by one embodiment of the present specification;

[0023] FIG. 3 is a processing schematic diagram of a dialogue state control large model in a voice processing method provided by one embodiment of the present specification;

[0024] FIG. 4 is a flowchart of a processing process of a voice processing method provided by one embodiment of the present specification;

[0025] FIG. 5 is a flowchart of a processing process of another voice processing method provided by one embodiment of the present specification;

[0026] FIG. 6 is a structural schematic diagram of a voice processing system provided by one embodiment of the present specification;

[0027] FIG. 7 is a structural block diagram of a computing device provided by one embodiment of the present specification. DETAILED DESCRIPTION

[0028] In the following description, many specific details are set forth in order to provide a thorough understanding of the present specification. However, the present specification can be practiced in many different ways beyond the specific embodiments described herein, and it is understood that persons having ordinary skill in the art can make similar modifications to the present specification without departing from the spirit of the present specification, so the present specification is not limited to the specific implementations disclosed below.

[0029] The terminology used in this disclosure, one or more embodiments of the present specification, is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present specification. As used in this disclosure and the appended claims herein, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "and / or," as used in this disclosure, refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0030] It will be understood that, although the terms first, second, etc. can be employed in this disclosure, one or more embodiments of the present specification, to describe various information, these information should not be limited to these terms. These terms are only used to differentiate one piece of information from another piece of information of the same type. For example, a first can also be referred to as a second, and similarly, a second can also be referred to as a first, without departing from the scope of one or more embodiments of the present specification. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining".

[0031] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present specification are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0032] First, the nomenclature involved in one or more embodiments of the present specification is explained.

[0033] Full-Duplex: Full-Duplex refers to a mode in which both parties in a communication system can send and receive data at the same time, similar to a telephone conversation in which both parties can talk and hear each other at the same time.

[0034] Multimodal Large Model: Multimodal Large Model is a large machine learning model that can process and correlate multiple modal (such as text, speech, image) information, commonly used to improve the understanding and interaction capabilities of machines.

[0035] Speech Encoding: Speech Encoding is the process of converting speech signals into a digital format that can be processed by computers, usually involving sampling, quantization and compression; in the present specification, it is used to convert audio into feature input for multimodal large models.

[0036] Turn-taking: Turn-Taking is the process of alternating speech in a conversation, which is crucial to effectively identify when one participant ends a speech and another participant starts a speech.

[0037] User barge in: User Barge In refers to the behavior of a user actively interrupting or barging in during the system voice output process, requiring the system to quickly identify and process the new input of the user.

[0038] AI: Artificial Intelligence, Artificial Intelligence.

[0039] In the present specification, a voice processing method is provided, and the present specification also relates to a voice processing system, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.

[0040] Referring to FIG. 1, FIG. 1 shows a scene schematic diagram of a voice processing method according to an embodiment of the present specification.

[0041] Specifically, the voice processing method is implemented by using a device end 102 and a voice processing server 104, the device end 102 is configured to send voice data of a user to the voice processing server 104; for example, the voice data is “wait a minute, why…”; in actual application, in the case that the user can input the voice data in the device end 102 in the form of text, if the text form is adopted, the device end 102 will also include a corresponding text processing part, such as a text analysis module, a text-to-speech module, a speech synthesis module, etc., which are configured to convert the text input by the user in the form of text into voice data, and the present specification does not limit this.

[0042] In the voice processing server 104, a multi-modal large model is trained, in the case that the device end 102 is in a voice output state and receives the voice data sent by the device end 102, the multi-modal large model performs data analysis on the voice data to obtain a data analysis result; in the case that it is determined according to the data analysis result that the voice data satisfies a first voice processing condition, a first state update instruction is sent to the device end 102 by a state management unit of the voice processing server 104, so that the device end 102 updates from the voice output state to a voice input state according to the first state update instruction; the multi-modal large model analyzes the voice data to obtain reply data corresponding to the voice data (for example, the reply data is “good, because…”), and sends the reply data to the device end 102; the voice processing server 104 sends a second state update instruction to the device end 102, and the device end 102 updates from the voice input state to the voice output state according to the second state update instruction, and outputs the reply data.

[0043] The device side 102 can include a browser, an APP (Application), or a web application such as an H5 (Hyper Text Markup Language 5) application, or a light application (also known as a small program, a lightweight application program), or a cloud application, and the like. The device side can be developed based on a software development kit (SDK) of a corresponding service provided by the server side, such as an RTC (Real Time Communication) SDK, and the like. The device side can be deployed in an electronic device, and needs to rely on the device or some APP in the device to run, and the like. The electronic device can have a display screen and support information browsing, and the like, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, and the like. Various other types of applications can also be configured in the electronic device, for example, human-computer dialogue type applications, model training type applications, text processing type applications, web browser applications, shopping type applications, search type applications, instant communication tools, mailbox clients, social platform software, and the like.

[0044] The voice processing server 104 can be understood as a server that provides various services, including a physical server, a cloud server, for example, a server that provides communication services for multiple clients, for example, a server for background training that supports models used on the client, for example, a server that processes data sent by the client, and the like. It should be noted that the voice processing server 104 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The voice processing server 104 can also be a server of a distributed system, or a server combined with a blockchain. The voice processing server 104 can also be a cloud server of cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDNs), and big data and artificial intelligence platforms, and the like. basic cloud computing services, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0045] It should be noted that the voice processing method provided in the embodiments of the present specification can be executed by the voice processing server 104. In other embodiments of the present specification, a multi-modal large model can be deployed in the device side 102, so that the device side 102 can also have similar functions as the voice processing server 104, thereby executing the voice processing method provided in the embodiments of the present specification. In other embodiments, the voice processing method provided in the embodiments of the present specification can also be executed by the device side 102 and the voice processing server 104 together.

[0046] Referring to FIG. 2, FIG. 2 shows a flowchart of a voice processing method according to an embodiment of the present specification, which specifically includes the following steps.

[0047] In step 202, in the case that the device is in a voice output state and receives input data sent by the device, data analysis is performed on the input data to obtain a data analysis result.

[0048] The voice output state can be understood as a state in which the device is outputting data, for example, in the case that the device is playing a voice, the voice output state is a "speaking" state.

[0049] The input data can be understood as the current voice data received by the device, that is, the audio stream of the user received by the device. Of course, the input data can also be voice data processed by the device on the audio stream of the user, such as noise reduction processing, cropping processing, etc.

[0050] In actual application, the voice processing server can store historical voice data associated with the current voice data and historical reply data corresponding to the historical voice data. Therefore, in the case that the voice processing server performs data analysis on the input data, the current voice data can be more accurately analyzed according to the association information between the historical voice data, the historical reply data and the current voice data. The historical voice data associated with the current voice data can be understood as the audio stream of the user in the previous k rounds (k can be set according to actual situation) before the current dialogue round. The dialogue data in the previous k rounds is determined by the historical voice data and the historical reply data.

[0051] Specifically, in the case that the device is in a voice output state and receives the current voice data sent by the device, because it is necessary to determine whether the user wants to interrupt the content being output, data analysis is performed on the received current voice data to obtain a data analysis result, so that correct processing is made by using the data analysis result.

[0052] In one or more embodiments of the present specification, the voice processing server includes a state control unit. The state control unit performs data interruption and data integrity analysis on the current voice data to obtain corresponding data interruption analysis result and data integrity analysis result, and uses different results to update the state of the device. The specific implementation is as follows:

[0053] In the case that the device is in a voice output state and receives input data sent by the device, data analysis is performed on the input data to obtain a data analysis result, including:

[0054] The state control unit, in a case where the input data sent by the device end is received and it is determined that the device end is in a voice output state, performs data interruption and data integrity analysis on the input data to obtain a data interruption analysis result and a data integrity analysis result.

[0055] The data interruption analysis can be understood as analyzing whether the input data meets a first voice processing condition. The data interruption analysis can be performed by analyzing the relevance, content length, and other information of the previous k rounds of dialogue data. The first voice processing condition is that the data interruption analysis result indicates that the input data is used to interrupt the device output of the device end. The data interruption analysis result includes a result that the input data is used to interrupt the device output of the device end, or a result that the input data is not used to interrupt the device output of the device end. That is, in a case where the data interruption analysis result is the result that the input data is used to interrupt the device output of the device end, it is determined that the input data meets the first voice processing condition. In a case where the data interruption analysis result is the result that the input data is not used to interrupt the device output of the device end, it is determined that the input data does not meet the first voice processing condition.

[0056] The data integrity analysis can be understood as analyzing whether the input data meets a second voice processing condition. The data integrity analysis can be performed by analyzing the relevance, pause length, and other information of the previous k rounds of dialogue data. The second input data processing condition is that the data integrity analysis result indicates that the input data meets a preset data integrity. The data integrity analysis result includes a result that the input data is complete data, or a result that the input data is incomplete data. That is, in a case where the data integrity analysis result is that the input data is complete data, it is determined that the input data meets the second voice processing condition. In a case where the data integrity analysis result is that the input data is incomplete data, it is determined that the input data does not meet the second voice processing condition.

[0057] Specifically, the state control unit can include a dialogue state control large model. The dialogue state control large model is used to perform data interruption analysis on the received input data. In a specific data interruption analysis case, the semantic, content length (time length), and relevance to historical dialogue data of the current voice data are analyzed to obtain a data interruption analysis result. The data interruption analysis result includes the semantic, content length (time length), and relevance to historical dialogue data of the current voice data.

[0058] The dialogue state control large model is used to perform data integrity analysis on the received input data. In the case of data integrity analysis, the semantic integrity, pause duration, and relevance to historical dialogue data of the current voice data are analyzed to obtain a data integrity analysis result, that is, the data integrity analysis result includes the semantic integrity, pause duration, and relevance to historical dialogue data of the current voice data.

[0059] The voice processing method provided by the embodiments of the present specification uses the dialogue state control large model of the state control unit to perform data interruption and data integrity analysis on voice data, thereby quickly and accurately obtaining data interruption analysis results and data integrity analysis results to provide a correct direction for subsequent processing according to the data interruption analysis results and the data integrity analysis results.

[0060] Step 204: In the case where the data analysis result indicates that the input data meets the first voice processing condition, a first state update instruction is sent to the device end, wherein the first state update instruction is used to update the device end from the voice output state to a voice input state.

[0061] In one or more embodiments of the present specification, the data analysis result includes a data interruption analysis result. Based on the data interruption analysis result, it is determined whether the input data is input data for interrupting the device output of the device end. If so, the state of the device end is changed by sending a first state update instruction to the device end. The specific implementation is as follows:

[0062] In the case where the data analysis result indicates that the input data meets the voice processing condition, a first state update instruction is sent to the device end, comprising:

[0063] In the case where the data interruption analysis result indicates that the input data is input data for interrupting the device output of the device end, a first state update instruction is sent to the device end.

[0064] Wherein, the device output can be understood as the reply data output by the device end in the last round for the current voice data of the user; the voice input state can be understood as a state of not outputting any data but processing the received voice data, for example, the voice input state is a "listen" state.

[0065] Specifically, when the device end is in a voice output state and is outputting device output, the user inputs current voice data, the voice processing server side obtains a data interruption analysis result by comprehensively analyzing the current voice data and historical dialogue data, judges whether the current voice data of the user is data for interrupting the device output being output by the device end, that is, whether the user wants to interrupt, if yes, it indicates that the user wants to interrupt the device output of the device end, therefore, a first state update instruction needs to be sent to the device end to make the device end update from the voice output state to the voice input state according to the first state update instruction, for example, update the “speak” state to the “listen” state, the device end ends and terminates the device output being output, and receives and forwards the current voice data of the user to the voice processing server side.

[0066] In actual application, when the device end is in the “speak” state, the device end still listens to the user audio input in real time, and whether the user needs to insert the dialogue is judged by the dialogue state control large model; if it is judged that the user needs to insert, the state of the device end is updated to the “listen” state.

[0067] The voice processing method provided by the embodiments of the present specification can accurately determine whether the input data is data for interrupting the device output of the device end according to the data interruption analysis result, and if yes, the state change of the device end is realized by timely sending the first state update instruction to the device end.

[0068] In one or more embodiments of the present specification, the voice processing server side includes a state control unit and a state management unit; the state control unit is used to make a judgment according to the data analysis result, and the state management unit is used to interact with the device end to send the first state update instruction to the device end. The specific implementation is as follows:

[0069] In the case where the data analysis result indicates that the voice data meets the first voice processing condition, the first state update instruction is sent to the device end, including:

[0070] The state control unit sends the first state update instruction to the state management unit in the case where the data analysis result indicates that the voice data meets the first voice processing condition;

[0071] The state management unit sends the first state update instruction to the device end.

[0072] The state management unit is a unit responsible for actually performing the state update operation, which receives the first state update instruction from the state control unit and forwards the first state update instruction to the device end, so as to change the current state of the device end.

[0073] Specifically, the state management unit receives a first state update instruction from the state control unit, sends the first state update instruction to the device end, and instructs the device end to update its state according to the first state update instruction, that is, the device end updates the device end from the current voice output state (such as the "speak" state) to the voice input state (such as the "listen" state) according to the first state update instruction, so that the device end can stop the current output and start receiving and processing new voice input of the user.

[0074] Referring to FIG. 3, FIG. 3 shows a processing schematic diagram of a dialogue state control large model in a voice processing method provided by an embodiment of the present specification.

[0075] In FIG. 3, the multi-modal large model includes a dialogue state control large model, the model input is the historical dialogue data of the last K rounds (including historical voice data and historical reply data), and the current voice data and the historical voice data are obtained by voice coding to obtain voice features. Through feature alignment, the dimension of the voice feature is converted to align the voice feature with the text feature space to obtain aligned voice features. The historical reply data can be voice modality reply data or text modality reply data. In the embodiment of the present specification, the historical reply data is taken as an example of text modality reply data, the historical reply data is encoded to obtain reply text features, and the current state of the device end (such as the voice output state) is input into the dialogue state control large model. By analyzing the aligned voice features and the reply text features, the state instruction of the next state is output, that is, the first state instruction is output, when it is determined that the current voice data is used to interrupt the output of the device end.

[0076] The voice processing method provided by the embodiment of the present specification, the voice processing service end includes a state control unit and a state management unit, the state control unit is used for judging whether the state of the device end needs to be updated according to the data analysis result of the input data, and the state management unit is used for being responsible for actually executing the state update operation. Through this kind of design, the voice processing service end can flexibly adjust the state of the device end according to the voice input of the user, and realize a more natural and smooth interaction experience.

[0077] Step 206: parsing the input data, obtaining reply data corresponding to the input data, and sending the reply data to the device end.

[0078] The reply data can be voice modality data, text modality data, a combination of text and voice data, or even video modality data, which is not limited here.

[0079] In one or more embodiments of the present specification, data analysis is performed on the input data, and it is determined whether the input data meets the second voice processing condition according to the data analysis result. If so, the input data is parsed to obtain corresponding reply data. The specific implementation is described as follows:

[0080] The parsing of the voice data to obtain the reply data corresponding to the voice data comprises:

[0081] In the case where the data analysis result indicates that the voice data meets the second voice processing condition, the voice data is parsed to obtain the reply data corresponding to the voice data.

[0082] In practical applications, the data analysis result includes a data integrity analysis result, and the second voice data processing condition is that the data integrity analysis result indicates that the input data meets a preset data integrity.

[0083] In the case where the data analysis result indicates that the voice data meets the second voice processing condition, the input data is parsed to obtain the reply data corresponding to the input data, comprising:

[0084] In the case where the data integrity analysis result indicates that the voice data meets the preset data integrity, the input data is parsed to obtain the reply data corresponding to the input data.

[0085] The preset data integrity can be understood as a condition that is preset and determines that the voice data is complete voice data, such as semantic integrity, pause duration, etc.

[0086] Specifically, in the case where the data integrity analysis result indicates that the voice data meets the preset data integrity, it is considered that the current input voice data of the user is coherent and complete, i.e., the user has finished speaking at this time and is waiting for the system to reply, i.e., the round switching is to be performed, the user ends the speech, and the device end needs to output the reply data.

[0087] In practical applications, when the device end is in the "listen" state, the voice input by the user is sent into the dialogue state control large model in real time in fragments until the model determines that the user has finished speaking. At this time, the text reply is generated for the voice data input by the user. In the case where the reply data is text modal data, the generated text reply can be output by the device end. In the case where the reply data is voice modal data, the generated text reply is sent into a streaming speech synthesis module for audio synthesis to obtain the reply data in the voice modal.

[0088] The voice processing method provided by the embodiments of the present specification can accurately determine whether the current voice data meets the preset data integrity according to the data integrity analysis result, correctly process based on the determination result, and only when the voice data meets the preset data integrity, analyze the voice data to obtain accurate reply data, thereby improving user experience.

[0089] In one or more embodiments of the present specification, the voice processing server comprises a state control unit and a voice reply unit; the state control unit is used to determine whether the input data meets the second voice processing condition, and the voice reply unit is used to analyze and process the input data to obtain reply data. The specific implementation is as follows:

[0090] The voice data is analyzed to obtain reply data corresponding to the voice data, comprising:

[0091] The state control unit, in a case where it is determined according to the data analysis result that the voice data meets the second voice processing condition, sends the voice data to the voice reply unit;

[0092] The voice reply unit analyzes the voice data to obtain reply data corresponding to the voice data.

[0093] The voice reply unit can be understood as a unit responsible for generating and sending reply information in the system. The voice reply unit receives voice data that meets the second voice processing condition from the state control unit, analyzes and generates replies based on the voice data.

[0094] Specifically, the voice reply unit comprises a voice semantic large model. The voice semantic large model receives voice data that meets the second voice processing condition, analyzes the received voice data, including but not limited to voice recognition and further semantic analysis, to accurately understand the user's intention and demand, and generates corresponding reply data based on the analysis result. These reply data can be in the form of text for subsequent text reply; or can be directly converted into voice format for voice reply. Of course, in actual application, text reply and voice reply can be output as reply data together.

[0095] The voice processing method provided by the embodiments of the present specification realizes the processing of voice data that meets the second voice processing condition through the close cooperation between the state control unit and the voice reply unit, obtains corresponding reply data, and ensures the accuracy of the reply data in the case of complete voice data.

[0096] In one or more embodiments of the present specification, in the case of reply data in the voice mode, the reply text of the input data is first obtained, and then the reply text is synthesized by voice to obtain the reply data in the voice mode. The specific implementation is described as follows:

[0097] The input data is parsed to obtain the reply data corresponding to the input data, including:

[0098] The input data is parsed to obtain the reply text of the input data;

[0099] The reply text is synthesized by voice to obtain the reply data corresponding to the input data.

[0100] Specifically, the voice reply unit includes a text reply module and a voice synthesis module; the text reply module is used to parse the voice data and generate corresponding reply text, and the voice synthesis module is responsible for converting the reply text generated by the text reply module into natural and fluent voice reply; wherein the voice synthesis module uses text-to-speech technology to convert text information into audible voice signals.

[0101] Specifically, the text reply module includes a voice semantic large model, which receives voice data satisfying the second voice processing condition from the state control unit and generates text reply in a streaming manner, and sends the generated reply text to the voice synthesis module for subsequent voice synthesis processing.

[0102] The voice synthesis module converts text into voice using text-to-speech technology, which includes but is not limited to multiple steps such as text analysis, voice coding, sound synthesis, etc., aiming to generate natural and fluent voice signals matching the original text content, and output the synthesized voice reply to the device end, which can be achieved by speaker playback, saving to file, network transmission, etc.

[0103] In one or more embodiments of the present specification, the voice synthesis module uses a voice synthesis model to obtain text features of the reply text, uses a vocoder to generate voice waveforms, and synthesizes reply data in the voice mode. The specific implementation is described as follows:

[0104] The reply text is synthesized by voice to obtain the reply data corresponding to the input data, including:

[0105] Using a voice synthesis model, the text features of the text reply are determined, and the text features are input into a vocoder;

[0106] Using the vocoder, the text features are generated by voice waveform to obtain the reply data corresponding to the input data.

[0107] Specifically, the speech synthesis process adopts a progressive synthesis strategy for streaming speech synthesis (Streaming TTS, Streaming Text-to-Speech), that is, starting speech synthesis while inputting text, using prediction and parallel processing technology to reduce the delay of generating speech; and in the case of streaming text reply generated by the voice semantic large model, the voice semantic large model outputs a predicted text token (the basic unit of model processing) one by one, and sends each output predicted text token to a speech synthesis model (such as a neural text-to-speech model, i.e., Neural TTS model). The Neural TTS model uses a deep learning language model (such as a generative pre-trained transformer, i.e., GPT, a general model in the field of pre-trained language models, i.e., Text to Text Transfer Transformer, i.e., T5) to convert text-to-speech intermediate representation for streaming vocoder input. The streaming vocoder uses a streaming neural vocoder (such as a wave recurrent neural network, i.e., WaveRNN, a parallel wave generative adversarial network, i.e., Parallel WaveGAN) to provide efficient real-time speech waveform generation, and synthesize the reply data corresponding to the input data.

[0108] The voice processing method provided by the embodiments of the present specification generates text reply by a voice semantic large model and performs streaming speech synthesis by a progressive synthesis strategy, which greatly reduces the time delay of speech synthesis, realizes fast response to user input voice, meets the characteristics of real interaction, and can respond with extremely low waiting time (theoretically, the time can reach within 100 milliseconds, and the response time of mainstream streaming TTS (Text-to-Speech) is within 500 milliseconds).

[0109] In one or more embodiments of the present specification, in the case of obtaining the reply data, a second state update instruction is sent to the device end to update the device end from the voice input state to the voice output state, so that the reply data is output in the voice output state. The specific implementation is as follows:

[0110] After analyzing the voice data and obtaining the reply data corresponding to the voice data, the method further includes:

[0111] sending a second state update instruction to the device end, wherein the second state update instruction is used to update the device end from the voice input state to the voice output state, and outputting the reply data.

[0112] Specifically, after the voice processing server obtains the reply data corresponding to the voice data, the reply data needs to be output to the user through the device end, therefore, at this time, the second state update instruction needs to be sent to the device end through the state management unit, and the device end updates from the voice input state (the "listen" state) to the voice output state (the "speak" state) and outputs the reply data in the case of receiving the second state update instruction.

[0113] In one or more embodiments of the present specification, the voice processing server can receive feedback information of the device end, such as a device state update instruction sent by the device end or a reply data output completion instruction, so as to update the current state of the device end according to the instruction of the device end. The specific implementation is as follows:

[0114] After sending the reply data to the device end, the method further comprises:

[0115] In the case of receiving the device state update instruction sent by the device end and determining that the device end is in the voice output state, a first state update instruction is sent to the device end, wherein the first state update instruction is used to update the device end from the voice output state to the voice input state; or

[0116] In the case of receiving the reply data output completion instruction sent by the device end and determining that the device end is in the voice output state, a first state update instruction is sent to the device end.

[0117] Specifically, the interactive operation of the user on the device end can trigger the device end to send a device state update instruction to the voice processing server, such as the existence of a break control and an exit control on the user interaction interface of the device end, in the case of the user clicking these controls, triggering the device end to send a device state update instruction to the voice processing server; in the case of the voice processing server receiving the device state update instruction sent by the device end and determining that the device end is in the voice output state, a first state update instruction is sent to the device end through the state management unit.

[0118] Or, in the case that the user does not interact on the device side and the user does not interrupt, the device side sends a reply data output completion instruction to the voice processing server side in the case that the reply data output is completed, to indicate that the device side has completed output and needs to switch from the "speak" state to the "listen" state. At this time, the voice processing server side sends a first state update instruction to the device side in the case that the voice processing server side receives the reply data output completion instruction sent by the device side and determines that the device side is in the voice output state.

[0119] The voice processing method provided by the embodiments of the present specification can accept feedback information of the device side, such as a device state update instruction corresponding to "user clicks interrupt" and "user exits", and a reply data output completion instruction corresponding to "broadcasting is completed", and accurately adjust the global state by using these instructions.

[0120] In one or more embodiments of the present specification, the voice processing method further comprises:

[0121] In the case that the device side is in the voice input state, receiving initial voice data sent by the device side;

[0122] Performing data integrity analysis on the initial voice data to obtain a data integrity analysis result;

[0123] In the case that the data integrity analysis result indicates that the initial voice data meets a second voice processing condition, sending a second state update instruction to the device side, wherein the second state update instruction is used to update the device side from the voice input state to the voice output state;

[0124] Analyzing the initial voice data to obtain initial reply data corresponding to the initial voice data, and sending the initial reply data to the device side.

[0125] Specifically, in the case that the device side is in the "listen" state, receiving initial voice data of the user, and performing data integrity analysis on the initial voice data, so as to process the initial voice data according to the data integrity analysis result. For details, refer to the above embodiments, which will not be repeated here.

[0126] The voice processing method provided by the embodiments of the present specification is based on the state control unit, which avoids the need to define a long silence time for judgment in the "round switching" detection in the prior art, naturally exists a fixed waiting time, and causes an unnatural effect of human-computer interaction. The state control unit is used to judge "round switching" and "user interruption" in different states, and the judgment is combined with current voice data and historical dialogue data, which conforms to the real interaction characteristics and can respond with extremely low waiting based on the voice reply unit.

[0127] Referring to FIG. 4, FIG. 4 shows a processing process flow diagram of a voice processing method provided by an embodiment of the present specification.

[0128] An embodiment of the present specification takes the device end in a voice input state as an example to explain the voice processing method in detail.

[0129] When the device end is in the "listen" state, the dialogue state control large model (for example, the model is modified and optimized based on the Audio-text Large Language Models audio reply voice model, and it is noted that the model and the subsequent voice semantic large model can share the model structure and parameters) receives the audio stream, analyzes the audio stream and the dialogue history, judges according to multiple factors such as semantic integrity, pause length, and previous association, and decides whether the AI end (i.e., the voice processing server in the above embodiment) needs to start responding to the user; 0 indicates that the user is in a speaking state or the user has not started speaking at all, and the AI side does not need to respond at this time; 1 indicates that the user's new round of input has been completed, and the AI side needs to respond at this time.

[0130] Specifically, the audio stream of the user's speech is sent in real time to the dialogue state control large model of the multi-modal large model, and the historical dialogue data of the user (including the historical audio stream of the user and the text reply of the multi-modal large model), the current audio stream, and the current state of the device end are input into the dialogue state control large model; the dialogue state control large model identifies whether the semantics of the received current audio stream is complete according to the "listen" state of the device end, judges according to multiple factors such as semantic integrity, pause length, and previous association, determines that the user has finished speaking, and outputs "1"; in actual application, the output <eos>, indicating that the user has finished speaking, and start to report into "speak" state (here, an over-parameter can be set to dynamically control the tolerance level); otherwise, output "0", in practical applications, output <listen>, indicating that the user continues to speak.

[0131] In the case that the dialogue state control large model outputs "1", that is, the user has finished speaking, the historical dialogue data of the user and the current audio stream are input into the voice semantic large model, and the voice semantic large model can generate a more accurate text reply for the current audio stream according to the context of the historical dialogue data; the text reply output by the voice semantic large model is input into the stream voice synthesis module. Specifically, the stream voice synthesis module converts the text reply into voice features by using a deep learning language model, and inputs the voice features into a vocoder to generate a broadcast voice. In the process of generating the broadcast voice by the stream voice synthesis module, a progressive synthesis strategy is adopted, that is, instead of inputting the complete text reply generated by the voice reply module into the stream voice synthesis module, the stream voice synthesis module generates voice every time a text Token is predicted by the voice reply module, so as to reduce the delay of generating voice.

[0132] In the case that the dialogue state control large model outputs "1", that is, the user has finished speaking, the output result of the dialogue state control large model is sent to the cloud dialogue state management module (that is, the state management unit in the above embodiment) to make the cloud dialogue state management module send a state switching instruction (that is, the second state update instruction in the above embodiment) to the device end, and the state switching instruction of the cloud dialogue state management module and the broadcast voice generated by the stream voice synthesis module are sent to the device end to make the device end switch to the "speak" state and start broadcasting the generated broadcast voice.

[0133] Referring to FIG. 5, FIG. 5 shows a process flow diagram of another voice processing method provided by an embodiment of the present specification.

[0134] The embodiments of the present specification take the device end in the voice output state as an example to explain the voice processing method in detail.

[0135] When the device end state is the "speak" state, the dialogue state control large model receives the audio stream, analyzes the audio stream and the dialogue history, and comprehensively judges whether the user's speech at this time constitutes an interruption to the AI end according to the current input audio semantics, content length and other factors; 0 indicates that the user's input audio stream has a relatively complete semantic to try to insert into the dialogue, and the AI end needs to stop speaking at this time; 1 indicates that the user's input audio stream does not have complete semantics or interruption ability, and the AI end continues to broadcast.

[0136] When the device is in "speaking" mode, it also listens for the user's audio input. When the device detects the user's audio stream, it inputs the user's audio stream, historical dialogue data, and the device's current state into the dialogue state control model. The dialogue state control model comprehensively judges whether the user's speech constitutes an interruption to the AI ​​based on factors such as the semantics and content length of the current input audio. If the judgment result is yes, that is, it is determined that the user wants to interrupt, it outputs "0". In practical applications, the output is... <stop>, indicating that the user is ready to intervene, interrupt the broadcast into the "listen" state (here allows to set a super parameter, for dynamic control tolerance); and in the case of determining that the user has no to break, output "1", the actual application, output <speak>, indicating to continue the broadcast.

[0137] In the case of the dialogue state control large model outputting "0", that is, the user is ready to intervene, the cloud dialogue state management module sends a state switching instruction (which is the first state update instruction in the above embodiment) to the device end according to the output of the dialogue state control large model, so that the device end switches the state of the device end to the "listen" state according to the state switching instruction.

[0138] The user can click on the device end user interaction interface to interrupt, and the device end can send a user interruption instruction to the cloud dialogue state management module according to the user's click interruption action, so that the cloud dialogue state management module returns a state switching response based on the user interruption instruction, and switches the state of the device end to the "listen" state.

[0139] In actual application, in the case that the device end completes the broadcast of the broadcast voice, the device end can send a broadcast completion instruction to the cloud dialogue state management module, so that the cloud dialogue state management module returns a state switching response based on the broadcast completion instruction, thereby switching the state of the device end to the "listen" state.

[0140] The voice processing method provided by the embodiments of the present specification can, in the case that the device end is in a voice output state and the input data sent by the device end is received, determine whether the input data meets the first voice processing condition by analyzing the input data, and in the case of meeting, change the voice output state of the device end to a voice input state by sending a first state update instruction to the device end, analyze the voice data to obtain reply data, and send the reply data to the device end. That is, the device end can still receive voice data and forward it to the voice processing service end in the voice output state, and the voice processing service end can update the state of the device end in a timely manner in the case that the voice data meets the first voice processing condition through analysis of the voice data. By receiving voice data in a timely manner and updating the state of the device end, the machine can speak and receive feedback at the same time, and obtain reply data quickly to give a reaction under very low latency, so that the human-computer interaction can achieve a very natural effect.

[0141] Corresponding to the above method embodiments, the present specification also provides voice processing system embodiments. FIG. 6 shows a structural schematic diagram of a voice processing system according to an embodiment of the present specification. As shown in FIG. 6, the system includes a voice processing service end 104 and a device end 102, wherein,

[0142] The device end 102 is configured to send the received input data to the voice processing service end 104.

[0143] The voice processing server 104 is configured to, in a case where the device end 102 is in a voice output state and the input data sent by the device end 102 is received, perform data analysis on the input data to obtain a data analysis result, and in a case where the data analysis result indicates that the voice data meets a voice processing condition, send a first state update instruction to the device end 102.

[0144] The device end 102 is further configured to update the voice output state to a voice input state according to the first state update instruction.

[0145] The voice processing server 104 is further configured to parse the input data to obtain reply data corresponding to the input data, and send the reply data to the device end 102.

[0146] Optionally, the voice processing server 104 is further configured to send a second state update instruction to the device end 102.

[0147] Optionally, the device end 102 is further configured to update the voice input state to the voice output state according to the second state update instruction, and output the reply data.

[0148] Optionally, the voice processing server 104 is further configured to:

[0149] In a case where the data analysis result indicates that the input data meets a second voice processing condition, parse the input data to obtain reply data corresponding to the input data.

[0150] Optionally, the voice processing server 104 is further configured to:

[0151] In a case where it is determined according to the data integrity analysis result that the input data meets a preset data integrity, parse the input data to obtain reply data corresponding to the input data.

[0152] Optionally, the voice processing server 104 is further configured to:

[0153] Parse the input data to obtain reply text of the input data.

[0154] Perform voice synthesis on the reply text to obtain the reply data corresponding to the input data.

[0155] Optionally, the voice processing server 104 is further configured to:

[0156] Determine a text feature of the text reply by using the voice synthesis model, and input the text feature into a vocoder.

[0157] The voice coder is used to generate a speech waveform for the text feature, and obtain reply data corresponding to the input data.

[0158] Optionally, the speech processing server 104 is further configured to send a first state update instruction to the device end 102 in a case where the speech processing server 104 receives a device state update instruction sent by the device end 102 and determines that the device end 102 is in the speech output state.

[0159] Optionally, the speech processing server 104 is further configured to send a first state update instruction to the device end 102 in a case where the speech processing server 104 receives a reply data output completion instruction sent by the device end 102 and determines that the device end 102 is in the speech output state.

[0160] Optionally, the speech processing server 104 is further configured to receive initial speech data sent by the device end 102 in a case where the speech processing server 104 determines that the device end 102 is in the speech input state, perform data integrity analysis on the initial speech data, and obtain a data integrity analysis result; in a case where the data integrity analysis result indicates that the initial speech data meets a second speech processing condition, send a second state update instruction to the device end 102, analyze the initial speech data, obtain initial reply data corresponding to the initial speech data, and send the initial reply data to the device end 102.

[0161] The above is a schematic scheme of the speech processing system of the embodiment. It should be noted that the technical scheme of the speech processing system belongs to the same concept as the technical scheme of the speech processing method described above, and the details of the technical scheme of the speech processing system that are not described in detail can be referred to the description of the technical scheme of the speech processing method.

[0162] FIG. 7 shows a structural block diagram of a computing device 700 according to an embodiment of the present specification. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 through a bus 730, and a database 750 is used to save data.

[0163] The computing device 700 also includes an access device 740 that enables the computing device 700 to communicate via one or more networks 760. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or combinations of such networks, such as the Internet. The access device 740 can include one or more of any type of network interface (for example, a network interface card (NIC)) such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC).

[0164] In one embodiment of the present specification, the above-mentioned components of the computing device 700 and other components not shown in FIG. 7 can also be connected to each other, for example, through a bus. It should be understood that the computing device structure block diagram shown in FIG. 7 is only for the purpose of example, and is not a limitation on the scope of the present specification. Other components can be added or replaced as needed by those skilled in the art.

[0165] The computing device 700 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (for example, a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (for example, a smartphone), a wearable computing device (for example, a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 700 can also be a mobile or stationary server.

[0166] Among them, the processor 720 is used to execute the following computer program / instruction, which is executed by the processor to realize the steps of the above voice processing method.

[0167] The various embodiments in the specification are described in progressive manner, and the same or similar parts among the various embodiments can be mutually referred to, and each embodiment focuses on the difference from other embodiments. In particular, for the computing device embodiments, since they are basically similar to the speech processing method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the speech processing method embodiments.

[0168] An embodiment of the specification further provides a computer readable storage medium storing computer programs / instructions, which, when executed by a processor, implement the steps of the above speech processing method.

[0169] The various embodiments in the specification are described in progressive manner, and the same or similar parts among the various embodiments can be mutually referred to, and each embodiment focuses on the difference from other embodiments. In particular, for the computer readable storage medium embodiments, since they are basically similar to the speech processing method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the speech processing method embodiments.

[0170] An embodiment of the specification further provides a computer program product, comprising computer programs / instructions, which, when executed by a processor, implement the steps of the above speech processing method.

[0171] The above is a schematic scheme of a computer program product of the embodiment. It should be noted that the technical scheme of the computer program product and the technical scheme of the above speech processing method belong to the same concept, and the details of the technical scheme of the computer program product which are not described in detail can be referred to the description of the technical scheme of the speech processing method.

[0172] The above describes specific embodiments of the specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than the order in which they are recited and still achieve desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.

[0173] The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, software distribution medium, etc. It should be noted that the computer readable medium can include appropriate additions or deletions according to the requirements of patent practice, for example, according to the patent practice in some regions, the computer readable medium does not include electrical carrier signals and telecommunication signals.

[0174] It should be noted that for the foregoing method embodiments, in order to facilitate description, they are all expressed as a combination of a series of actions, but those skilled in the art should know that the embodiments of the present specification are not limited by the order of the described actions, because according to the embodiments of the present specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of the present specification.

[0175] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0176] The preferred embodiments of the present specification disclosed above are only used to help explain the present specification. The alternative embodiments do not describe all the details and limit the invention to the specific embodiments described. Obviously, according to the content of the embodiments of the present specification, many modifications and changes can be made. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present specification, so that those skilled in the art can well understand and use the present specification. The present specification is limited only by the claims and their entire scope and equivalents.< / speak> < / stop> < / listen> < / eos>

Claims

1. A voice processing method applied to a voice processing server, comprising: in a case that a device end is in a voice output state and input data sent by the device end is received, performing data analysis on the input data to obtain a data analysis result; in a case that the data analysis result indicates that the input data satisfies a first voice processing condition, sending a first state update instruction to the device end, wherein the first state update instruction is used to update the device end from the voice output state to a voice input state; parsing the input data to obtain reply data corresponding to the input data, and sending the reply data to the device end. 2.The voice processing method of claim 1, wherein the sending the reply data to the device end comprises: sending a second state update instruction and the reply data to the device end, wherein the second state update instruction is used to update the device end from the voice input state to the voice output state, and output the reply data. 3.The voice processing method of claim 1 or 2, wherein the data analysis result comprises a data interruption analysis result, and the first voice processing condition is that the data interruption analysis result indicates that the input data is used to interrupt device output of the device end. 4.The voice processing method of claim 3, wherein in a case that the data interruption analysis result is a result that the input data is used to interrupt device output of the device end, it is determined that the input data satisfies the first voice processing condition; and in a case that the data interruption analysis result is a result that the input data is not used to interrupt device output of the device end, it is determined that the input data does not satisfy the first voice processing condition. 5.The voice processing method of claim 4, further comprising: determining whether the input data is input data used to interrupt device output of the device end based on the data interruption analysis result; and in a case that the input data is input data used to interrupt device output of the device end, changing a state of the device end by sending a first state update instruction to the device end. 6.The voice processing method of claim 5, wherein the sending the first state update instruction to the device end in a case that the data analysis result indicates that the input data satisfies the first voice processing condition comprises: in a case that it is determined that the input data is input data used to interrupt device output of the device end according to the data interruption analysis result, sending the first state update instruction to the device end. 7.The voice processing method of any one of claims 1 to 6, wherein the parsing the input data to obtain reply data corresponding to the input data comprises: in a case that the data analysis result indicates that the input data satisfies a second voice processing condition, parsing the input data to obtain reply data corresponding to the input data. ​ 8. The voice processing method of claim 7, wherein the data analysis result comprises a data integrity analysis result, and the second voice processing condition is that the data integrity analysis result indicates that the input data satisfies a preset data integrity. The input data is parsed to obtain reply data corresponding to the input data in a case where the data analysis result indicates that the input data satisfies a second voice processing condition. The input data is parsed to obtain reply data corresponding to the input data in a case where the data integrity analysis result indicates that the input data satisfies a preset data integrity.

9. The voice processing method of any one of claims 1 to 6, wherein the input data is parsed to obtain reply data corresponding to the input data, comprising: The input data is parsed to obtain reply text of the input data. The reply text is subjected to voice synthesis to obtain the reply data corresponding to the input data.

10. The voice processing method of claim 9, wherein the reply text is subjected to voice synthesis to obtain the reply data corresponding to the input data, comprising: A text feature of the text reply is determined using a voice synthesis model, and the text feature is input into a vocoder. The text feature is subjected to voice waveform generation using the vocoder to obtain the reply data corresponding to the input data.

11. The voice processing method of any one of claims 1 to 10, further comprising, after the reply data is sent to the device end: In a case where a device state update instruction sent by the device end is received and it is determined that the device end is in the voice output state, a first state update instruction is sent to the device end, wherein the first state update instruction is used to update the device end from the voice output state to the voice input state.

12. The voice processing method of any one of claims 1 to 11, further comprising, after the reply data is sent to the device end: In a case where a reply data output completion instruction sent by the device end is received and it is determined that the device end is in the voice output state, a first state update instruction is sent to the device end.

13. The voice processing method of any one of claims 1 to 12, wherein the input data comprises current voice data received by the device end.

14. The voice processing method of any one of claims 1 to 13, further comprising: In a case where it is determined that the device end is in the voice input state, initial voice data sent by the device end is received. The initial voice data is subjected to data integrity analysis to obtain a data integrity analysis result. In a case where the data integrity analysis result indicates that the initial voice data satisfies a second voice processing condition, a second state update instruction is sent to the device end, wherein the second state update instruction is used to update the device end from the voice input state to the voice output state. analyzing the initial voice data to obtain initial reply data corresponding to the initial voice data, and sending the initial reply data to the device end. 15.A voice processing method applied to a device end, comprising: sending received input data to a voice processing server, so that the voice processing server performs data analysis on voice data in a voice output state of the device end, obtains a data analysis result, and sends a first state update instruction to the device end in a case where the data analysis result indicates that the input data satisfies a first voice processing condition; updating from the voice output state to a voice input state according to the first state update instruction. 16.A voice processing system comprising a voice processing server and a device end, wherein the device end is configured to send received input data to the voice processing server; the voice processing server is configured to perform data analysis on voice data in a voice output state of the device end, obtain a data analysis result, and send a first state update instruction to the device end in a case where the data analysis result indicates that the input data satisfies a first voice processing condition; the device end is further configured to update from the voice output state to a voice input state according to the first state update instruction; the voice processing server is further configured to analyze the input data to obtain reply data corresponding to the input data, and send the reply data to the device end. 17.A voice processing method system according to claim 16, wherein the voice processing server is further configured to send a second state update instruction to the device end; the device end is further configured to update from the voice input state to the voice output state according to the second state update instruction, and output the reply data. 18.A computing device, comprising: a memory and a processor; the memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the voice processing method of any one of claims 1 to 15. 19.A computer readable storage medium storing computer programs / instructions, which, when executed by a processor, implement the steps of the voice processing method of any one of claims 1 to 15. 20.A computer program product comprising computer programs / instructions, which, when executed by a processor, implement the steps of the voice processing method of any one of claims 1 to 15.

Citation Information

Patent Citations

  • Voice interaction method and device and computer readable storage medium

    CN112637431A

  • Voice dialogue processing method and device based on multi-modal features and electronic equipment

    CN114078474A

  • Voice interaction method and system, electronic equipment and storage medium

    CN115148205A

  • Real-time semantic understanding method and system for spoken dialogue and electronic equipment

    CN116052664A

  • Semantic integrity judgment method and system based on multi-head attention mechanism

    CN118520879A