Streaming Voice Interaction Method and Related Devices, Equipment, and Storage Media

Through endpoint detection and sliding window operation combined with voice recognition and intelligent dialogue model, the problem of untimely streaming voice interaction is solved, accurate processing and timely interaction of voice frames are realized, and the timeliness and efficiency of streaming voice interaction is improved.

CN119479620BActive Publication Date: 2025-05-30IFLYTEK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510026410.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-30
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The existing streaming voice interaction technology is difficult to interact in a timely manner during voice input, and there is a problem of early or too late interaction intervention, which affects the applicability.

Method used

The start endpoint of streaming voice is obtained through endpoint detection, and the speech frame is extracted through sliding window operation, and the speech recognition system and intelligent dialogue model are used to perform feature extraction and classification prediction, judge the semantic end, generate reply text, and realize the process closed loop of streaming voice interaction.

Benefits of technology

Improve the timeliness of streaming voice interaction, ensure that the interaction is carried out at the end of semantics, avoiding early or too late intervention, and improving the accuracy and efficiency of the interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479620B_ABST
    Figure CN119479620B_ABST
Patent Text Reader

Abstract

The present application discloses a streaming voice interaction method and related devices, equipment, and storage media. Among them, the streaming voice interaction method includes: performing endpoint detection on the streaming voice, and in response to detecting the start endpoint of the streaming voice, performing a sliding window operation on the streaming voice to obtain speech frames, extracting features based on the speech frames to obtain the speech features of the speech frames; inputting the speech features of the speech frames into a speech recognition system for recognizing the streaming voice to obtain the recognition results of the speech frames, and performing classification prediction based on the encoding features of the speech frames to obtain the classification results of the speech frames; in response to the classification result indicating the end of semantics, obtaining the recognition text based on the recognition results of each speech frame from the start endpoint to the end endpoint, and at least processing the recognition text by an intelligent dialogue model to generate a reply text; in response to the classification result indicating that the semantics have not ended, continue to return and perform the sliding window operation. The above solution can improve the timeliness of streaming voice interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of voice interaction technology, and in particular to a streaming voice interaction method and related devices, equipment and storage media. Background Art

[0002] Streaming voice interaction is widely used in many scenarios, including customer service and business office. Unlike non-streaming voice interaction, which requires voice interaction to be performed after the voice input is completed, streaming voice interaction requires timely voice interaction during the voice input process.

[0003] Existing voice interaction technologies primarily focus on non-streaming speech and are difficult to apply to streaming speech. Furthermore, while a small number of streaming speech interaction technologies exist, these technologies suffer from flaws such as inappropriate timing (e.g., premature or late interaction) during the streaming process, severely hindering their widespread application. Therefore, improving the timeliness of streaming speech interaction has become a pressing issue. Summary of the Invention

[0004] The main technical problem solved by this application is to provide a streaming voice interaction method and related devices, equipment and storage media, which can improve the timeliness of streaming voice interaction.

[0005] In order to solve the above technical problems, the first aspect of the present application provides a streaming voice interaction method, including: performing endpoint detection on the streaming voice, and in response to detecting the starting endpoint of the streaming voice, performing a sliding window operation on the streaming voice starting from the starting endpoint to obtain a voice frame, and performing feature extraction based on the voice frame to obtain the voice features of the voice frame; inputting the voice features of the voice frame into a voice recognition system for identifying streaming voice to obtain a recognition result of the voice frame, and performing classification prediction based on the encoding features of the voice frame to obtain a classification result of the voice frame; wherein the encoding features are extracted by the encoding network in the intelligent dialogue model after the voice features are input into the intelligent dialogue model, and the classification result includes whether the voice frame represents an ending endpoint of a semantic end; in response to the classification result indicating the semantic end, obtaining a recognition text based on the recognition results of each voice frame from the starting endpoint to the ending endpoint, and generating a reply text based on at least the recognition text by the intelligent dialogue model; in response to the classification result indicating that the semantics have not ended, continuing to return to perform the sliding window operation.

[0006] To solve the above technical problems, the second aspect of the present application provides a streaming voice interaction device, comprising: a preparatory processing module, a recognition and classification module, a first response module, and a second response module. The preparatory processing module is configured to perform endpoint detection on the streaming voice and, in response to detecting the starting endpoint of the streaming voice, perform a sliding window operation on the streaming voice starting from the starting endpoint to obtain voice frames, and perform feature extraction based on the voice frames to obtain voice features of the voice frames. The recognition and classification module is configured to input the voice features of the voice frames into a voice recognition system for recognizing the streaming voice to obtain recognition results of the voice frames, and perform classification prediction based on the encoding features of the voice frames to obtain classification results of the voice frames. The encoding features are extracted by an encoding network in the intelligent dialogue model after the voice features are input into the intelligent dialogue model, and the classification result includes whether the voice frame represents an ending endpoint indicating a semantic end. The first response module is configured to, in response to the classification result indicating a semantic end, obtain recognition text based on the recognition results of each voice frame from the starting endpoint to the ending endpoint, and generate a reply text by processing the recognition text based on at least the recognition text by the intelligent dialogue model. The second response module is configured to, in response to the classification result indicating that the semantic end has not been completed, continue to return to perform the sliding window operation.

[0007] In order to solve the above technical problems, the third aspect of the present application provides an electronic device, which at least includes a memory and a processor coupled to each other, wherein the memory stores at least program instructions, and the processor is used to execute the program instructions to implement the streaming voice interaction method in the above first aspect.

[0008] In order to solve the above technical problems, the fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be executed by a processor, and the program instructions are used to implement the streaming voice interaction method of the above first aspect.

[0009] The above scheme performs endpoint detection on the streaming speech, and in response to detecting the starting endpoint of the streaming speech, performs a sliding window operation on the streaming speech starting from the starting endpoint to obtain a speech frame, and performs feature extraction based on the speech frame to obtain the speech features of the speech frame, and then inputs the speech features of the speech frame into a speech recognition system for identifying the streaming speech to obtain the recognition result of the speech frame, and performs classification prediction based on the coding features of the speech frame to obtain the classification result of the speech frame, and the coding features are extracted by the coding network in the intelligent dialogue model after being input by the speech features into the intelligent dialogue model, and the classification result includes whether the speech frame represents the ending endpoint of the semantic end, based on this, in response to the classification result representing the semantic end, based on the recognition results of each speech frame from the starting endpoint to the ending endpoint, a recognition text is obtained, and at least based on the recognition text, it is processed by the intelligent dialogue model to generate The reply text is obtained, and in response to the classification result representing that the semantics have not ended, the sliding window operation is continued to be performed. Therefore, on the one hand, in the streaming voice process, the sliding window operation can be continuously performed after the starting endpoint is detected to obtain the voice frame, and each voice frame is processed in turn until the voice frame representing the semantic end is detected, and then the reply text is generated, which can realize the closed loop process of streaming voice interaction. On the other hand, the voice features of the voice frame are sent to the voice recognition system for voice recognition on one side and to the intelligent dialogue model on the other side to be encoded by the encoding network in the intelligent dialogue model first and then classified and predicted based on whether the semantics have ended. With the help of the semantic understanding ability of the intelligent dialogue model, it can be judged as accurately as possible whether the voice frame has ended to determine whether to intervene in the interaction, and to avoid intervening in the interaction too early or too late as much as possible, which helps to improve the timeliness of streaming voice interaction. Therefore, the timeliness of streaming voice interaction can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 This is a flow chart of an embodiment of the streaming voice interaction method of the present application;

[0011] Figure 2 This is a process diagram of an embodiment of the streaming voice interaction method of the present application;

[0012] Figure 3 This is a schematic diagram of the framework of an embodiment of the streaming voice interaction device of the present application;

[0013] Figure 4 This is a schematic diagram of the framework of an embodiment of the electronic device of the present application;

[0014] Figure 5 It is a schematic diagram of a framework of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0015] The following describes the embodiments of the present application in detail with reference to the accompanying drawings.

[0016] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0017] The terms "system" and "network" are often used interchangeably in this document. The term "and / or" is simply a description of an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the fragment " / " generally indicates that the related objects are in an "or" relationship. Furthermore, "multiple" in this document refers to two or more than two.

[0018] See also Figure 1 , Figure 1 This is a flow chart of an embodiment of the streaming voice interaction method of the present application. Specifically, it may include the following steps:

[0019] Step S11: performing endpoint detection on the streaming speech, and in response to detecting the starting endpoint of the streaming speech, performing a sliding window operation on the streaming speech starting from the starting endpoint to obtain a speech frame, and performing feature extraction based on the speech frame to obtain speech features of the speech frame.

[0020] In the disclosed embodiments, streaming voice refers to an audio stream that is input or collected in real time, and the audio stream is processed during the real-time input or real-time collection process to perform voice interaction. Taking a business office scenario as an example, the real-time speech of a speaker in a meeting is streaming voice, such as "I think this project (short pause L seconds) is still under discussion"; or, taking an everyday life scenario as an example, the real-time voice input of a user during a human-computer dialogue is streaming voice, such as "(short pause M seconds) We have now arrived in City A (long pause N seconds) Please tell the children about the local weather conditions tomorrow in the tone of a kindergarten teacher." Of course, the above example is only one possible example in actual application, and other possible situations will not be given examples one by one here.

[0021] In one implementation scenario, the specific process of endpoint detection (Voice Activity Detection, VAD) can be found in the technical details of VAD and will not be elaborated on here. For example, consider the streaming voice "(short pause M seconds) We have now arrived in City A (long pause N seconds) Please tell the children about tomorrow's local weather conditions in the voice of a kindergarten teacher." When the current speech progress is in the "short pause M seconds", endpoint detection does not detect valid speech, so endpoint detection must continue on the streaming voice until the current speech progress reaches the point where the pronunciation of "I" begins, at which point the starting endpoint of the streaming voice can be detected. Of course, the above example is only one possible example in actual application, and other possible scenarios will not be given one by one here.

[0022] In one implementation scenario, the window length and window shift of the sliding window operation can be set according to application needs. For example, if a larger future view is required, the window shift can be set to a smaller one, or if a smaller future view is required, the window shift can be set to a larger one. For example, if the window length is set to 200ms, the window shift can be set to 120ms, meaning the future view is 80ms. In this case, with 0s of the streaming voice as the starting endpoint, the first voice frame is from 0 to 200ms of the streaming voice, and the second voice frame is from 120 to 320ms of the streaming voice, meaning the last 80ms of the first voice frame is the future view. Of course, the above example is only one possible example of a sliding window operation in actual application, and other possible scenarios will not be given here one by one. It should be noted that after determining the starting endpoint, the sliding window operation can be continuously performed on the streaming voice according to the window length and window shift to obtain individual voice frames, and subsequent operations can be performed on each voice frame in sequence. For details, please refer to the relevant description below and will not be repeated here.

[0023] In one implementation scenario, after the speech frame is obtained through the sliding window operation, feature extraction can be performed on the speech frame to obtain the speech features of the speech frame. As a possible example, acoustic features such as MFCC, FBank, etc. can be extracted from the speech frame as the speech features of the speech frame; or, as another possible example, feature extraction can also be performed on the speech frame based on the speech coding network to obtain the speech features of the speech frame. The coding network can include but is not limited to convolutional neural networks, long short-term memory networks, recurrent neural networks, etc., and the specific structure of the coding network is not limited here. Of course, the above examples are only a few possible examples of extracting speech features, and do not limit other possible ways of extracting speech features, so they will not be given one by one here.

[0024] Step S12: inputting the speech features of the speech frame into a speech recognition system for recognizing streaming speech to obtain a recognition result of the speech frame, and performing classification prediction based on the coding features of the speech frame to obtain a classification result of the speech frame.

[0025] In one implementation scenario, after extracting the speech features of a speech frame, one of the processing branches feeds the features into a speech recognition system for recognizing streaming speech, thereby obtaining a recognition result for the speech frame. It should be noted that the recognition result may include a number of characters, or a number of phonemes, and the specific content of the recognition result is not limited herein. Furthermore, the speech recognition system may include, but is not limited to, wenet, etc., and the specific structure of the speech recognition system is not limited herein.

[0026] In one implementation scenario, after extracting the speech features of a speech frame, another processing branch feeds the speech features into an intelligent dialogue model for processing. It should be noted that the intelligent dialogue model may include an encoding network, and after the speech features are encoded by the encoding network, encoding features may be obtained. On this basis, classification prediction can be performed based on the encoding features of the speech frame to obtain a classification result for the speech frame. It should be noted that the classification result may include whether the speech frame represents an end endpoint of a semantic end, where the end endpoint is the end moment of the speech frame representing the semantic end. Furthermore, the intelligent dialogue model may include, but is not limited to, a large language model, and the encoding network may include, but is not limited to, a Transformer, etc. The specific structures of the intelligent dialogue model and the encoding network are not limited herein.

[0027] In a specific implementation scenario, classification prediction can be performed by a neural network model. For example, the neural network model can include but is not limited to network layers such as convolutional layers and fully connected layers, and the specific structure of the neural network model is not limited here. After the encoding features are input into the neural network model, the speech frame can be classified into a probability value representing the end of semantics and a probability value representing the end of semantics. On this basis, it is possible to determine whether the speech frame represents the end of semantics based on the larger probability value. If the probability value representing the end of semantics is greater than the probability value representing the end of semantics, it can be determined that the speech frame represents the end of semantics. On the contrary, if the probability value representing the end of semantics is less than the probability value representing the end of semantics, it can be determined that the speech frame represents the end of semantics; or, it is possible to determine whether the speech frame represents the end of semantics by taking a probability value greater than a probability threshold. If the probability value representing the end of semantics is greater than the probability threshold, it can be determined that the speech frame represents the end of semantics. On the contrary, if the probability value representing the end of semantics is greater than the probability threshold, it can be determined that the speech frame represents the end of semantics. Of course, the above examples are only several possible examples of determining whether a speech frame represents a semantic end, and do not limit other possible ways of determining whether a speech frame represents a semantic end. Other possible ways will not be given examples one by one here.

[0028] In a specific implementation scenario, for ease of understanding, let's take the example of the streaming speech "(short pause M seconds) We have now arrived in City A (long pause N seconds) Please tell the children about tomorrow's local weather conditions in the tone of a kindergarten teacher." If the speech frame contains the word "la" (specifically, it can be the last phoneme of "la") in the above streaming speech, the meaning of "arrived in City A" has been fully expressed, so the speech frame represents the end of the semantics. Alternatively, take the example of the streaming speech "I think this project (short pause L seconds) is still under discussion." If the speech frame contains the word "ne" (specifically, it can be the last phoneme of "ne") in the above streaming speech, the meaning of "project attitude" has not been fully expressed, so the speech frame represents the end of the semantics. Of course, the above examples are merely specific examples to facilitate understanding of the specific meaning of semantic end and do not limit streaming speech in actual application scenarios. Therefore, we will not give more examples here.

[0029] Step S13: In response to the semantic end represented by the classification result, a recognition text is obtained based on the recognition results of each speech frame from the starting endpoint to the ending endpoint, and the intelligent dialogue model is used to process at least based on the recognition text to generate a reply text.

[0030] In one implementation scenario, when the semantic representation of the classification result ends, a control signal indicating the end of recognition can be transmitted to the speech recognition system. This causes the speech recognition system to pause speech recognition after receiving the control signal until a new round of interaction begins, at which point speech recognition is restarted. This approach, when the semantic representation of the classification result ends, transmits a control signal indicating the end of recognition to the speech recognition system. This can minimize the processing load on the speech recognition system and ensure that speech recognition, speech interaction, and other functions are synchronized.

[0031] In one implementation scenario, when the classification result represents a semantic end, the recognition results of each speech frame from the starting endpoint to the ending endpoint may be obtained first, and the above recognition results may be combined to obtain the recognized text. For example, still taking the streaming voice "(short pause M seconds) We have now arrived in City A (long pause N seconds) Please tell the children about tomorrow's local weather conditions in the tone of a kindergarten teacher" as an example, as mentioned above, when the voice frame contains "la" in the above streaming voice (specifically, it can be the last phoneme containing "la"), the voice frame representation semantics ends, and then the recognition results of each voice frame starting from the voice frame containing "I" (specifically, it can be the first phoneme containing "I") to the voice frame containing "la" (specifically, it can be the last phoneme containing "la") can be obtained. Taking the recognition result containing phonemes as an example, the recognition results of these voice frames can be listed in sequence as {wo men xian zai yi jing di daA shi la}, based on which the recognized text "We have now arrived in City A" can be obtained; or, taking the recognition result containing characters as an example, the recognition results of these voice frames can be listed in sequence as {we have now arrived in City A}, based on which the recognized text "we have now arrived in City A" can be obtained. Of course, the above examples are only a few possible examples in actual application, and other possible methods will not be given one by one here.

[0032] In one implementation scenario, the recognized text can be directly input into the intelligent dialogue model, and the output text of the intelligent dialogue model can be obtained as the reply text. For example, still taking the intelligent dialogue model as a large language model, the large language model can generate a reply text for the recognized text based on its configured system instructions. Still taking the aforementioned recognized text "We have now arrived in City A" as an example, if the configuration information obtained by the system instruction includes "cute" and "energetic", then the large language model can generate a reply text "That's great!" based on the configuration information. Of course, the above example is only one possible example in actual application, and other possible scenarios will not be given one by one here.

[0033] In another implementation scenario, different from the aforementioned implementation, after obtaining the recognition text and before generating the reply text, a semantic analysis can be performed based on the recognition text to obtain an analysis result, and the analysis result indicates whether a query retrieval is required before replying to the recognition text. In response to the analysis result indicating that a query retrieval is required first, a query retrieval can be performed based on the recognition text to obtain a retrieval result, so that the intelligent dialogue model can process the recognition text and the retrieval result to obtain a reply text. In the above method, after obtaining the recognition text and before replying to the recognition text, a semantic analysis is first performed on the recognition text to determine whether a query retrieval is required before replying to the recognition text, and if a query retrieval is required first, a query retrieval is performed based on the recognition text to obtain a retrieval result, so that the recognition text and the retrieval result are combined and processed by the intelligent dialogue model to generate a reply text, which helps to improve the accuracy of the interaction.

[0034] In a specific implementation scenario, the semantic analysis of the recognized text can be specifically performed by a semantic analysis model. It should be noted that the semantic analysis model may include but is not limited to Transformer, etc., and the network structure of the semantic analysis model is not limited here. After the recognized text is semantically analyzed by the semantic analysis model, the analysis result can be obtained. Exemplarily, the semantic analysis model can be a two-category model, and the analysis result can include: the probability value that the reply recognition text needs to be searched through a query first and the probability value that the reply recognition text does not need to be searched through a query first. Based on this probability value, it can be determined whether the reply recognition text needs to be searched through a query first. For details, please refer to the aforementioned description of determining whether the speech frame represents the end of semantics based on the probability value, which will not be repeated here. Let's take the streaming speech "(short pause M seconds) We have now arrived in City A. (Long pause N seconds) Please tell the children about tomorrow's local weather conditions in the voice of a kindergarten teacher." The recognized text "We have now arrived in City A" can be represented through semantic analysis without requiring a query. However, the second half of the recognized text, "Please tell the children about tomorrow's local weather conditions in the voice of a kindergarten teacher," requires a query because it requires semantic analysis and requires a representation of tomorrow's weather conditions. Of course, the above example is just one possible example in actual application, and other possible scenarios will not be listed here.

[0035] In a specific implementation scenario, as a possible example, before performing semantic analysis, the recognized text can be semantically regularized based on the context of the recognized text during the interaction with the streaming voice to obtain regularized text. Still taking the aforementioned recognized text "Please report tomorrow's local weather conditions to the children in the tone of a kindergarten teacher" as an example, combined with the context "We have now arrived in City A", the recognized text can be semantically regularized to obtain the regularized text "Please report tomorrow's local weather conditions to the children in the tone of a kindergarten teacher", so that the semantically regularized recognized text (i.e., regularized text) is semantically clear and unambiguous. On this basis, semantic analysis can be performed based on the regular text to obtain analysis results. For example, if semantic analysis is performed on the regular text "Please report the weather in City A tomorrow to the children in the tone of a kindergarten teacher," it can be determined that a query search is required first. Then, a query search is performed based on the regular text to obtain search results. For example, if a query search is performed based on the regular text "Please report the weather in City A tomorrow to the children in the tone of a kindergarten teacher," a search result containing the weather conditions in City A tomorrow is obtained, such as "The weather is sunny." Based on this, the intelligent dialogue model can process the regular text "Please report the weather in City A tomorrow to the children in the tone of a kindergarten teacher" and the search result "The weather is sunny" to generate a response text, such as "Children, tomorrow is a sunny day!" Of course, the above example is only one possible example in actual application, and other possible scenarios will not be given one by one here. The above method first recognizes the context of the text during the interaction with the streaming voice, semantically regularizes the recognized text to obtain regularized text, then performs semantic analysis based on the regularized text to obtain analysis results, and performs query retrieval based on the regularized text to obtain retrieval results. The regularized text and retrieval results are processed by the intelligent dialogue model to generate a reply text, which can improve the accuracy and pertinence of reply generation.

[0036] In one specific implementation scenario, after obtaining search results and before generating a reply, a pre-text generated by the intelligent dialogue model can be obtained and output to alleviate waiting anxiety. Taking the aforementioned recognized text "Please report tomorrow's local weather conditions to the children in the voice of a kindergarten teacher" as an example, after obtaining the search result "The weather is sunny," the intelligent dialogue model can generate pre-text: "Query in progress, please wait," and so on. Other possible scenarios are not listed here. As a specific example, when the intelligent dialogue model utilizes a large language model, a prompt instruction can be constructed based on the recognized text (or regularized text). The prompt instruction instructs the large language model to generate pre-text matching the style of the reply text with reference to the recognized text (or regularized text). For example, using the aforementioned recognized text "Please report tomorrow's local weather conditions to the children in the voice of a kindergarten teacher" as an example, a pre-text matching the style of the reply text can be generated: "Well, let me check this carefully!" It should be noted that the pre-conversational text can be output in text form, or it can be synthesized using speech synthesis and then output as speech. This approach, after obtaining search results and before generating reply text, can help alleviate waiting anxiety by first obtaining and outputting the pre-conversational text generated by the intelligent dialogue model.

[0037] In a specific implementation scenario, unlike the aforementioned scenario, in response to the analysis result representation, a reply text can be generated based on the recognized text without first undergoing a query retrieval. Still taking the aforementioned recognized text "We have now arrived in City A" as an example, semantic analysis can determine that replying to the recognized text does not require a query retrieval, so it can be directly processed based on the intelligent dialogue model to generate a reply text, such as "That's really great!" It should be noted that, similar to the aforementioned example, the reply text can be output in either text or voice form, which is not limited here. In addition, the above example is only one possible example in actual application and does not limit other possible scenarios. Other scenarios will not be given one by one here. The above approach, in response to the analysis result representation, a reply text is generated based on the recognized text without first undergoing a query retrieval, which can improve the efficiency of voice interaction.

[0038] In another implementation scenario, as described above, the intelligent dialogue model can be a large language model. In this case, when the classification result represents a semantic end, before generating the reply text, intent recognition can also be performed based on the encoded features of each speech frame from the starting endpoint to the ending endpoint to determine whether there is a control intent regarding the voice interaction method. In response to the existence of the control intent, the system instructions of the intelligent dialogue model are updated based on the control intent. Then, based on at least the recognized text, the intelligent dialogue model after the updated system instructions is processed to generate the reply text. The above method, which determines whether there is a control intent regarding the voice interaction method through intent recognition and, if there is a control intent, updates the system instructions of the intelligent dialogue model accordingly to generate the reply text based on the intelligent dialogue model after the updated system instructions, helps to improve the quality of voice interaction and interaction satisfaction.

[0039] In a specific implementation scenario, taking the aforementioned streaming speech "(short pause M seconds) We have now arrived in City A (long pause N seconds) Please tell the children about tomorrow's local weather conditions in the tone of a kindergarten teacher" as an example, when the speech frame contains "la" (specifically, it can be the last phoneme of "la"), the speech frame can be considered to indicate the semantic end, and the process can return to the aforementioned detection of the starting endpoint of the streaming speech. When the speech frame contains "please" (specifically, it can be the first phoneme of "please"), the starting endpoint can be considered to have been detected, and when the speech frame contains "bar" (specifically, it can be the last phoneme of "bar"), the speech frame can be considered to indicate the semantic end. At this time, based on the encoding features from the speech frame containing "please" to the speech frame containing "bar", intent recognition can be performed to determine whether there is a control intent regarding the voice interaction method. It should be noted that the voice interaction method includes but is not limited to: the tone of voice interaction, the voice interaction tone, the language of voice interaction, etc. The specific connotations of the voice interaction method will not be given one by one here. In the above example, there is a control intent regarding the voice interaction method "in the tone of a kindergarten teacher". Of course, the above example is only one possible example in actual application, and other possible situations will not be given one by one here.

[0040] In a specific implementation scenario, as mentioned above, the system instructions of the intelligent dialogue model can contain a lot of configuration information. When it is determined that there is a control intention, the configuration information in the system instructions can be updated based on the control intention. For example, in the above example, the configuration information related to the tone of voice in the system instructions can be updated to "kindergarten teacher". On this basis, the intelligent dialogue model after updating the system instructions can process the recognized text "Please report the local weather conditions tomorrow to the children in the tone of a kindergarten teacher" to generate a reply text, such as "Children, tomorrow morning in City A, the sun will rise from the east with a smile and give us a big, warm hug!". Of course, the above example is only one possible example in the actual application process, and other possible situations will not be given one by one here.

[0041] In one implementation scenario, when the classification result represents the end of semantics, after generating the reply text, it is also possible to generate control parameters for speech synthesis based on the reply text, and then perform speech synthesis on the reply text based on the control parameters to obtain a reply voice. It should be noted that the control parameters may include but are not limited to pronunciation, pauses, etc., and the specific connotations of the control parameters will not be given examples one by one here. In addition, the specific process of speech synthesis can refer to the technical details of speech synthesis, and will not be given examples one by one here. The above method, after generating the reply text, further generates control parameters for speech synthesis based on the reply text, and then performs speech synthesis on the reply text based on the control parameters to obtain a reply voice, which can enhance the fun and immersion of voice interaction.

[0042] In one implementation scenario, if the classification result indicates a semantic end, after generating a response text, the process can return to the endpoint detection step for the streaming speech, starting from the end endpoint, to initiate a new round of voice interaction. Taking the aforementioned streaming speech "(short pause M seconds) We have now arrived in City A. (Long pause N seconds) Please tell the children about tomorrow's local weather conditions in the voice of a kindergarten teacher" as an example, when the speech frame contains the word "la" (specifically, the last phoneme of "la"), the speech frame can be considered to indicate a semantic end. After generating a response text (e.g., "That's great!") for the recognized text "We have now arrived in City A," the process can return to the start endpoint of the streaming speech detection process. When the speech frame contains the word "please" (specifically, the first phoneme of "please"), the start endpoint can be considered detected. This cycle can be repeated, and examples will not be given here one by one. In the above method, when the classification result represents the end of semantics, after generating the reply text, starting from the end endpoint, return to the step of performing endpoint detection on the streaming voice to start a new round of voice interaction, so that voice interaction can be automatically and timely performed during the input or collection process of streaming voice.

[0043] Step S14: In response to the classification result representation semantics not being completed, continue to return to perform the sliding window operation.

[0044] In one implementation scenario, when the semantic representation of the classification result has not ended, the sliding window operation can continue to be performed on the streaming speech to perform a series of operations such as feature extraction, speech recognition, and classification prediction on the continuously generated speech frames until it is detected that the semantic representation of the speech frame has ended. For details, please refer to the above-mentioned relevant description and will not be repeated here.

[0045] In an implementation scenario, as a special example, please refer to Figure 2 , Figure 2 This is a process diagram of an embodiment of the streaming voice interaction method of the present application. Figure 2As shown, the streaming speech is firstly detected by endpoint detection. When the starting endpoint of the streaming speech is detected, a sliding window operation can be performed on the streaming speech starting from the starting endpoint of the streaming speech to continuously obtain speech frames. The speech frames are firstly subjected to feature extraction to obtain the speech features of the speech frames. The speech features of the speech frames are processed separately in two ways. One way is sent to the speech recognition system for identifying the streaming speech to obtain the recognition results, and the other way is sent to the intelligent dialogue model. The encoding features obtained by encoding the speech features by the encoding network in the intelligent dialogue model will be classified and predicted to determine whether the speech frame represents the semantic end. If it represents the semantic end, the control signal representing the stop of recognition will be transmitted to the speech recognition system to control the speech recognition system to stop speech recognition until a new round of voice interaction is started and then the speech recognition is restarted. When determining the semantic end of the speech frame representation, on the one hand, the recognition results of each speech frame from the starting endpoint to the semantic end can be combined to obtain the recognized text, and semantic regularization can be performed in combination with the context to obtain the regularized text. The regularized text is then semantically analyzed to obtain the analysis result, and the analysis result represents whether the reply to the regularized text needs to be queried first. If no query retrieval is required, the query retrieval step can be skipped and the intelligent dialogue model generates the reply. If query retrieval is required first, the query retrieval can be performed in the database, search engine, etc. based on the regularized text to obtain the retrieval result. At this time, the intelligent dialogue model can first generate the padding text to alleviate the waiting anxiety, and then the retrieval results and the regularized text are handed over to the intelligent dialogue model for reply generation. On the other hand, the encoding features of each speech frame from the starting endpoint to the semantic end can be combined to perform intention recognition to determine whether there is a control intention regarding the voice interaction method. If so, the intelligent dialogue model first updates the system instructions of the intelligent dialogue model based on the control intention before generating a reply, and then the intelligent dialogue model after the updated system instructions generates a reply. After obtaining the reply text through the aforementioned process steps, speech synthesis can be performed based on the reply text to obtain the reply speech, completing a round of voice interaction. After that, the process can return to the aforementioned endpoint detection step for the streaming speech to start a new round of voice interaction. In addition, if the speech frame representation semantics has not yet been completed, the sliding window operation can continue to be performed on the streaming speech and the subsequent sliding window operation process can be executed until the speech frame representation semantics are completed.

[0046] The above scheme performs endpoint detection on the streaming speech, and in response to detecting the starting endpoint of the streaming speech, performs a sliding window operation on the streaming speech starting from the starting endpoint to obtain a speech frame, and performs feature extraction based on the speech frame to obtain the speech features of the speech frame, and then inputs the speech features of the speech frame into a speech recognition system for identifying the streaming speech to obtain the recognition result of the speech frame, and performs classification prediction based on the coding features of the speech frame to obtain the classification result of the speech frame, and the coding features are extracted by the coding network in the intelligent dialogue model after being input by the speech features into the intelligent dialogue model, and the classification result includes whether the speech frame represents the ending endpoint of the semantic end, based on this, in response to the classification result representing the semantic end, based on the recognition results of each speech frame from the starting endpoint to the ending endpoint, a recognition text is obtained, and at least based on the recognition text, it is processed by the intelligent dialogue model to generate The reply text is obtained, and in response to the classification result representing that the semantics have not ended, the sliding window operation is continued to be performed. Therefore, on the one hand, in the streaming voice process, the sliding window operation can be continuously performed after the starting endpoint is detected to obtain the voice frame, and each voice frame is processed in turn until the voice frame representing the semantic end is detected, and then the reply text is generated, which can realize the closed loop process of streaming voice interaction. On the other hand, the voice features of the voice frame are sent to the voice recognition system for voice recognition on one side and to the intelligent dialogue model on the other side to be encoded by the encoding network in the intelligent dialogue model first and then classified and predicted based on whether the semantics have ended. With the help of the semantic understanding ability of the intelligent dialogue model, it can be judged as accurately as possible whether the voice frame has ended to determine whether to intervene in the interaction, and to avoid intervening in the interaction too early or too late as much as possible, which helps to improve the timeliness of streaming voice interaction. Therefore, the timeliness of streaming voice interaction can be improved.

[0047] See also Figure 3 , Figure 3The figure is a schematic diagram of a framework of an embodiment of the streaming voice interaction device of the present application. The streaming voice interaction device 30 includes: a preparatory processing module 31, a recognition and classification module 32, a first response module 33 and a second response module 34. The preparatory processing module 31 is used to perform endpoint detection on the streaming voice, and in response to detecting the starting endpoint of the streaming voice, perform a sliding window operation on the streaming voice from the starting endpoint to obtain a voice frame, and perform feature extraction based on the voice frame to obtain the voice feature of the voice frame; the recognition and classification module 32 is used to input the voice feature of the voice frame to the voice recognition system for recognizing the streaming voice, obtain the recognition result of the voice frame, and perform classification based on the coding feature of the voice frame. Classification prediction is performed to obtain a classification result of the speech frame; wherein, the coding feature is extracted by the coding network in the intelligent dialogue model after being inputted into the intelligent dialogue model by the speech feature, and the classification result includes whether the speech frame represents the end endpoint of the semantic end; the first response module 33 is used to respond to the classification result representing the semantic end, obtain the recognition text based on the recognition results of each speech frame from the starting endpoint to the ending endpoint, and generate a reply text based on at least the recognition text processed by the intelligent dialogue model; the second response module 34 is used to respond to the classification result representing that the semantics have not ended, and continue to return to perform the sliding window operation.

[0048] In the above scheme, the streaming voice interaction device 30 performs endpoint detection on the streaming voice, and in response to detecting the starting endpoint of the streaming voice, performs a sliding window operation on the streaming voice starting from the starting endpoint to obtain a voice frame, and extracts features based on the voice frame to obtain the voice features of the voice frame, and then inputs the voice features of the voice frame into a voice recognition system for identifying streaming voice to obtain a recognition result of the voice frame, and performs classification prediction based on the coding features of the voice frame to obtain a classification result of the voice frame, and the coding features are extracted by the coding network in the intelligent dialogue model after being input by the voice features into the intelligent dialogue model, and the classification result includes whether the voice frame represents the ending endpoint of the semantic end, based on this, in response to the classification result representing the semantic end, based on the recognition results of each voice frame from the starting endpoint to the ending endpoint, a recognition text is obtained, and at least based on the recognition text, the intelligent dialogue model performs classification prediction. The processing is performed to generate a reply text. In response to the classification result indicating that the semantics have not ended, the sliding window operation is continued to be performed. Therefore, on the one hand, after the starting endpoint is detected during the streaming voice process, the sliding window operation can be continuously performed to obtain voice frames, and each voice frame is processed in turn until the semantics of the voice frame is detected to end, and then the reply text is generated. This can achieve a closed loop process for streaming voice interaction. On the other hand, the voice features of the voice frame are sent to the voice recognition system for voice recognition and to the intelligent dialogue model for encoding by the encoding network in the intelligent dialogue model. Based on this, a classification prediction is made as to whether the semantics have ended. With the help of the semantic understanding ability of the intelligent dialogue model, it is possible to judge as accurately as possible whether the voice frame has ended to determine whether to intervene in the interaction, thereby avoiding intervening in the interaction too early or too late as much as possible, which helps to improve the timeliness of streaming voice interaction. Therefore, the timeliness of streaming voice interaction can be improved.

[0049] In some disclosed embodiments, the streaming voice interaction device 30 also includes a semantic analysis module for performing semantic analysis based on the recognized text to obtain an analysis result; wherein the analysis result indicates whether a query retrieval is required before replying to the recognized text; the streaming voice interaction device 30 also includes a query retrieval module for performing a query retrieval based on the recognized text in response to the analysis result indicating that a query retrieval is required to obtain a retrieval result; the first response module 33 is also specifically used to generate a reply text by processing the recognized text and the retrieval result by an intelligent dialogue model.

[0050] In some disclosed embodiments, the streaming voice interaction device 30 further includes a filler control module for acquiring and outputting filler text generated by the intelligent dialogue model.

[0051] In some disclosed embodiments, the streaming voice interaction device 30 also includes a semantic regularization module, which is used to semantically regularize the recognized text based on the context of the recognized text during the interaction with the streaming voice to obtain regularized text; the semantic analysis module is specifically used to perform semantic analysis based on the regularized text to obtain analysis results; the query retrieval module is specifically used to perform query retrieval based on the regularized text to obtain retrieval results; the first response module 33 is also specifically used to generate a reply text based on the regularized text and the retrieval results processed by the intelligent dialogue model.

[0052] In some disclosed embodiments, the first response module 33 is further specifically configured to generate a reply text based on the recognition text processed by the intelligent dialogue model in response to the analysis result representation without first undergoing query retrieval.

[0053] In some disclosed embodiments, the intelligent dialogue model is a large language model, and the streaming voice interaction device 30 also includes an intention recognition module for performing intention recognition based on the encoding features of each voice frame from the starting endpoint to the ending endpoint when the classification result represents the semantic end, and determining whether there is a control intention regarding the voice interaction method; the streaming voice interaction device 30 also includes an instruction update module for updating the system instructions of the intelligent dialogue model based on the control intention in response to the existence of the control intention; the first response module 33 is also specifically used to generate a reply text based on at least the recognition text processed by the intelligent dialogue model after the system instructions are updated.

[0054] In some disclosed embodiments, the streaming voice interaction device 30 also includes a parameter generation module for generating control parameters for voice synthesis based on the reply text when the classification result represents the end of semantics; the streaming voice interaction device 30 also includes a voice synthesis module for performing voice synthesis on the reply text based on the control parameters to obtain a reply voice.

[0055] In some disclosed embodiments, the streaming voice interaction device 30 also includes a signal transmission module for returning to the step of performing endpoint detection on the streaming voice starting from the end endpoint when the classification result represents the end of semantics, so as to start a new round of voice interaction.

[0056] See also Figure 4 , Figure 4This is a schematic diagram of the framework of an embodiment of an electronic device of the present application. The electronic device 40 includes at least a memory 41 and a processor 42 coupled to each other. The memory 41 stores at least program instructions, and the processor 42 is used to execute the program instructions to implement the steps of any of the above-mentioned streaming voice interaction method embodiments. For details, please refer to the aforementioned disclosed embodiments and will not be repeated here. As a possible example, the electronic device 40 may include but is not limited to a server, a camera, a smartphone, a tablet computer, etc., and the specific type of the electronic device 40 is not limited here.

[0057] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps in any of the above-mentioned streaming voice interaction method embodiments. The processor 42 can also be called a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 42 can be implemented by an integrated circuit chip.

[0058] In the above scheme, the electronic device 40 performs endpoint detection on the streaming voice, and in response to detecting the starting endpoint of the streaming voice, performs a sliding window operation on the streaming voice starting from the starting endpoint to obtain a voice frame, and performs feature extraction based on the voice frame to obtain the voice features of the voice frame, and then inputs the voice features of the voice frame into a voice recognition system for identifying the streaming voice to obtain the recognition result of the voice frame, and performs classification prediction based on the coding features of the voice frame to obtain the classification result of the voice frame, and the coding features are extracted by the coding network in the intelligent dialogue model after being input by the voice features into the intelligent dialogue model, and the classification result includes whether the voice frame represents the ending endpoint of the semantic end, based on this, in response to the classification result representing the semantic end, based on the recognition results of each voice frame from the starting endpoint to the ending endpoint, a recognized text is obtained, and at least based on the recognized text, the intelligent dialogue model performs processing The system generates a reply text, and in response to the classification result indicating that the semantics have not ended, it continues to return to perform the sliding window operation. Therefore, on the one hand, during the streaming voice process, the sliding window operation can be continuously performed after the starting endpoint is detected to obtain voice frames, and each voice frame is sequentially processed until the semantics of the voice frame representation is detected to end, and then the reply text is generated, which can achieve a closed loop process for streaming voice interaction. On the other hand, the voice features of the voice frame are sent to the voice recognition system for voice recognition, and the other way is sent to the intelligent dialogue model to be encoded by the encoding network in the intelligent dialogue model first and then classified and predicted based on the semantics. With the help of the semantic understanding ability of the intelligent dialogue model, it can judge as accurately as possible whether the voice frame is semantically ended to determine whether to intervene in the interaction, and avoid intervening in the interaction too early or too late as much as possible, which helps to improve the timeliness of streaming voice interaction. Therefore, it can improve the timeliness of streaming voice interaction.

[0059] See also Figure 5 , Figure 5 1 is a schematic diagram of a framework of an embodiment of a computer-readable storage medium 50 of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be executed by a processor, and the program instructions 51 are used to implement the steps of any of the above-mentioned streaming voice interaction method embodiments.

[0060] In the above scheme, the computer-readable storage medium 50 performs endpoint detection on the streaming speech, and in response to detecting the starting endpoint of the streaming speech, performs a sliding window operation on the streaming speech starting from the starting endpoint to obtain a speech frame, and extracts features based on the speech frame to obtain the speech features of the speech frame, and then inputs the speech features of the speech frame into a speech recognition system for identifying the streaming speech to obtain the recognition result of the speech frame, and performs classification prediction based on the coding features of the speech frame to obtain the classification result of the speech frame, and the coding features are extracted by the coding network in the intelligent dialogue model after being input by the speech features into the intelligent dialogue model, and the classification result includes whether the speech frame represents the ending endpoint of the semantic end, based on this, in response to the classification result representing the semantic end, based on the recognition results of each speech frame from the starting endpoint to the ending endpoint, a recognized text is obtained, and at least based on the recognized text, the intelligent dialogue model performs classification prediction. The processing is performed to generate a reply text. In response to the classification result indicating that the semantics have not ended, the sliding window operation is continued to be performed. Therefore, on the one hand, after the starting endpoint is detected during the streaming voice process, the sliding window operation can be continuously performed to obtain voice frames, and each voice frame is processed in turn until the semantics of the voice frame is detected to end, and then the reply text is generated. This can achieve a closed loop process for streaming voice interaction. On the other hand, the voice features of the voice frame are sent to the voice recognition system for voice recognition and to the intelligent dialogue model for encoding by the encoding network in the intelligent dialogue model. Based on this, a classification prediction is made as to whether the semantics have ended. With the help of the semantic understanding ability of the intelligent dialogue model, it is possible to judge as accurately as possible whether the voice frame has ended to determine whether to intervene in the interaction, thereby avoiding intervening in the interaction too early or too late as much as possible, which helps to improve the timeliness of streaming voice interaction. Therefore, the timeliness of streaming voice interaction can be improved.

[0061] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0062] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.

[0063] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation methods described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0064] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0065] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0066] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the various implementation methods of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0067] If the technical solution of this application involves personal information, the product that applies the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing personal information. If the technical solution of this application involves sensitive personal information, the product that applies the technical solution of this application has obtained the individual's separate consent before processing sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information; among which, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

Claims

1. A streaming voice interaction method, characterized in that: include: Performing endpoint detection on the streaming voice, and in response to detecting the starting endpoint of the streaming voice, performing a sliding window operation on the streaming voice from the starting endpoint to obtain a voice frame, and performing feature extraction based on the voice frame to obtain a voice feature of the voice frame; Inputting the speech features of the speech frame into a speech recognition system for identifying streaming speech to obtain a recognition result of the speech frame, and performing classification prediction based on the coding features of the speech frame to obtain a classification result of the speech frame; wherein the coding features are extracted by the coding network in the intelligent dialogue model after the speech features are input into the intelligent dialogue model, and the classification result includes whether the speech frame represents an end endpoint of a semantic end; In response to the classification result representing the semantic end, based on the recognition results of each of the speech frames from the starting endpoint to the ending endpoint, a recognition text is obtained, and based on at least the recognition text, the intelligent dialogue model is processed to generate a reply text, and starting from the ending endpoint, the step of performing endpoint detection on the streaming speech is returned to start a new round of voice interaction; In response to the classification result representing that the semantics are not finished, continue to return to execute the sliding window operation.

2. The method according to claim 1, characterized in that After obtaining the recognition text based on the recognition results of each of the speech frames from the starting endpoint to the ending endpoint, and before generating a reply text by processing the intelligent dialogue model at least based on the recognition text, the method further includes: Perform semantic analysis based on the recognized text to obtain an analysis result; wherein the analysis result indicates whether a query retrieval is required before replying to the recognized text; In response to the analysis result indicating that a query search is required first, a query search is performed based on the recognized text to obtain a search result; The step of processing the intelligent dialogue model at least based on the recognition text to generate a reply text includes: The intelligent dialogue model processes the recognized text and the search results to generate the reply text.

3. The method according to claim 2, characterized in that After performing a query search based on the recognition text to obtain a search result, and before the intelligent dialogue model processes the recognition text and the search result to generate the reply text, the method further includes: Acquire and output the pad text generated by the intelligent dialogue model.

4. The method according to claim 2, characterized in that: Before performing semantic analysis based on the recognized text to obtain an analysis result, the method further includes: Based on the context of the recognized text in the process of interaction with the streaming voice, semantically regularize the recognized text to obtain a regularized text; The performing semantic analysis based on the recognized text to obtain analysis results includes: Performing semantic analysis based on the regularized text to obtain the analysis result; The query retrieval based on the recognized text to obtain the retrieval results includes: Performing a query search based on the regular text to obtain the search result; The processing based on the recognition text and the search result by the intelligent dialogue model to generate the reply text includes: The intelligent dialogue model is used to process at least the regular text and the search results to generate the reply text.

5. The method according to claim 2, characterized in that: The method further comprises: In response to the analysis result representation, the reply text is generated by processing the recognized text by the intelligent dialogue model without the need for query retrieval.

6. The method according to claim 1, characterized in that The intelligent dialogue model is a large language model. When the classification result represents a semantic end, before the intelligent dialogue model processes the recognized text at least based on the recognition text to generate a reply text, the method further includes: Performing intention recognition based on the coding features of each of the voice frames from the starting endpoint to the ending endpoint to determine whether there is a control intention regarding the voice interaction mode; In response to the existence of the control intention, updating the system instructions of the intelligent dialogue model based on the control intention; The step of processing the intelligent dialogue model at least based on the recognition text to generate a reply text includes: The reply text is generated based on at least the recognition text being processed by the intelligent dialogue model after updating the system instructions.

7. The method according to claim 1, characterized in that In the case where the classification result represents a semantic end, after the intelligent dialogue model processes the recognized text at least based on the generated reply text, the method further includes: Based on the reply text, generating control parameters for speech synthesis; The reply text is subjected to speech synthesis based on the control parameters to obtain a reply speech.

8. The method according to claim 1, characterized in that In the case where the classification result represents a semantic end, before obtaining the recognized text based on the recognition results of each of the speech frames from the starting endpoint to the ending endpoint, the method further includes: A control signal indicating that recognition is to be stopped is transmitted to the speech recognition system.

9. A streaming voice interaction device, characterized in that: include: A preparatory processing module, configured to perform endpoint detection on the streaming voice, and in response to detecting the starting endpoint of the streaming voice, perform a sliding window operation on the streaming voice from the starting endpoint to obtain a voice frame, and perform feature extraction based on the voice frame to obtain a voice feature of the voice frame; A recognition and classification module, used for inputting the speech features of the speech frame into a speech recognition system for identifying streaming speech, obtaining a recognition result of the speech frame, and performing classification prediction based on the coding features of the speech frame to obtain a classification result of the speech frame; wherein the coding features are extracted by the coding network in the intelligent dialogue model after the speech features are input into the intelligent dialogue model, and the classification result includes whether the speech frame represents an end endpoint of a semantic end; A first response module is used for obtaining a recognition text based on the recognition results of each of the speech frames from the starting endpoint to the ending endpoint in response to the classification result representing the semantic end, and generating a reply text by processing the recognition text by the intelligent dialogue model at least based on the recognition text, and returning to the step of performing endpoint detection on the streaming speech from the ending endpoint to start a new round of voice interaction; The second response module is used for returning to execute the sliding window operation in response to the semantic representation of the classification result not being completed.

10. An electronic device, characterized in that: The method comprises at least a memory and a processor coupled to each other, wherein the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the streaming voice interaction method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the streaming voice interaction method described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Audio stream processing method and device, storage medium and electronic device

    CN117672188A