Streaming voice interaction method and related devices, equipment and storage media
By hierarchically classifying streaming voices, semantic contents are solved, and the error triggering problem in streaming voice interaction is achieved, achieving higher interaction accuracy and experience quality.
Patent Information
- Application Number
- CN202510202865.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-24
AI Technical Summary
The existing streaming voice interaction technology has problems such as false triggering, which leads to problems such as not responding but not responding, which affects the interactive experience.
By hierarchically classifying the collected streaming voice, first checking for other people's voice, noise or muting, further checking whether other people's voice is real or background voice, then judging the semantic integrity and relevance to the context content, and then performing target interaction operations.
It effectively reduces the false triggering of streaming voice, improves the accuracy of streaming voice interaction, and improves the interactive experience.
Smart Images

Figure CN119694304B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice interaction technologies, and particularly to a streaming voice interaction method and related devices, equipment, and storage media. Background Art
[0002] With the rapid development of related technologies such as natural language processing, streaming voice interaction has been widely applied in many scenarios such as intelligent assistants and customer services. Through the streaming voice interaction technology, intelligent functions such as the ability to interrupt a conversation at any time during the conversation can be achieved, making human-machine conversations similar to human-human conversations.
[0003] However, existing streaming voice interaction technologies still have the technical defect of mis-triggering, resulting in problems such as responding when not supposed to respond and not responding when supposed to respond, thereby affecting the interaction experience. In view of this, how to minimize the mis-triggering of streaming voice as much as possible and improve the accuracy of streaming voice interaction has become an urgent problem to be solved. Summary of the Invention
[0004] The main technical problem to be solved by this application is to provide a streaming voice interaction method and related devices, equipment, and storage media, which can minimize the mis-triggering of streaming voice as much as possible and improve the accuracy of streaming voice interaction.
[0005] To solve the above technical problem, in the first aspect of this application, a streaming voice interaction method is provided, including: performing a first classification on the currently collected first streaming voice to obtain a first predicted category of the first streaming voice; where the first predicted category is human voice, noise, or silence; in response to the first predicted category being human voice, performing a second classification on at least the first streaming voice to obtain a second predicted category of the first streaming voice; where the second predicted category is real human voice or background human voice; in response to the second predicted category being real human voice, performing a third classification on at least the first streaming voice to obtain a third predicted category of the first streaming voice; where the third predicted category relates to whether the semantics is complete and whether it is relevant to the context content; based on the third predicted category, performing a target interaction operation on the currently output machine conversation content; where the target interaction operation is any one of several preset interaction operations, and the several preset interaction operations include maintaining the current topic, switching to a new topic, and continuing the original topic from the interrupted conversation.
[0006] To solve the above technical problems, a second aspect of the present application provides a streaming voice interaction device, including: a first classification module, a second classification module, a third classification module, and an interaction operation module. The first classification module is configured to perform a first classification on the currently collected first streaming voice to obtain a first predicted category of the first streaming voice; wherein, the first predicted category is human voice, noise, or silence. The second classification module is configured to, in response to the first predicted category being human voice, perform a second classification on at least the first streaming voice to obtain a second predicted category of the first streaming voice; wherein, the second predicted category is real human voice or background human voice. The third classification module is configured to, in response to the second predicted category being real human voice, perform a third classification on at least the first streaming voice to obtain a third predicted category of the first streaming voice; wherein, the third predicted category relates to whether the semantics is complete and whether it is relevant to the context content. The interaction operation module is configured to perform a target interaction operation on the currently output machine dialogue content based on the third predicted category; wherein, the target interaction operation is any one of a number of preset interaction operations, and the number of preset interaction operations includes maintaining the current topic, switching to a new topic, and continuing the original topic from the interrupted dialogue.
[0007] To solve the above technical problems, a third aspect of the present application provides an electronic device, at least including a memory and a processor coupled to each other. The memory stores at least program instructions, and the processor is configured to execute the program instructions to implement the streaming voice interaction method in the first aspect above.
[0008] To solve the above technical problems, a fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor, and the program instructions are used to implement the streaming voice interaction method in the first aspect above.
[0009] The above solution performs a first classification on the currently collected first streaming voice to obtain a first predicted category of the first streaming voice, and the first predicted category is human voice, noise, or silence. In response to the first predicted category being human voice, at least a second classification is performed based on the first streaming voice to obtain a second predicted category of the first streaming voice, and the second predicted category is real human voice or background human voice. In response to the second predicted category being real human voice, at least a third classification is performed based on the first streaming voice to obtain a third predicted category of the first streaming voice, and the third predicted category relates to whether the semantics is complete and whether it is relevant to the context content. Then, based on the third predicted category, a target interaction operation is performed on the machine dialogue content currently being output. The target interaction operation is any one of several preset interaction operations, and the several preset interaction operations include maintaining the current topic, switching to a new topic, and continuing the original topic from the interrupted dialogue. Therefore, a hierarchical classification of the first classification, second classification, and third classification is sequentially performed on the currently collected first streaming voice. During the first classification process, human voice, noise, or silence is distinguished. During the second classification process, it is further distinguished whether the human voice belongs to real human voice or background human voice. And during the third classification process, it is further distinguished whether the semantics of the real human voice is complete and whether the content is relevant, so as to gradually reduce the mis-triggering of the streaming voice as much as possible through hierarchical progression. Therefore, it is possible to reduce the mis-triggering of the streaming voice as much as possible and improve the accuracy of the streaming voice interaction. Description of the Drawings
[0010] Figure 1 is a schematic flowchart of an embodiment of the streaming voice interaction method of the present application;
[0011] Figure 2 is a schematic process diagram of an embodiment of the streaming voice interaction method of the present application;
[0012] Figure 3 is a schematic framework diagram of an embodiment of the streaming voice interaction device of the present application;
[0013] Figure 4 is a schematic framework diagram of an embodiment of the electronic device of the present application;
[0014] Figure 5 is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. Detailed Embodiments
[0015] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.
[0016] In the following description, specific details such as specific system structures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.
[0017] The terms "system" and "network" are often used interchangeably in this document. The term " / or" in this document is merely a description of the relationship between associated objects, indicating that three relationships can exist. For example, A / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, in this document, the segment " / " generally indicates that the associated objects before and after are in an "or" relationship. Furthermore, "plurality" in this document means two or more than two.
[0018] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the streaming voice interaction method of this application. Specifically, it can include the following steps:
[0019] Step S11: Perform a first classification on the currently collected first streaming voice to obtain the first predicted category of the first streaming voice.
[0020] In the embodiments of the present disclosure, the streaming voice is voice data in the form of a stream. Different from non-streaming voice, which usually uses the start of voice interaction (e.g., clicking and holding the voice button) as the starting point of the voice and the end of voice interaction (e.g., releasing the voice button) as the end point of the entire segment of voice data for voice interaction, the streaming voice interaction can imitate the voice interaction mode of human-to-human interaction and perform human-computer interaction timely and real-time during the voice collection process. Additionally, the first predicted category can specifically be human voice, noise, or silence. Exemplarily, at t = 0 s, the first classification of the first streaming voice can be started. When t = 1 s, since no one has spoken yet, the first predicted category is silence at this time (or when there is background noise, the first predicted category can also be noise). When t = 2 s, someone starts speaking, and the first predicted category is human voice at this time. Of course, the above examples are merely possible examples in the actual application process and do not limit other possible situations that may exist in the actual application process, and no further examples will be given here.
[0021] In an implementation scenario, as a possible example, the first prediction category can be predicted by a first classification network. The first classification network can be trained based on the first sample speech, and the first sample speech can be labeled with an annotation category, such as human voice, noise, or silence. Specifically, the first sample speech can be classified by the first classification network to obtain the prediction category of the first sample speech, and then the network parameters of the first classification network can be adjusted based on the difference between the prediction category of the first sample speech and the annotation category. It should be noted that when the first classification network processes the first sample speech, it can specifically obtain the prediction probability values of the first sample speech belonging to human voice, noise, and silence respectively. Based on this, loss functions such as cross-entropy can be used to process the prediction probability values of the first sample speech belonging to human voice, noise, and silence respectively and the annotation category of the first sample speech to obtain the training loss of the first classification network, and the network parameters of the first classification network can be adjusted based on the training loss of the first classification network. In addition, the first classification network can specifically include, but is not limited to, convolutional neural network, recurrent neural network, long short-term memory network, Transformer, etc. The network structure of the first classification network is not limited here.
[0022] In another implementation scenario, different from the foregoing implementation manner, as another possible example, during the process of training the first classification network based on the first sample speech, the first sample speech can be first classified by the first classification network to obtain the prediction category of the first sample speech, and the sample original speech can be classified and predicted by a voice endpoint detection network to obtain the voice classification of each sample speech segment in the sample original speech. The sample original speech is the complete speech where the first sample speech is located, the first sample speech is one of the sample speech segments, and the voice classification is human voice, noise, or silence. Then, the network parameters of the first classification network can be adjusted based on the difference between the prediction category of the first sample speech and the annotation category, and the difference between the prediction category of the first sample speech and the voice classification. In the above manner, by combining the difference between the prediction category of the first sample speech and the annotation category, and the difference between the prediction category of the first sample speech and the voice classification during the training process of the first classification network, the first classification network can be forced to learn the true annotation of the sample speech on the one hand, so that the output result of the first classification network is as close as possible to the true classification of the input speech, and on the other hand, the first classification network can be forced to learn the voice endpoint detection model to achieve knowledge distillation based on the voice endpoint detection model, and then the knowledge / performance of the voice endpoint detection model can be transferred to the first classification network as much as possible. Furthermore, combining the above two aspects helps to improve the accuracy of the first classification network in performing the first classification.
[0023] In a specific implementation scenario, the voice endpoint detection network is used to implement voice activity detection (VAD). For the specific implementation, refer to the technical details of VAD, which will not be elaborated here.
[0024] In a specific implementation scenario, the first classification network processes the first sample voice, and specifically obtains the predicted probability values of the first sample voice belonging to human voice, noise, and silence respectively. Based on this, loss functions such as cross-entropy can be used to process the predicted probability values of the first sample voice belonging to human voice, noise, and silence respectively, as well as the labeled category of the first sample voice, so as to obtain the first loss of the first classification network.
[0025] In a specific implementation scenario, for the sake of description, the predicted probability values of the first sample voice belonging to human voice, noise, and silence respectively obtained by the first classification network processing the first sample voice can be called the first probability distribution, and the predicted probability values of the sample voice segments belonging to human voice, noise, and silence respectively obtained by the voice endpoint detection network processing the original sample voice can be called the second probability distribution. Then, loss functions such as KL divergence can be used to measure the first probability distribution and the second probability distribution of the first sample voice, so as to measure the difference between the predicted category of the first sample voice and the voice classification based on this, and obtain the second loss of the first classification network.
[0026] In a specific implementation scenario, based on the first loss and the second loss, the training loss of the first classification network can be obtained, and then based on the classification loss of the first classification network, the network parameters of the first classification network can be adjusted. Exemplarily, the training loss of the first classification network can be obtained through fusion methods such as weighting, averaging, and summing based on the first loss and the second loss.
[0027] Step S12: In response to the first predicted category being human voice, perform a second classification based on at least the first streaming voice to obtain the second predicted category of the first streaming voice.
[0028] In the embodiments of the present disclosure, the second predicted category is real human voice or background human voice. It should be noted that real human voice specifically refers to human voice with a real human-computer interaction intention, and background human voice can include player human voice (such as TV human voice, etc.), human voice of other speakers (such as when the speaker in a human-computer interaction is talking to others, the human voice emitted by others), etc. The possible situations of background human voice will not be exemplified one by one here.
[0029] In an implementation scenario, as a possible implementation method, the second classification can be directly performed on the first streaming voice to obtain the second prediction category of the first streaming voice. Exemplarily, the feature extraction can be performed based on the first streaming voice to obtain the voice features of the first streaming voice, and then the classification prediction can be performed based on the voice features of the first streaming voice to obtain the second prediction category of the first streaming voice.
[0030] In another implementation scenario, different from the foregoing implementation method, as another possible implementation method, the voice collected at the start of the voice interaction can also be obtained as the reference voice before performing the second classification. Exemplarily, when a click operation on the voice button is detected, the voice data of a preset duration (such as 5 seconds, 10 seconds, etc.) can be collected as the reference voice. On this basis, the second classification can be performed based on the first streaming voice and the reference voice to obtain the second prediction category of the first streaming voice. Exemplarily, the feature extraction can be performed based on the first streaming voice to obtain the voice features of the first streaming voice, and the feature extraction can be performed based on the reference voice to obtain the voice features of the reference voice. Then, the voice features of the first streaming voice and the voice features of the reference voice can be fused to obtain the fusion features, and the classification prediction can be performed based on the fusion features to obtain the second prediction category of the first streaming voice. By the above method, the voice collected during the voice interaction is obtained as the reference voice first, and then the second classification is performed based on the first streaming voice and the reference voice to obtain the second prediction category of the first streaming voice, which can assist the second classification of the first streaming voice through the reference voice and help improve the accuracy of the second classification.
[0031] In an implementation scenario, similar to the foregoing first classification, the second prediction category can be predicted by a second classification network, which can be trained based on second sample voices. The labeled category of the second sample voices is real human voice or background human voice. For the specific meanings of real human voice and background human voice, reference can be made to the foregoing related descriptions and will not be elaborated here. For the specific training process, reference can be made to the related descriptions of training the first classification network and will not be elaborated here either. In addition, the second classification network can include, but is not limited to, convolutional neural network, recurrent neural network, Transformer, etc. The network structure of the second classification network is not limited here. For the second sample voices with the labeled category of real human voice, they can be obtained by means such as real recording in a recording studio, voice collection of real human voices in a real interaction scenario, etc., which are not limited here. For the second sample voices with the labeled category of background human voice, as a possible example, voice collection can be performed on the sample source at a target distance from the location where the voice interaction occurs, and the second sample voices with the labeled category of background human voice are obtained, and the sample source is a speaker or a player. For example, when having a voice interaction, the speaker is generally about 50 centimeters away from the machine device. Then a sample source can be set at a target distance of 1 meter, 2 meters, etc. away from the machine device by the speaker, so that the machine device can perform voice collection on the sample source as the second sample voices with the labeled category of background human voice. Or, as another possible example, the second sample voices with the labeled category of background human voice can also be obtained by at least attenuating the volume of the second sample voices with the labeled category of real human voice. Of course, the above examples are only several possible ways to obtain the second sample voices with the labeled category of background human voice, and will not be listed one by one here. For example, the second sample voices can also be obtained by voice collection of background human voices in a real interaction scenario. By the above methods, voice collection is performed on the sample source and the volume of real speech is attenuated to obtain the second sample voices with the labeled category of background human voice, which can achieve sample amplification as much as possible for training the second classification network.
[0032] Step S13: In response to the second prediction category being real human voice, perform a third classification based on at least the first streaming voice to obtain the third prediction category of the first streaming voice.
[0033] In the embodiments of the present disclosure, the third prediction category relates to whether the semantics is complete and whether it is relevant to the context content. It should be noted that whether the first streaming speech is semantically complete is used to measure whether the first streaming speech contains relatively complete semantic information. For example, expressions such as "Today...", "Today the...", "Today's weather..." have incomplete semantics. Among these three, the first two have relatively short semantics (or close to none), and the last one has relatively long semantics (i.e., close to complete). Expressions such as "Today's weather is fine", "What's the weather like today", "Is today's weather suitable for outdoor sports" have complete semantics. In addition, whether the first streaming speech is relevant to the context content is used to measure whether the first streaming speech has a semantic correlation with the context (in the streaming speech, it can exactly be the previous text / foregoing text) at the semantic level. Taking the previous text "What's the weather like today" as an example, when the first streaming speech is "Is it suitable for outdoor sports today", the first streaming speech is relevant to the context content, while when the first streaming speech is "Is the C Mountain in City B suitable for the elderly to climb", the first streaming speech is irrelevant to the context content. Of course, the above examples are only several possible examples of the third prediction category in the actual application process, and other possible situations will not be exemplified one by one here.
[0034] In one implementation scenario, as a possible implementation manner, the third classification can be directly performed based on the first streaming speech to obtain the third prediction category of the first streaming speech. Exemplarily, feature extraction can be performed based on the first streaming speech to obtain the speech features of the first streaming speech, and then classification prediction can be performed based on the speech features of the first streaming speech to obtain the third prediction category of the first streaming speech.
[0035] In another implementation scenario, different from the foregoing implementation, as another possible implementation, before performing the third classification, it is also possible to first obtain the recognized text of the second streaming voice as the historical text, where the second streaming voice is the real human voice in the voice interaction before the first streaming voice. Exemplarily, still taking the first streaming voice as "Is it suitable for outdoor sports today?" as an example, the voice data "What's the weather like today?" in the voice interaction before it can be used as the second streaming voice. Of course, the above example is only a possible example in the actual application process, and the specific content of the first streaming voice and the second streaming voice is not limited here. On this basis, the third classification can be performed based on the first streaming voice and the historical text to obtain the third prediction category of the first streaming voice. Exemplarily, feature extraction can be performed based on the first streaming voice to obtain the voice features of the first streaming voice, and feature extraction can be performed based on the historical text to obtain the text features of the historical text. Then, the voice features of the first streaming voice and the text features of the historical text can be fused to obtain the fused features, so as to perform classification prediction based on the fused features to obtain the third prediction category of the first streaming voice. The above method first obtains the recognized text of the second streaming voice as the historical text, and then performs the third classification based on the first streaming voice and the historical text to obtain the third prediction category of the first streaming voice, which can assist in performing the third classification of the first streaming voice through the historical text and help improve the accuracy of the third classification.
[0036] In one implementation scenario, similar to the relevant implementations of the foregoing first classification and second classification, the third classification can be performed by a third classification network. It should be noted that the third classification network can include, but is not limited to, convolutional neural network, recurrent neural network, long short-term memory network, Transformer, etc., and the network structure of the third classification network is not limited here. In addition, the third classification network can be trained based on the third sample voice, and the labeled categories of the third sample voice can involve whether the semantics is complete and whether it is relevant to the context content. The third sample voice can also be labeled with the text corresponding to the voice data in the voice interaction before it (for the convenience of description, it can be called the sample historical text) and the text corresponding to the third sample voice itself (for the convenience of description, it can be called the sample real text). On this basis, the third classification network can be used to process the third sample voice and the sample historical text to obtain the prediction category of the third sample voice and the sample recognized text. Thus, the first loss can be obtained based on the difference between the prediction category of the third sample voice and the labeled category, and the second loss can be obtained based on the difference between the sample real text and the sample recognized text of the third sample voice. Furthermore, the first loss and the second loss can be fused (such as weighted, averaged, summed, etc.) to obtain the training loss of the third classification network, and the network parameters of the third classification network can be adjusted based on the training loss of the third classification network.
[0037] Step S14: Based on the third predicted category, perform a target interaction operation on the machine dialogue content currently being output.
[0038] In the embodiments of the present disclosure, the target interaction operation is any one of several preset interaction operations, and the several preset interaction operations include maintaining the current topic, switching to a new topic, and continuing the original topic from the interrupted dialogue. Of course, the above examples are only several possible examples of the several preset interaction operations, and other possible interaction operations will not be exemplified one by one here.
[0039] In an implementation scenario, in response to the third predicted category being semantically complete, the machine dialogue content currently being output can be interrupted and switched to a new topic. It should be noted that the new topic is the topic to be interacted with by the first streaming voice. Taking the first streaming voice "What's the weather like today?" as an example, if the machine dialogue content currently being output is the machine dialogue content for the previous voice data "Please help me plan a trip to City B" (i.e., the dialogue content generated and output in response to this voice data), since the third predicted category of the first streaming voice represents semantic completeness, the machine dialogue content currently being output can be interrupted. For example, if the current dialogue content has been generated and output to "... Mountain C located in the north of City B is one of the famous scenic spots in the country", it can be interrupted here and switched to the new topic of "What's the weather like today?", and at this time, new machine dialogue content can be generated and output, such as "It's sunny all day today, the temperature is XX degrees, the wind force is XX level, and the humidity is XX%...". In addition, when the third predicted category represents semantic completeness, as a possible example, regardless of whether the third predicted category represents relevance to the context content, the machine dialogue content currently being output can be interrupted and switched to a new topic.
[0040] In an implementation scenario, in response to the third predicted category being semantically incomplete and unable to determine whether it is relevant to the context content, the machine dialogue content currently being output can continue the original topic from the interrupted point. It should be noted that the original topic is the topic to which the machine dialogue content currently being output belongs. In addition, in the case where it is unable to determine whether it is relevant to the context content, it can be considered that the degree of semantic incompleteness of the first streaming voice is relatively high, so that it is unable to determine whether it is relevant to the context content, such as a short semantic (e.g., "Today the weather..."), approaching none (e.g., "Today..."), or even having no semantics (e.g., the first streaming voice is a modal particle such as "um", "oh", etc.). Taking the first streaming voice "Today the weather..." as an example, if the machine dialogue content currently being output is the machine dialogue content for the previous voice data "Please help me plan the itinerary to City B" (i.e., the dialogue content generated and output in response to this voice data), then since the third predicted category of the first streaming voice is semantically incomplete and unable to determine whether it is relevant to the context content, the machine dialogue content currently being output can continue the original topic from the interrupted dialogue point. For example, if the current dialogue content is generated and output to "... Mountain C located in the north of City B is one of the famous scenic spots in the country", then the original topic can be continued at this interruption point (i.e., "one of the famous scenic spots in the country"), such as continuing to generate and output machine dialogue content like "It ranks among the four famous mountains together with Mountain D, Mountain E, and Mountain F. In addition,...", "I was just interrupted. Let's continue. Just now we were talking about Mountain C located in the north of City B being one of the famous scenic spots in the country, and it ranks among the four famous mountains together with Mountain D, Mountain E, and Mountain F. In addition,..." and so on.
[0041] In an implementation scenario, in response to the third predicted category being semantically incomplete and related to the context content, the machine dialogue content currently being output can be paused and preparations can be made to switch to a new topic. It should be noted that the new topic is the topic to be interacted with in the first streaming voice. In addition, in the case where it can be determined to be related to the context content, it can be considered that although the first streaming voice is semantically incomplete, it has a relatively long semantics, such as "The weather today..." etc., and it can already express a certain degree of semantic information. If the machine dialogue content currently being output is the machine dialogue content for the previous voice data "Please help me plan the itinerary to City B" (that is, the dialogue content generated and output in response to this voice data), and the current dialogue content is generated and output to "... Mountain C located in the north of City B is one of the famous scenic spots in the country", taking the first streaming voice "Is Mountain C in City B suitable for the elderly..." as an example, since the third predicted category of the first streaming voice is semantically incomplete (for example, in this example, there are the following possible semantic situations: Is Mountain C in City B suitable for the elderly to climb, Is Mountain C in City B suitable for the elderly to live in, Is Mountain C in City B suitable for the elderly to avoid the heat, etc.) but related to the context content (both involve City B, and the latter "Mountain C in City B" is also one of the links in the former "itinerary to City B"), it can be paused at "one of" in the previously generated and output machine dialogue content, and preparations can be made to switch to the new topic (that is, "Is Mountain C in City B suitable for the elderly..."). At this time, wait for the first streaming voice to continue to be collected over time, and when the semantics are complete at a certain moment, the currently output machine dialogue content can be interrupted and switched to the new topic. For example, if the first streaming voice is completely expressed as "Is Mountain C in City B suitable for the elderly to climb" at a certain moment, then the new topic can be switched to and new machine dialogue content can be generated and output, such as "The annual temperature of Mountain C in City B is relatively low, and the lowest even drops below minus 20 degrees. For the elderly and people with weak constitutions, it is very unsuitable for climbing...".
[0042] In an implementation scenario, in response to the third predicted category being semantically incomplete and irrelevant to the context content, the machine dialogue content currently being output can continue to output dialogue content while maintaining the current topic. It should be noted that in the case of being irrelevant to the context content, the first streaming voice collected at this time may be that the speaker in the current human-machine interaction intends to have a conversation with others around him / her using the first streaming voice collected at this time instead of having a human-machine conversation. If the machine dialogue content currently being output is the machine dialogue content for the previous voice data "Please help me plan my itinerary to City B" (i.e., the dialogue content generated and output in response to this voice data), and the current dialogue content is generated and output to "... Mountain C located in the north of City B is one of the famous scenic spots in the country", taking the first streaming voice "The stock market trend today..." as an example, its semantics is incomplete and irrelevant to the context content, so it can be considered that the first streaming voice collected at this time may be that the speaker is having a person-to-person conversation with others instead of a human-machine conversation. Then, the machine dialogue content currently being output can continue to output dialogue content while maintaining the current topic. For example, the dialogue content "Together with Mountain D, Mountain E, and Mountain F, they are among the four famous mountains. In addition,..." can be continued to be output.
[0043] It should be noted that the above examples are only possible examples of the target interaction operation in several situations where the third predicted category is semantically complete, semantically incomplete and it is impossible to determine whether it is relevant to the context content, semantically incomplete and relevant to the context content, semantically incomplete and irrelevant to the context content, etc. Other possible situations will not be exemplified one by one here. In addition, please refer to Figure 2 , Figure 2 is a schematic process diagram of an embodiment of the streaming voice interaction method of this application. As Figure 2As shown, the target interaction operation can be triggered and executed by a streaming voice interaction model. The streaming voice interaction model can at least include a first classification network, a second classification network, a third classification network, and a large language model. The first classification network can be used to perform the first classification, the second classification network can be used to perform the second classification, the third classification network can be used to perform the third classification, and the large language model can be used to perform the target interaction operation. Among them, the relevant connotations of the first classification network, the second classification network, and the third classification network can refer to the foregoing relevant descriptions and will not be elaborated here. In addition, the large language model can perform the corresponding target interaction operation according to the third prediction category. Exemplarily, the prompt instruction can be constructed by combining the third prediction category and the recognized text of the first streaming voice, and the prompt instruction is used to instruct the large language model to perform the target interaction operation, and then the prompt instruction is input into the large language model so that the large language model can respond to the first streaming voice in real time. For example, when the third prediction category is semantically incomplete and it is impossible to determine whether it is relevant to the context content, the prompt instruction can include, but is not limited to, "continue the original topic from the interrupted conversation of the current machine conversation content being output"; when the third prediction category is semantically incomplete and relevant to the context content, the prompt instruction can include, but is not limited to, "pause the current machine conversation content being output and prepare to switch to a new topic [recognized text of the first streaming voice]"; when the third prediction category is semantically incomplete and irrelevant to the context content, the prompt instruction can include, but is not limited to, "maintain the current topic of the current machine conversation content and continue to output the conversation content"; when the third prediction category is semantically complete, the prompt instruction can include, but is not limited to, "interrupt the current machine conversation content being output and switch to a new topic [recognized text of the first streaming voice]". Of course, the above examples are only several possible examples in the actual application process, and the specific content of the prompt instruction for the large language model is not limited here.
[0044] Based on the currently collected first streaming voice, the above solution performs a first classification to obtain the first predicted category of the first streaming voice, and the first predicted category is human voice, noise, or silence. In response to the first predicted category being human voice, at least a second classification is performed based on the first streaming voice to obtain the second predicted category of the first streaming voice, and the second predicted category is real human voice or background human voice. In response to the second predicted category being real human voice, at least a third classification is performed based on the first streaming voice to obtain the third predicted category of the first streaming voice, and the third predicted category relates to whether the semantics is complete and whether it is relevant to the context content. Then, based on the third predicted category, a target interaction operation is performed on the currently output machine dialogue content. The target interaction operation is any one of several preset interaction operations, and the several preset interaction operations include maintaining the current topic, switching to a new topic, and continuing the original topic from the interrupted dialogue. Therefore, a hierarchical classification of the first classification, the second classification, and the third classification is sequentially performed on the currently collected first streaming voice. During the first classification process, human voice, noise, or silence is distinguished. During the second classification process, it is further distinguished whether the human voice belongs to real human voice or background human voice. And during the third classification process, it is further distinguished whether the semantics of the real human voice is complete and whether the content is relevant, so as to gradually reduce the mis-triggering of the streaming voice as much as possible through hierarchical progression. Therefore, it is possible to reduce the mis-triggering of the streaming voice as much as possible and improve the accuracy of the streaming voice interaction.
[0045] Please refer to Figure 3 , Figure 3 which is a schematic framework diagram of an embodiment of the streaming voice interaction device of the present application. The streaming voice interaction device 30 includes: a first classification module 31, a second classification module 32, a third classification module 33, and an interaction operation module 34. The first classification module 31 is configured to perform a first classification based on the currently collected first streaming voice to obtain the first predicted category of the first streaming voice; wherein, the first predicted category is human voice, noise, or silence. The second classification module 32 is configured to, in response to the first predicted category being human voice, perform at least a second classification based on the first streaming voice to obtain the second predicted category of the first streaming voice; wherein, the second predicted category is real human voice or background human voice. The third classification module 33 is configured to, in response to the second predicted category being real human voice, perform at least a third classification based on the first streaming voice to obtain the third predicted category of the first streaming voice; wherein, the third predicted category relates to whether the semantics is complete and whether it is relevant to the context content. The interaction operation module 34 is configured to perform a target interaction operation on the currently output machine dialogue content based on the third predicted category; wherein, the target interaction operation is any one of several preset interaction operations, and the several preset interaction operations include maintaining the current topic, switching to a new topic, and continuing the original topic from the interrupted dialogue.
[0046] In the above solution, the streaming voice interaction device 30 performs a first classification based on the currently collected first streaming voice to obtain a first predicted category of the first streaming voice, and the first predicted category is human voice, noise, or silence. In response to the first predicted category being human voice, a second classification is performed based at least on the first streaming voice to obtain a second predicted category of the first streaming voice, and the second predicted category is real human voice or background human voice. In response to the second predicted category being real human voice, a third classification is performed based at least on the first streaming voice to obtain a third predicted category of the first streaming voice, and the third predicted category relates to whether the semantics is complete and whether it is relevant to the context content. Then, based on the third predicted category, a target interaction operation is performed on the machine dialogue content currently being output. The target interaction operation is any one of several preset interaction operations, and the several preset interaction operations include maintaining the current topic, switching to a new topic, and continuing the original topic from the interrupted dialogue. Therefore, a hierarchical classification of the first classification, the second classification, and the third classification is sequentially performed on the currently collected first streaming voice. During the first classification process, human voice, noise, or silence is distinguished. During the second classification process, it is further distinguished whether the human voice belongs to real human voice or background human voice. And during the third classification process, it is further distinguished whether the semantics of the real human voice is complete and whether the content is relevant, so as to gradually reduce the mis-triggering of the streaming voice as much as possible through hierarchical progression. Therefore, it is possible to reduce the mis-triggering of the streaming voice as much as possible and improve the accuracy of the streaming voice interaction.
[0047] In some disclosed embodiments, the first predicted category is predicted by a first classification network. The streaming voice interaction device 30 includes an endpoint prediction module for performing a first classification on a first sample voice based on the first classification network to obtain a predicted category of the first sample voice, and performing a classification prediction on the original sample voice based on a voice endpoint detection network to obtain the voice classification of each sample voice segment in the original sample voice; wherein, the original sample voice is the complete voice where the first sample voice is located, the first sample voice is one of the sample voice segments, and the voice classification is human voice, noise, or silence; the streaming voice interaction device 30 includes a parameter adjustment module for adjusting the network parameters of the first classification network based on the difference between the predicted category and the labeled category of the first sample voice, and the difference between the predicted category and the voice classification of the first sample voice.
[0048] In some disclosed embodiments, the streaming voice interaction device 30 includes a voice acquisition module for acquiring the voice collected when the voice interaction is turned on as a reference voice; the second classification module 32 is specifically configured to perform a second classification based on the first streaming voice and the reference voice to obtain a second predicted category of the first streaming voice.
[0049] In some disclosed embodiments, the second predicted category is predicted by a second classification network, which is trained based on second sample voices. The labeled category of the second sample voices is real human voice or background human voice. The streaming voice interaction device 30 includes a data augmentation module for performing at least one of the following: at a target distance from the location where the voice interaction occurs, collecting a sample source voice to obtain second sample voices with a labeled category of background human voice, where the sample source is a speaker or a player; and at least attenuating the volume of the second sample voices with a labeled category of real human voice to obtain second sample voices with a labeled category of background human voice.
[0050] In some disclosed embodiments, the streaming voice interaction device 30 includes a text acquisition module for acquiring the recognized text of the second streaming voice as historical text, where the second streaming voice is a real human voice in a voice interaction before the first streaming voice. The third classification module 33 is specifically configured to perform a third classification based on the first streaming voice and the historical text to obtain the third predicted category of the first streaming voice.
[0051] In some disclosed embodiments, the interaction operation module 34 includes a first response sub-module for resuming the original topic from the interrupted conversation for the machine dialogue content currently being output in response to the third predicted category being semantically incomplete and unable to determine whether it is related to the context content; the interaction operation module 34 includes a second response sub-module for pausing the machine dialogue content currently being output and preparing to switch to a new topic in response to the third predicted category being semantically incomplete and related to the context content; the interaction operation module 34 includes a third response sub-module for maintaining the current topic and continuing to output the dialogue content for the machine dialogue content currently being output in response to the third predicted category being semantically incomplete and unrelated to the context content.
[0052] In some disclosed embodiments, the target interaction operation is triggered and executed by a streaming voice interaction model, which at least includes a first classification network, a second classification network, a third classification network, and a large language model. The first classification network is used to perform a first classification, the second classification network is used to perform a second classification, the third classification network is used to perform a third classification, and the large language model is used to perform the target interaction operation.
[0053] Please refer to Figure 4 , Figure 4It is a schematic diagram of the framework of an embodiment of the electronic device of the present application. The electronic device 40 at least includes a memory 41 and a processor 42 that are coupled to each other. At least program instructions are stored in the memory 41, and the processor 42 is configured to execute the program instructions to implement the steps in any of the foregoing embodiments of the flow voice interaction method. For details, reference may be made to the foregoing disclosed embodiments, which will not be elaborated herein. As a possible example, the electronic device 40 may include, but is not limited to, an office notebook, a smart phone, a tablet computer, a learning machine, a server, etc. The specific type of the electronic device 40 is not limited herein.
[0054] Specifically, the processor 42 is configured to control itself and the memory 41 to implement the steps in any of the foregoing embodiments of the flow voice interaction method. The processor 42 may also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 42 may be implemented jointly by integrated circuit chips.
[0055] In the above solution, the electronic device 40 performs a first classification based on the currently collected first streaming voice to obtain a first predicted category of the first streaming voice, and the first predicted category is human voice, noise, or silence. In response to the first predicted category being human voice, a second classification is performed based at least on the first streaming voice to obtain a second predicted category of the first streaming voice, and the second predicted category is real human voice or background human voice. In response to the second predicted category being real human voice, a third classification is performed based at least on the first streaming voice to obtain a third predicted category of the first streaming voice, and the third predicted category relates to whether the semantics is complete and whether it is relevant to the context content. Then, based on the third predicted category, a target interaction operation is performed on the currently output machine dialogue content. The target interaction operation is any one of several preset interaction operations, and the several preset interaction operations include maintaining the current topic, switching to a new topic, and continuing the original topic from the interrupted dialogue. Therefore, a hierarchical classification of the first classification, the second classification, and the third classification is sequentially performed on the currently collected first streaming voice. During the first classification process, human voice, noise, or silence is distinguished. During the second classification process, it is further distinguished whether the human voice belongs to real human voice or background human voice. And during the third classification process, it is further distinguished whether the semantics of the real human voice is complete and whether the content is relevant, so as to gradually reduce the mis-triggering of the streaming voice as much as possible through hierarchical progression. Therefore, it is possible to reduce the mis-triggering of the streaming voice as much as possible and improve the accuracy of the streaming voice interaction.
[0056] Please refer to Figure 5 , Figure 5 FIG. is a schematic framework diagram of an embodiment of the computer-readable storage medium 50 of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be run by a processor, and the program instructions 51 are used to implement the steps in any of the above embodiments of the streaming voice interaction method.
[0057] In the above solution, the computer-readable storage medium 50 performs a first classification based on the currently collected first streaming voice to obtain a first predicted category of the first streaming voice, and the first predicted category is human voice, noise, or silence. In response to the first predicted category being human voice, a second classification is performed based at least on the first streaming voice to obtain a second predicted category of the first streaming voice, and the second predicted category is real human voice or background human voice. In response to the second predicted category being real human voice, a third classification is performed based at least on the first streaming voice to obtain a third predicted category of the first streaming voice, and the third predicted category relates to whether the semantics is complete and whether it is relevant to the context content. Then, based on the third predicted category, a target interaction operation is performed on the machine dialogue content currently being output. The target interaction operation is any one of several preset interaction operations, and the several preset interaction operations include maintaining the current topic, switching to a new topic, and continuing the original topic from the interrupted dialogue. Therefore, a hierarchical classification of the first classification, the second classification, and the third classification is sequentially performed on the currently collected first streaming voice. During the first classification process, human voice, noise, or silence is distinguished. During the second classification process, it is further distinguished whether the human voice belongs to real human voice or background human voice. And during the third classification process, it is further distinguished whether the semantics of the real human voice is complete and whether the content is relevant, so as to gradually reduce the mis-triggering of the streaming voice as much as possible through hierarchical progression. Therefore, it is possible to reduce the mis-triggering of the streaming voice as much as possible and improve the accuracy of the streaming voice interaction.
[0058] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0059] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated in this article.
[0060] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation manners described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0061] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0062] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0063] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
[0064] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
Claims
1. A streaming voice interaction method, characterized in that: include: Performing a first classification based on the currently collected first stream speech to obtain a first predicted category of the first stream speech; wherein the first predicted category is human voice, noise, or silence; In response to the first predicted category being human voice, performing a second classification based at least on the first stream speech to obtain a second predicted category of the first stream speech; wherein the second predicted category is a real human voice or a background human voice; In response to the second predicted category being a real human voice, performing a third classification based at least on the first stream speech to obtain a third predicted category of the first stream speech; wherein the third predicted category involves whether the first stream speech is semantically complete and whether it is related to contextual content; Based on the third prediction category, a target interaction operation is performed on the machine dialogue content currently being output; wherein the target interaction operation is any one of several preset interaction operations, and the several preset interaction operations include maintaining the current topic, switching to a new topic, and continuing the original topic from the interrupted dialogue; The first predicted category is predicted by a first classification network, and the training steps of the first classification network include: Performing a first classification on the first sample speech based on the first classification network to obtain a predicted category of the first sample speech, and performing classification prediction on the sample original speech based on the speech endpoint detection network to obtain a speech classification of each sample speech segment in the sample original speech; wherein the sample original speech is the complete speech where the first sample speech is located, the first sample speech is one of the sample speech segments, and the speech classification is human voice, noise, or silence; Based on the difference between the predicted category and the marked category of the first sample speech, and the difference between the predicted category and the speech classification of the first sample speech, the network parameters of the first classification network are adjusted.
2. The method according to claim 1, characterized in that Before performing a second classification based at least on the first stream speech to obtain a second predicted category of the first stream speech, the method further includes: Get the voice collected when voice interaction is enabled as the reference voice; The performing a second classification based at least on the first stream speech to obtain a second predicted category of the first stream speech includes: A second classification is performed based on the first stream speech and the reference speech to obtain a second predicted category of the first stream speech.
3. The method according to claim 1, characterized in that: The second predicted category is predicted by a second classification network, the second classification network is trained based on a second sample speech, the labeled category of the second sample speech is a real human voice or a background human voice, and the step of obtaining the second sample speech includes at least one of the following: At a target distance from the location at the time of voice interaction, voice is collected from a sample source to obtain a second sample voice with the labeled category being the background human voice; wherein the sample source is a speaker or a player; Based on the second sample speech whose labeled category is the real human voice, at least volume attenuation is performed to obtain the second sample speech whose labeled category is the background human voice.
4. The method according to claim 1, characterized in that: Before performing a third classification based at least on the first stream speech to obtain a third predicted category of the first stream speech, the method further includes: Acquire the recognized text of the second stream voice as the historical text; wherein the second stream voice is a real human voice that interacts with the first stream voice before; The performing a third classification based at least on the first stream speech to obtain a third predicted category of the first stream speech includes: A third classification is performed based on the first stream speech and the historical text to obtain a third predicted category of the first stream speech.
5. The method according to claim 1, characterized in that The performing a target interactive operation on the machine dialogue content currently being output based on the third prediction category includes at least one of the following: In response to the third prediction category being semantically incomplete and unable to determine whether it is related to the context content, continuing the original topic of the machine dialogue content currently being output from the interrupted dialogue; In response to the third prediction category being semantically incomplete and relevant to the context content, pausing the machine dialogue content currently being output and preparing to switch to a new topic; In response to the third prediction category being semantically incomplete and irrelevant to contextual content, the current topic of the machine dialogue content currently being output is maintained to continue outputting the dialogue content.
6. The method according to any one of claims 1 to 5, characterized in that: The target interaction operation is triggered and executed by a streaming voice interaction model, and the streaming voice interaction model includes at least a first classification network, a second classification network, a third classification network and a large language model, wherein the first classification network is used to perform the first classification, the second classification network is used to perform the second classification, the third classification network is used to perform the third classification, and the large language model is used to perform the target interaction operation.
7. A streaming voice interaction device, characterized in that: include: A first classification module, configured to perform a first classification based on the currently collected first stream speech to obtain a first predicted category of the first stream speech; wherein the first predicted category is human voice, noise, or silence; a second classification module, configured to, in response to the first predicted category being human voice, perform a second classification based at least on the first stream speech to obtain a second predicted category of the first stream speech; wherein the second predicted category is a real human voice or a background human voice; A third classification module is configured to, in response to the second predicted category being a real human voice, perform a third classification based at least on the first stream speech to obtain a third predicted category of the first stream speech; wherein the third predicted category involves whether the first stream speech is semantically complete and whether the first stream speech is related to contextual content; An interactive operation module, configured to perform a target interactive operation on the machine dialogue content currently being output based on the third prediction category; wherein the target interactive operation is any one of a number of preset interactive operations, and the number of preset interactive operations includes maintaining the current topic, switching to a new topic, and continuing the original topic from the interrupted dialogue; The first predicted category is predicted by a first classification network, and the streaming voice interaction device further includes: An endpoint prediction module, configured to perform a first classification on the first sample speech based on the first classification network to obtain a predicted category of the first sample speech, and to perform classification prediction on the sample original speech based on the speech endpoint detection network to obtain a speech classification of each sample speech segment in the sample original speech; wherein the sample original speech is the complete speech in which the first sample speech is located, the first sample speech is one of the sample speech segments, and the speech classification is human voice, noise, or silence; The parameter adjustment module is used to adjust the network parameters of the first classification network based on the difference between the predicted category and the marked category of the first sample speech, and the difference between the predicted category and the speech classification of the first sample speech.
8. An electronic device, characterized in that: The method comprises at least a memory and a processor coupled to each other, wherein the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the streaming voice interaction method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the streaming voice interaction method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Man-machine conversation interruption method, electronic equipment and computer readable storage medium
CN113488047A
Voice processing method, voice processing device, electronic equipment and storage medium
CN114464204A