Streaming voice interaction method and related devices, equipment and storage media
By detecting the semantic integrity of streaming speech in the voice interaction system and adjusting the response delay, the problem of interruption of speech objects caused by response time optimization is solved, and the response time is shortened and the interruption risk is reduced.
Patent Information
- Application Number
- CN202510291994.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-12
AI Technical Summary
In voice interaction systems, optimization of response time may lead to an increased risk of the speech subject being interrupted by unexpectedly, and how to reduce this risk while shortening the response time becomes a challenge.
By detecting whether the streaming voice is semantic intact at the current moment, if it is incomplete, the target delay is obtained based on the basic delay plus the additional delay, and the waiting time is extended to analyze whether the new valid voice is added; if it is complete, the basic delay is waiting and the same analysis is performed. If no new valid voice is added, reply data is generated.
It effectively reduces the possibility that the speaker is interrupted unexpectedly, and at the same time optimizes the response time to ensure that it responds as soon as possible if necessary.
Smart Images

Figure CN119811378B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice interaction technologies, and particularly to a streaming voice interaction method and related devices, equipment, and storage media. Background Art
[0002] Currently, voice interaction technologies have been widely applied in many scenarios such as intelligent driving and smart home. Thanks to the rapid development of voice interaction technologies, the response time of voice interaction has been greatly optimized, even reaching the millisecond level, which has greatly improved the entire voice interaction system.
[0003] However, while the response time of voice interaction has been greatly optimized, new problems have emerged. That is, if the response is too fast, it may frequently interrupt the speaker, and even very likely interrupt the speaker unexpectedly. In view of this, how to reduce the possibility of the speaker being interrupted unexpectedly on the premise of shortening the response time of voice interaction has become an urgent problem to be solved. Summary of the Invention
[0004] The main technical problem to be solved by this application is to provide a streaming voice interaction method and related devices, equipment, and storage media, which can reduce the possibility of the speaker being interrupted unexpectedly on the premise of shortening the response time of voice interaction.
[0005] To solve the above technical problem, in the first aspect of this application, a streaming voice interaction method is provided, including: detecting whether the streaming voice is semantically complete at the current moment; in response to the detection result being that the semantics is incomplete, obtaining a target delay based on the basic delay plus an additional delay, and analyzing whether there is new valid voice in the streaming voice from the current moment until waiting for the target delay. If so, return to detect whether the streaming voice is semantically complete at the current moment. If not, generate reply data; in response to the detection result being that the semantics is complete, analyze whether there is new valid voice in the streaming voice from the current moment until waiting for the basic delay. If so, return to detect whether the streaming voice is semantically complete at the current moment. If not, generate reply data.
[0006] To solve the above technical problems, the second aspect of the present application provides a streaming voice interaction device, including: a semantic detection module, a first response module, and a second response module. The semantic detection module is used to detect whether the streaming voice is semantically complete at the current moment; the first response module is used to, in response to the detection result indicating that the semantics is incomplete, obtain a target delay based on the basic delay plus an additional delay, and analyze whether there is any new valid voice in the streaming voice from the current moment until waiting for the target delay. If so, return to detect whether the streaming voice is semantically complete at the current moment. If not, generate reply data; the second response module is used to, in response to the detection result indicating that the semantics is complete, analyze whether there is any new valid voice in the streaming voice from the current moment until waiting for the basic delay. If so, return to detect whether the streaming voice is semantically complete at the current moment. If not, generate reply data.
[0007] To solve the above technical problems, the third aspect of the present application provides an electronic device, at least including a memory and a processor coupled to each other. The memory stores at least program instructions, and the processor is used to execute the program instructions to implement the streaming voice interaction method in the first aspect above.
[0008] To solve the above technical problems, the fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor. The program instructions are used to implement the streaming voice interaction method in the first aspect above.
[0009] In the above solution, it is detected whether the streaming voice is semantically complete at the current moment. In response to the detection result indicating that the semantics is incomplete, a target delay is obtained based on the basic delay plus an additional delay, and whether there is any new valid voice in the streaming voice is analyzed from the current moment until waiting for the target delay. If so, return to detect whether the streaming voice is semantically complete at the current moment. Otherwise, generate reply data. In response to the detection result indicating that the semantics is complete, whether there is any new valid voice in the streaming voice is analyzed from the current moment until waiting for the basic delay. If so, return to detect whether the streaming voice is semantically complete at the current moment. If not, generate reply data. On the one hand, when it is detected at the current moment that the streaming voice semantics is incomplete, compared with detecting that the semantics is complete and then adding an additional delay on the basis of the basic delay, it is possible to extend the delay when the speaker may be about to continue speaking, which helps to reduce the possibility of the speaker being unexpectedly interrupted. On the other hand, regardless of whether the semantics is complete, as long as new valid voice is detected during the waiting process, the semantic integrity detection is continued and no reply is made temporarily, which also helps to reduce the possibility of the speaker being unexpectedly interrupted. On the other hand, regardless of whether the semantics is complete, if no new valid voice is detected during the waiting process, a reply is made, which can also shorten the response time of the voice interaction as much as possible. Therefore, it is possible to reduce the possibility of the speaker being unexpectedly interrupted on the premise of shortening the response time of the voice interaction. Brief Description of the Drawings
[0010] Figure 1 is a schematic flowchart of an embodiment of the streaming voice interaction method of the present application;
[0011] Figure 2 is a schematic diagram of the effect of an embodiment of the streaming voice interaction method of the present application;
[0012] Figure 3 is a schematic framework diagram of an embodiment of the streaming voice interaction device of the present application;
[0013] Figure 4 is a schematic framework diagram of an embodiment of the electronic device of the present application;
[0014] Figure 5 is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. Detailed Embodiment
[0015] The following will combine the accompanying drawings of the specification to elaborate on the solutions of the embodiments of the present application in detail.
[0016] In the following description, specific details such as specific system structures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.
[0017] In this article, the terms "system" and "network" are often used interchangeably. The term " / and" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the fragment " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more than two.
[0018] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the streaming voice interaction method of the present application. Specifically, it may include the following steps:
[0019] Step S11: Detect whether the streaming voice is semantically complete at the current moment.
[0020] In the embodiments of the present disclosure, during the interaction process, the streaming voice representation continuously collects audio signals. In the actual application process, the continuously collected audio signals (i.e., streaming voice) may contain valid speech (i.e., audio signals containing substantial speech content), may also contain invalid speech (i.e., audio signals without substantial speech content such as pauses), or valid speech and invalid speech may intermittently and alternately exist in the streaming voice. For example, the streaming voice may be "[Short pause] (i.e., this segment is invalid speech) [What's the weather like today] (i.e., this segment is valid speech)..." Of course, the above examples are only several possible examples of the streaming voice, and the specific content of the streaming voice is not limited herein, nor will they be enumerated one by one.
[0021] In an implementation scenario, as a possible implementation example, in order to detect semantic integrity, a deep learning model can be selected as the first artificial intelligence model, and based on the first artificial intelligence model, it can be detected whether the streaming voice is semantically complete at the current moment, that is, it can be detected whether the streaming voice is semantically complete up to the current moment by using the deep learning model. It should be noted that the deep learning model may include, but is not limited to: convolutional neural network, recurrent neural network, long short-term memory network, Transformer, etc. The network structure of the deep learning model is not limited herein.
[0022] In a specific implementation scenario, sample speech can be obtained in advance, and the sample speech can be labeled with a sample mark indicating whether it is actually semantically complete. Then, based on the deep learning model, the sample speech is detected to obtain the predicted mark of the sample speech, and the predicted mark indicates whether the sample speech is semantically complete after prediction. Then, based on the difference between the sample mark and the predicted mark, the network parameters of the deep learning model are adjusted. Through the above process steps, one round of training of the deep learning model can be achieved. In the actual application process, the above process steps can be repeatedly cycled to perform iterative training on the deep learning model until the training converges, and then it can be detected whether the streaming voice is semantically complete based on the deep learning model.
[0023] In a specific implementation scenario, when using the deep learning model to perform semantic integrity detection on the streaming voice, specifically, the voice segment from the last effective interaction to the current moment in the streaming voice can be sent into the deep learning model, and the deep learning model can perform semantic integrity detection on the above voice segment to obtain the detection result, such as the probability value of semantic completeness and the probability value of semantic incompleteness. When the former is higher than the latter, it can be considered semantically complete; on the contrary, when the former is lower than the latter, it can be considered semantically incomplete.
[0024] In a specific implementation scenario, for the sake of easy understanding, taking the streaming voice "[Short pause] [What's the weather like today] [Long pause] [Check for me if there are any flights to City A tomorrow]..." as an example, the last valid interaction in the above streaming voice is "What's the weather like today". If the speaker at the current moment says "Check for me", the deep learning model can perform semantic integrity detection on the voice segment from after "What's the weather like today" to "Check for me" (alternatively, the long pause can also be ignored and only the voice segment of "Check for me" can be used for semantic integrity detection), and determine that the semantics are "incomplete" at the current moment; conversely, if the speaker at the current moment has said "If there are any flights to City A tomorrow", the deep learning model can perform semantic integrity detection on the voice segment from after "What's the weather like today" to "If there are any flights to City A tomorrow" (alternatively, the long pause can also be ignored and only the voice segment of "Check for me if there are any flights to City A tomorrow" can be used for semantic integrity detection), and determine that the semantics are "complete" at the current moment. Of course, the above example is only one possible example in the actual application process, and other possible situations will not be exemplified one by one here.
[0025] In another implementation scenario, as another possible implementation example, different from the foregoing implementation manner, in order to detect semantic integrity, a large language model can also be selected as the first artificial intelligence model, and based on the first artificial intelligence model, it can be detected whether the streaming speech is semantically complete at the current moment, that is, the large language model can be used to detect whether the streaming speech is semantically complete up to the current moment. It should be noted that the large language model can include, but is not limited to, open-source large models such as Llama and Bloom, or large language models obtained by fine-tuning parameters based on a specific corpus, or, it can also be a custom large model, etc. The network structure of the large language model is not limited here. Exemplarily, a prompt instruction can be constructed based on the speech segment from the last effective interaction to the current moment in the streaming speech, and the prompt instruction is used to instruct the large language model to perform semantic integrity detection on the speech segment, and the prompt instruction is input into the large language model to obtain the output result of the large language model, which can be used as the detection result of the streaming speech regarding semantic integrity at the current moment. For the sake of easy understanding, taking the streaming speech "
Short pause
What's the weather like today
Long pause
Check if there are any flights to City A tomorrow for me again
[0026] In yet another implementation scenario, as yet another possible implementation example, different from the foregoing implementation manner, in order to detect semantic integrity, a deep learning model and a large language model can also be selected as the first artificial intelligence model respectively. Based on this, the detection result of the deep learning model regarding the semantic integrity of the streaming speech at the current moment (including whether the semantics is complete and the corresponding confidence level) can be obtained, and the detection result of the large language model regarding the semantic integrity of the streaming speech at the current moment (including whether the semantics is complete and the corresponding confidence level) can be obtained, and then the one with the higher confidence level is selected to determine whether the streaming speech is semantically complete at the current moment.
[0027] In an implementation scenario, different from directly detecting whether the semantics is complete for the streaming voice at the current moment, it is also possible to first perform voice activity detection (VAD) based on the streaming voice before that, and in response to the streaming voice being detected as the end point of speech (such as a short pause, etc.) at the current moment, the step of detecting whether the semantics of the streaming voice is complete at the current moment can be executed. Conversely, in response to the streaming voice being detected as not being the end point of speech (such as still being valid speech) at the current moment, the step of performing voice activity detection based on the streaming voice can be returned. In the above manner, before performing semantic integrity detection, voice activity detection is first carried out, and when it is detected that the current moment is the end point of speech, semantic integrity detection is performed. On the one hand, it can reduce the detection burden of semantic integrity, and on the other hand, it can also improve the detection accuracy of whether the semantics is complete at the current moment by combining the two dimensions of voice activity detection and semantic integrity detection.
[0028] Step S12: In response to the detection result being that the semantics is incomplete, obtain the target delay based on the basic delay plus the additional delay, and during the process from the current moment until waiting for the target delay, analyze whether there is any newly added valid speech in the streaming voice. If so, return to detect whether the semantics of the streaming voice is complete at the current moment. If not, generate the response data.
[0029] In an implementation scenario, as a possible implementation example, the basic delay can be custom-set according to the actual application scenario. For example, it can be set to 200ms, 300ms, 400ms, 500ms, etc. The specific value of the basic delay is not limited here. In addition, the additional delay can also be custom-set according to the actual application scenario. For example, it can be set to 200ms, 300ms, 400ms, 500ms, etc. The specific value of the additional delay is not limited here.
[0030] In another implementation scenario, as another possible implementation example, different from the foregoing implementation manner, the basic time delay can also be adaptively obtained. Specifically, the basic time delay can be determined based on the historical speech of the speaker to whom the streaming speech belongs before the current moment, and there is a correlation between the basic time delay and the speech state represented by the historical speech, and the speech state at least includes the speech rate. It should be noted that the historical speech can specifically be several rounds of speech that the speaker has already interacted with before the current moment, such as five rounds of speech, ten rounds of speech, etc. that the speaker has already interacted with before the current moment. The number of rounds of speech that have been interacted with is not limited herein. In the above manner, the basic time delay is determined based on the historical speech of the speaker to whom the streaming speech belongs before the current moment, and there is a correlation between the basic time delay and the speech state represented by the historical speech, and the speech state at least includes the speech rate. Therefore, the basic time delay can be adaptively determined according to the actual situation of the speaker, which helps to improve the adaptability of the basic time delay to the actual situation.
[0031] In a specific implementation scenario, as a possible implementation example, in order to adaptively determine the basic delay based on historical speech, statistics can be performed on the historical speech to obtain the speech rate of the speaker. The speech rate is expressed as the number of words spoken per standard duration. It should be noted that the standard duration can be set according to the statistical requirements of actual applications. For example, it can be set to 500 ms, etc. The specific value of the standard duration is not limited here. Based on this, the unit delay is adjusted proportionally based on the comparison result between the speech rate and the reference speech rate to obtain the reference delay. It should be noted that the reference speech rate can be a custom normal speech rate, which can also be expressed as the number of words spoken per standard duration. For example, specifically, the number of words spoken per 500 ms can be 3. In addition, the unit delay is expressed as the adjusted delay per unit of text. For example, it can be set to 10 ms per 0.1 words spoken, 15 ms per 0.1 words spoken, etc. The specific value of the unit delay is not limited here. The comparison result can specifically be the number of words difference between the speech rate of the speaker and the reference speech rate per standard delay. The ratio of the difference in the number of words to the unit of text in the unit delay can be set as the adjustment ratio. Then, the adjusted delay in the unit delay is scaled using this adjustment ratio to obtain the reference delay. Still taking the example where the speech rate of the speaker is 2 words spoken per 500 ms, if the reference speech rate is 3 words spoken per 500 ms, the difference in the number of words is 1. If the unit delay represents an adjusted delay of 50 ms per 0.1 words spoken, the adjustment ratio is 1 / 0.1 = 10. Using this adjustment ratio of 10 to scale the adjusted delay of 50 ms, the reference delay of 500 ms can be obtained. Of course, the above example is only one possible example in actual applications, and other possible situations are not exemplified one by one here. After obtaining the reference delay, the reference delay can be further constrained based on the delay interval to obtain the basic delay. It should be noted that the delay interval can be set according to the actual application requirements. For example, it can be set to 400 ms to 800 ms, 300 ms to 700 ms, etc. The specific range of the delay interval is not limited here. In addition, after obtaining the delay interval, the reference delay can be constrained. For example, if the reference delay is within the delay interval, the reference delay can be directly determined as the basic delay. If the reference delay is lower than the lower limit of the delay interval, the lower limit of the delay interval can be determined as the basic delay. If the reference delay is higher than the upper limit of the delay interval, the upper limit of the delay interval can be determined as the basic delay. The above method, which performs statistics on historical speech to obtain the speech rate of the speaker, where the speech rate is expressed as the number of words spoken per standard duration, adjusts the unit delay proportionally based on the comparison result between the speech rate and the reference speech rate to obtain the reference delay, and then constrains the reference delay based on the delay interval to obtain the basic delay, can analyze historical speech through relevant rules to determine the basic delay.
[0032] In a specific implementation scenario, as another possible implementation example, in order to adaptively determine the base delay according to historical speech, the historical speech can also be predicted based on a second artificial intelligence model to obtain the base delay. The second artificial intelligence model at least includes a deep learning model, and the deep learning model is trained based on sample speech annotated with sample delays. It should be noted that the deep learning model can include, but is not limited to: convolutional neural network, recurrent neural network, long short-term memory network, Transformer, etc. The network structure of the deep learning model is not limited here. Of course, in order to improve the accuracy of the deep learning model in predicting the base delay, the deep learning model can be pre-trained. For example, the deep learning model can be supervised-trained based on sample speech annotated with sample base delays. The training process can refer to the relevant description of using the deep learning model in the aforementioned first artificial intelligence model and will not be elaborated here. In the above manner, the historical speech is predicted based on the second artificial intelligence model to obtain the base delay. The second artificial intelligence model at least includes a deep learning model, and the deep learning model is trained based on sample speech annotated with sample delays, which can predict the historical speech through a network model to determine the base delay.
[0033] In a specific implementation scenario, the base delay can also be adaptively updated. Specifically, the recorded data of the current device for streaming voice interaction can be analyzed to obtain the usage pattern of the current device for streaming voice interaction. It should be noted that the recorded data can include the total number of times of streaming voice interaction on the current device at each time period. Based on this, the usage pattern of the current device for streaming voice interaction can be analyzed, such as the possibility of streaming voice interaction on the current device at each time period. Based on this, the update timing of the base delay can be determined according to the usage pattern. For example, when the usage pattern includes the possibility of streaming voice interaction on the current device at each time period, if it is found that the possibility of streaming voice interaction at a certain time period is relatively high, then this time period or a time period before this time period can be determined as the update timing of the base delay. Of course, the above example is only a possible example of the update timing, and other possible situations will not be exemplified one by one here. After determining the update timing, the aforementioned steps for obtaining the base delay can be executed according to the update timing to adaptively update the base delay. That is to say, once the update timing of the base delay is reached, the aforementioned steps for obtaining the base delay can be executed. In the above manner, the recorded data of the current device for streaming voice interaction is analyzed to obtain the usage pattern of the current device for streaming voice interaction, and then the update timing of the base delay is determined according to the usage pattern, so that the steps for obtaining the base delay can be executed according to the update timing to adaptively update the base delay, thereby improving the matching of the base delay.
[0034] In an implementation scenario, after obtaining the basic delay, an additional delay can be added to the basic delay to obtain the target delay, and starting from the current moment, wait for new valid speech in the streaming speech. If no new valid speech is added in the streaming speech until the target delay is reached, reply data for responding to the streaming speech can be generated. On the contrary, if new valid speech is added in the streaming speech at a certain moment before waiting until the target delay, the aforementioned step of detecting whether the streaming speech is semantically complete at the current moment can be returned.
[0035] In a specific implementation scenario, in order to generate reply data, the speech segment of the streaming speech from the completion of the last interaction until the current moment can be obtained, and based on this speech segment, a prompt instruction can be constructed. The prompt instruction is used to instruct the large language model to generate reply data for this speech segment. Then, input the prompt instruction into the large language model, and obtain the output data of the large language model as the reply data. It should be noted that the required format of the reply data, such as text, speech, video, picture, etc., can be included in the prompt instruction, which is not limited here. In addition, after generating the reply data, the reply data can be output on the current device.
[0036] In a specific implementation scenario, as mentioned above, before detecting whether the streaming speech is semantically complete at the current moment, speech activity detection can also be performed based on the streaming speech. In this case, in response to new valid speech in the streaming speech, the step of performing speech activity detection based on the streaming speech can be returned.
[0037] Step S13: In response to the detection result being semantically complete, during the process from the current moment until waiting for the basic delay, analyze whether new valid speech is added in the streaming speech. If so, return to detect whether the streaming speech is semantically complete at the current moment. If not, generate reply data.
[0038] In an implementation scenario, the method for obtaining the basic delay can refer to the aforementioned relevant description and will not be elaborated here. In the case where the detection result is semantically complete, new valid speech in the streaming speech can be waited for starting from the current moment. If no new valid speech is added in the streaming speech until the basic delay is reached, reply data for responding to the streaming speech can be generated. On the contrary, if new valid speech is added in the streaming speech at a certain moment before waiting until the basic delay, the aforementioned step of detecting whether the streaming speech is semantically complete at the current moment can be returned. It should be noted that as mentioned above, before detecting whether the streaming speech is semantically complete at the current moment, speech activity detection can also be performed based on the streaming speech. In this case, in response to new valid speech in the streaming speech, the step of performing speech activity detection based on the streaming speech can be returned. In addition, the specific method for generating reply data can refer to the aforementioned relevant description and will not be elaborated here.
[0039] In addition, as an implementation example in the actual application process, after generating and outputting the response data, it is also possible to obtain the feedback data of the speaking object on the response data. The feedback data indicates whether the response data unexpectedly interrupts the speaking object. The feedback data can be aggregated to the server, and the server can then select, from the user set composed of each speaking object, the speaking objects that have not given feedback on being unexpectedly interrupted by the response data as reference objects. Then, based on the basic delay and speaking speed of the reference objects, a speaking speed-delay mapping relationship is formed. Exemplarily, the basic delay of the reference object (e.g., the average value of the basic delays during each interaction) and the speaking speed (e.g., the average value of the speaking speeds during each interaction) can be fitted to obtain a mathematical function between the basic delay and the speaking speed as the speaking speed-delay mapping relationship. On this basis, in response to the speaking object being a newly registered user, a prompt message for guiding the newly registered user to conduct speaking interactions in a normal state can be output, and the speaking speed of the newly registered user can be obtained by analyzing the voice data input by the newly registered user triggered by the prompt message. Then, based on the speaking speed of the newly registered user and the speaking speed-delay mapping relationship, the initial basic delay of the newly registered user can be obtained. After the newly registered user starts streaming voice interactions, the basic delay can be recalculated and updated through the process steps described above. For specific details, reference can be made to the relevant descriptions of the basic delay calculation and update above, which will not be elaborated here. In the above manner, in response to the speaking object being a newly registered user, a prompt message for guiding the newly registered user to conduct speaking interactions in a normal state is output, the speaking speed of the newly registered user is obtained by analyzing the voice data input by the newly registered user triggered by the prompt message, and then based on the speaking speed of the newly registered user and the speaking speed-delay mapping relationship, the initial basic delay of the newly registered user is obtained. Moreover, the speaking speed-delay mapping relationship is obtained based on the basic delays and speaking speeds of each reference object, and the reference objects are the speaking objects in the user set that have not given feedback on being unexpectedly interrupted by the response data. Therefore, an initial basic delay that is as adaptable as possible can also be initialized for the newly registered user.
[0040] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the effect of an embodiment of the streaming voice interaction method of this application. As Figure 2As shown, in one implementation example, at the current time t1, since it is detected that the semantics are complete, it is possible to analyze whether there is any new valid speech in the streaming voice from the current time t1 until the waiting basic delay. Since it is detected that there is no new valid speech during this period, the response data "The weather is nice today" can be obtained after waiting for the generation delay; or, in another implementation example, at the current time t2, since it is detected that the semantics are incomplete, an additional delay can be added to the basic delay to obtain the target delay, and then analyze whether there is any new valid speech in the streaming voice from the current time t2 until the waiting target delay. Since it is detected that new valid speech appears in the streaming voice after the current time t2, the step of detecting whether the streaming voice is semantically complete at the current time can be returned, and this process can be repeated until the speaker says "How about it" and is detected as semantically complete. Of course, the above examples are only two possible examples in the actual application process, and other possible situations will not be listed one by one here.
[0041] In the above solution, it is detected whether the streaming voice is semantically complete at the current time. In response to the detection result that the semantics are incomplete, based on the basic delay plus an additional delay, the target delay is obtained, and during the process from the current time until the waiting target delay, it is analyzed whether there is any new valid speech in the streaming voice. If so, the step of detecting whether the streaming voice is semantically complete at the current time is returned; otherwise, the response data is generated. In response to the detection result that the semantics are complete, during the process from the current time until the waiting basic delay, it is analyzed whether there is any new valid speech in the streaming voice. If so, the step of detecting whether the streaming voice is semantically complete at the current time is returned; if not, the response data is generated. On the one hand, when it is detected at the current time that the semantics of the streaming voice are incomplete, compared with detecting that the semantics are complete and then adding an additional delay to the basic delay, it is possible to extend the delay when the speaker may continue to speak, which helps to reduce the possibility of the speaker being interrupted unexpectedly. On the other hand, regardless of whether the semantics are complete, as long as new valid speech is detected during the waiting process, the semantic integrity detection is continued and no response is made temporarily, which also helps to reduce the possibility of the speaker being interrupted unexpectedly. On the other hand, regardless of whether the semantics are complete, if no new valid speech is detected during the waiting process, a response is made, which can also shorten the response time of the voice interaction as much as possible. Therefore, it is possible to reduce the possibility of the speaker being interrupted unexpectedly on the premise of shortening the response time of the voice interaction.
[0042] Please refer to Figure 3 , Figure 3It is a schematic framework diagram of an embodiment of the streaming voice interaction device of the present application. The streaming voice interaction device 30 includes: a semantic detection module 31, a first response module 32, and a second response module 33. The semantic detection module 31 is used to detect whether the streaming voice is semantically complete at the current moment; the first response module 32 is used to, in response to the detection result being that the semantics is incomplete, obtain a target delay based on the base delay plus an additional delay, and analyze whether there is new valid voice in the streaming voice from the current moment until waiting for the target delay. If so, return to detect whether the streaming voice is semantically complete at the current moment. If not, generate response data; the second response module 32 is used to, in response to the detection result being that the semantics is complete, analyze whether there is new valid voice in the streaming voice from the current moment until waiting for the base delay. If so, return to detect whether the streaming voice is semantically complete at the current moment. If not, generate response data.
[0043] In the above solution, the streaming voice interaction device 30 detects whether the streaming voice is semantically complete at the current moment. In response to the detection result being that the semantics is incomplete, a target delay is obtained based on the base delay plus an additional delay, and whether there is new valid voice in the streaming voice is analyzed from the current moment until waiting for the target delay. If so, return to detect whether the streaming voice is semantically complete at the current moment. Otherwise, generate response data. In response to the detection result being that the semantics is complete, whether there is new valid voice in the streaming voice is analyzed from the current moment until waiting for the base delay. If so, return to detect whether the streaming voice is semantically complete at the current moment. If not, generate response data. On the one hand, when it is detected at the current moment that the streaming voice semantics is incomplete, compared with detecting that the semantics is complete and then adding an additional delay on the basis of the base delay, it can extend the delay when the speaker may continue to speak, which helps to reduce the possibility of the speaker being interrupted unexpectedly. On the other hand, regardless of whether the semantics is complete, as long as new valid voice is detected during the waiting process, the semantic integrity detection continues and no reply is made temporarily, which also helps to reduce the possibility of the speaker being interrupted unexpectedly. On the other hand, regardless of whether the semantics is complete, if no new valid voice is detected during the waiting process, a reply is made, which can also shorten the response time of the voice interaction as much as possible. Therefore, it is possible to reduce the possibility of the speaker being interrupted unexpectedly on the premise of shortening the response time of the voice interaction.
[0044] In some disclosed embodiments, the streaming voice interaction device 30 further includes an activity detection module, which is used to perform voice activity detection based on the streaming voice before detecting whether the streaming voice is semantically complete at the current moment; the semantic detection module 31 is specifically used to, in response to the streaming voice being detected as the voice end point at the current moment, execute the detection of whether the streaming voice is semantically complete at the current moment.
[0045] In some disclosed embodiments, the streaming voice interaction device 30 further includes a loop iteration module, which is configured to, in response to the streaming voice being detected as not being the end point of speech at the current moment, return to perform voice activity detection based on the streaming voice.
[0046] In some disclosed embodiments, the first response module 32 and the second response module 33 are further specifically configured to, in response to new valid voice in the streaming voice, return to perform voice activity detection based on the streaming voice.
[0047] In some disclosed embodiments, the semantic detection module 31 is specifically configured to detect whether the streaming voice is semantically complete at the current moment based on a first artificial intelligence model; wherein, the first artificial intelligence model includes at least one of a deep learning model and a large language model.
[0048] In some disclosed embodiments, the streaming voice interaction device 30 further includes a delay determination module, which is configured to determine a basic delay based on the historical voice of the speaking object to which the streaming voice belonged before the current moment; wherein, the basic delay is correlated with the speaking state represented by the historical voice, and the speaking state at least includes the speaking speed.
[0049] In some disclosed embodiments, the delay determination module includes a historical analysis sub-module, which is configured to perform statistics based on the historical voice to obtain the speaking speed of the speaking object; wherein, the speaking speed is expressed as the number of words spoken per standard duration; the delay determination module includes a ratio adjustment sub-module, which is configured to perform ratio adjustment on the unit delay based on the comparison result between the speaking speed and the reference speaking speed to obtain a reference delay; the delay determination module includes an interval constraint sub-module, which is configured to constrain the reference delay based on a delay interval to obtain a basic delay.
[0050] In some disclosed embodiments, the delay determination module includes a model prediction sub-module, which is configured to predict the historical voice based on a second artificial intelligence model to obtain a basic delay; wherein, the second artificial intelligence model at least includes a deep learning model, and the deep learning model is trained based on sample voices annotated with sample delays.
[0051] In some disclosed embodiments, the delay determination module includes a record analysis sub-module, which is configured to analyze the record data of the streaming voice interaction on the current device to obtain the usage pattern of the streaming voice interaction on the current device; the delay determination module includes an opportunity determination sub-module, which is configured to determine the update opportunity of the basic delay based on the usage pattern; the delay determination module includes a delay update sub-module, which is configured to execute the step of obtaining the basic delay according to the update opportunity to adaptively update the basic delay.
[0052] In some disclosed embodiments, the streaming voice interaction device 30 further includes a speech guidance module for outputting a prompt message for guiding a newly registered user to interact in a normal state in response to the speech object being a newly registered user. The streaming voice interaction device 30 further includes a speech rate analysis module for analyzing the speech data input by the newly registered user triggered by the prompt message to obtain the speech rate of the newly registered user. The streaming voice interaction device 30 further includes a delay initial module for obtaining the initial basic delay of the newly registered user based on the speech rate of the newly registered user and the speech rate-delay mapping relationship. The speech rate-delay mapping relationship is obtained based on the basic delays and speech rates of each reference object, and the reference object is a speech object in the user set who has not fed back an unexpected interruption to the reply data.
[0053] Please refer to Figure 4 , Figure 4 is a schematic framework diagram of an embodiment of the electronic device of the present application. The electronic device 40 at least includes a memory 41 and a processor 42 coupled to each other. At least program instructions are stored in the memory 41, and the processor 42 is configured to execute the program instructions to implement the steps in any of the above embodiments of the streaming voice interaction method. Specifically, reference may be made to the foregoing disclosed embodiments, which will not be elaborated herein. As a possible example, the electronic device 40 may include, but is not limited to, a smart phone, a tablet computer, a learning machine, an office book, a translator, a smart watch, a server, etc. The specific type of the electronic device 40 is not limited herein.
[0054] Specifically, the processor 42 is configured to control itself and the memory 41 to implement the steps in any of the above embodiments of the streaming voice interaction method. The processor 42 may also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 42 may be implemented jointly by integrated circuit chips.
[0055] In the above solution, the electronic device 40 detects whether the streaming voice is semantically complete at the current moment. In response to the detection result indicating semantic incompleteness, a target delay is obtained based on the basic delay plus an additional delay. During the process from the current moment until waiting for the target delay, it is analyzed whether there is any newly added valid voice in the streaming voice. If so, it returns to detect whether the streaming voice is semantically complete at the current moment; otherwise, reply data is generated. In response to the detection result indicating semantic completeness, during the process from the current moment until waiting for the basic delay, it is analyzed whether there is any newly added valid voice in the streaming voice. If so, it returns to detect whether the streaming voice is semantically complete at the current moment; if not, reply data is generated. On the one hand, when it is detected at the current moment that the streaming voice is semantically incomplete, compared with the case of detecting semantic completeness and adding an additional delay on the basis of the basic delay, it can extend the delay when the speaker may be about to continue speaking, which helps to reduce the possibility of the speaker being unexpectedly interrupted. On the other hand, regardless of whether the semantics is complete or not, as long as newly added valid voice is detected during the waiting process, the semantic completeness detection continues and no reply is made temporarily, which also helps to reduce the possibility of the speaker being unexpectedly interrupted. On the other hand, regardless of whether the semantics is complete or not, if no newly added valid voice is detected during the waiting process, a reply is made, which can also shorten the response time of the voice interaction as much as possible. Therefore, it is possible to reduce the possibility of the speaker being unexpectedly interrupted on the premise of shortening the response time of the voice interaction.
[0056] Please refer to Figure 5 , Figure 5 which is a framework schematic diagram of an embodiment of the computer-readable storage medium 50 of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be run by a processor, and the program instructions 51 are used to implement the steps in any of the above embodiments of the streaming voice interaction method.
[0057] In the above solution, the computer-readable storage medium 50 detects whether the streaming voice is semantically complete at the current moment. In response to the detection result indicating semantic incompleteness, based on the base delay plus an additional delay, a target delay is obtained. During the process from the current moment until waiting for the target delay, it is analyzed whether there is any newly added valid voice in the streaming voice. If so, it returns to detect whether the streaming voice is semantically complete at the current moment; otherwise, reply data is generated. In response to the detection result indicating semantic completeness, during the process from the current moment until waiting for the base delay, it is analyzed whether there is any newly added valid voice in the streaming voice. If so, it returns to detect whether the streaming voice is semantically complete at the current moment; if not, reply data is generated. On the one hand, when it is detected at the current moment that the streaming voice is semantically incomplete, compared with detecting semantic completeness and then adding an additional delay on the basis of the base delay, the delay can be extended when the speaker may be about to continue speaking, which helps to reduce the possibility of the speaker being interrupted unexpectedly. On the other hand, regardless of whether the semantics is complete, as long as newly added valid voice is detected during the waiting process, the semantic completeness detection continues and no reply is made temporarily, which also helps to reduce the possibility of the speaker being interrupted unexpectedly. On the further hand, regardless of whether the semantics is complete, if no newly added valid voice is detected during the waiting process, a reply is made, which can also shorten the response time of the voice interaction as much as possible. Therefore, it is possible to reduce the possibility of the speaker being interrupted unexpectedly on the premise of shortening the response time of the voice interaction.
[0058] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be elaborated here.
[0059] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be elaborated in this article.
[0060] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation manners described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0061] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0062] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0063] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
[0064] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
Claims
1. A streaming voice interaction method, characterized in that: include: Detect whether the streaming speech is semantically complete at the current moment; In response to the detection result being semantically incomplete, a target delay is obtained based on the basic delay plus the additional delay, and from the current moment until the target delay is reached, whether valid speech is newly added to the streaming speech is analyzed, and if so, whether the detected streaming speech is semantically complete at the current moment is returned, and if not, reply data is generated; In response to the detection result being semantically complete, from the current moment until the process of waiting for the basic delay, analyzing whether valid speech is newly added to the streaming speech, and if so, returning whether the detected streaming speech is semantically complete at the current moment, and if not, generating reply data; The method for determining the basic delay includes: Based on the historical speech of the speaker to which the streaming speech belongs before the current moment, statistics are performed to obtain the speech speed of the speaker; wherein the speech speed is expressed as the number of speech words per standard time length; Based on the comparison result between the speech speed and the target speech speed, the unit delay is proportionally adjusted to obtain a reference delay; wherein, based on the difference in the number of speech words per standard delay between the speech speed and the target speech speed, the ratio of the difference in the number of speech words to the unit word in the unit delay is obtained as an adjustment ratio for proportional adjustment; Constraining the reference delay based on the delay interval to obtain the basic delay; The method further comprises: In response to the speech object being a newly registered user, outputting a prompt message for guiding the newly registered user to perform speech interaction in a normal state, and analyzing voice data input by the newly registered user triggered by the prompt message to obtain the speech speed of the newly registered user; Based on the mapping relationship between the speaking speed and speaking speed delay of the newly registered user, the initial basic delay of the newly registered user is obtained; wherein the speaking speed delay mapping relationship is obtained based on the basic delay and speaking speed of each reference object, and the reference object is a speaking object in the user set whose response to the reply data feedback was not unexpectedly interrupted.
2. The method according to claim 1, characterized in that Before detecting whether the streaming speech is semantically complete at the current moment, the method further includes: performing voice activity detection based on the streaming voice; The detecting whether the streaming speech is semantically complete at the current moment includes: In response to the streaming voice being detected as a voice end point at the current moment, the detecting of whether the streaming voice is semantically complete at the current moment is performed.
3. The method according to claim 2, characterized in that The method further comprises: In response to the streaming voice being detected as not being the voice end endpoint at the current moment, returning to perform the voice activity detection based on the streaming voice; And / or, the method further comprises: In response to valid voice being newly added to the streaming voice, returning to the voice activity detection based on the streaming voice.
4. The method according to claim 1, characterized in that: The detecting whether the streaming speech is semantically complete at the current moment includes: Based on a first artificial intelligence model, it is detected whether the streaming speech is semantically complete at the current moment; wherein the first artificial intelligence model includes: at least one of a deep learning model and a large language model.
5. The method according to claim 1, characterized in that The determining the basic delay based on the historical speech of the speaker to which the streaming speech belongs before the current moment includes: The historical speech is predicted based on a second artificial intelligence model to obtain the basic delay; wherein the second artificial intelligence model at least includes a deep learning model, and the deep learning model is trained based on sample speech marked with sample delays.
6. The method according to claim 1, characterized in that The method further comprises: Analyze the recorded data of the streaming voice interaction on the current device to obtain the usage rules of the streaming voice interaction on the current device; Based on the usage pattern, determining an update timing of the basic delay; According to the update timing, the step of acquiring the basic delay is performed to adaptively update the basic delay.
7. A streaming voice interaction device, characterized in that: include: A semantic detection module, used to detect whether the streaming speech is semantically complete at the current moment; A first response module is used for, in response to the detection result that the semantics are incomplete, obtaining a target delay based on the basic delay plus the additional delay, and analyzing whether a valid voice is added to the streaming voice from the current moment until the target delay is waited, and if so, returning whether the detected streaming voice is semantically complete at the current moment, and generating reply data if otherwise; A second response module is used for analyzing whether a valid voice is newly added in the streaming voice in response to the detection result being semantically complete, from the current moment to the process of waiting for the basic delay, if so, returning whether the detected streaming voice is semantically complete at the current moment, and generating reply data if otherwise; A history analysis submodule, configured to obtain a speech speed of the speaker based on statistics of the historical speech of the speaker to which the streaming speech belongs before the current moment; wherein the speech speed is expressed as the number of speech words per standard time length; A delay determination module, configured to adjust the unit delay proportionally based on the comparison result between the speech speed and the target speech speed to obtain a reference delay; wherein, based on the difference in the number of speech words per standard delay between the speech speed and the target speech speed, a ratio of the difference in the number of speech words to the unit word in the unit delay is obtained as an adjustment ratio for proportional adjustment; An interval constraint submodule, used to constrain the reference delay based on the delay interval to obtain the basic delay; a speech guiding module, for outputting a prompt message for guiding the newly registered user to perform speech interaction in a normal state in response to the speech object being a newly registered user, and analyzing voice data input by the newly registered user triggered by the prompt message to obtain the speech speed of the newly registered user; A speech rate analysis module is used to obtain the initial basic delay of the newly registered user based on the speech rate and speech rate delay mapping relationship of the newly registered user; wherein the speech rate delay mapping relationship is obtained based on the basic delay and speech rate of each reference object, and the reference object is a speaking object in the user set whose response to the reply data feedback was not unexpectedly interrupted.
8. An electronic device, characterized in that: The method comprises at least a memory and a processor coupled to each other, wherein the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the streaming voice interaction method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the streaming voice interaction method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Tail point detection method and device, equipment and storage medium
CN114203204A
Speech recognition method and device, electronic equipment and storage medium
CN114582333A
Human-computer interaction method, intelligent robot and storage medium
CN114936560A
Speech recognition method, speech recognition device and vehicle
CN118800283A