Voice interaction method and related devices, equipment, systems and storage media

By using streaming audio-based speech activity and semantic end detection methods in complex interactive environments, the audio ending end point is determined, and the problem of low accuracy of speech recognition in the prior art is solved, and the quality and efficiency of speech interactions are improved.

CN119785791BActive Publication Date: 2025-06-27IFLYTEK CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510278029.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

In the prior art, the speech detection technology has a low accuracy rate of speech recognition in complex interactive environments, resulting in insufficient interaction quality.

Method used

By performing voice activity detection based on streaming audio, in response to detecting a speech start endpoint, a semantic end detection of the streaming audio is performed from the speech start endpoint, and an audio end endpoint is determined in combination with at least one of the speech endpoint and the semantic end endpoint, thereby generating target content for responding to the target audio.

Benefits of technology

The quality of speech recognition and interaction is improved, the impact of multiple factors on target audio detection during the interaction process is reduced, and the delay in target content generation is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785791B_ABST
    Figure CN119785791B_ABST
Patent Text Reader

Abstract

The present application discloses a voice interaction method and related devices, equipment, systems, and storage media. The method includes: performing voice activity detection based on streaming audio; in response to detecting a voice start endpoint, performing semantic end detection on the streaming audio from the voice start endpoint to detect a semantic end endpoint after the voice start endpoint, and continuing to perform voice activity detection on the streaming audio from the voice start endpoint to detect a voice end endpoint after the voice start endpoint; determining an audio end endpoint based on at least one of the voice end endpoint and the semantic end endpoint; and generating target content for responding to the target audio based on the target audio from the voice start endpoint to the audio end endpoint. The above solution can improve the quality of voice recognition and interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of human-computer interaction technologies, and particularly to a voice interaction method and related devices, equipment, systems, and storage media. Background Art

[0002] With the development of artificial intelligence technologies, intelligent human-computer interaction functions can be realized based on technologies such as speech recognition, speech synthesis, and natural language understanding.

[0003] In the prior art, usually based on the Voice Activity Detection (VAD), the end position of the voice activity in the audio signal is detected, and then the response generation is performed based on the separated voice segment. However, the voice activity detection technology needs to detect whether there is a preset duration of silence after the audio activity to determine whether the voice activity ends, resulting in a low accuracy of speech recognition in a complex interaction environment, and further resulting in the interaction quality not being sufficient to meet the user's needs. In view of this, how to improve the quality of speech recognition and interaction has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide a voice interaction method and related devices, equipment, systems, and storage media, which can improve the quality of speech recognition and interaction.

[0005] To solve the above technical problem, in the first aspect of this application, a voice interaction method is provided, including performing voice activity detection based on streaming audio; in response to detecting the start endpoint of the voice, performing semantic end detection on the streaming audio from the start endpoint of the voice to detect the semantic end endpoint after the start endpoint of the voice, and continuing to perform voice activity detection on the streaming audio from the start endpoint of the voice to detect the voice end endpoint after the start endpoint of the voice; determining the audio end endpoint based on at least one of the voice end endpoint and the semantic end endpoint; and generating a target content for responding to the target audio based on the target audio from the start endpoint of the voice to the audio end endpoint.

[0006] To solve the above technical problems, a second aspect of the present application provides a voice interaction device, including a first detection module, a second detection module, an endpoint determination module, and a content generation module. The first detection module is configured to perform voice activity detection based on streaming audio; the second detection module is configured to, in response to detecting a voice start endpoint, perform semantic end detection on the streaming audio from the voice start endpoint to detect a semantic end endpoint after the voice start endpoint, and continue to perform voice activity detection on the streaming audio from the voice start endpoint to detect a voice end endpoint after the voice start endpoint; the endpoint determination module is configured to determine an audio end endpoint based on at least one of the voice end endpoint and the semantic end endpoint; and the content generation module is configured to generate target content for responding to the target audio based on the target audio from the voice start endpoint to the audio end endpoint.

[0007] To solve the above technical problems, a third aspect of the present application provides an electronic device, including a memory and a processor coupled to each other. Program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the voice interaction method in the first aspect above.

[0008] To solve the above technical problems, a fourth aspect of the present application provides a voice interaction system, including an audio acquisition device, a data processing device, and a data output device. The audio acquisition device is configured to acquire streaming audio, the data output device is configured to output the target content generated by the data processing device, and the data processing device is the electronic device described in the third aspect above.

[0009] To solve the above technical problems, a fifth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor, and the program instructions are used to implement the voice interaction method in the first aspect above.

[0010] Based on the above solution, speech activity detection is performed on streaming audio. In response to detecting the start endpoint of speech, semantic end detection is performed on the streaming audio from the start endpoint of speech to detect the semantic end endpoint after the start endpoint of speech, and speech activity detection is continued on the streaming audio from the start endpoint of speech to detect the speech end endpoint after the start endpoint of speech. Based on at least one of the speech end endpoint and the semantic end endpoint, the audio end endpoint is determined, and based on the target audio from the start endpoint of speech to the audio end endpoint, the target content for responding to the target audio is generated. Since the determination of the semantic end endpoint does not require detecting whether there is a preset duration of silence, the detection of the speech end endpoint does not depend on the integrity of the speech activity semantics, and both the speech end endpoint and the semantic end endpoint represent the end of the speech activity. Therefore, semantic end detection and speech end detection are respectively performed on the streaming audio. In different scenarios, the audio end endpoint determined based on at least one of the speech end endpoint and the semantic end endpoint can minimize the influence of various factors on the detection of the target audio during the interaction process, which helps to determine the audio end endpoint based on the end endpoint obtained by a better detection method in the current scenario, so as to minimize the delay in generating the target content. Therefore, the quality of speech recognition and interaction can be improved. Description of the Drawings

[0011] Figure 1 is a schematic flowchart of an embodiment of the speech interaction method of the present application;

[0012] Figure 2 is a schematic diagram of an embodiment of endpoint determination of the speech interaction method of the present application;

[0013] Figure 3 is a schematic framework diagram of an embodiment of the speech interaction device of the present application;

[0014] Figure 4 is a schematic framework diagram of an embodiment of the electronic device of the present application;

[0015] Figure 5 is a schematic framework diagram of an embodiment of the speech interaction system of the present application;

[0016] Figure 6 is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. Detailed Embodiments

[0017] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0018] The terms "system" and "network" are often used interchangeably in this document. The term "and / or" in this document is merely a description of the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, both A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after. Furthermore, "plurality" in this document means two or more than two.

[0019] Please refer to Figure 1 , Figure 1 is a schematic flowchart of an embodiment of the voice interaction method of this application. Specifically, it may include the following steps:

[0020] Step S10: Perform voice activity detection based on streaming audio.

[0021] In the embodiments of the present disclosure, streaming audio has an audio format that divides audio data into small chunks for transmission and processing, allowing the audio data to be gradually loaded and played in the form of a stream during transmission. Regarding the specific data nature of streaming audio, more technical details of streaming audio can be referred to and will not be elaborated here.

[0022] In one implementation scenario, the voice activity detection of streaming audio can be implemented based on a VAD module (Voice Activity Detection module). The VAD module identifies the valid voice part from continuous streaming audio and separates it from the non-voice part or noise. Specifically, the VAD module can calculate the zero-crossing rate of each frame of audio in the streaming audio and compare it with a set threshold to determine whether the frame belongs to voice activity, or determine the presence of voice by detecting the fundamental period of the voice signal, etc. It should be noted that the specific method for the VAD module to perform voice activity detection is not limited in this application.

[0023] In another implementation scenario, voice activity detection of streaming audio can be performed based on a pre-trained acoustic detection model, spectral filter, etc. For the sake of brevity, the specific implementation manner of voice activity detection will not be elaborated here.

[0024] In a specific implementation scenario, as a possible implementation method, an acoustic detection model can be pre-trained. The acoustic detection model can include, but is not limited to, network models with convolutional neural network and recurrent neural network architectures. The streaming audio is input into the acoustic detection model, and the output result of the acoustic detection model is used as the result of acoustic activity detection. To ensure the prediction accuracy of the acoustic detection model as much as possible, sample audio can be collected, and the true start endpoint and true end endpoint are marked on the sample audio. On this basis, the sample audio can be processed by the acoustic detection model to obtain the sample predicted start endpoint and predicted end endpoint of the sample audio. Thus, based on the difference between the true start endpoint and the predicted start endpoint, and the difference between the true end endpoint and the predicted end endpoint, the network parameters of the acoustic detection model can be adjusted until the acoustic detection model training converges. Then, the streaming audio collected during the interaction can be processed based on the trained and converged acoustic detection model to obtain the result of speech activity detection. It should be noted that for the specific processing process of the acoustic detection model, the technical details of network models such as neural network architectures can be referred to and will not be elaborated here. The above solution can process the streaming audio during the interaction based on the pre-trained acoustic detection model, improving the accuracy of audio activity determination.

[0025] Step S20: In response to detecting the speech start endpoint, perform semantic end detection on the streaming audio from the speech start endpoint to detect the semantic end endpoint after the speech start endpoint, and continue to perform speech activity detection on the streaming audio from the speech start endpoint to detect the speech end endpoint after the speech start endpoint.

[0026] In an implementation scenario, when the start endpoint of speech is detected, that is, there is speech activity in the streaming audio, semantic end detection is performed on the streaming audio from the start endpoint of speech to detect the semantic end endpoint after the start endpoint of speech, and speech activity detection is continued on the streaming audio from the start endpoint of speech to detect the end endpoint of speech after the start endpoint of speech. Since the detection of the end endpoint of speech is related to the duration of the audio features in the speech activity, and the detection of the semantic end endpoint is related to the semantics represented by the audio features in the speech activity, when different types of endpoint detections are performed on the same streaming audio after the start endpoint of speech, the order of the moment when the end endpoint of speech is detected and the moment when the semantic end endpoint is detected is different in different scenarios. For example, without considering hardware resources, it can be understood that since semantic detection has predictability, that is, based on the semantic information represented by the input streaming audio, the semantic information represented by the uninput streaming audio can be predicted. Therefore, the moment when the semantic end endpoint is detected is earlier than the moment when the end endpoint of speech is detected. Another example is when the semantics of the speech activity output by the target object is incomplete. Since the detection of the end endpoint of speech is only related to the duration of the audio features in the speech activity, the moment when the end endpoint of speech is detected is earlier than the moment when the semantic end endpoint is detected.

[0027] Please refer to Figure 2 , Figure 2 which is a schematic diagram of an embodiment of endpoint determination of the voice interaction method of the present application. As Figure 2 shown, in a specific implementation scenario, speech activity detection is implemented based on a speech detection module, and semantic end detection is implemented based on a semantic detection module. Taking the time axis as a reference, streaming audio is continuously input. It can be understood that although the moment when the end endpoint of speech is detected is inconsistent with the moment when the semantic end endpoint is detected due to the different processing methods of the speech detection module and the semantic detection module, both the semantic end endpoint and the end endpoint of speech are used to represent the end of the speech activity, that is, the audio end endpoint. Therefore, the semantic end endpoint and the end endpoint of speech are represented as the same moment on the time axis, or there is a slight difference between the semantic end endpoint and the end endpoint of speech on the time axis due to the different detection accuracies of the speech detection module and the semantic detection module.

[0028] In an implementation scenario, semantic end detection is performed based on a large language model, which can understand human multi-modal natural language input, obtain keyword information about the input content and the logical relationships between various elements according to the input content, and generate semantically relevant outputs. For the specific principle, more technical details of the large language model can be referred to and will not be elaborated here. It should be noted that the large language model can include, but is not limited to, open source large models such as LLAMA and Bloom. The network architecture of the large language model is not limited here. The large language model in this application can be an open source large model or a finished large model after parameter adjustment, which is not specifically limited in this application.

[0029] In a specific implementation scenario, based on the voice start endpoint, a prompt instruction is constructed to indicate that the large language model understands the semantic information of the streaming audio from the voice start endpoint to determine the semantic end endpoint of the streaming audio. For example, the constructed prompt instruction is "Please understand the semantic information represented by the voice activity after (the corresponding moment of the voice start endpoint) of the continuously input streaming audio, and predict the semantic end endpoint of this voice activity according to the semantic information", and the prompt instruction is input to the large language model to obtain the output endpoint of the large language model as the semantic end endpoint.

[0030] Step S30: Determine the audio end endpoint based on at least one of the voice end endpoint and the semantic end endpoint.

[0031] In an implementation scenario, in response to the voice end endpoint being detected earlier than the semantic end endpoint, select the voice end endpoint as the audio end endpoint; in response to the semantic end endpoint being detected earlier than the voice end endpoint, select the semantic end endpoint as the audio end endpoint.

[0032] In a specific implementation scenario, when the voice end endpoint is detected earlier than the semantic end endpoint, that is, when the voice end endpoint is used as the audio end endpoint, for example, when the voice activity semantics is incomplete, the resource consumption required for voice activity detection is much lower than that required for semantic detection, etc., forcibly end the currently executing semantic end detection.

[0033] In a specific implementation scenario, when the voice end endpoint is detected, start the detection operation of the voice activity for the latest voice start endpoint.

[0034] In a specific implementation scenario, when the semantic end endpoint is detected earlier than the voice end endpoint, forcibly end the detection operation of the voice activity for the voice end endpoint, and while generating the target content, start the detection operation of the voice activity for the latest voice start endpoint.

[0035] Step S40: Generate target content for responding to the target audio based on the target audio from the voice start endpoint to the audio end endpoint.

[0036] In an implementation scenario, since the determination of the semantic end endpoint does not require detecting whether there is a preset duration of silence, the detection of the voice end endpoint does not depend on the integrity of the voice activity semantics, and both the voice end endpoint and the semantic end endpoint represent the end of the voice activity. Therefore, semantic end detection and voice end detection are respectively performed on the streaming audio. In different scenarios, the audio end endpoint determined based on at least one of the voice end endpoint and the semantic end endpoint can minimize the influence of various factors on the detection of the target audio during the interaction process, which helps to determine the audio end endpoint based on the end endpoint obtained by a better detection method in the current scenario, so as to minimize the delay in generating the target content. Therefore, the quality of speech recognition and interaction can be improved.

[0037] In an implementation scenario, after detecting the voice start endpoint of the target audio and before generating the target content for responding to the target audio based on the target audio from the voice start endpoint to the audio end endpoint, select the previous audio end endpoint as the reference end endpoint, and obtain the comparison result between the interval duration between the reference end endpoint and the current voice start endpoint and the preset duration. It should be noted that the specific setting method of the preset duration is not limited in this application. For example, a fixed duration can be set according to the generation efficiency of the target content. Generate the target content for responding to the target audio based on the data processing method matching the comparison result. In the above solution, using the previous audio end endpoint as the reference end endpoint and determining the generation method of the target content based on the interval duration between the reference end endpoint and the current voice start endpoint can minimize the data processing pressure and improve the interaction efficiency when the target object continuously inputs multiple voice messages to be interacted.

[0038] In a specific implementation scenario, when the comparison result indicates that the interval duration is less than the preset duration, regard the target audio ending at the reference end endpoint as the historical audio, and stop generating the target content for responding to the historical audio. Generate the target content for responding to the historical audio and the current target audio based on the historical audio and the current target audio.

[0039] In another specific implementation scenario, when the comparison result indicates that the interval duration is not less than the preset duration, generate the target content for responding to the current target audio based on the current target audio.

[0040] In a specific implementation scenario, when network resources are relatively abundant, both voice activity detection and semantic end detection have a faster processing speed. Therefore, when the voice end endpoint serves as the audio end endpoint, the target audio from the voice start endpoint to the audio end endpoint has the defect of unclear semantics. Therefore, when the voice end endpoint serves as the audio end endpoint, the target content used to respond to the target audio at least includes content used to guide the target to supplement the relevant interaction, such as "Please supplement your interaction needs", "I don't quite understand what you mean, can you say it again", etc.

[0041] In one implementation scenario, after generating target content for responding to the target audio based on the target audio from the voice start endpoint to the audio end endpoint, the target content is output, and the step of performing voice activity detection based on the streaming audio is returned to start a new round of voice interaction. Specifically, the target content can be played based on an audio playback device, the corresponding text can be displayed based on a display device, etc. The output method of the target content is not limited in this application.

[0042] In a specific implementation scenario, speech synthesis is performed based on the target content to obtain playback audio for responding to the target audio. When the target content is played based on an audio playback device, in order to improve the accuracy of voice interaction and minimize the impact of the playback content on voice activity detection, echo cancellation is performed on the playback content based on the target content, and when the step of performing voice activity detection based on streaming audio is performed, the streaming audio does not contain the playback audio. For the specific implementation of speech synthesis and echo cancellation, please refer to the relevant technical details of speech synthesis, echo cancellation, etc., which will not be repeated here for the sake of brevity.

[0043] In a specific implementation scenario, when the target content of the previous target audio is output at the detection moment of the latest voice start endpoint, for example, when the audio about the interactive content is being played and interrupted by the target object, the target content of the previous target audio is stopped from being output, and the target content of the previous target audio is used as reference data for assisting in generating the target content for responding to the latest target audio in the latest round of voice interaction. The above scheme can realize the rapid interruption of the interaction process and improve the interaction efficiency, and use the target content that has been processed and generated as reference data for generating the target content for responding to the latest target audio, which can reduce the data processing pressure as much as possible and improve the interaction efficiency while improving the integrity of the target content generation.

[0044] In the above solution, voice activity detection is performed based on streaming audio. In response to detecting the start endpoint of speech, semantic end detection is performed on the streaming audio from the start endpoint of speech to detect the semantic end endpoint after the start endpoint of speech, and voice activity detection is continued on the streaming audio from the start endpoint of speech to detect the voice end endpoint after the start endpoint of speech. Based on at least one of the voice end endpoint and the semantic end endpoint, the audio end endpoint is determined. Based on the target audio from the start endpoint of speech to the audio end endpoint, target content for responding to the target audio is generated. Since the determination of the semantic end endpoint does not require detecting whether there is a preset duration of silence, the detection of the voice end endpoint does not depend on the integrity of the voice activity semantics, and both the voice end endpoint and the semantic end endpoint represent the end of the voice activity, semantic end detection and voice end detection are respectively performed on the streaming audio. In different scenarios, the audio end endpoint determined based on at least one of the voice end endpoint and the semantic end endpoint can minimize the influence of various factors on the detection of the target audio during the interaction process, which helps to determine the audio end endpoint based on the end endpoint obtained by a better detection method in the current scenario, so as to minimize the delay in generating the target content. Therefore, the quality of speech recognition and interaction can be improved.

[0045] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of an embodiment of the voice interaction device 30 of the present application. As Figure 3 shown, the voice interaction device 30 includes a first detection module 31, a second detection module 32, an endpoint determination module 33, and a content generation module 34. The first detection module 31 is configured to perform voice activity detection based on streaming audio; the second detection module 32 is configured to, in response to detecting the start endpoint of speech, perform semantic end detection on the streaming audio from the start endpoint of speech to detect the semantic end endpoint after the start endpoint of speech, and continue to perform voice activity detection on the streaming audio from the start endpoint of speech to detect the voice end endpoint after the start endpoint of speech; the endpoint determination module 33 is configured to determine the audio end endpoint based on at least one of the voice end endpoint and the semantic end endpoint; the content generation module 34 is configured to generate target content for responding to the target audio based on the target audio from the start endpoint of speech to the audio end endpoint.

[0046] Therefore, the voice interaction device 30 performs voice activity detection based on the streaming audio. In response to detecting the voice start endpoint, it performs semantic end detection on the streaming audio from the voice start endpoint to detect the semantic end endpoint after the voice start endpoint, and continues to perform voice activity detection on the streaming audio from the voice start endpoint to detect the voice end endpoint after the voice start endpoint. Based on at least one of the voice end endpoint and the semantic end endpoint, it determines the audio end endpoint, and generates the target content for responding to the target audio based on the target audio from the voice start endpoint to the audio end endpoint. Since the determination of the semantic end endpoint does not require detecting whether there is a preset duration of silence, the detection of the voice end endpoint does not depend on the integrity of the voice activity semantics, and both the voice end endpoint and the semantic end endpoint represent the end of the voice activity. Therefore, performing semantic end detection and voice end detection on the streaming audio respectively, in different scenarios, the audio end endpoint determined based on at least one of the voice end endpoint and the semantic end endpoint can minimize the influence of various factors on the target audio detection during the interaction process, which helps to determine the audio end endpoint based on the end endpoint obtained by a better detection method in the current scenario, so as to minimize the delay in generating the target content. Thus, the quality of voice recognition and interaction can be improved.

[0047] In some publicly disclosed embodiments, after responding to detecting the voice start endpoint of the target audio and before generating the target content for responding to the target audio based on the target audio from the voice start endpoint to the audio end endpoint, the voice interaction device 30 further includes a reference end module (not shown) for selecting the previous audio end endpoint as the reference end endpoint; the voice interaction device 30 further includes a comparison and acquisition module (not shown) for obtaining the comparison result of the interval duration between the reference end endpoint and the current voice start endpoint and the preset duration; the content generation module 34 further includes a selection and processing module (not shown) for generating the target content based on the data processing method matching the comparison result.

[0048] In some publicly disclosed embodiments, when the comparison result indicates that the interval duration is less than the preset duration, the voice interaction device 30 further includes a generation stop module (not shown) for stopping generating the target content for responding to the historical audio; wherein the historical audio is the target audio ending at the reference end endpoint; the selection and processing module (not shown) further includes a first generation module (not shown) for generating the target content for responding to the historical audio and the current target audio based on the historical audio and the current target audio.

[0049] In some publicly disclosed embodiments, when the comparison result indicates that the interval duration is not less than the preset duration, the selection and processing module (not shown) further includes a second generation module (not shown) for generating the target content for responding to the current target audio based on the current target audio.

[0050] In some disclosed embodiments, the endpoint determination module 33 further includes a first selection module (not shown) for selecting the voice end endpoint as the audio end endpoint in response to the voice end endpoint being detected earlier than the semantic end endpoint; the endpoint determination module 33 further includes a second selection module (not shown) for selecting the semantic end endpoint as the audio end endpoint in response to the semantic end endpoint being detected earlier than the voice end endpoint.

[0051] In some disclosed embodiments, when the voice end endpoint is detected earlier than the semantic end endpoint, the voice interaction device 30 further includes a first forced module (not shown) for forcibly ending the currently executing semantic end detection; and / or, when the semantic end endpoint is detected earlier than the voice end endpoint, the voice interaction device 30 further includes a second forced module (not shown) for forcibly ending the detection operation of the voice end endpoint by the voice activity detection, and starting the detection operation of the latest voice start endpoint by the voice activity detection while generating the target content.

[0052] In some disclosed embodiments, the second detection module 32 further includes an instruction construction module (not shown) for constructing a prompt instruction based on the voice start endpoint; wherein, the prompt instruction is used to instruct the large language model to understand the semantic information of the streaming audio from the voice start endpoint to determine the semantic end endpoint of the streaming audio; the second detection module 32 further includes a large model module (not shown) for inputting the prompt instruction to the large language model to obtain the output endpoint of the large language model as the semantic end endpoint.

[0053] In some disclosed embodiments, after generating the target content for responding to the target audio based on the target audio from the voice start endpoint to the audio end endpoint, the voice interaction device 30 further includes a content output module (not shown) for outputting the target content and returning to execute the step of performing voice activity detection based on the streaming audio to start a new round of voice interaction.

[0054] In some disclosed embodiments, when there is an output of the target content of the previous target audio at the detection moment of the latest voice start endpoint, the voice interaction device 30 further includes a data processing module (not shown) for stopping the output of the target content of the previous target audio and using the target content of the previous target audio as reference data; wherein, the reference data is used to assist in generating the target content for responding to the latest target audio in the latest round of voice interaction.

[0055] Please refer to Figure 4 , Figure 4 which is a schematic framework diagram of an embodiment of the electronic device 40 of the present application. As Figure 4As shown in the figure, the electronic device 40 includes a memory 41 and a processor 42 that are coupled to each other. Program instructions are stored in the memory 41, and the processor 42 is configured to execute the program instructions to implement the steps in any of the above-described embodiments of the voice interaction method. Specifically, the electronic device 40 may include, but is not limited to: a server, a desktop computer, a laptop computer, a tablet computer, a smart phone, etc., which are not limited herein. Specifically, the processor 42 is configured to control itself and the memory 41 to implement the steps in any of the above-described embodiments of the voice interaction method. The processor 42 may also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 42 may be implemented jointly by integrated circuit chips.

[0056] Therefore, the electronic device 40 performs voice activity detection based on streaming audio. In response to detecting a voice start endpoint, semantic end detection is performed on the streaming audio from the voice start endpoint to detect a semantic end endpoint after the voice start endpoint, and voice activity detection is continued on the streaming audio from the voice start endpoint to detect a voice end endpoint after the voice start endpoint. Based on at least one of the voice end endpoint and the semantic end endpoint, an audio end endpoint is determined. Based on the target audio from the voice start endpoint to the audio end endpoint, target content for responding to the target audio is generated. Since the determination of the semantic end endpoint does not require detecting whether there is a preset duration of silence, the detection of the voice end endpoint does not depend on the integrity of the voice activity semantics, and both the voice end endpoint and the semantic end endpoint represent the end of the voice activity, semantic end detection and voice end detection are respectively performed on the streaming audio. In different scenarios, the audio end endpoint determined based on at least one of the voice end endpoint and the semantic end endpoint can minimize the influence of various factors on the detection of the target audio during the interaction process, and helps to determine the audio end endpoint based on the end endpoint obtained by a better detection method in the current scenario, so as to minimize the delay in generating the target content. Therefore, the quality of speech recognition and interaction can be improved.

[0057] Please refer to Figure 5 , Figure 5It is a schematic framework diagram of an embodiment of the voice interaction system 50 of the present application. As Figure 5 shown, the voice interaction system 50 includes an audio acquisition device 51, a data processing device 52, and a data output device 53. The audio acquisition device 51 is used to acquire streaming audio, and the data output device 53 is used to output the target content generated by the data processing device 52, and the data processing device 52 is the electronic device 40 described in any of the foregoing embodiments.

[0058] In an implementation scenario, the data processing device 52 runs at least a voice detection model and a semantic detection model. The voice detection model is at least used to perform voice activity detection on the streaming audio, and the semantic detection model is at least used to perform semantic end detection on the streaming audio.

[0059] It should be noted that the specific network structures of the voice detection model and the semantic detection model are not limited in the present application. For details, reference can be made to the detailed descriptions of the voice activity detection and semantic end detection technical solutions in the foregoing embodiments. For the sake of brevity, they will not be elaborated here.

[0060] In an implementation scenario, the data processing device 52 also runs a response generation model, which is used to generate target content for responding to the target audio based on the target audio from the voice start endpoint to the audio end endpoint. For details, reference can be made to the detailed descriptions of the technical solutions in the foregoing embodiments. For the sake of brevity, they will not be elaborated here.

[0061] In an implementation scenario, the data output device 53 runs a content synthesis model, which is used to synthesize the target content for responding to the target audio into data in a preset form and output it. For example, the target content is synthesized into audio, text, etc. For details, reference can be made to the detailed descriptions of the technical solutions in the foregoing embodiments. For the sake of brevity, they will not be elaborated here.

[0062] In the above solution, the voice interaction system 50 performs voice activity detection based on streaming audio. In response to detecting the start endpoint of the voice, semantic end detection is performed on the streaming audio from the start endpoint of the voice to detect the semantic end endpoint after the start endpoint of the voice, and voice activity detection is continued on the streaming audio from the start endpoint of the voice to detect the voice end endpoint after the start endpoint of the voice. Based on at least one of the voice end endpoint and the semantic end endpoint, the audio end endpoint is determined. Based on the target audio from the start endpoint of the voice to the audio end endpoint, the target content for responding to the target audio is generated. Since the determination of the semantic end endpoint does not require detecting whether there is a preset duration of silence, the detection of the voice end endpoint does not depend on the integrity of the voice activity semantics, and both the voice end endpoint and the semantic end endpoint represent the end of the voice activity, semantic end detection and voice end detection are respectively performed on the streaming audio. In different scenarios, the audio end endpoint determined based on at least one of the voice end endpoint and the semantic end endpoint can minimize the influence of various factors on the detection of the target audio during the interaction process, which helps to determine the audio end endpoint based on the end endpoint obtained by a better detection method in the current scenario, so as to minimize the latency of target content generation. Therefore, the quality of speech recognition and interaction can be improved.

[0063] Please refer to Figure 6 , Figure 6 is a schematic framework diagram of an embodiment of the computer-readable storage medium 60 of the present application. The computer-readable storage medium 60 includes program instructions 61 that can be run by a processor. The program instructions 61 are used to implement the steps in any of the above-described embodiments of the voice interaction method.

[0064] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0065] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated here.

[0066] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0067] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, and may be located in one place or distributed over multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0068] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0069] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to execute all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.

[0070] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

Claims

1. A voice interaction method, characterized in that: include: Voice activity detection based on streaming audio; In response to detecting a speech start endpoint, performing semantic end detection on the streaming audio from the speech start endpoint to detect a semantic end endpoint after the speech start endpoint, and continuing to perform the voice activity detection on the streaming audio from the speech start endpoint to detect a speech end endpoint after the speech start endpoint; Based on at least one of the speech end endpoint and the semantic end endpoint, determine the audio end endpoint; wherein, in response to the speech end endpoint being detected before the semantic end endpoint, select the speech end endpoint as the audio end endpoint, and forcibly end the semantic end detection currently being executed; in response to the semantic end endpoint being detected before the speech end endpoint, select the semantic end endpoint as the audio end endpoint, forcibly end the detection operation of the speech end endpoint by the voice activity detection, and while generating the target content, start the detection operation of the voice activity detection on the latest speech start endpoint; Based on the target audio from the speech start endpoint to the audio end endpoint, a target content for responding to the target audio is generated.

2. The method according to claim 1, characterized in that After the response to detecting the speech start endpoint of the target audio and before the target content for responding to the target audio is generated based on the target audio from the speech start endpoint to the audio end endpoint, the method further includes: Select the audio end endpoint mentioned above as the reference end endpoint; Obtain a comparison result of the interval time between the reference end endpoint and the current speech start endpoint and a preset time; The step of generating target content for responding to the target audio based on the target audio from the speech start endpoint to the audio end endpoint includes: The target content is generated based on a data processing method that matches the comparison result.

3. The method according to claim 2, characterized in that When the comparison result indicates that the interval duration is less than the preset duration, the method further includes: Stop generating target content for responding to historical audio; wherein the historical audio is the target audio ending at the reference end endpoint; The generating the target content based on the data processing method matching the comparison result includes: Based on the historical audio and the current target audio, the target content for responding to the historical audio and the current target audio is generated.

4. The method according to claim 2, characterized in that: In the case where the comparison result indicates that the interval duration is not less than the preset duration, generating the target content based on a data processing method matching the comparison result includes: Based on the current target audio, the target content for responding to the current target audio is generated.

5. The method according to claim 1, characterized in that The performing semantic end detection on the streaming audio from the voice start endpoint includes: Based on the speech start endpoint, construct a prompt instruction; wherein the prompt instruction is used to instruct the large language model to understand the semantic information of the streaming audio from the speech start endpoint to determine the semantic end endpoint of the streaming audio; The prompt instruction is input to the large language model to obtain an output endpoint of the large language model as the semantic end endpoint.

6. The method according to any one of claims 1 to 5, characterized in that: After generating target content for responding to the target audio based on the target audio from the speech start endpoint to the audio end endpoint, the method further includes: The target content is output, and the step of performing voice activity detection based on streaming audio is returned to start a new round of voice interaction.

7. The method according to claim 6, characterized in that In the case where there is a target content output of the previous target audio at the detection time of the latest voice start endpoint, the method further includes: Stop outputting the target content of the last target audio, and use the target content of the last target audio as reference data; wherein the reference data is used to assist in generating the target content for responding to the latest target audio in the latest round of voice interaction.

8. A voice interaction device, characterized in that: include: A first detection module, configured to perform voice activity detection based on streaming audio; A second detection module is used for, in response to detecting a speech start endpoint, performing semantic end detection on the streaming audio from the speech start endpoint to detect a semantic end endpoint after the speech start endpoint, and continuing to perform the voice activity detection on the streaming audio from the speech start endpoint to detect a speech end endpoint after the speech start endpoint; An endpoint determination module is used to determine an audio end endpoint based on at least one of the voice end endpoint and the semantic end endpoint; wherein, in response to the voice end endpoint being detected prior to the semantic end endpoint, the voice end endpoint is selected as the audio end endpoint, and the semantic end detection currently being executed is forcibly terminated; in response to the semantic end endpoint being detected prior to the voice end endpoint, the semantic end endpoint is selected as the audio end endpoint, and the detection operation of the voice activity detection on the voice end endpoint is forcibly terminated, and while generating the target content, the voice activity detection is started to detect the latest voice start endpoint; The content generation module is used to generate target content for responding to the target audio based on the target audio from the voice start endpoint to the audio end endpoint.

9. An electronic device, characterized in that: It comprises a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the voice interaction method according to any one of claims 1 to 7.

10. A voice interaction system, characterized in that: It comprises an audio acquisition device, a data processing device and a data output device, wherein the audio acquisition device is used to acquire streaming audio, the data output device is used to output target content generated by the data processing device, and the data processing device is the electronic device described in claim 9.

11. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the voice interaction method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Voice endpoint judgment method and device, equipment, storage medium and product

    CN114495981A

  • Voice interaction method, device, equipment, medium and product

    CN119181361A

  • Streaming voice interaction method, related device, equipment and storage medium

    CN119479620A