Human-computer voice interaction method and device, storage medium and electronic device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD
- Filing Date
- 2024-12-20
- Publication Date
- 2026-06-23
AI Technical Summary
[0005]本申请实施例提供了一种人机语音交互方法、装置、存储介质及电子装置,以至少解决相关技术中,在人机语音交互过程中,自动语音识别不够准确的技术问题
[0017]在本申请实施例中,对目标对象的语音交互信息进行自动语音识别,分别得到第一中间识别文本和第二中间识别文本,其中,所述第一中间识别文本是使用第一语音活性检测策略对所述语音交互信息进行自动语音识别时得到的,所述第二中间识别文本是使用第二语音活性检测策略对所述语音交互信息进行自动语音识别时得到的,所述第一语音活性检测策略中用于检测所述目标对象停止发声的时刻所设置的停顿时长阈值低于所述第二语音活性检测策略中用于检测所述目标对象停止发声的时刻所设置的停顿时长阈值;比较所述第一中间识别文本和所述第二中间识别文本,直至所述第二中间识别文本与所述第一中间识别文本一致,基于监测到的所述语音交互信息的语音识别结束标识获取所述语音交互信息的目标识别结果;对所述目标识别结果进行自然语言处理,得到第一交互意图。基于上述技术方案,通过使用两种不同停顿时长阈值的语音活性检测策略进行自动语音识别,能够更准确地判断语音的结束点,减少误识别和漏识别的情况,解决了在人机语音交互过程中,自动语音识别不够准确的技术问题,进而提高了自动语音识别的准确性和实时性。
Smart Images

Figure CN122266348A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart home technology, and more specifically, to a human-computer voice interaction method, device, storage medium, and electronic device. Background Technology
[0002] With the development of technology, voice interaction, as an important method of human-computer interaction, facilitates communication between users and machines through natural language, greatly improving the user's interactive experience. In practical voice interaction scenarios, Automatic Speech Recognition (ASR) technology is generally used to process users' voice interaction requests. ASR technology can convert speech information into text information, and then use natural language understanding technology to recognize the semantic information corresponding to the text information, and then convert the semantic information into response information. However, speech may be affected by factors such as environmental noise, sound source location, and sound source distance. For example, for a speech segment, the result of automatic speech recognition may contain redundant or erroneous information, leading to errors, reducing the accuracy of semantic information, and also affecting the accuracy of the response information.
[0003] Therefore, in related technologies, there is a technical problem that automatic speech recognition is not accurate enough during human-computer voice interaction.
[0004] No effective solution has yet been proposed to address the technical problem of inaccurate automatic speech recognition during human-computer voice interaction. Summary of the Invention
[0005] This application provides a human-computer voice interaction method, apparatus, storage medium, and electronic device to at least solve the technical problem of inaccurate automatic speech recognition during human-computer voice interaction in related technologies.
[0006] According to one embodiment of this application, a human-computer voice interaction method is provided, comprising: performing automatic speech recognition on voice interaction information of a target object to obtain a first intermediate recognition text and a second intermediate recognition text, wherein the first intermediate recognition text is obtained when performing automatic speech recognition on the voice interaction information using a first speech activity detection strategy, and the second intermediate recognition text is obtained when performing automatic speech recognition on the voice interaction information using a second speech activity detection strategy, wherein the pause duration threshold set in the first speech activity detection strategy for detecting the moment when the target object stops speaking is lower than the pause duration threshold set in the second speech activity detection strategy for detecting the moment when the target object stops speaking; comparing the first intermediate recognition text and the second intermediate recognition text until the second intermediate recognition text is consistent with the first intermediate recognition text; obtaining a target recognition result of the voice interaction information based on a detected speech recognition end marker of the voice interaction information; and performing natural language processing on the target recognition result to obtain a first interaction intent.
[0007] In an exemplary embodiment, comparing the first intermediate recognized text and the second intermediate recognized text until the second intermediate recognized text matches the first intermediate recognized text includes: determining the recognition time of the automatic speech recognition, determining the recognition duration to which the recognition time belongs, obtaining a first sub-text recognized within the recognition duration from the first intermediate recognized text, and obtaining a second sub-text recognized within the recognition duration from the second intermediate recognized text; comparing the first sub-text and the second sub-text to obtain a comparison result; and determining that the second intermediate recognized text matches the first intermediate recognized text if any one of the multiple comparison results is used to indicate that the first sub-text and the second sub-text are consistent, for multiple comparison results corresponding to different recognition times.
[0008] In one exemplary embodiment, obtaining the first intermediate recognition text includes: acquiring audio words obtained by recognizing the voice interaction information, and obtaining the audio pause duration based on the time difference between different audio words; generating an end identifier of the first voice activity detection strategy when it is determined that the audio pause duration is less than the pause duration threshold of the first voice activity detection strategy; and obtaining the first intermediate recognition text obtained by automatically recognizing the voice interaction information of the target object within a first duration based on the end identifier of the first voice activity detection strategy, wherein the start time of the first duration represents the time when automatic voice recognition begins, and the end time of the first duration represents the time when the end identifier of the first voice activity detection strategy is generated.
[0009] In one exemplary embodiment, the method further includes: performing natural language processing on the first intermediate identified text to obtain a second interaction intent; and determining whether the first interaction intent matches the second interaction intent. Figure 1 In the event of a conflict, the response information corresponding to the first interaction intent will be sent to the target object.
[0010] In one exemplary embodiment, obtaining the second intermediate recognition text includes: acquiring audio words obtained from recognizing the voice interaction information, and obtaining the audio pause duration based on the time difference between different audio words; if it is determined that the audio pause duration is less than the pause duration threshold of the second voice activity detection strategy, generating an end identifier of the second voice activity detection strategy; and acquiring the second intermediate recognition text obtained by automatically recognizing the voice interaction information of the target object within a second duration based on the end identifier of the second voice activity detection strategy, wherein the start time of the second duration represents the time when automatic voice recognition begins, and the end time of the second duration represents the time when the end identifier of the second voice activity detection strategy is generated.
[0011] In an exemplary embodiment, the method further includes: when the second intermediate identified text is inconsistent with the first intermediate identified text, performing natural language processing on the first intermediate identified text to obtain a third interaction intent, and performing natural language processing on the target identification result to obtain a fourth interaction intent; if it is determined that the third interaction intent and the fourth interaction intent match, and if it is determined that the fourth interaction intent and the first interaction intent also match, then sending a reply message of the fourth interaction intent to the target object.
[0012] In an exemplary embodiment, the method further includes: obtaining the speech rate of the target object based on the voice interaction information, and setting a first pause duration threshold of the first speech activity detection strategy and a second pause duration threshold of the second speech activity detection strategy based on the speech rate; wherein, setting the first pause duration threshold of the first speech activity detection strategy and the second pause duration threshold of the second speech activity detection strategy based on the speech rate includes: when it is determined that the speech rate is lower than a preset speech rate, setting a third duration corresponding to the preset speech rate as the first pause duration threshold, and setting a fourth duration corresponding to the preset speech rate as the second pause duration threshold, wherein the third duration is less than the fourth duration; when it is determined that the speech rate is higher than the preset speech rate, determining an adjustment duration corresponding to the speech rate, setting the sum of the adjustment duration and the third duration as the first pause duration threshold, and setting the adjustment duration and the fourth duration as the second pause duration threshold.
[0013] In an exemplary embodiment, the method further includes: setting a first pause duration threshold for the first speech activity detection strategy and a second pause duration threshold for the second speech activity detection strategy based on the ambient noise of the environment in which the target object is located, wherein setting the first pause duration threshold for the first speech activity detection strategy and the second pause duration threshold for the second speech activity detection strategy based on the ambient noise of the environment in which the target object is located includes: when the current noise value of the ambient noise is higher than a preset threshold, setting a first duration corresponding to the current noise value as the first pause duration threshold, and setting a second duration corresponding to the current noise value as the second pause duration threshold, wherein the first duration is less than the second duration.
[0014] According to another aspect of the embodiments of this application, a human-computer voice interaction device is also provided, comprising: a first acquisition module, configured to automatically recognize the voice interaction information of a target object, and obtain a first intermediate recognition text and a second intermediate recognition text, wherein the first intermediate recognition text is obtained when the voice interaction information is automatically recognized using a first voice activity detection strategy, and the second intermediate recognition text is obtained when the voice interaction information is automatically recognized using a second voice activity detection strategy, wherein the pause duration threshold set in the first voice activity detection strategy for detecting the moment when the target object stops speaking is lower than the pause duration threshold set in the second voice activity detection strategy for detecting the moment when the target object stops speaking; a second acquisition module, configured to compare the first intermediate recognition text and the second intermediate recognition text until the second intermediate recognition text is consistent with the first intermediate recognition text, and obtain the target recognition result of the voice interaction information based on the voice recognition end marker of the monitored voice interaction information; and a sending module, configured to perform natural language processing on the target recognition result to obtain a first interaction intent.
[0015] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, and the computer program is configured to execute the above-described human-computer voice interaction method when it is run.
[0016] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-described human-computer voice interaction method through the computer program.
[0017] In this embodiment, automatic speech recognition is performed on the voice interaction information of the target object to obtain a first intermediate recognition text and a second intermediate recognition text. The first intermediate recognition text is obtained when the voice interaction information is automatically recognized using a first speech activity detection strategy, and the second intermediate recognition text is obtained when the voice interaction information is automatically recognized using a second speech activity detection strategy. The pause duration threshold set in the first speech activity detection strategy for detecting the moment when the target object stops speaking is lower than the pause duration threshold set in the second speech activity detection strategy for detecting the moment when the target object stops speaking. The first intermediate recognition text and the second intermediate recognition text are compared until the second intermediate recognition text matches the first intermediate recognition text. Based on the detected speech recognition end marker of the voice interaction information, the target recognition result of the voice interaction information is obtained. Natural language processing is performed on the target recognition result to obtain a first interaction intent. Based on the above technical solution, by using two different pause duration thresholds for speech activity detection, automatic speech recognition can more accurately determine the end point of speech, reduce misidentification and missed identification, solve the technical problem of insufficient accuracy of automatic speech recognition in human-computer voice interaction, and thus improve the accuracy and real-time performance of automatic speech recognition. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of the hardware environment for a human-computer voice interaction method according to an embodiment of this application;
[0021] Figure 2 This is a flowchart of a human-computer voice interaction method according to an embodiment of this application;
[0022] Figure 3 This is a schematic diagram illustrating the principle of the human-computer voice interaction method according to an embodiment of this application;
[0023] Figure 4 This is a structural block diagram of a human-computer voice interaction device according to an embodiment of this application. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present application.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] According to one aspect of the embodiments of this application, a human-computer voice interaction method is provided. This human-computer voice interaction method is widely used in whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligence house ecosystems. Optionally, in this embodiment, the above-mentioned human-computer voice interaction method can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.
[0027] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.
[0028] This embodiment provides a human-computer voice interaction method, applied to the aforementioned terminal device. Figure 2 This is a flowchart of a human-computer voice interaction method according to an embodiment of this application, which includes the following steps:
[0029] Step S202: Automatic speech recognition is performed on the voice interaction information of the target object to obtain a first intermediate recognition text and a second intermediate recognition text. The first intermediate recognition text is obtained when the voice interaction information is automatically recognized using a first speech activity detection strategy, and the second intermediate recognition text is obtained when the voice interaction information is automatically recognized using a second speech activity detection strategy. The pause duration threshold set in the first speech activity detection strategy for detecting when the target object stops speaking is lower than the pause duration threshold set in the second speech activity detection strategy for detecting when the target object stops speaking.
[0030] It's important to note that in speech recognition, speech pause duration refers to the time period during which the speaker stops speaking but hasn't yet finished the entire speech input. This pause appears as silence or background noise in the speech signal. The VAD (Voice Activity Detection) algorithm can use speech pause duration to determine whether the speaker has completed transmitting a segment of speech.
[0031] Specifically, the VAD algorithm analyzes features of the audio signal, such as energy, zero-crossing rate, and spectrum, to determine whether the signal contains valid speech. When the characteristic value of the speech signal is detected to be continuously below a set threshold after a certain period of time, the VAD algorithm considers that the speaker has paused, and the recorded duration is the speech pause duration. This duration is crucial for speech recognition and interaction systems because it affects when the system considers a speech input to have ended, thus enabling it to begin processing and generate a response.
[0032] Therefore, the threshold setting for speech pause duration will vary in different speech recognition scenarios. For example, in quiet environments, a lower threshold can be set according to the first speech activity detection strategy, meaning that shorter pauses may be recognized as the end of speech, which can speed up the system's response. In noisy environments, a higher threshold can be set according to the second speech activity detection strategy to ensure that the system does not misjudge the end of speech due to a brief decrease in noise, thus ensuring the accuracy of speech recognition. It is evident that appropriately setting the preset speech pause duration is one of the key factors in speech recognition design, directly affecting recognition real-time performance and user experience. By dynamically adjusting the VAD strategy, the response speed and efficiency of voice interaction can be improved while maintaining recognition accuracy.
[0033] By using VAD (Voice Amplifier) to detect the presence of a speech signal during the ASR (Automatic Speech Reduction) process, it can be determined whether human voice activity has ended after a certain period of time. This process includes determining whether the signal contains human voice based on characteristics such as energy, zero-crossing rate, and spectrum of the audio signal, and then analyzing the extracted signal features by setting a pause duration threshold. If the feature value exceeds the pause duration threshold, the system considers the current signal to be speech; if it does not exceed it, it is considered silence or noise.
[0034] The end of voice activity detection is indicated by "VAD end" (corresponding to the end indicators of the second and first voice activity detection strategies mentioned above). This indicates that the audio data segment has ended at this time point, and subsequent audio at this time point is either silence or noise. This time point can be used as the basis for the end of valid audio for subsequent logical processing. The use and threshold configuration of VAD end vary in different scenarios and applications, and there may also be some errors.
[0035] The process of extracting features from an audio signal and determining whether it contains human voice may include the following steps:
[0036] 1. Audio signal preprocessing. Before feature extraction, the raw audio signal needs to be preprocessed. Preprocessing includes:
[0037] Denoising: Use filters to remove background noise.
[0038] Normalization: Adjusting the volume of a signal to a uniform volume level.
[0039] 2. Feature Extraction. Next, key features are extracted from the preprocessed audio signal. Commonly used features include:
[0040] Energy: The energy of a signal is usually measured using the sum of squares of the signal.
[0041] Zero-crossing rate: Calculates the number of times a signal waveform crosses zero, which can help distinguish between speech and non-speech signals.
[0042] Spectral characteristics: The spectral information of the signal is extracted using methods such as Fourier transform, including the spectral envelope, Mel frequency cepstral coefficients, etc.
[0043] 3. Feature Analysis and Threshold Setting. Extracted features need to be compared with preset thresholds to determine whether the signal contains human voice. The threshold settings are as follows:
[0044] Energy threshold: If the energy of a signal exceeds a certain threshold, it may indicate that the signal contains human voice.
[0045] Zero-crossing rate threshold: The zero-crossing rate of a speech signal is usually higher than that of noise or silence.
[0046] Spectral feature threshold: Speech signals have specific distribution characteristics in the spectrum, such as fundamental frequency and harmonics.
[0047] 4. Decision Logic. Based on the comparison results between the feature value and the threshold, the following decision can be made:
[0048] If the feature value exceeds the threshold, the system considers the current signal to be speech.
[0049] If the feature value does not exceed the threshold, the system considers the current signal to be either silent or noise.
[0050] 5. Implementation and Optimization. In practical applications, the above process is adjusted and optimized to adapt to different environments and needs. Specifically:
[0051] Dynamically adjust threshold: The threshold is dynamically adjusted according to the ambient noise level.
[0052] Multi-feature fusion: Combining multiple features for comprehensive judgment improves the accuracy of recognition.
[0053] Machine learning: Using machine learning methods, such as support vector machines or neural networks, to perform more complex analysis and classification of features.
[0054] By following the steps above, features can be effectively extracted from audio signals, and it can be determined whether the signal contains human voice. This method has wide applications in speech recognition, speech enhancement, and speech separation.
[0055] Step S204: Compare the first intermediate recognition text and the second intermediate recognition text until the second intermediate recognition text matches the first intermediate recognition text; obtain the target recognition result of the voice interaction information based on the voice recognition end marker of the monitored voice interaction information.
[0056] In this step, the target recognition result includes at least the first intermediate recognition text.
[0057] Optionally, if the second intermediate recognition text is inconsistent with the first intermediate recognition text, the target recognition result of the voice interaction information is obtained based on the voice recognition end marker of the monitored voice interaction information. In this case, the target recognition result includes at least the first intermediate recognition text and the second intermediate recognition text.
[0058] Step S206: Perform natural language processing on the target recognition result to obtain the first interactive intent.
[0059] Natural Language Processing (NLP) is a term used to describe the processing of natural language.
[0060] Through the above steps, automatic speech recognition is performed on the voice interaction information of the target object, resulting in a first intermediate recognition text and a second intermediate recognition text. The first intermediate recognition text is obtained when the voice interaction information is automatically recognized using a first voice activity detection strategy, and the second intermediate recognition text is obtained when the voice interaction information is automatically recognized using a second voice activity detection strategy. The pause duration threshold set in the first voice activity detection strategy for detecting the moment when the target object stops speaking is lower than the pause duration threshold set in the second voice activity detection strategy for detecting the moment when the target object stops speaking. The first intermediate recognition text and the second intermediate recognition text are compared until the second intermediate recognition text matches the first intermediate recognition text. Based on the detected speech recognition end marker of the voice interaction information, the target recognition result of the voice interaction information is obtained. Natural language processing is performed on the target recognition result to obtain the first interaction intent. By using the above technical solution and employing two voice activity detection strategies with different pause duration thresholds for automatic speech recognition, the end point of speech can be determined more accurately, reducing misrecognition and missed recognition. This solves the technical problem of insufficient accuracy in automatic speech recognition during human-computer voice interaction, thereby improving the accuracy and real-time performance of automatic speech recognition.
[0061] Optionally, in this embodiment, a triggering strategy is configured by setting thresholds for both strong and weak VADs. This means that two VAD thresholds can be configured within the ASR recognition process, supporting two VAD end callbacks. First, the threshold of the weak VAD strategy (corresponding to the first speech activity detection strategy) is used to preprocess most interaction scenarios. Second, the threshold of the strong VAD strategy (corresponding to the second speech activity detection strategy) is used as a secondary check and fallback. Then, semantic consistency verification is performed in the final ASR end stage, which can improve ASR recognition accuracy and reduce voice interaction latency. If the recognition results are consistent, only the intermediate recognition result of the weak VAD is used; otherwise, the intermediate recognition result of the strong VAD can be used, or the recognized text can be obtained again.
[0062] Optionally, idempotency can be implemented for voice interaction information within the same session request. This involves avoiding duplicate parsing by detecting the uniqueness of the NLP parsing results, ensuring consistency in the parsing results within the same session. For example, if the semantics of the NLP intent parsing results are consistent, parsing will not be repeated. This allows subsequent stages of the voice interaction process to support idempotency. For instance, when NLP intent parsing is called multiple times, it first checks whether the semantics of the same session are consistent. If they are consistent, there is no need to trigger business logic again, and NLP intent parsing will not be repeated. This ensures the integrity and consistency of subsequent voice interaction logic after ASR recognition is completed, reducing the number of repetitive processing steps.
[0063] Optionally, voice interaction information and corresponding interaction intents can be stored in the interaction record. For new voice interaction information, it can be searched from the interaction record first, which improves the parsing efficiency of instructions with the same semantics.
[0064] In an exemplary embodiment, the process of comparing the first intermediate recognized text and the second intermediate recognized text until the second intermediate recognized text matches the first intermediate recognized text includes: determining the recognition time of the automatic speech recognition, determining the recognition duration to which the recognition time belongs, obtaining a first sub-text recognized within the recognition duration from the first intermediate recognized text, and obtaining a second sub-text recognized within the recognition duration from the second intermediate recognized text; comparing the first sub-text and the second sub-text to obtain a comparison result; and determining that the second intermediate recognized text matches the first intermediate recognized text if any one of the multiple comparison results is used to indicate that the first sub-text and the second sub-text are consistent, for multiple comparison results corresponding to different recognition times. This embodiment provides a way to verify the consistency of intermediate results in the speech recognition process. By analyzing and comparing intermediate recognition results under weak VAD strategies and strong VAD strategies in real time, the real-time performance and accuracy of speech recognition are improved, while optimizing the resource utilization of the system, enabling the voice interaction system to provide a smoother and more accurate user experience in different environments and scenarios.
[0065] The different identification times have a sequential order of identification, and the process of determining any one of the multiple comparison results is carried out in accordance with the sequential order of identification.
[0066] In an exemplary embodiment, the process of obtaining the first intermediate recognized text obtained by automatically performing speech recognition on the voice interaction information of the target object specifically includes: obtaining audio words obtained by recognizing the voice interaction information, and obtaining the audio pause duration based on the time difference between different audio words; generating an end marker of the first voice activity detection strategy when it is determined that the audio pause duration is less than the pause duration threshold of the first voice activity detection strategy; and obtaining the first intermediate recognized text obtained by automatically performing speech recognition on the voice interaction information of the target object within a first duration based on the end marker of the first voice activity detection strategy, wherein the start time of the first duration represents the time when automatic speech recognition begins, and the end time of the first duration represents the time when the end marker of the first voice activity detection strategy is generated. This embodiment can generate a weak VAD end marker in a quiet or low-background-noise environment. Using a weak VAD strategy can reduce the probability of false truncation and improve the accuracy of recognized text.
[0067] The above embodiments can cover most interaction scenarios, reduce the probability of accidental truncation and loss of audio data, improve the accuracy of ASR recognition, shorten the voice interaction time, reduce the latency of voice interaction, effectively improve the real-time performance and efficiency of voice interaction, and enhance the user experience of voice interaction.
[0068] In one exemplary embodiment, further, the specific scheme for sending the response information corresponding to the first interaction intent to the target object may include: performing natural language processing on the first intermediate identified text to obtain a second interaction intent; determining the relationship between the first interaction intent and the second interaction intent... Figure 1 If the two intentions are consistent, the response information corresponding to the first interaction intent is sent to the target object. This embodiment ensures consistency between speech recognition and NLP processing by comparing the interaction intents (first interaction intent and second interaction intent) obtained from the first intermediate recognized text and the final ASR recognition result. Figure 1 The system can directly send response information, avoiding unnecessary processing and delays. This not only improves the system's response speed but also reduces resource consumption, as there is no need to re-parse NLP intents, thus saving computing resources and reducing overall interaction latency.
[0069] When the end of the strong and weak VAD is triggered, in addition to obtaining the next intermediate ASR recognition result, the solution provided in this embodiment can also be implemented, that is, to asynchronously trigger the NLP intent parsing step required for subsequent voice interaction in advance, so that the subsequent NLP parsing step and the current ASR recognition link can be processed asynchronously in parallel.
[0070] In an exemplary embodiment, the process of obtaining the second intermediate recognized text is described through the following steps: acquiring audio words obtained from recognizing the voice interaction information, and obtaining the audio pause duration based on the time difference between different audio words; if the audio pause duration is determined to be less than the pause duration threshold of the second voice activity detection strategy, generating an end marker of the second voice activity detection strategy; and obtaining the second intermediate recognized text obtained by automatically recognizing the voice interaction information of the target object within a second duration based on the end marker of the second voice activity detection strategy, wherein the start time of the second duration represents the time when automatic voice recognition begins, and the end time of the second duration represents the time when the end marker of the second voice activity detection strategy is generated. This embodiment can use a strong VAD strategy (i.e., the second voice activity detection strategy) to generate a strong VAD end marker in noisy or high-noise environments. This helps ensure that even in complex environments, the system can accurately acquire complete voice information, thereby obtaining a more accurate second intermediate recognized text. The strong VAD strategy improves the accuracy of voice recognition in noisy environments, avoids instruction truncation due to misjudgment of the end of voice, and enhances the robustness and adaptability of the system.
[0071] In an exemplary embodiment, further, if the second intermediate identified text is inconsistent with the first intermediate identified text, natural language processing can be performed on the first intermediate identified text to obtain a third interaction intent, and natural language processing can be performed on the target identification result to obtain a fourth interaction intent; if it is determined that the third interaction intent and the fourth interaction intent match, and if it is determined that the fourth interaction intent and the first interaction intent also match, then a reply to the fourth interaction intent is sent to the target object.
[0072] In this embodiment, when the second intermediate recognized text is inconsistent with the first intermediate recognized text, the system further analyzes the two intermediate texts and the final ASR result to obtain the third and fourth interaction intents. If the third and fourth interaction intents match, and the fourth interaction intent also matches the first interaction intent, the system will directly use the intent corresponding to the final ASR result to respond. This strategy avoids repetitive processing caused by multiple text changes, reduces unnecessary NLP intent parsing times, and ensures the accuracy and consistency of the final decision. This embodiment optimizes the processing flow while ensuring speech recognition quality, further shortening the overall voice interaction latency and improving system efficiency and user experience.
[0073] Optionally, if it is determined that the fourth interaction intent and the first interaction intent also match, the response information of the first interaction intent can be sent to the target object.
[0074] Optionally, if the fourth interaction intent does not match the first interaction intent, the text can be re-parsed and recognized, or an alarm message can be generated for manual processing.
[0075] In an exemplary embodiment, further, before obtaining the first intermediate recognized text and the second intermediate recognized text, the following steps are implemented: obtaining the speech rate of the target object based on the voice interaction information, and setting a first pause duration threshold of the first speech activity detection strategy and a second pause duration threshold of the second speech activity detection strategy based on the speech rate, wherein setting the first pause duration threshold of the first speech activity detection strategy and the second pause duration threshold of the second speech activity detection strategy based on the speech rate includes: when it is determined that the speech rate is lower than a preset speech rate, setting a third duration corresponding to the preset speech rate as the first pause duration threshold, and setting a fourth duration corresponding to the preset speech rate as the second pause duration threshold, wherein the third duration is less than the fourth duration; when it is determined that the speech rate is higher than the preset speech rate, determining an adjustment duration corresponding to the speech rate, setting the sum of the adjustment duration and the third duration as the first pause duration threshold, and setting the adjustment duration and the fourth duration as the second pause duration threshold.
[0076] In an exemplary embodiment, further, a first pause duration threshold for the first speech activity detection strategy and a second pause duration threshold for the second speech activity detection strategy can be set based on the environmental noise of the environment in which the target object is located. The setting of the first pause duration threshold for the first speech activity detection strategy and the second pause duration threshold for the second speech activity detection strategy based on the environmental noise of the environment in which the target object is located includes: when the current noise value of the environmental noise is higher than a preset threshold, setting a first duration corresponding to the current noise value as the first pause duration threshold, and setting a second duration corresponding to the current noise value as the second pause duration threshold, wherein the first duration is less than the second duration.
[0077] Through the above embodiments, by monitoring environmental noise and the speech rate of the target audience, the pause duration threshold can be dynamically adjusted to adapt to the speech recognition needs in different environments. This optimizes speech recognition performance in various usage scenarios, thereby improving the human-computer interaction experience and enhancing the practicality of smart devices.
[0078] To better understand the process of the above-described human-computer voice interaction method, the implementation flow of the above-described human-computer voice interaction method will be described below in conjunction with optional embodiments, but this is not intended to limit the technical solution of the embodiments of this application.
[0079] This embodiment provides a human-computer voice interaction method, the specific implementation steps of which are as follows:
[0080] Step 1: After submitting the voice interaction information to the cloud, perform ASR recognition. Taking the command to turn on the air conditioner as an example, the ASR recognition request configures two VAD thresholds, namely weak VAD and strong VAD.
[0081] Dynamically adjusting the VAD strategy to adapt to the speech recognition needs of different environments is key to improving the flexibility and robustness of voice interaction systems. The following examples illustrate how to dynamically adjust the VAD strategy based on the environment:
[0082] In a home environment with relatively low background noise, user interaction with the device is typically smooth, without significant pauses in speech. In this quiet setting, a weak VAD (Voice-Only Decision) strategy can be implemented for rapid voice recognition, reducing user wait time and improving the interactive experience. A weak VAD threshold is low, allowing for even short silence intervals before the speech is considered finished, thus immediately triggering NLP intent parsing for a rapid response.
[0083] In office environments, there may be more background noise, such as from printers and colleagues talking, and users may pause longer during commands. In such cases, a strong VAD (Voice Over-Driving) strategy is needed to more accurately determine the end of the voice and avoid accidentally truncating important voice information. A strong VAD has a higher threshold, allowing for longer pauses and ensuring that the user's complete commands are accurately captured even in noisy environments.
[0084] For users who speak quickly, the VAD strategy can be dynamically adjusted to lower the threshold of weak VAD, reduce waiting time, ensure timely response to user commands, and prevent users from feeling that the device is "sluggish".
[0085] Conversely, for users who speak slowly, the threshold for strong VAD can be increased to allow for longer silence times, ensuring that the device can fully receive the user's voice and avoid misinterpreting the command as the end of the voice.
[0086] In low-noise environments, a lower VAD threshold can be used to reduce waiting time and speed up the recognition process. In such environments, the risk of false positives is lower, therefore a weak VAD strategy can effectively improve the user experience.
[0087] In high-noise environments, the probability of misjudgment increases, so a higher VAD threshold should be used to increase the accuracy of the judgment. Through secondary verification with strong VAD, the end of the speech signal can be accurately identified in complex noise, ensuring recognition accuracy.
[0088] By dynamically adjusting the VAD (Voice over Action) strategy to suit the intonation, rhythm, and pause habits of different languages, it can better adapt to the speech recognition needs of different languages. For example, a higher VAD threshold can be used for languages with more pauses, while a lower VAD threshold can be set for speech with stronger continuity.
[0089] Through the above examples of dynamically adjusting the VAD strategy, it can be seen that flexibly adjusting the VAD threshold and strategy according to different environmental conditions and user habits can significantly improve the accuracy, response speed, and user experience of speech recognition, achieving more intelligent and efficient human-computer voice interaction. This dynamic adjustment mechanism is a key innovation of this application, enabling the speech recognition system to maintain optimal performance in various application scenarios.
[0090] Step 2: During the ASR recognition process, callbacks of intermediate recognition result text will be continuously received, for example... Figure 3 The phrase "open, open empty" in the text;
[0091] During ASR recognition, the system first receives an "end" flag indicating weak VAD detection has ended. Then, it waits for the callback of the next intermediate recognition result, for example, if the next callback's recognized text is: "Turn on the air conditioner." Upon receiving this recognition result, it immediately triggers subsequent asynchronous streaming NLP intent parsing.
[0092] Step 3: In the subsequent ASR recognition process, intermediate recognition results will be continuously received. For example, if an "end" flag for strong VAD is received, it indicates that the strong VAD detection has ended. Then, wait for the callback of the next intermediate recognition result. For example, the recognition text of the next callback is: "Turn on the air conditioner." Determine whether the current intermediate result is consistent with the intermediate recognition result after the previous "end" flag for weak VAD. If they are consistent, proceed to step 5; otherwise, proceed to step 4.
[0093] If the intermediate recognition results after the end of the strong and weak VADs do not match, the subsequent logic will be triggered again, such as asynchronously triggering streaming NLP intent parsing again, and proceeding to step 6.
[0094] Step 5: If the intermediate recognition results after the two VAD ends match, the subsequent logic will not be triggered asynchronously again. Wait for the final result of ASR recognition and proceed to step 6.
[0095] Step 6: After the ASR final recognition is completed, an ASR end marker will be received, along with the final ASR recognition text result, such as "Turn on the air conditioner." At this point, subsequent NLP intent parsing will be triggered, proceeding to Step 7.
[0096] Step 7: With ASR recognition completed in steps 1-6, the subsequent NLP intent parsing process supports idempotent operations. This means that after the streaming NLP call following both strong and weak VAD ends and the final NLP call following the ASR end, the consistency of the text semantics is determined. For example, "turn on the air conditioner" and "turn on the air conditioner bar" have the same semantics and will not trigger NLP intent parsing again. When the semantics of multiple NLP calls are consistent, the preprocessing result of the previous NLP call will be directly used for subsequent processing, such as issuing commands. Semantic parsing will only be performed again when the semantics are inconsistent.
[0097] In one embodiment, such as Figure 3 As shown, for example, the time difference between the end of a weak VAD and the end of a strong VAD is set to approximately 300ms, and the time difference between the end of a strong VAD and the final end of an ASR is approximately 100ms. The entire ASR speech recognition process can include: Step 1, triggering streaming NLP intent parsing 400ms in advance (the end of the weak VAD relative to the end of the ASR). Step 2, determining whether the recognized texts of the two VADs have the same semantics. If they do, proceed to Step 3. Otherwise, proceed to Step 4. Step 3, obtaining the recognized text. Further, in Step 3.1, storing the recognized text. Step 4, when the recognized texts of the ends of the two VADs are inconsistent, triggering the control logic of streaming NLP intent parsing again (triggering once 100ms in advance for the end of the strong VAD relative to the end of the ASR). Step 5, sending a response message based on the semantics corresponding to the recognized text.
[0098] In this embodiment, the triggering timing of these two streaming NLP operations is earlier than the end of the ASR. If the recognized text at the end of the two strong and weak VAD operations is semantically consistent with the recognized text at the end of the final ASR, the intent parsing of the subsequent NLP can be advanced by up to 400ms, which shortens the overall voice interaction time and improves the interactive experience.
[0099] Furthermore, subsequent NLP intent resolution is idempotent even with multiple calls within a single session. That is, if the NLP triggered by the end of a strong and weak VAD call and the end of an ASR call are semantically identical, the NLP intent resolution will not be triggered again; the previous semantic intent resolution will prevail.
[0100] It is evident that this application can trigger the subsequent NLP intent parsing process in parallel before the final ASR speech recognition result is obtained (i.e., at the end time points of the two strong and weak VADs), and trigger the NLP matching verification once after the final ASR recognition result is obtained; thus, the triggering time of the subsequent NLP intent parsing is effectively advanced, and the overall speech interaction time is effectively reduced.
[0101] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0102] Figure 4 This is a structural block diagram of a human-computer voice interaction device according to an embodiment of this application; as shown below. Figure 4 As shown, it includes:
[0103] The first acquisition module 42 is used to automatically recognize the voice interaction information of the target object and obtain a first intermediate recognition text and a second intermediate recognition text respectively. The first intermediate recognition text is obtained when the voice interaction information is automatically recognized using a first voice activity detection strategy, and the second intermediate recognition text is obtained when the voice interaction information is automatically recognized using a second voice activity detection strategy. The pause duration threshold set in the first voice activity detection strategy for detecting the moment when the target object stops speaking is lower than the pause duration threshold set in the second voice activity detection strategy for detecting the moment when the target object stops speaking.
[0104] The second acquisition module 44 is used to compare the first intermediate recognition text and the second intermediate recognition text until the second intermediate recognition text is consistent with the first intermediate recognition text, and to acquire the target recognition result of the voice interaction information based on the voice recognition end marker of the monitored voice interaction information.
[0105] The sending module 46 is used to perform natural language processing on the target recognition result to obtain the first interactive intent.
[0106] Using the aforementioned device, automatic speech recognition is performed on the voice interaction information of the target object to obtain a first intermediate recognition text and a second intermediate recognition text. The first intermediate recognition text is obtained when the voice interaction information is automatically recognized using a first speech activity detection strategy, and the second intermediate recognition text is obtained when the voice interaction information is automatically recognized using a second speech activity detection strategy. The pause duration threshold set in the first speech activity detection strategy for detecting when the target object stops speaking is lower than the pause duration threshold set in the second speech activity detection strategy for detecting when the target object stops speaking. The first intermediate recognition text and the second intermediate recognition text are compared until the second intermediate recognition text matches the first intermediate recognition text. Based on the detected speech recognition end marker of the voice interaction information, the target recognition result of the voice interaction information is obtained. Natural language processing is performed on the target recognition result to obtain a first interaction intent. By adopting the above technical solution and using two different pause duration thresholds for speech activity detection, automatic speech recognition can more accurately determine the end point of speech, reduce misidentification and missed recognition, solve the technical problem of insufficient accuracy of automatic speech recognition in human-computer voice interaction, and thus improve the accuracy and real-time performance of automatic speech recognition.
[0107] In an exemplary embodiment, the second acquisition module is further configured to: determine the recognition time of the automatic speech recognition, determine the recognition duration to which the recognition time belongs, acquire a first sub-text recognized within the recognition duration from the first intermediate recognized text, and acquire a second sub-text recognized within the recognition duration from the second intermediate recognized text; compare the first sub-text and the second sub-text to obtain a comparison result; and, for multiple comparison results corresponding to different recognition times, determine that the second intermediate recognized text is consistent with the first intermediate recognized text if any one of the multiple comparison results is used to indicate that the first sub-text and the second sub-text are consistent.
[0108] In an exemplary embodiment, the first acquisition module is further configured to: acquire audio words obtained by recognizing the voice interaction information, and obtain audio pause duration based on the time difference between different audio words; generate an end identifier of the first voice activity detection strategy when it is determined that the audio pause duration is less than the pause duration threshold of the first voice activity detection strategy; and acquire a first intermediate recognition text obtained by automatically recognizing the voice interaction information of the target object within a first duration based on the end identifier of the first voice activity detection strategy, wherein the start time of the first duration represents the time when automatic voice recognition begins, and the end time of the first duration represents the time when the end identifier of the first voice activity detection strategy is generated.
[0109] In an exemplary embodiment, the first acquisition module is further configured to: perform natural language processing on the first intermediate identified text to obtain a second interaction intent; and determine whether the first interaction intent matches the second interaction intent. Figure 1 In the event of a conflict, the response information corresponding to the first interaction intent will be sent to the target object.
[0110] In an exemplary embodiment, the first acquisition module is further configured to: acquire audio words obtained by recognizing the voice interaction information, and obtain audio pause duration based on the time difference between different audio words; generate an end identifier of the second voice activity detection strategy when it is determined that the audio pause duration is less than the pause duration threshold of the second voice activity detection strategy; and acquire second intermediate recognition text obtained by automatically recognizing the voice interaction information of the target object within a second duration based on the end identifier of the second voice activity detection strategy, wherein the start time of the second duration represents the time when automatic voice recognition begins, and the end time of the second duration represents the time when the end identifier of the second voice activity detection strategy is generated.
[0111] In an exemplary embodiment, the second acquisition module is further configured to: perform natural language processing on the first intermediate identification text to obtain a third interaction intent when the second intermediate identification text is inconsistent with the first intermediate identification text, and perform natural language processing on the target identification result to obtain a fourth interaction intent; and if it is determined that the third interaction intent and the fourth interaction intent match, and if it is determined that the fourth interaction intent and the first interaction intent also match, then send a reply message of the fourth interaction intent to the target object.
[0112] In an exemplary embodiment, the first acquisition module is further configured to: obtain the speech rate of the target object based on the voice interaction information, and set a first pause duration threshold of the first voice activity detection strategy and a second pause duration threshold of the second voice activity detection strategy based on the speech rate, wherein setting the first pause duration threshold of the first voice activity detection strategy and the second pause duration threshold of the second voice activity detection strategy based on the speech rate includes: when it is determined that the speech rate is lower than a preset speech rate, setting a third duration corresponding to the preset speech rate as the first pause duration threshold, and setting a fourth duration corresponding to the preset speech rate as the second pause duration threshold, wherein the third duration is less than the fourth duration; when it is determined that the speech rate is higher than the preset speech rate, determining an adjustment duration corresponding to the speech rate, setting the sum of the adjustment duration and the third duration as the first pause duration threshold, and setting the adjustment duration and the fourth duration as the second pause duration threshold.
[0113] In an exemplary embodiment, the first acquisition module is further configured to: set a first pause duration threshold of the first speech activity detection strategy and a second pause duration threshold of the second speech activity detection strategy based on the environmental noise of the environment in which the target object is located, specifically including: when the current noise value of the environmental noise is higher than a preset threshold, setting a first duration corresponding to the current noise value as the first pause duration threshold, and setting a second duration corresponding to the current noise value as the second pause duration threshold, wherein the first duration is less than the second duration.
[0114] Embodiments of this application also provide a storage medium including a stored program, wherein the program executes any of the methods described above when it is run.
[0115] Optionally, in this embodiment, the storage medium may be configured to store program code for performing the following steps:
[0116] S1, Automatic speech recognition is performed on the voice interaction information of the target object to obtain a first intermediate recognition text and a second intermediate recognition text respectively. The first intermediate recognition text is obtained when the voice interaction information is automatically recognized using a first speech activity detection strategy, and the second intermediate recognition text is obtained when the voice interaction information is automatically recognized using a second speech activity detection strategy. The pause duration threshold set in the first speech activity detection strategy for detecting the moment when the target object stops speaking is lower than the pause duration threshold set in the second speech activity detection strategy for detecting the moment when the target object stops speaking.
[0117] S2, compare the first intermediate recognition text and the second intermediate recognition text until the second intermediate recognition text matches the first intermediate recognition text, and obtain the target recognition result of the voice interaction information based on the voice recognition end marker of the monitored voice interaction information;
[0118] S3, perform natural language processing on the target recognition result to obtain the first interactive intent.
[0119] Embodiments of this application also provide an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0120] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0121] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0122] S1, Automatic speech recognition is performed on the voice interaction information of the target object to obtain a first intermediate recognition text and a second intermediate recognition text respectively. The first intermediate recognition text is obtained when the voice interaction information is automatically recognized using a first speech activity detection strategy, and the second intermediate recognition text is obtained when the voice interaction information is automatically recognized using a second speech activity detection strategy. The pause duration threshold set in the first speech activity detection strategy for detecting the moment when the target object stops speaking is lower than the pause duration threshold set in the second speech activity detection strategy for detecting the moment when the target object stops speaking.
[0123] S2, compare the first intermediate recognition text and the second intermediate recognition text until the second intermediate recognition text matches the first intermediate recognition text, and obtain the target recognition result of the voice interaction information based on the voice recognition end marker of the monitored voice interaction information;
[0124] S3, perform natural language processing on the target recognition result to obtain the first interactive intent.
[0125] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0126] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.
[0127] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0128] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method of human-to-computer speech interaction, the method comprising: include: Automatic speech recognition is performed on the voice interaction information of the target object to obtain a first intermediate recognition text and a second intermediate recognition text. The first intermediate recognition text is obtained when the voice interaction information is automatically recognized using a first speech activity detection strategy, and the second intermediate recognition text is obtained when the voice interaction information is automatically recognized using a second speech activity detection strategy. The pause duration threshold set in the first speech activity detection strategy for detecting the moment when the target object stops speaking is lower than the pause duration threshold set in the second speech activity detection strategy for detecting the moment when the target object stops speaking. Compare the first intermediate recognition text and the second intermediate recognition text until the second intermediate recognition text matches the first intermediate recognition text, and obtain the target recognition result of the voice interaction information based on the voice recognition end marker of the monitored voice interaction information; The target recognition result is processed by natural language to obtain the first interactive intent.
2. The human-to-computer voice interaction method of claim 1, wherein, Comparing the first intermediate identified text and the second intermediate identified text until the second intermediate identified text matches the first intermediate identified text includes: Determine the recognition time for the automatic speech recognition, determine the recognition duration to which the recognition time belongs, obtain the first sub-text recognized within the recognition duration from the first intermediate recognized text, and obtain the second sub-text recognized within the recognition duration from the second intermediate recognized text; Compare the first subtext and the second subtext to obtain the comparison result; For multiple comparison results corresponding to different recognition times, if any one of the multiple comparison results is determined to indicate that the first sub-text and the second sub-text are consistent, then the second intermediate recognition text is determined to be consistent with the first intermediate recognition text.
3. The human-computer voice interaction method according to claim 1, characterized in that, The first intermediate recognized text is obtained, including: The audio words obtained from recognizing the voice interaction information are acquired, and the audio pause duration is obtained based on the time difference between different audio words; If the audio pause duration is determined to be less than the pause duration threshold of the first speech activity detection strategy, an end identifier of the first speech activity detection strategy is generated. The first intermediate recognized text is obtained by automatically recognizing the voice interaction information of the target object within a first duration based on the end marker of the first voice activity detection strategy. The start time of the first duration represents the time when automatic voice recognition begins, and the end time of the first duration represents the time when the end marker of the first voice activity detection strategy is generated.
4. The human-computer voice interaction method according to claim 3, characterized in that, The method further includes: performing natural language processing on the first intermediate identified text to obtain a second interactive intent; If it is determined that the first interaction intent is consistent with the second interaction intent, the response information corresponding to the first interaction intent is sent to the target object.
5. The human-computer voice interaction method according to claim 1, characterized in that, The second intermediate recognized text is obtained, including: The audio words obtained from recognizing the voice interaction information are acquired, and the audio pause duration is obtained based on the time difference between different audio words; If the audio pause duration is determined to be less than the pause duration threshold of the second speech activity detection strategy, an end marker for the second speech activity detection strategy is generated. The second intermediate recognized text is obtained by automatically recognizing the voice interaction information of the target object within a second duration based on the end marker of the second voice activity detection strategy. The start time of the second duration represents the time when automatic voice recognition begins, and the end time of the second duration represents the time when the end marker of the second voice activity detection strategy is generated.
6. The human-computer voice interaction method according to claim 1, characterized in that, The method further includes: when the second intermediate identified text is inconsistent with the first intermediate identified text, performing natural language processing on the first intermediate identified text to obtain a third interactive intent, and performing natural language processing on the target identification result to obtain a fourth interactive intent; If it is determined that the third interaction intent and the fourth interaction intent match, and if it is determined that the fourth interaction intent and the first interaction intent also match, then the reply message of the fourth interaction intent is sent to the target object.
7. The human-computer voice interaction method according to claim 1, characterized in that, The method further includes: obtaining the speech rate of the target object based on the voice interaction information, and setting a first pause duration threshold of the first voice activity detection strategy and a second pause duration threshold of the second voice activity detection strategy based on the speech rate; The step of setting a first pause duration threshold for the first speech activity detection strategy and a second pause duration threshold for the second speech activity detection strategy based on the speech rate includes: If it is determined that the speech rate is lower than the preset speech rate, the third duration corresponding to the preset speech rate is set as the first pause duration threshold, and the fourth duration corresponding to the preset speech rate is set as the second pause duration threshold, wherein the third duration is less than the fourth duration; If the speech rate is determined to be higher than the preset speech rate, the adjustment duration corresponding to the speech rate is determined, the sum of the adjustment duration and the third duration is set as the first pause duration threshold, and the adjustment duration and the fourth duration are set as the second pause duration threshold.
8. The human-computer voice interaction method according to claim 1, characterized in that, The method further includes: setting a first pause duration threshold for the first speech activity detection strategy and a second pause duration threshold for the second speech activity detection strategy based on the environmental noise of the environment in which the target object is located; The step of setting the first pause duration threshold of the first speech activity detection strategy and the second pause duration threshold of the second speech activity detection strategy based on the environmental noise of the target object's environment includes: When the current noise value of the ambient noise is higher than a preset threshold, the first duration corresponding to the current noise value is set as the first pause duration threshold, and the second duration corresponding to the current noise value is set as the second pause duration threshold, wherein the first duration is less than the second duration.
9. A human-computer voice interaction device, characterized in that, include: The first acquisition module is used to automatically recognize the voice interaction information of the target object and obtain a first intermediate recognition text and a second intermediate recognition text respectively. The first intermediate recognition text is obtained when the voice interaction information is automatically recognized using a first voice activity detection strategy, and the second intermediate recognition text is obtained when the voice interaction information is automatically recognized using a second voice activity detection strategy. The pause duration threshold set in the first voice activity detection strategy for detecting the moment when the target object stops speaking is lower than the pause duration threshold set in the second voice activity detection strategy for detecting the moment when the target object stops speaking. The second acquisition module is used to compare the first intermediate recognition text and the second intermediate recognition text until the second intermediate recognition text is consistent with the first intermediate recognition text, and to acquire the target recognition result of the voice interaction information based on the voice recognition end marker of the monitored voice interaction information. The sending module is used to perform natural language processing on the target recognition result to obtain the first interaction intent.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method described in any one of claims 1 to 8.