A voice wake-up interactive response method and system

By dividing the wake-up time window into multiple regions and employing different detection methods and techniques, the problem of poor detection performance in voice wake-up interaction was solved, achieving highly accurate user speech detection and improving user experience.

CN119541487BActive Publication Date: 2025-11-14PACHIRA TIMES (ZHUHAI HENGQIN) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411699261.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-11-14
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

In existing technologies, the methods for detecting whether a user has spoken immediately during voice wake-up interaction are easily affected by the ending sound of the wake-up word, and are prone to missed detection in the latter half of the 400ms time window, resulting in poor detection performance and easy to cause missed or false detection.

Method used

The time window after wake-up is divided into false trigger zone, normal zone and blind zone, and different technical means are used for processing, including false trigger detection, normal speech detection and blind zone base frequency calculation. Combined with wake-up model and correlation analysis, it is determined whether the user has spoken.

Benefits of technology

It effectively reduces the proportion of false detections and missed detections. In actual applications, the false detection rate is less than 0.01%, which improves the accuracy of voice wake-up interaction and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119541487B_ABST
    Figure CN119541487B_ABST
Patent Text Reader

Abstract

A voice wake-up interactive response method and system are disclosed, primarily used to detect whether a user has spoken after a voice interaction system has been woken up. The given time window after wake-up is divided into different regions, and different techniques are used to process each region. Specifically, in the false trigger region immediately following the wake-up word detection, the system detects whether the user has spoken and determines whether the spoken words are confused with the ending sound of the wake-up word. In the normal speech detection region following the false trigger region, the system detects whether the user has spoken. In the blind zone detection region following the normal speech detection region, the system detects whether the user has spoken. If no spoken words are detected in the blind zone detection region, the system calculates the fundamental frequency. If the fundamental frequency calculation indicates the presence of spoken words for a certain duration, then spoken words are considered detected. Using the method and system of this invention can significantly reduce the proportion of false detections and false negatives.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent voice interaction technology, and specifically relates to a voice wake-up interactive response method and system. Background Technology

[0002] With the development of society and electronic information technology, artificial intelligence products have become indispensable necessities in people's lives, such as smart speakers, smart cars, smart TVs, and smart air conditioners. Many other devices and applications are also rapidly developing towards intelligence. At the same time, users have diverse functional requirements for artificial intelligence products, demanding not only basic functions but also audio-visual entertainment, life services, and interconnectivity.

[0003] With the proliferation of smart applications and the development of AI technology, consumers' understanding of smart products is constantly evolving, and their expectations for product experience are shifting from single-function to more intelligent scenarios. This necessitates a redefinition of many product functions. Simultaneously, significant advancements in technologies such as artificial intelligence, 5G, human-computer interaction devices, and operating systems are driving the rapid development of smart applications to meet users' ever-increasing awareness and needs.

[0004] In the interaction between smart voice devices and humans, voice wake-up is a fundamental capability used to initiate interaction with the smart device. The purpose of voice wake-up is to activate the smart device from a dormant state to an active state, so the wake word should be detected immediately after it is spoken for a better user experience. Apple's "Hey Siri" and Baidu's "Xiaodu Xiaodu" are existing wake words.

[0005] Normally, after a user is activated, the system will respond with a message such as "I'm here." However, if the user speaks immediately after being activated, it may cause the user's voice to overlap with the system's announcement, thus affecting the interactive experience and effectiveness.

[0006] In existing technologies, one form of interaction involves the system detecting whether the user immediately speaks a voice command after the system is activated by voice, i.e., a "one-time" interaction. After activation, the system checks whether the user has any subsequent immediate speech, using this as a basis to decide whether to play a welcome message. If the system detects that the user has spoken immediately, it will not play a welcome message or similar response to avoid interfering with the user's input and affecting the experience and effectiveness. The above-mentioned voice activation interaction process in existing technologies is shown in the attached diagram of the specification. Figure 1 As shown.

[0007] According to the attached drawings in the instruction manual Figure 1The common practice in existing technologies is to detect whether the user has spoken within a given time window (e.g., 400ms) after wake-up, using voice endpoint detection. This can lead to two problems: first, the detection process is easily affected by the ending sound of the wake-up word; second, the latter half of the 400ms period is prone to missed detections. This results in poor detection performance and a high risk of missed or false detections. Summary of the Invention

[0008] To address the aforementioned deficiencies in existing technologies, this invention proposes a novel method for detecting whether a user has spoken after being woken up. This method divides a given time window after wake-up into different regions and employs different technical means to process each region, thereby improving the interactive response of voice wake-up.

[0009] This invention provides a voice wake-up interactive response method, comprising the following steps:

[0010] Get user speech;

[0011] Detect whether the user's speech contains a wake word;

[0012] After detecting the wake word, check if a user is speaking within the detection area following the wake word;

[0013] The detection within the detection area includes detecting whether a user is speaking within the normal speaking detection area;

[0014] If no user speech is detected within the detection area, a welcome message will be played.

[0015] The detection within the detection area also includes:

[0016] Within the false trigger area of ​​the time zone immediately following the wake word detection, it is detected whether a user has spoken. If a speech is detected, it is further determined whether the speech is confused with the ending sound of the wake word; if there is confusion, it is considered that no speech has been detected.

[0017] and / or

[0018] The system detects whether a user is speaking in the blind zone detection area after the normal speaking detection area. If no speaking content is detected in the blind zone detection area, the system calculates the base frequency for the blind zone detection area. If the base frequency calculation shows that there is a speaking session of a certain duration, then the system considers that a speaking session has been detected.

[0019] Furthermore, a wake-up model is used to detect user speech within the detection area.

[0020] Furthermore, a correlation analysis method is used to determine whether the utterance is confused with the ending sound of the wake-up word.

[0021] This invention provides a voice wake-up interactive response system, the system comprising:

[0022] The acquisition module is used to acquire user messages.

[0023] The wake-up word detection module is used to detect whether a wake-up word exists in the user's speech acquired by the acquisition unit;

[0024] The user speech detection module is used to detect the user's second speech after the wake word;

[0025] The user speech detection module includes a normal speech detection module, and also includes a false trigger detection module and / or a blind spot detection module;

[0026] The false trigger detection module is used to detect whether the second speech exists within a certain period of time immediately following the wake word. If the second speech is detected, it is necessary to further determine whether the second speech is confused with the ending sound of the wake word.

[0027] The normal speech detection module is used to detect whether the second speech exists within a certain period of time after the wake-up word or after the false trigger detection.

[0028] The blind zone detection module is used to detect whether the second voice call exists within a certain period of time after the normal voice call detection. If the second voice call is still not detected, the baseband detection is performed during the time period to determine whether the user voice call exists.

[0029] The voice output module is used to output a welcome message when the user speech detection module does not detect the second speech after the system is woken up.

[0030] Furthermore, the user speech detection module uses a wake-up model for detection when detecting the second speech.

[0031] Furthermore, the false trigger detection module uses a correlation analysis method to determine whether the second utterance is confused with the ending sound of the wake-up word.

[0032] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method.

[0033] This invention provides an electronic device, comprising:

[0034] One or more processors; and

[0035] A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method.

[0036] Compared with the prior art, the present invention has the following advantages and positive effects: The method provided by the present invention can effectively improve the problems existing in the prior art, use actual data in the normal interaction process for verification, the false detection rate is less than 0.01%, and no false detection occurs in the actual process. Attached Figure Description

[0037] To gain a more complete understanding of the invention, reference will now be made to the following description taken in conjunction with the accompanying drawings, wherein:

[0038] Figure 1 This is a schematic diagram of the voice wake-up interaction process in the existing technology;

[0039] Figure 2 A schematic diagram of the interactive response method for voice wake-up provided by the present invention;

[0040] Figure 3 This is a schematic diagram illustrating the division of the user speech detection area according to the present invention;

[0041] Figure 4 A schematic diagram of the interactive response system for voice wake-up provided by the present invention. Detailed Implementation

[0042] To make the problems to be solved, the technical solutions, and the technical effects of this invention clearer, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments and methods obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0043] Before describing exemplary embodiments of the present invention in more detail, it should be noted that although the flowcharts of the present invention describe the operations as sequential processes, many of the operation steps can be performed in parallel or simultaneously, and the order of the operations can be rearranged as needed. Furthermore, the process can be terminated when an operation is completed, but additional steps not included in the drawings may also be included.

[0044] In the following description, the terms "first", "second", "third", etc. only distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first", "second", "third", etc. can be interchanged in a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0045] Embodiment 1

[0046] The present invention provides an interactive response method for voice wake-up. As shown in the accompanying drawings of the specification Figure 2 shown, mainly including the following steps:

[0047] Step S01: Obtain the first utterance of the user and detect whether the first utterance of the user contains a wake-up word.

[0048] For example, the wake-up word is "Hello, Xiaobing"; collect the user's audio and detect the similarity between the user's audio and the wake-up word "Hello, Xiaobing". If the similarity is greater than the set threshold, it is considered that the first utterance of the user contains the wake-up word.

[0049] Step S02: Detect the second utterance of the user after the wake-up word in the detection area. Divide the detection area into a false trigger area, a normal area, and a blind area, as shown in the accompanying drawings of the specification Figure 3 shown, and perform the following steps:

[0050] Step S201: Detect the user's utterance in the false trigger area after wake-up. If an utterance is detected, further determine whether the utterance is confused with the ending sound of the wake-up word; if there is confusion, it is considered that no utterance is detected.

[0051] The false trigger area is the time area immediately adjacent to the wake-up word detection area, for example, within a duration of 100 ms.

[0052] When detecting an utterance in the false trigger area, it is necessary to determine whether the current utterance is caused by the ending sound of the wake-up word. For example, the wake-up word is "Hello, Xiaobing". If the pronunciation of the word "bing" is long and the wake-up is triggered before "bing" is finished, the ending sound of "bing" is likely to be falsely triggered and misinterpreted as a new voice command. Therefore, after detecting the wake-up word, in the false trigger area after wake-up, for example, within a duration of 100 ms, detect whether the content of the utterance is confused with "bing".

[0053] When detecting whether the voice command is confused with the ending sound of the wake-up word, a correlation analysis method can be used to analyze whether the detected content is correlated with the ending sound of the wake-up word.

[0054] When performing correlation analysis, a confusion table that records the confusion or similarity relationships of each speech content can be used to effectively solve this type of problem.

[0055] Step S202: In the normal area following the false triggering area, detect user speech.

[0056] Step S203: Detect speech in the blind zone region after the normal region. If no speech content is detected in the blind zone region, perform baseband calculation on the blind zone region. If the baseband calculation can determine that there is speech for a certain duration, it is also considered that there is speech after wake-up.

[0057] Following the normal region, a blind zone is defined, for example, with a duration of 120ms. Since the blind zone is the final segment of the second call detection, the final result is given immediately after detection and processing. Therefore, excessive delay is not allowed in the blind zone. In other words, a short detection time in the blind zone may prevent timely detection of user calls. To address this, low-latency techniques, such as F0 (baseband), can be used for auxiliary judgment.

[0058] When a person speaks, the sound is divided into voiced and unvoiced sounds. Voiced sounds are produced by the vibration of the vocal cords and have a distinct periodicity. All voiced sounds possess a fundamental frequency that can be extracted. The basic principle of fundamental frequency detection is to utilize the periodicity of voiced sounds. By estimating the period of the voiced sound signal, the fundamental frequency, i.e., the frequency of vocal cord vibration, is obtained. The fundamental frequency of the current audio is extracted, and this is used to determine whether speech has occurred. If the fundamental frequency calculation shows a certain continuous speech duration, such as 50ms, then speech is also considered to have occurred after wake-up.

[0059] In step S02, the detection area can also be divided into two areas, such as the tail sound detection area (false trigger area) and the real-time speech area (normal area), and the corresponding detection steps can be performed.

[0060] Speech content detection typically uses wake-up models or Automatic Speech Recognition (ASR) models. This invention directly uses a wake-up model as the content detection technique. It leverages the different characteristics of the false-trigger region, normal region, and blind zone region, sharing these techniques with other detection methods. For example, the false-trigger region contains normal interference, requiring further determination of whether it was falsely triggered by a wake-up word after detection, in conjunction with the model's detection results; the blind zone region, after detection using the wake-up model, requires secondary confirmation using low-latency technology.

[0061] Step S03: Based on the detection results in step S02, determine whether there is a user speaking within the detection area. If no user speaking is detected, play a welcome message; if a user speaking is detected, perform an interactive response based on the user speaking.

[0062] For example, if a user says, "Hello Xiaoyi, please open the car window for me," and the system detects the user's message "please open the car window for me" within the detection area following the wake word "Hello Xiaoyi," then the system will interact with the user based on that message. If the user says, "Hello Xiaoyi <pause 600ms> please open the car window for me," and no user message is detected within the detection area following the wake word "Hello Xiaoyi," then a welcome message will be played.

[0063] Example 2

[0064] Corresponding to the method described in this invention, this invention provides a voice wake-up interactive response system, see [link to relevant documentation]. Figure 4 The system includes:

[0065] Module 100 is used to acquire user messages.

[0066] The wake-up word detection module 200 is used to detect whether a wake-up word exists in the user speech acquired by the acquisition unit 101.

[0067] The user speech detection module 300 is used to detect the user's second speech after the wake word.

[0068] The user speech detection module 300 includes a false trigger detection module 301, a normal speech detection module 302, and a blind spot detection module 303.

[0069] The false trigger detection module 301 is used to detect whether the second utterance exists within a certain period of time after the wake-up word. If the second utterance is detected, it is necessary to further determine whether the second utterance is confused with the ending sound of the wake-up word.

[0070] The normal speech detection module 302 is used to detect whether the second speech exists within a certain period of time after the wake-up word or after the false trigger detection.

[0071] The blind spot detection module 303 is used to detect whether the second voice call exists within a certain period of time after the normal voice call detection. If the second voice call is still not detected, the baseband detection is performed during that period of time to further determine whether the user voice call exists.

[0072] The voice output module 400 is used to output a welcome message when no user speech is detected after the system is woken up.

[0073] In addition, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the foregoing method embodiments.

[0074] The present invention also provides an electronic device, comprising:

[0075] One or more processors; and

[0076] A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method described in the method embodiment of the preceding embodiment one.

[0077] It should be noted that a portion of this invention can be applied as a computer program product, such as computer program instructions. When executed by a smart electronic device (such as a smartphone or tablet computer), these instructions can invoke or provide the methods and / or technical solutions according to this invention through the operation of the smart electronic device. The program instructions invoking the methods of this invention may be stored in a fixed or removable recording medium, and / or transmitted via data streams in broadcast or other signal-carrying media, and / or stored in the working memory of the smart electronic device operating according to the program instructions. Here, one embodiment of the invention includes an apparatus comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the apparatus is triggered to operate the methods and / or technical solutions based on the foregoing embodiments of the invention.

[0078] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the description of the embodiments above. Therefore, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in the system claims may also be implemented by a single unit or device in software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any particular order.

Claims

1. A voice wake-up interactive response method, comprising the following steps: Get user speech; Detect whether the user's speech contains a wake word; After detecting the wake word, check if a user is speaking within the detection area following the wake word; The detection within the detection area includes detecting whether a user is speaking within the normal speaking detection area; If no user speech is detected within the detection area, a welcome message will be played. The detection within the detection area also includes: Within the false trigger area of ​​the time zone immediately following the wake word detection, it is detected whether a user has spoken. If a speech is detected, it is further determined whether the speech is confused with the ending sound of the wake word; if there is confusion, it is considered that no speech has been detected. and / or The system detects whether a user is speaking in the blind zone detection area after the normal speaking detection area. If no speaking content is detected in the blind zone detection area, the system calculates the base frequency for the blind zone detection area. If the base frequency calculation shows that there is a speaking session of a certain duration, then the system considers that a speaking session has been detected.

2. The method according to claim 1, characterized in that, The detection of user speech within the detection area employs a wake-up model.

3. The method according to claim 1, characterized in that, Correlation analysis was used to determine whether the spoken words were confused with the ending sound of the wake-up word.

4. A voice-activated interactive response system, the system comprising: The acquisition module is used to acquire user messages; The wake-up word detection module is used to detect whether a wake-up word exists in the user's speech acquired by the acquisition module; The user speech detection module is used to detect the user's second speech after the wake word; The user speech detection module includes a normal speech detection module, and also includes a false trigger detection module and / or a blind spot detection module; The false trigger detection module is used to detect whether the second speech exists within a certain period of time immediately following the wake word. If the second speech is detected, it is necessary to further determine whether the second speech is confused with the ending sound of the wake word. The normal speech detection module is used to detect whether the second speech exists within a certain period of time after the wake-up word or after the false trigger detection. The blind zone detection module is used to detect whether the second voice call exists within a certain period of time after the normal voice call detection. If the second voice call is still not detected, the baseband detection is performed during the time period to determine whether the user voice call exists. The voice output module is used to output a welcome message when the user speech detection module does not detect the second speech after the system is woken up.

5. The system according to claim 4, characterized in that, When detecting the second user speech, the user speech detection module uses a wake-up model for detection.

6. The system according to claim 4, characterized in that, The false trigger detection module uses a correlation analysis method to determine whether the second utterance is confused with the ending sound of the wake-up word.

7. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method according to any one of claims 1-3.

8. An electronic device, comprising: One or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Voice wake-up method, electronic equipment and storage medium

    CN114155857A

  • Wake-up word recognition method and device combined with dynamic time warping, equipment and medium

    CN118173094A