A voice wake-up method, device and system with dual protection mechanism

The voice wake-up technology with a dual protection mechanism utilizes sound source localization and timestamp arbitration to solve the problem of false triggering of voice wake-up under background noise, improves the accuracy of wake-up word recognition and system security, and reduces the risk of false wake-up.

CN119207396BActive Publication Date: 2025-10-03FIBERHOME TELECOMMUNICATION TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411127758.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2025-10-03
Estimated Expiration
2044-08-16

AI Technical Summary

Technical Problem

Existing voice wake-up technologies have difficulty distinguishing between real people speaking and device playback when background noise is a strong interference source, leading to false wake-ups. This is especially prone to flooding attacks in live broadcast scenarios.

Method used

A dual protection mechanism is adopted, through the collaborative work of the voice wake-up device and the voice service server, using sound source positioning and timestamp arbitration to identify and filter interfering audio, confirm the authenticity of the wake-up word, and avoid false triggering.

Benefits of technology

Effectively reduce the interference of external strong sound source devices on voice wake-up devices, reduce false triggering of wake-up words, reduce system response costs, and avoid flooding attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119207396B_ABST
    Figure CN119207396B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of voice wake-up technology, and provides a voice wake-up method, device, and system with a dual protection mechanism, including: a playback device superimposes interference audio onto original audio to obtain superimposed audio, then performs wake-up word recognition on the original audio or superimposed audio, obtains a wake-up word interference mark, and sends it to a voice service server; a voice wake-up device monitors ambient sound, performs noise processing on sounds carrying interference audio to obtain pure audio, then performs wake-up word recognition on the pure audio, and when a wake-up word is recognized, generates a wake-up word arbitration and sends it to a voice service server; the voice service server matches the wake-up word interference mark with the wake-up word arbitration, detects whether there is a false triggering of the wake-up word, and generates a detection result, and the voice wake-up device executes a wake-up word response or cancels the wake-up word response based on the detection result. The present invention can reduce interference from external strong sound source devices and avoid false triggering of wake-up words.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice wake-up technology, and in particular to a voice wake-up method, device, and system with a dual protection mechanism. Background Art

[0002] Voice wake-up is a human-computer interaction technology that allows users to activate smart devices by speaking a preset wake-up word, switching them from standby or dormant mode to active mode, ready to receive and execute further voice commands. This technology is widely used in various smart devices, including smartphones, smart speakers, and laptops, improving user convenience and user experience. Existing voice wake-up technologies primarily rely on microphone arrays to receive speech, recognize the wake-up word, and perform corresponding functions.

[0003] Because the accuracy and interaction efficiency of speech recognition are limited or affected by the performance of microphone sensors, ambient noise, and the quality of algorithms, existing technologies use multiple microphones arranged in a specific layout to synchronously collect sound. By analyzing the subtle time differences between the microphone signals, beamforming and spatial filtering are implemented to construct directional beams that enhance the speech signal in the target direction while suppressing noise, improving speech recognition. Some existing technologies even incorporate adaptive algorithms and deep learning techniques to dynamically optimize processing effects, enhancing speech recognition in noisy environments.

[0004] The existing technology has the following technical problems:

[0005] 1. When background noise is strong, such as from a TV or other audio equipment, the voice wake-up device's sound source localization is easily directed toward the strong sound source, making it unable to identify the actual speaker in a timely manner. Especially when playing human voices, the microphone array is easily directed toward the TV and mistakenly woken up by keywords in the program, making it difficult to distinguish between real people speaking and the device itself.

[0006] 2. In a live broadcast scenario, when the presenter speaks the wake-up word, it is easy to wake up a large number of user devices at the same time to access the voice service provider's server, causing a flood attack. Summary of the Invention

[0007] The present invention provides a voice wake-up method, device, and system with a dual protection mechanism, aiming to solve the technical problems in the above-mentioned prior art, such as the difficulty in distinguishing between real people speaking and device playback when background noise is a strong interference sound source, and the problem that in live broadcast scenarios, when the presenter speaks the wake-up word, it is easy to wake up a large number of user devices at the same time to access the voice service provider's server, causing a flood attack.

[0008] The technical solution of the present invention to solve the above technical problems is as follows:

[0009] In a first aspect, the present invention provides a voice wake-up method with a dual protection mechanism, which is used in a voice wake-up device. The method includes:

[0010] Monitor ambient sound, identify interference audio from the ambient sound, and perform noise processing on the sound carrying the interference audio to obtain pure audio;

[0011] Performing wake-up word recognition on the clean audio, generating a wake-up word arbitration when a wake-up word is recognized and sending it to the voice service server, wherein the wake-up word arbitration is used to trigger the voice service server to send a detection result to the voice wake-up device after detecting whether the wake-up word is a wake-up word false trigger;

[0012] The detection result is received, and a wake-up word response is performed or a wake-up word response is canceled based on the detection result.

[0013] Furthermore, the interference audio is pre-set in the voice wake-up device, or received from the voice service server.

[0014] Furthermore, obtaining the pure audio specifically includes:

[0015] Continuously monitor ambient sounds and identify the source direction of various sounds through sound source localization;

[0016] When interference audio is monitored, the direction of the interference sound source of the interference audio is obtained, the sound source in the direction of the interference sound source is used as noise interference, and the voice wake-up device is used to continue monitoring the ambient sounds in other directions to obtain pure audio.

[0017] Furthermore, the above-mentioned wake-up word arbitration includes a second wake-up word identifier and a second timestamp, and the second wake-up word identifier is a wake-up word identifier used to activate the voice wake-up device.

[0018] In a second aspect, the present invention provides a voice wake-up method with a dual protection mechanism, which is used in a voice service server, and the method includes:

[0019] Receive a wake-up word interference mark sent by a playback device and a wake-up word arbitration sent by a voice wake-up device, match the wake-up word interference mark with the wake-up word arbitration, detect whether there is a wake-up word false trigger, and generate a detection result, where the wake-up word arbitration is generated when the voice wake-up device recognizes a wake-up word, and the wake-up word interference mark is generated when the voice recognition module built into the playback device recognizes a wake-up word;

[0020] The detection result is sent to the voice wake-up device, where the detection result is used to trigger the voice wake-up device to execute a wake-up word response or cancel a wake-up word response.

[0021] Furthermore, the above also includes:

[0022] Sending interference audio to the playback device, where the interference audio is used to superimpose the original audio of the playback device to obtain superimposed audio;

[0023] Interference audio is sent to the voice wake-up device, where the interference audio is used for filtering the superimposed audio when the voice recognition device recognizes the superimposed audio through encoding of the interference audio.

[0024] Furthermore, the wake-up word interference mark includes a first wake-up word identifier and a first timestamp, the wake-up word arbitration includes a second wake-up word identifier and a second timestamp, and detecting whether there is a wake-up word false trigger specifically includes:

[0025] Determining whether the first wake-up word identifier and the second wake-up word identifier are the same wake-up word;

[0026] If yes, calculate the time difference between the first timestamp and the second timestamp, and determine whether the time difference is within the statistical confidence interval;

[0027] If so, a detection result indicating the presence of a false trigger is generated; if not, a detection result indicating the absence of a false trigger is generated.

[0028] Furthermore, the calculation formula for the statistical confidence interval is as follows:

[0029]

[0030] Among them, z represents a constant, n represents the number of statistics, and w i represents the observed value of the i-th statistic, Represents the mean of all observations.

[0031] In a third aspect, the present invention provides a voice wake-up method with a dual protection mechanism, which is used in a playback device, and the method includes:

[0032] Superimposing the interference audio onto the original audio to obtain superimposed audio;

[0033] Performing wake-up word recognition on the original audio or the superimposed audio to obtain a wake-up word interference mark;

[0034] A wake-up word interference mark is sent to the voice service server, and the wake-up word interference mark is used by the voice service server to match the wake-up word interference mark with the wake-up word arbitration, detect whether there is a wake-up word false trigger, and generate a detection result. The detection result is used to trigger the voice wake-up device to execute a wake-up word response or cancel the wake-up word response. The wake-up word arbitration is generated when the voice wake-up device recognizes the wake-up word.

[0035] Furthermore, the above-mentioned wake-up word interference mark includes a first wake-up word identifier and a first timestamp.

[0036] Furthermore, the interference audio is pre-set in the playback device, or received from the voice service server.

[0037] Furthermore, the above also includes:

[0038] The superimposed audio is played, and the superimposed audio is used to identify and filter the superimposed audio through the interference audio when the voice wake-up device performs noise filtering.

[0039] In a fourth aspect, the present invention provides a voice wake-up device with a dual protection mechanism, which is used in a voice wake-up device, and the device includes:

[0040] A signal processing module is used to monitor ambient sound, identify interference audio from the ambient sound, and perform noise processing on the sound carrying the interference audio to obtain pure audio;

[0041] A second wake-up word recognition module is configured to perform wake-up word recognition on the clean audio, generate a wake-up word arbitration when a wake-up word is recognized, and send the generated wake-up word arbitration to the voice service server, wherein the wake-up word arbitration is configured to trigger the voice service server to send a detection result to the voice wake-up device after detecting whether the wake-up word is a false trigger of the wake-up word;

[0042] The wake-up word response module is used to receive the detection result and execute the wake-up word response or cancel the wake-up word response based on the detection result.

[0043] In a fifth aspect, the present invention provides a voice wake-up device with a dual protection mechanism, which is used in a voice service server, and the device includes:

[0044] a false trigger detection module, configured to receive a wake-up word interference mark sent by a playback device and a wake-up word arbitration mark sent by a voice wake-up device, match the wake-up word interference mark with the wake-up word arbitration mark, detect whether there is a wake-up word false trigger, and generate a detection result, wherein the wake-up word arbitration mark is generated when the voice wake-up device recognizes a wake-up word, and the wake-up word interference mark is generated when the voice recognition module built into the playback device recognizes a wake-up word;

[0045] The first sending module is used to send a detection result to the voice wake-up device, where the detection result is used to trigger the voice wake-up device to execute a wake-up word response or cancel a wake-up word response.

[0046] In a sixth aspect, the present invention provides a voice wake-up device with a dual protection mechanism, for use in a playback device, the device comprising:

[0047] A superposition module, used for superimposing the interference audio onto the original audio to obtain superimposed audio;

[0048] a first wake-up word recognition module, configured to perform wake-up word recognition on the original audio or the superimposed audio to obtain a wake-up word interference mark;

[0049] The second sending module is used to send a wake-up word interference mark to the voice service server. The wake-up word interference mark is used by the voice service server to match the wake-up word interference mark with the wake-up word arbitration, detect whether there is a wake-up word false trigger, and generate a detection result. The detection result is used to trigger the voice wake-up device to execute a wake-up word response or cancel the wake-up word response. The wake-up word arbitration is generated when the voice wake-up device recognizes the wake-up word.

[0050] In a seventh aspect, the present invention provides a voice wake-up system with a dual protection mechanism, the system comprising a voice wake-up device, a voice service server, and a playback device;

[0051] The voice wake-up device includes the voice wake-up device with a dual protection mechanism as described in the fourth aspect, the voice service server includes the voice wake-up device with a dual protection mechanism as described in the fifth aspect, and the playback device includes the voice wake-up device with a dual protection mechanism as described in the sixth aspect. Compared with the prior art, the present invention has the following advantages:

[0052] 1. Effectively reduce the interference of strong external sound source devices on the microphone array of the voice wake-up device;

[0053] 2. Effectively reduce false trigger interference caused by wake-up words in audio streams played by external devices;

[0054] 3. The solution has low implementation cost and does not require increasing algorithm complexity or hardware costs.

[0055] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0057] Figure 1 A schematic flow chart of a voice wake-up method with a dual protection mechanism according to an embodiment of the present invention is shown;

[0058] Figure 2 A schematic flow chart showing a voice wake-up method with a dual protection mechanism according to another embodiment of the present invention is shown;

[0059] Figure 3 A schematic diagram of a voice wake-up device according to an embodiment of the present invention is shown;

[0060] Figure 4 A schematic diagram of a voice service server according to an embodiment of the present invention is shown;

[0061] Figure 5 A schematic diagram of a playback device according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0063] The implementation environment involved in each embodiment of the present invention includes: a voice service server, a playback device, and a voice wake-up device.

[0064] The voice service server can be a single server, a server cluster consisting of several servers, or a cloud computing service center.

[0065] Playback devices can be portable music players, smart speakers, Bluetooth speakers, car audio systems, home theater systems, and TVs.

[0066] Voice-activated devices can be smart assistants, smartphones, car navigation systems, smart home controllers, and wearable devices.

[0067] Figure 1 FIG. 4 shows a flow chart of a voice wake-up method with a dual protection mechanism according to an embodiment of the present invention. Figure 1 As shown, the voice wake-up method of an embodiment of the present invention includes:

[0068] S1: The playback device recognizes the wake-up word on the original audio, obtains a wake-up word interference mark, and sends the wake-up word interference mark to the voice service server;

[0069] S2: The playback device superimposes the interference audio onto the original audio to obtain superimposed audio;

[0070] S3: The voice wake-up device monitors the ambient sound, identifies the interference audio of the ambient sound, and performs noise processing on the sound carrying the interference audio to obtain pure audio;

[0071] S4: The voice wake-up device recognizes the wake-up word on the clean audio, generates a wake-up word arbitration when a wake-up word is recognized, and sends it to the voice service server;

[0072] S5: The voice service server receives the wake-up word interference mark and the wake-up word arbitration, matches the wake-up word interference mark with the wake-up word arbitration, detects whether there is a wake-up word false trigger, generates a detection result, and sends the detection result to the voice wake-up device;

[0073] S6: The voice wake-up device receives the detection result and executes a wake-up word response or cancels the wake-up word response based on the detection result.

[0074] In this embodiment, the voice wake-up device first localizes the sound source using a multi-microphone array, using the signal's time difference of arrival (TDOA) and intensity differences to determine the direction of the sound. This process not only helps focus on the target speech but also serves as a foundation for identifying interference sources. Interference audio is encoded using specific frequency coding rules, which the audio signal processing unit in the voice wake-up device uses to identify the interference audio. When the voice wake-up device detects an audio signal carrying interference audio, it immediately identifies the input from the relevant direction as noise interference and reduces or eliminates it in subsequent processing, while prioritizing clean speech signals from other directions. Even if noise interference penetrates the front-end defenses, the voice wake-up device still implements a second line of defense. Before confirming the wake-up word response, the voice wake-up device sends an arbitration request to the voice service server to further verify the authenticity of the wake-up event. The voice service server compares the second timestamp with the first timestamp to obtain the time difference and analyzes the time difference using statistical confidence intervals. If the two timestamps are closely aligned, it indicates that the two recognitions may have originated from the same noise event. The voice service server deems this a false trigger and promptly notifies the voice wake-up device to cancel the wake-up action, avoiding unnecessary system responses.

[0075] Optionally, the wake-up word interference mark includes a first wake-up word identifier and a first timestamp, the wake-up word arbitration includes a second wake-up word identifier and a second timestamp, and the voice service server detects whether there is a wake-up word false trigger specifically including:

[0076] Determining whether the first wake-up word identifier and the second wake-up word identifier are the same wake-up word;

[0077] If yes, calculate the time difference between the first timestamp and the second timestamp, and determine whether the time difference is within the statistical confidence interval;

[0078] If yes, a detection result indicating false triggering is generated; if no, a detection result indicating no false triggering is generated;

[0079] Optionally, the calculation method of the statistical confidence interval includes: a confidence interval based on a normal distribution, a confidence interval based on a t distribution, and a confidence interval of a proportion;

[0080] When the number of statistics is large enough (n>30), the statistical confidence interval adopts the confidence interval based on the normal distribution. The calculation formula of the statistical confidence interval is shown as follows:

[0081]

[0082] Among them, z represents a constant, n represents the number of statistics, and w i represents the observed value of the i-th statistic, Represents the mean of all observations.

[0083] Among them, when the number of statistics is small (n≤30), the statistical confidence interval adopts the confidence interval based on t distribution;

[0084] Since the detection results are two types: the time difference is within the statistical confidence interval and the time difference is not within the statistical confidence interval, the statistical confidence interval can also use the proportion confidence interval.

[0085] Optionally, the interference audio is pre-set in the voice wake-up device, or received from the voice service server.

[0086] Optionally, the interference audio is pre-set in the playback device, or received from the voice service server.

[0087] Optionally, there are multiple encoding settings for the interference audio, and different playback devices receive interference audio with different encodings; or different interference audios are preset in different playback devices;

[0088] The voice wake-up device receives the codes of all interfering audios and is able to identify all interfering audios; or the codes of all interfering audios are built into the voice wake-up device.

[0089] Optionally, in S1, on the playback device side, the currently received data stream is first identified as audio content through the built-in hardware or software mechanism. At this stage, the playback device will check the format, encoding type and other information of the data packet to ensure the accuracy of subsequent processing. After confirming that it is an audio stream, the decoder in the playback device is started, which is responsible for decoding the compressed audio data (such as MP3, AAC and other formats) into a format that can be directly processed by the first speech recognition module, such as uncompressed PCM (Pulse Code Modulation) format.

[0090] Optionally, in S2, a section of interference audio with a specific frequency coding rule is superimposed on the decoded data. The frequency coding rule can be implemented at a low frequency or high frequency end to which the human ear is insensitive. For example, three frequencies of 16.5KHz, 17KHz, and 17.5KHz are recorded as A, B, and C, and an AC-B style coded sound wave signal is generated every t milliseconds. Among them, the AC-B style coded sound wave signal is only a coding of an interference audio. The coding of n types of interference audio can be distinguished by swapping the order of A, B, and C, such as AB-C, BC-A and other combination codings, and the frequency D can also be increased to increase the combination coding of the interference audio.

[0091] To facilitate subsequent interference source identification and management, a segment of interfering audio encoded at a specific frequency is superimposed on the decoded audio stream. This special interfering sound wave signal does not affect the user experience, but serves as a recognizable signature for voice wake-up devices, allowing them to track and analyze interference in the audio stream.

[0092] Optional, such as Figure 2 As shown, S1 and S2 can be swapped in order. The playback device first superimposes the interference audio on the original audio through the superposition module, and then continuously monitors the superimposed audio output by the superposition module to find the preset wake-up word. When the wake-up word is recognized, a mark, namely the wake-up word interference mark, is sent to the voice service server.

[0093] After receiving the wake-up word interference mark, the voice service server records it in a log or database. This step is helpful for subsequent wake-up word false triggering judgment.

[0094] Optionally, in S3, based on the frequency coding rules of the interference audio, the voice wake-up device is used to analyze the input signal and perform the sound source positioning function. By comparing the audio responses from different directions, the direction of the sound source is identified. During this process, when the voice wake-up device detects the coding frequency of the interference audio (such as the frequency of the above-mentioned AC-B-AC-B or B-AC-B-AC style coding) within 5t consecutive milliseconds, it is considered to be interference audio, and the sound source in this direction is automatically determined to be noise interference. The signal from this direction is temporarily ignored, and the focus is instead on monitoring the pure audio input from other directions.

[0095] This embodiment uses a preset frequency coding rule to analyze the input signal to detect interfering audio. It also uses the sound source localization function to ignore signals from the direction of the interfering audio and instead focus on listening to pure voice input from other directions to improve the accuracy of wake-up word recognition.

[0096] Optionally, in S4, when interfering audio breaks through the front-end defenses and enters the wake-up word response process, the voice wake-up device still has a redundant mechanism to prevent false triggering of the wake-up word. Before the voice wake-up device recognizes the wake-up word and responds to the wake-up word event, it immediately sends a wake-up word arbitration request to the voice service server.

[0097] This embodiment sends a wake-up word arbitration to the voice service server to initiate an additional verification process to avoid false activation caused by noise.

[0098] Optionally, in S5, the calculation formula for the statistical confidence interval is:

[0099] Statistical confidence interval = sample mean + (critical value × standard error)

[0100] Among them, when the sample size is large enough (in this embodiment, it is greater than 30), the critical value can be calculated based on the normal distribution function and the confidence level of 95%, and the obtained value is 1.96. Assume that the sample size is n and the sample mean is If the sample standard deviation is s, then The specific formula for obtaining the statistical confidence interval is:

[0101]

[0102] Among them, w i represents the observed value of the i-th sample, represents the mean of all observations, and n represents the number of observations in the sample.

[0103] Among them, in this embodiment, after the voice service server receives the wake-up word arbitration sent by the voice wake-up device, it will start a verification process, compare the previously recorded first timestamp with the newly received second timestamp, and calculate the time difference between the two. And compare the obtained time difference with the pre-set statistical confidence interval. If the time difference falls within the statistical confidence interval, it means that the two recognized wake-up words are very likely to come from the same event, and the second recognition may be a misjudgment caused by noise interference. In this case, a detection result of false triggering is generated; otherwise, a detection result of no false triggering is generated.

[0104] like Figure 2FIG. 1 shows a schematic diagram of the structure of a voice wake-up device with a dual protection mechanism provided by an embodiment of the present invention. The voice wake-up device can be applied to a voice wake-up device, including:

[0105] A signal processing module is used to monitor ambient sound, identify interference audio from the ambient sound, and perform noise processing on the sound carrying the interference audio to obtain pure audio;

[0106] A second wake-up word recognition module is configured to perform wake-up word recognition on the clean audio, generate a wake-up word arbitration when a wake-up word is recognized, and send the generated wake-up word arbitration to the voice service server, wherein the wake-up word arbitration is configured to trigger the voice service server to send a detection result to the voice wake-up device after detecting whether the wake-up word is a false trigger of the wake-up word;

[0107] The wake-up word response module is used to receive the detection result and execute the wake-up word response or cancel the wake-up word response based on the detection result.

[0108] like Figure 3 As shown, it shows a voice wake-up device with a dual protection mechanism provided by an embodiment of the present invention, which is used in a voice service server, and the device includes:

[0109] a false trigger detection module, configured to receive a wake-up word interference mark sent by a playback device and a wake-up word arbitration mark sent by a voice wake-up device, match the wake-up word interference mark with the wake-up word arbitration mark, detect whether there is a wake-up word false trigger, and generate a detection result, wherein the wake-up word arbitration mark is generated when the voice wake-up device recognizes a wake-up word, and the wake-up word interference mark is generated when the voice recognition module built into the playback device recognizes a wake-up word;

[0110] The first sending module is used to send a detection result to the voice wake-up device, where the detection result is used to trigger the voice wake-up device to execute a wake-up word response or cancel a wake-up word response.

[0111] like Figure 4 As shown, it shows a voice wake-up device with a dual protection mechanism provided by an embodiment of the present invention, which is used in a playback device, and the device includes:

[0112] A superposition module, used for superimposing the interference audio onto the original audio to obtain superimposed audio;

[0113] a first wake-up word recognition module, configured to perform wake-up word recognition on the original audio or the superimposed audio to obtain a wake-up word interference mark;

[0114] The second sending module is used to send a wake-up word interference mark to the voice service server. The wake-up word interference mark is used by the voice service server to match the wake-up word interference mark with the wake-up word arbitration, detect whether there is a wake-up word false trigger, and generate a detection result. The detection result is used to trigger the voice wake-up device to execute a wake-up word response or cancel the wake-up word response. The wake-up word arbitration is generated when the voice wake-up device recognizes the wake-up word.

[0115] The present invention provides a voice wake-up system with a dual protection mechanism, the system comprising a voice wake-up device, a voice service server and a playback device;

[0116] A playback device, configured to perform wake-up word recognition on the original audio, obtain a wake-up word interference mark, send the wake-up word interference mark to the voice service server, and superimpose the interference audio on the original audio to obtain superimposed audio;

[0117] The voice wake-up device is used to monitor the ambient sound, identify the interference audio of the ambient sound, perform noise processing on the sound carrying the interference audio to obtain pure audio, recognize the wake-up word on the pure audio, and generate a wake-up word arbitration when the wake-up word is recognized and send it to the voice service server;

[0118] A voice service server is configured to match the wake-up word interference mark with the wake-up word arbitration, detect whether there is a wake-up word false trigger, generate a detection result, and send the detection result to the voice wake-up device;

[0119] The voice wake-up device is also used to receive the detection result and execute a wake-up word response or cancel the wake-up word response based on the detection result.

[0120] Optionally, the playback device is further used to replace the original audio with the superimposed audio to perform wake-up word recognition and obtain a wake-up word interference mark.

[0121] Optionally, also include:

[0122] The switch module is used to construct a global voice protection switch. When the global voice protection switch is in the on state, the voice wake-up system with the dual protection mechanism is enabled.

[0123] The above description is merely a preferred embodiment of the present invention and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present invention is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the present invention.

Claims

1. A voice wake-up method with a dual protection mechanism, characterized in that: Used in a voice wake-up device, the method includes: Monitor ambient sound, identify interference audio from the ambient sound, and perform noise processing on the sound carrying the interference audio to obtain pure audio; Performing wake-up word recognition on the clean audio, generating a wake-up word arbitration when a wake-up word is recognized and sending it to the voice service server, wherein the wake-up word arbitration is used to trigger the voice service server to send a detection result to the voice wake-up device after detecting whether the wake-up word is a wake-up word false trigger; receiving the detection result, and executing a wake-up word response or canceling the wake-up word response based on the detection result; Among them, after the voice service server detects whether the wake-up word is a false trigger of the wake-up word, the detection result sent to the voice wake-up device is specifically: Receive the wake-up word interference mark sent by the playback device and the wake-up word arbitration sent by the voice wake-up device, match the wake-up word interference mark with the wake-up word arbitration, detect whether there is a wake-up word false trigger, and generate a detection result. The wake-up word arbitration is generated when the voice wake-up device recognizes the wake-up word, and the wake-up word interference mark is generated when the voice recognition module built into the playback device recognizes the wake-up word.

2. A voice wake-up method with a dual protection mechanism according to claim 1, characterized in that: The interference audio is pre-set in the voice wake-up device, or received from the voice service server.

3. The voice wake-up method with a dual protection mechanism according to claim 1, characterized in that: Getting pure audio specifically includes: Continuously monitor ambient sounds and identify the source direction of various sounds through sound source localization; When interference audio is monitored, the direction of the interference sound source of the interference audio is obtained, the sound source in the direction of the interference sound source is used as noise interference, and the voice wake-up device is used to continue monitoring the ambient sounds in other directions to obtain pure audio.

4. The voice wake-up method with a dual protection mechanism according to claim 1, characterized in that: The wake-up word arbitration includes a second wake-up word identifier and a second timestamp, where the second wake-up word identifier is a wake-up word identifier used to activate a voice wake-up device.

5. A voice wake-up method with a dual protection mechanism, characterized in that: Used in a voice service server, the method includes: Receive a wake-up word interference mark sent by a playback device and a wake-up word arbitration sent by a voice wake-up device, match the wake-up word interference mark with the wake-up word arbitration, detect whether there is a wake-up word false trigger, and generate a detection result, where the wake-up word arbitration is generated when the voice wake-up device recognizes a wake-up word, and the wake-up word interference mark is generated when the voice recognition module built into the playback device recognizes a wake-up word; The detection result is sent to the voice wake-up device, where the detection result is used to trigger the voice wake-up device to execute a wake-up word response or cancel a wake-up word response.

6. A voice wake-up method with a dual protection mechanism according to claim 5, characterized in that: Also includes: Sending interference audio to the playback device, where the interference audio is used to superimpose the original audio of the playback device to obtain superimposed audio; Interference audio is sent to the voice wake-up device, where the interference audio is used for filtering the superimposed audio when the voice recognition device recognizes the superimposed audio through encoding of the interference audio.

7. The voice wake-up method with a dual protection mechanism according to claim 5, characterized in that: The wake-up word interference mark includes a first wake-up word identifier and a first timestamp, the wake-up word arbitration includes a second wake-up word identifier and a second timestamp, and detecting whether there is a wake-up word false trigger specifically includes: Determine whether the first wake-up word identifier and the second wake-up word identifier are the same wake-up word; if so, calculate the time difference between the first timestamp and the second timestamp, and determine whether the time difference is within a statistical confidence interval; If so, a detection result indicating the presence of a false trigger is generated; if not, a detection result indicating the absence of a false trigger is generated.

8. A voice wake-up method with a dual protection mechanism according to claim 7, characterized in that: The calculation formula of the statistical confidence interval is shown as follows: Statistical confidence interval = Among them, z represents a constant, n represents the number of statistics, represents the observed value of the i-th statistic, Represents the mean of all observations.

9. A voice wake-up method with a dual protection mechanism, characterized in that: Used in a playback device, the method includes: Superimposing the interference audio onto the original audio to obtain superimposed audio; Performing wake-up word recognition on the original audio or the superimposed audio to obtain a wake-up word interference mark; A wake-up word interference mark is sent to the voice service server, and the wake-up word interference mark is used by the voice service server to match the wake-up word interference mark with the wake-up word arbitration, detect whether there is a wake-up word false trigger, and generate a detection result. The detection result is used to trigger the voice wake-up device to execute a wake-up word response or cancel the wake-up word response. The wake-up word arbitration is generated when the voice wake-up device recognizes the wake-up word.

10. A voice wake-up method with a dual protection mechanism according to claim 9, characterized in that: The wake-up word interference mark includes a first wake-up word identifier and a first timestamp.

11. The voice wake-up method with a dual protection mechanism according to claim 9, characterized in that: The interference audio is pre-set in the playback device, or received from the voice service server.

12. A voice wake-up method with a dual protection mechanism according to claim 11, characterized in that: Also includes: The superimposed audio is played, and the superimposed audio is used to identify and filter the superimposed audio through the interference audio when the voice wake-up device performs noise filtering.

13. A voice wake-up device with a dual protection mechanism, characterized in that: Used in a voice wake-up device, the device includes: A signal processing module is used to monitor ambient sound, identify interference audio from the ambient sound, and perform noise processing on the sound carrying the interference audio to obtain pure audio; The second wake-up word recognition module is used to recognize the wake-up word of the clean audio, and when the wake-up word is recognized, a wake-up word arbitration is generated and sent to the voice service server. The wake-up word arbitration is used to trigger the voice service server to send a detection result to the voice wake-up device after detecting whether the wake-up word is a false trigger of the wake-up word; wherein, after the voice service server detects whether the wake-up word is a false trigger of the wake-up word, the detection result sent to the voice wake-up device is specifically: Receive a wake-up word interference mark sent by a playback device and a wake-up word arbitration sent by a voice wake-up device, match the wake-up word interference mark with the wake-up word arbitration, detect whether there is a wake-up word false trigger, and generate a detection result, where the wake-up word arbitration is generated when the voice wake-up device recognizes a wake-up word, and the wake-up word interference mark is generated when the voice recognition module built into the playback device recognizes a wake-up word; The wake-up word response module is used to receive the detection result and execute the wake-up word response or cancel the wake-up word response based on the detection result.

14. A voice wake-up device with a dual protection mechanism, characterized in that: Used in a voice service server, the device includes: a false trigger detection module, configured to receive a wake-up word interference mark sent by a playback device and a wake-up word arbitration mark sent by a voice wake-up device, match the wake-up word interference mark with the wake-up word arbitration mark, detect whether there is a wake-up word false trigger, and generate a detection result, wherein the wake-up word arbitration mark is generated when the voice wake-up device recognizes a wake-up word, and the wake-up word interference mark is generated when the voice recognition module built into the playback device recognizes a wake-up word; The first sending module is used to send a detection result to the voice wake-up device, where the detection result is used to trigger the voice wake-up device to execute a wake-up word response or cancel a wake-up word response.

15. A voice wake-up device with a dual protection mechanism, characterized in that: Used in a playback device, the device includes: A superposition module, used for superimposing the interference audio onto the original audio to obtain superimposed audio; a first wake-up word recognition module, configured to perform wake-up word recognition on the original audio or the superimposed audio to obtain a wake-up word interference mark; The second sending module is used to send a wake-up word interference mark to the voice service server. The wake-up word interference mark is used by the voice service server to match the wake-up word interference mark with the wake-up word arbitration, detect whether there is a wake-up word false trigger, and generate a detection result. The detection result is used to trigger the voice wake-up device to execute a wake-up word response or cancel the wake-up word response. The wake-up word arbitration is generated when the voice wake-up device recognizes the wake-up word.

16. A voice wake-up system with a dual protection mechanism, characterized in that: The system includes a voice wake-up device, a voice service server and a playback device; The voice wake-up device includes a voice wake-up device with a dual protection mechanism as described in claim 13, the voice service server includes a voice wake-up device with a dual protection mechanism as described in claim 14, and the playback device includes a voice wake-up device with a dual protection mechanism as described in claim 15.

Citation Information

Patent Citations

  • Voice awakening method and device

    CN108335696A

  • Enhanced wake-up method and device for Internet of Things terminal, and storage medium

    CN114944153A