An intelligent device false wake-up interception method and system, and a storage medium
By acquiring a list of similar words and dynamically adjusting the wake-up threshold, the problem of false wake-up of intelligent voice devices is solved, the recognition accuracy and stability of the devices are improved, the user experience is optimized, and the characteristics of voice signals in complex environments are adapted.
Patent Information
- Application Number
- CN202510116565.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2045-01-24
AI Technical Summary
False wake-up is a common problem in existing intelligent voice control devices. Existing solutions cannot accurately control the wake-up frequency and prevent false wake-ups, which affects the accuracy of device response and user experience.
By acquiring a list of similar words, setting a list of absorbable wake words, and adjusting the wake threshold of each similar word based on the wake threshold of the target wake word, combined with speech feature analysis and score calculation, the target wake word and similar words are accurately distinguished, and the wake threshold is dynamically adjusted to reduce false wake-ups.
It improves the recognition accuracy and stability of smart devices, optimizes the user experience, reduces useless wake-ups, ensures that the target wake word triggers the device at the right time, and adapts to the voice signal characteristics under different environmental conditions.
Smart Images

Figure CN120048257B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech recognition, and in particular to a smart device false wake-up interception method and system and a storage medium. BACKGROUND
[0002] At present, the false wake-up problem in smart voice control devices is widespread. Existing solutions mainly reduce the false wake-up probability by dynamically adjusting the threshold of the wake-up word. For example, some methods automatically adjust the wake-up threshold according to the user's wake-up habits, and the threshold gradually decreases with the increase of the number of uses, or the threshold is adjusted over time. However, these methods have defects. In the case of long time without wake-up, the increase of the wake-up word threshold may cause difficulty in the first wake-up, and the too low threshold after frequent wake-up may easily cause false wake-up, affecting the response accuracy of the device. In addition, some technologies avoid false wake-up by detecting the do-not-disturb period, such as not responding to the wake-up instruction at night to avoid false wake-up during this period. However, this method has a narrow application range and is only limited to a specific time period, and cannot meet the demand of starting the device normally during the do-not-disturb period, affecting the user experience.
[0003] Therefore, the existing technology still faces a series of challenges such as how to more accurately control and optimize the frequency of false wake-up, improve the stability of the device under all environmental conditions, and balance the sensitivity of wake-up and the prevention of false wake-up. SUMMARY
[0004] Therefore, the present application is committed to providing a smart device false wake-up interception method and system and a storage medium to solve the technical problems of high wake-up probability, inaccurate threshold adjustment, and inability to balance sensitivity and false wake-up prevention in the prior art.
[0005] In a first aspect, the present application provides a smart device false wake-up interception method, which comprises the following steps:
[0006] obtaining an approximate word list;
[0007] adding the approximate words in the approximate word list to an absorption wake-up word list;
[0008] setting the wake-up threshold of a target wake-up word through testing;
[0009] setting the wake-up threshold corresponding to each approximate word in the absorption wake-up word list based on the wake-up threshold of the target wake-up word;
[0010] obtaining a human voice audio and scoring the human voice audio to obtain the score of the human voice audio for each approximate word and the score of the human voice audio for the target wake-up word;
[0011] When the score of the human voice audio for the target wake-up word first reaches the wake-up threshold of the target wake-up word, the received human voice audio signal is cleared, and the device is woken up; when the score of the human voice audio for the approximate word first reaches the wake-up threshold of the approximate word, the received human voice audio signal is cleared, and no response is made.
[0012] Optionally, the obtaining of the approximate word list comprises the following steps:
[0013] Voice samples of the target wake-up word uttered by a plurality of users are collected, and corresponding audio signals are obtained, the users including sample groups with different pronunciation habits and dialects;
[0014] Speech recognition is performed on the audio signals to generate text results corresponding to the audio signals, the text results including correctly recognized target wake-up word texts and approximate word texts generated due to unclear pronunciation of the users or recognition bias;
[0015] According to the text results, approximate words that are easily misrecognized as target wake-up words are screened out through analysis of the phonetic feature similarity and the close proximity of the pinyin unit sequence of the text, and an approximate word list is generated.
[0016] Optionally, the setting of the wake-up threshold of the target wake-up word through testing comprises the following steps:
[0017] Audio samples in multiple different scenarios are collected, including correct pronunciation samples of the target wake-up word and samples of background noise or other irrelevant speech;
[0018] Acoustic feature analysis is performed on the audio samples to calculate the score of the target wake-up word and the score of the background noise or other irrelevant speech in each sample;
[0019] Based on the score analysis result, an initial wake-up threshold of the target wake-up word is selected, and the initial wake-up threshold of the target wake-up word is tested through audio samples in multiple different environments;
[0020] The number of false wake-ups is observed, and the initial wake-up threshold of the target wake-up word is adjusted until a preset false wake-up number standard is reached within a specified time;
[0021] A final wake-up threshold of the target wake-up word is determined.
[0022] Optionally, the setting of the wake-up threshold corresponding to each approximate word in the absorption wake-up word list based on the wake-up threshold of the target wake-up word comprises the following steps:
[0023] On the basis of the set wake-up threshold of the target wake-up word, human voice samples are collected;
[0024] Calculate the score of the human voice audio sample for the target wake-up word and each approximate word respectively, and draw the score distribution curve of the target wake-up word and each approximate word;
[0025] Based on the wake-up threshold of the target wake-up word, adjust the wake-up threshold of each approximate word from high to low to ensure that the wake-up point of the target wake-up word is not later than the wake-up point of each approximate word.
[0026] Optionally, the adjusting the initial wake-up threshold of the target wake-up word according to the number of false wake-ups until the preset false wake-up number standard is reached within a specified time comprises the following steps:
[0027] Performing speech recognition on the audio segment of each false wake-up point to generate the corresponding false wake-up text, and comparing it with the approximate words in the absorption wake-up word list;
[0028] When the false wake-up text belongs to the approximate words in the absorption wake-up word list, ignore this false wake-up; when the false wake-up text does not belong to the approximate words in the absorption wake-up word list, mark it as a valid false wake-up and record it;
[0029] When the number of valid false wake-ups exceeds the preset false wake-up number, increase the wake-up threshold of the target wake-up word until the number of valid false wake-ups reaches the preset false wake-up number standard within a specified time.
[0030] Optionally, the obtaining human voice audio and scoring the human voice audio to obtain the score of the human voice audio for each approximate word and the score of the human voice audio for the target wake-up word comprises the following steps:
[0031] Obtain the audio features of the target wake-up word and each approximate word, and calculate the probability distribution of all pinyin units in each frame of audio based on the acoustic model;
[0032] Based on the probability of each frame of pinyin unit, use a combination algorithm to calculate the comprehensive score of the target wake-up word and each approximate word at the word level.
[0033] Optionally, it also includes:
[0034] Obtain multiple audio signals;
[0035] Based on the energy distribution comparison analysis of the multiple audio signals, determine the target audio and non-target audio;
[0036] Normally absorb and intercept the approximate words of the target audio, and ignore the wake-up results of the non-target audio;
[0037] When there is an overlap between the target audio and the non-target audio processing, preferentially use the wake-up result of the target audio.
[0038] In a second aspect, the present application also provides an intelligent device false wake-up interception system, the system comprising:
[0039] an approximate word generation module configured to generate an approximate word list according to a voice sample of a target wake-up word;
[0040] a threshold setting module configured to set a wake-up threshold of the target wake-up word by testing, and set a corresponding wake-up threshold for each approximate word based on the wake-up threshold of the target wake-up word;
[0041] an audio analysis module configured to obtain human voice audio and calculate a score of the human voice audio for the target wake-up word and each approximate word;
[0042] a signal processing module configured to clear a received human voice signal and wake up a device when a score rate of the human voice audio for the target wake-up word first reaches the wake-up threshold of the target wake-up word, and clear the received human voice signal and make no response when a score rate of the human voice audio for the approximate word first reaches the wake-up threshold of the approximate word.
[0043] In a third aspect, the present application also provides a device comprising a processor, a memory, a communication interface and a communication bus, the processor, the memory and the communication interface being in communication with each other through the communication bus;
[0044] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the above-mentioned intelligent device false wake-up interception method.
[0045] In a fourth aspect, the present application also provides a computer readable storage medium, characterized in that the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the above-mentioned intelligent device false wake-up interception method.
[0046] According to the first aspect of the present application, by treating the approximate word that is easy to cause false wake-up as an absorption wake-up word for special processing, the accuracy of the intelligent device is effectively improved. By setting an independent wake-up threshold for the absorption wake-up word and clearing the false wake-up signal, when the absorption wake-up word reaches the wake-up condition before the target wake-up word, the device makes no response, thereby avoiding false wake-up caused by different voice pronunciation or voice environment. This scheme ensures that the target wake-up word can trigger the device at the right time, thereby optimizing the user experience and reducing unnecessary wake-up.
[0047] Further, the present application improves the accuracy and flexibility of the system in intercepting false wake-up by dynamically adjusting the wake-up threshold of the absorption wake-up word. By adjusting the wake-up threshold of the absorption wake-up word from high to low, the system can adapt to the characteristics of voice signals in different environmental conditions and effectively avoid false interception. On the premise of ensuring the accuracy of target wake-up word recognition, the system can trigger an interception response in time when the absorption wake-up word reaches the preset wake-up threshold, preventing false wake-up. At the same time, based on the optimization of score calculation and processing logic, the system can accurately judge and eliminate false recognition of the absorption wake-up word, avoiding false response. The technical solution of the present application enables the intelligent device to dynamically adjust and optimize the recognition strategy of the wake-up word, thereby improving the intelligence, refinement and user experience of the system. Especially in complex or changing use environments, the system can more stably and accurately execute wake-up instructions.
[0048] The above description is only a summary of the technical solutions of the present application. In order to more clearly understand the technical means of the present application and to implement the content of the description, the following will describe the preferred embodiments of the present application in detail. BRIEF DESCRIPTION OF DRAWINGS
[0049] Figure 1 A schematic flowchart of a false wake-up interception method of an intelligent device according to an embodiment of the present application is shown;
[0050] Figure 2 A schematic flowchart of a method for setting a wake-up threshold of a target wake-up word in step S300 is shown; Figure 1
[0051] Figure 3 A schematic flowchart of a method for setting a wake-up threshold of a target wake-up word in step S300 is shown; Figure 1
[0052] Figure 4 A schematic flowchart of a method for setting a wake-up threshold of a target wake-up word in step S400 is shown; Figure 1
[0053] Figure 5 A schematic diagram of scores of different human voice audios on the target wake-up word "Hello" according to an embodiment of the present application is shown;
[0054] Figure 6 A schematic diagram of scores of human voice audio "Hello" on the target wake-up word and the absorption wake-up word according to an embodiment of the present application is shown;
[0055] Figure 7 A structural block diagram of a false wake-up interception system of an intelligent device according to an embodiment of the present application is shown;
[0056] Figure 8 A structural block diagram of an intelligent device false wake-up interception device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0057] In order to make the above objectives, features and advantages of the present application more obvious and understandable, specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. It can be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application. In addition, it should be noted that only the parts related to the present application are shown in the drawings for the convenience of description, rather than all the structures. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the present application.
[0058] The terms "comprising" and "having" and any variations thereof in the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to the process, method, product or device.
[0059] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of the phrase in various places in the specification does not necessarily all refer to the same embodiment, nor is it necessarily mutually exclusive of other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein are combinable.
[0060] Based on the problems in the above background art, the inventors considered setting a wake-up threshold for each word of the target wake-up word during the design process. This method aims to distinguish the target wake-up word from the similar words by refining the requirements for voice features. However, due to the influence of user pronunciation habits and environmental noise in actual use, each word of the target wake-up word may not always be clear and distinguishable. This method has limited effect in intercepting similar words, and too strict word-level threshold setting significantly reduces the wake-up rate, making it difficult for the device to effectively respond to the user's instructions. This limitation prompted further research to explore innovative solutions to improve false wake-up interception while ensuring wake-up rate.
[0061] Finally, after many research experiments, the inventors proposed the intelligent device false wake-up interception method of the present application. Figure 1 A schematic flowchart of an intelligent device false wake-up interception method according to an embodiment of the present application is shown. As shown in Figure 1 The intelligent device false wake-up interception method includes:
[0062] Step S100, obtaining a list of approximate words.
[0063] Collect all approximate words similar to the target wake-up word in terms of voice features, which are usually misrecognized as the target wake-up word in the voice recognition process.
[0064] Step S200, adding the approximate words in the approximate word list to the absorption wake-up word list.
[0065] The identified approximate words are added to the "absorption wake-up word list" of the device, that is, these words are filtered as non-wake-up instructions. The absorption wake-up word list is used to store all approximate words that may cause false wake-up.
[0066] Step S300, setting the wake-up threshold of the target wake-up word through testing.
[0067] Audio samples under different scenarios are collected, the acoustic features of the target wake-up word are analyzed, and appropriate wake-up thresholds are selected according to the influence of background noise, environmental noise, etc. The most appropriate threshold is determined through multi-scenario testing, so that the device can achieve accurate wake-up in different noise environments and improve the stability of the device.
[0068] Step S400, setting the wake-up threshold corresponding to each approximate word in the absorption wake-up word list based on the wake-up threshold of the target wake-up word.
[0069] The wake-up threshold of the target wake-up word is used as a reference to set a corresponding threshold for each approximate word, which can ensure that the approximate words that cause false wake-up do not wake up the device prematurely or repeatedly. Ensure that each approximate word has an appropriate threshold to avoid false wake-up. This step further accurately manages the wake-up threshold, enabling the device to have higher accuracy when processing approximate words.
[0070] Step S500, obtaining human voice audio and scoring the human voice audio to obtain the score of the human voice audio for each approximate word and the score of the human voice audio for the target wake-up word.
[0071] After the device receives the human voice signal, it performs voice recognition and compares the scores of each approximate word and the target wake-up word to determine whether the human voice audio is a valid instruction. Accurate calculation of the response score of each sound source to the wake-up word can effectively distinguish the difference between the target wake-up word and noise or approximate words, thereby reducing the possibility of false wake-up.
[0072] Step S600, when the score of the human voice audio for the target wake-up word first reaches the wake-up threshold of the target wake-up word, the received human voice signal is cleared and the device is woken up; when the score of the human voice audio for the approximate word first reaches the wake-up threshold of the approximate word, the received human voice signal is cleared and no response is made.
[0073] The identified target wake-up word score is compared with the approximate word score. If the target wake-up word meets the threshold condition first, the device is triggered to wake up. If the approximate word score exceeds the threshold, the instruction is ignored and no response is given. This ensures that the device responds to valid wake-up words and does not respond to received approximate words, minimizing the occurrence of false wake-ups. At the same time, the efficiency and user experience of the device are maintained.
[0074] According to the above embodiment, by treating the approximate word that is easy to cause false wake-up as an absorption wake-up word and performing special processing, the accuracy of the intelligent device is effectively improved. By setting an independent wake-up threshold for the absorption wake-up word and clearing the false wake-up signal, when the absorption wake-up word meets the wake-up condition before the target wake-up word, the device does not respond, thereby avoiding false wake-up caused by different pronunciation or speech environment. This scheme ensures that the target wake-up word can trigger the device at the right time, thereby optimizing the user experience and reducing unnecessary wake-up.
[0075] Figure 2 A schematic flowchart of the method for obtaining the approximate word list in step S100 is shown. Figure 1 As shown, step S100 includes: Figure 2 As shown, step S100 includes:
[0076] Step S110, collect voice samples of target wake-up words uttered by multiple users to obtain corresponding audio signals, the users including sample groups with different pronunciation habits and dialects.
[0077] In this step, in order to cover the diversity of different user groups, the target wake-up words uttered by multiple users in different environments are collected. The user groups may include different accents, dialects, pronunciation habits and voice characteristics. A special voice collection task can be set up using an intelligent device or a voice collection platform to ensure the diversity of the samples. By collecting a wide range of voice samples, the device can adapt to different user pronunciation, dialect or accent, so that the wake-up system can better understand the voice input of different groups and improve the recognition accuracy of the device in a diversified user environment.
[0078] Step S120, performing voice recognition on the audio signal to generate a text result corresponding to the audio signal, the text result including a correctly recognized target wake-up word text and an approximate word text generated due to unclear pronunciation or recognition bias of the user.
[0079] The collected audio signal is converted into text by a speech recognition system (such as an ASR system). The focus of this step is whether each target wake-up word in the audio is correctly recognized in the speech recognition process. Sometimes, due to unclear pronunciation or different accents of the user, the speech recognition system may misrecognize the target wake-up word as a word that is similar but not exactly the same in pronunciation, and these misrecognized words are called "approximate words". The generated text includes correctly recognized target wake-up words and approximate words misrecognized by the system due to unclear pronunciation or dialect differences. By converting the speech signal into text through speech recognition technology, the possible bias in speech recognition can be revealed, the difference between the target wake-up word and the approximate word can be captured, and effective data support can be provided for further approximate word screening.
[0080] Step S130, according to the text result, by analyzing the phonetic feature similarity and the close proximity of the pinyin unit sequence of the text, screening out approximate words that are easy to be misrecognized as target wake-up words, and generating an approximate word list.
[0081] This step S130 performs a detailed phonetic feature similarity analysis on the text result, compares the audio features of the target wake-up word and other words in the text, such as pitch, speed, and speech rhythm, to identify words that are similar in pronunciation. In addition, through the analysis of the close proximity of the pinyin unit sequence, words with similar pinyin unit sequences are detected, which may produce errors in speech recognition and be misjudged as target wake-up words. By combining the analysis of both phonetic features and pinyin sequences, the system can more accurately identify potential approximate words and include them in the approximate word list. This method can effectively reduce the probability of misrecognition and ensure that the distinction between the target wake-up word and the false wake-up word is more clear, thereby improving the accuracy of speech recognition and ensuring that the user's wake-up instruction is not disturbed by other sounds or similar words, optimizing the user experience.
[0082] Figure 3 A schematic flowchart of the method for setting the wake-up threshold of the target wake-up word in step S300 is shown. Figure 1 As shown, this step S300 includes: Figure 3 As shown, this step S300 includes:
[0083] Step S310, collect audio samples in multiple different scenarios, including correct pronunciation samples of the target wake-up word and samples of background noise or other irrelevant speech.
[0084] The step S310 collects audio samples in multiple different scenarios, which can include noisy environments, quiet environments, and normal conversation environments, etc. In addition, the audio samples include correct pronunciation samples of the target wake-up word and background noise or other irrelevant voice samples. By covering multiple environments and audio categories, different situations in actual use can be effectively simulated to provide comprehensive basic data for subsequent acoustic feature analysis. The advantage is that through diversified sample collection, it is helpful to comprehensively evaluate the performance of the target wake-up word in different environments, ensure that the device can be accurately awakened in multiple environments, and enhance the applicability and anti-interference ability of the device.
[0085] The step S320 performs acoustic feature analysis on the audio samples to calculate the score of the target wake-up word and the score of the background noise or other irrelevant voice in each sample.
[0086] The acoustic feature analysis is performed on the collected audio samples to calculate the score of the target wake-up word and the score of the background noise or irrelevant voice in each sample. By performing detailed analysis on the audio samples, the acoustic features of the target wake-up word (such as pitch, duration, tone quality, etc.) and the feature differences of the background noise or irrelevant voice are identified and quantified. The advantage of this step S320 is that it can accurately evaluate the distinguishability of the target wake-up word and the background noise, and provide data support for subsequent threshold setting, ensuring that the wake-up threshold can be reasonably adjusted under different environmental conditions to improve the accuracy of wake-up.
[0087] The step S330 selects an initial wake-up threshold for the target wake-up word based on the score analysis results, and tests the initial wake-up threshold for the target wake-up word through audio samples in multiple different environments.
[0088] Based on the score analysis results in step S320, an initial wake-up threshold for the target wake-up word is selected, and the effectiveness of this threshold is tested through audio samples in multiple different environments. In this process, the preliminary selection of the threshold needs to consider the obtained score distribution to ensure that the selected threshold can accurately identify the target wake-up word in most environments and avoid false wake-up. The advantage is that this step selects a reasonable initial threshold by considering the environment and sound scores comprehensively, and verifies its performance through multi-scene testing, which provides strong data support for further adjusting and optimizing the wake-up sensitivity.
[0089] The step S340 observes the number of false wake-ups and adjusts the initial wake-up threshold for the target wake-up word until the preset false wake-up number standard is reached within a specified time.
[0090] This step S340 mainly optimizes the threshold setting of the target wake-up word by pre-recording audio data of the real environment. In this stage, the device does not have actual wake-up function, mainly for collecting audio samples for algorithm model analysis. By inputting these audio data into the algorithm model in the development environment, the system generates a corresponding score graph showing the matching degree of audio at each time point with the target wake-up word. In the analysis process, the staff observes the false wake-up peaks that may occur in the identification process of the model, and by adjusting the score threshold, the staff can control the sensitivity of the system to false wake-up and accurately calibrate the accuracy of the device response. This adjustment process ensures that the model can identify and filter out potential false wake-up points in the actual environment without the target wake-up word, thereby optimizing the performance of the system.
[0091] This step can effectively reduce the probability of false wake-up by fine-tuning the threshold, while ensuring the reliability and stability of the system in real use. Finally, only the most accurate matching result will trigger wake-up, thereby improving the overall recognition ability of the system and user experience.
[0092] Step S350, determine the final wake-up threshold of the target wake-up word.
[0093] After multiple adjustments in step S340, when the number of false wake-ups of the target wake-up word meets the preset standard, the final threshold can be determined. This threshold is fully tested in the actual environment, ensuring that the device can effectively wake up under accurate identification while avoiding false wake-up. Through adjustment and verification, the most suitable wake-up threshold is finally obtained to ensure the performance stability of the intelligent voice device in various environments, while improving the accuracy and reliability of the device.
[0094] In some embodiments, the above step S340 includes:
[0095] Step S341, performing speech recognition on the audio segment of each false wake-up point to generate a corresponding false wake-up text, and comparing it with the similar words in the absorption wake-up word list.
[0096] This step S341 performs speech recognition on the audio segment of each false wake-up point to generate a corresponding false wake-up text, and compares the text with the similar words in the absorption wake-up word list. In this process, the false wake-up text is first processed by the speech recognition system and converted into text form, and then compared with the similar words in the preset absorption wake-up word list to determine whether the text belongs to the similar word range. The advantage is that this step can effectively distinguish whether the false wake-up is caused by similar words, thereby providing a basis for subsequent judgment of whether to record false wake-up. This detailed comparison mechanism helps to avoid misjudgment of non-true false wake-up due to language accent, non-standard pronunciation, etc.
[0097] Step S342, when the false wake-up text belongs to the approximate words in the absorption wake-up word list, ignore this false wake-up; when the false wake-up text does not belong to the approximate words in the absorption wake-up word list, mark as valid false wake-up and record.
[0098] After voice recognition and comparison, it is determined whether the false wake-up text belongs to the approximate words in the absorption wake-up word list. If the false wake-up text belongs to one of these approximate words, the system will ignore this false wake-up and not record it; if the false wake-up text does not belong to the approximate words in the absorption wake-up word list, it is considered that this false wake-up is a valid false wake-up and is recorded. The advantage of this step is that by processing the approximate words, irrelevant errors can be reduced, and false wake-ups caused by pronunciation or other non-standard voice characteristics can be avoided to be counted as valid false wake-ups, thereby improving the accuracy of false wake-up statistics. This process provides an effective tolerance mechanism for intelligent devices to interfere with interference sounds or non-target wake-up words in the environment.
[0099] Step S343, when the number of valid false wake-ups exceeds the preset false wake-up number, the wake-up threshold of the target wake-up word is increased until the number of valid false wake-ups reaches the preset false wake-up number standard within a specified time.
[0100] When the number of valid false wake-ups exceeds the preset false wake-up number standard, the system will raise the wake-up threshold of the target wake-up word. By raising the threshold, the system will reduce the frequency of false wake-ups until the number of valid false wake-ups no longer exceeds the preset standard. This step S343 can dynamically adjust the wake-up threshold according to the actual situation, and flexibly cope with false wake-up problems in different environments and use scenarios. If the false wake-up frequency is too high, the tolerance of the system to environmental noise and approximate words can be increased by increasing the threshold, thereby reducing unnecessary false responses and ensuring accurate response of the device to the target wake-up word. This way provides a gradual optimization process for the device, ensuring the long-term effectiveness and accuracy of the system.
[0101] Figure 4 A schematic flowchart of a method for setting a wake-up threshold of a target wake-up word is shown Figure 1 As shown in step S400, the method for setting a wake-up threshold of a target wake-up word includes Figure 4 As shown, this step S400 includes:
[0102] Step S410, based on the set wake-up threshold of the target wake-up word, collect human voice audio samples.
[0103] On the basis of the set wake-up threshold of the target wake-up word, human voice audio samples are collected. In this process, the system first sets a preliminary target wake-up word threshold, and then collects audio data containing the target wake-up word and other related language features. The advantage of this step is to provide the required real audio data for the subsequent test and optimization process, ensuring that the adjustment and optimization work is based on the actual collected data. Collecting audio samples can cover a variety of situations, which helps to more comprehensively analyze the score distribution of the target wake-up word and the approximate word.
[0104] Step S420, respectively calculating the score of the human voice audio sample for the target wake-up word and each approximate word, and drawing the score distribution curve of the target wake-up word and each approximate word.
[0105] In this step S420, the system will evaluate each audio sample collected, respectively calculate the recognition score of the audio under the target wake-up word and each approximate word, and draw the score distribution graph. Through the score distribution curve, the system can analyze the performance of the wake-up word and the approximate word under different audio conditions, intuitively reflect their relative score situation, and provide objective data support for subsequent threshold adjustment. It can clearly see the score distribution of different wake-up words, so as to evaluate the necessity and rationality of threshold adjustment.
[0106] Step S430, based on the wake-up threshold of the target wake-up word, adjusting the wake-up threshold of each approximate word from high to low to ensure that the wake-up point of the target wake-up word is not later than the wake-up point of each approximate word.
[0107] Based on the set threshold of the target wake-up word, the wake-up threshold of each approximate word is adjusted from high to low step by step to ensure that the wake-up point of the target wake-up word is not later than the wake-up point of any approximate word. That is, by adjusting the threshold of each approximate word, when a voice signal is received, the target wake-up word can be preferentially recognized as a wake-up signal, and the false triggering as a certain approximate word is avoided. The advantage of this adjustment process is that it can effectively prevent the false wake-up delay problem of the target wake-up word. By strictly controlling the wake-up time of the approximate word and the target wake-up word, it is ensured that the target wake-up word can accurately and timely respond to the user's wake-up request, improve the user experience, reduce the false triggering phenomenon, and enhance the stability and accuracy of the system.
[0108] In some embodiments, the obtaining human voice audio and scoring the human voice audio to obtain the score of the human voice audio corresponding to each approximate word and the score of the human voice audio for the target wake-up word in the step S500 comprises:
[0109] Step S510, obtaining the audio features of the target wake-up word and each approximate word, and calculating the probability distribution of all pinyin units in each frame of audio in combination with the acoustic model.
[0110] In this step, the system performs accurate probability calculation for each phonetic unit in the audio signal by obtaining the audio features of the target wake-up word and the approximate words and combining the acoustic model. Specifically, this process involves analysis of the phonemes, syllables, and their corresponding pinyins in the audio signal, and assigning a corresponding probability value to each phonetic unit through the acoustic model. This basic audio recognition work ensures that the system can capture and quantify the audio characteristics in the speech in detail, including the pronunciation patterns of phonemes and pinyins. The advantage is that, with the help of the acoustic model to analyze the probability of the phonetic unit, the system can accurately identify and process various speech signal features, and effectively distinguish the target wake-up word from the approximate words in clear or fuzzy speech environment. This not only improves the recognition accuracy of the wake-up word, but also provides reliable data support for subsequent score calculation, thereby enhancing the robustness and flexibility of the speech recognition system, ensuring its stable operation and accurate wake-up response in dynamic and complex environments.
[0111] Step S520, based on the probability of each frame of phonetic unit, a combination algorithm is used to calculate the comprehensive score of the target wake-up word and each approximate word at the word level.
[0112] Based on the probability of each frame of phonetic unit, the system uses a combination algorithm (such as a weighted or ranking algorithm) to calculate the comprehensive score of the target wake-up word and each approximate word. This algorithm not only considers the matching degree of the phonetic unit, but also combines the speech features and semantic context to accurately integrate multi-dimensional information, thereby obtaining a more accurate score. Through this multi-dimensional feature fusion calculation method, the system can efficiently process the comprehensive score of each word, significantly improving the reliability and accuracy of recognition. This algorithm effectively improves the discrimination between the target wake-up word and the approximate word, reduces the probability of misrecognition and false wake-up, and enhances the adaptability of the system in complex speech environments. Through this flexible and efficient algorithm, the device can provide more accurate and reliable wake-up response under varying audio conditions.
[0113] In some embodiments, the above-mentioned intelligent device false wake-up interception method further comprises the following steps:
[0114] Obtaining multiple audio signals.
[0115] The system captures audio data from different locations or devices in parallel through multiple audio signal sources (such as multiple microphones or audio channels), thereby realizing multi-channel audio acquisition. This multi-channel audio recording provides richer and more comprehensive sound input, greatly enhancing the adaptability of the system in complex audio environments, especially in the presence of noise interference, echo reflection, or multiple sound sources. By comprehensively processing multiple audio signals, the system can significantly improve the recognition accuracy of the target speech, while effectively distinguishing background noise from meaningful speech signals, thereby optimizing the accuracy and reliability of wake-up word recognition.
[0116] Based on the comparative analysis of the energy distribution of the multi-channel audio signals, the target channel audio and the non-target channel audio are determined.
[0117] By analyzing the energy distribution of multiple audio channel signals, based on the intensity difference of each channel signal, the target channel audio (i.e. the audio from the sound area where the test person is located) and the non-target channel audio (i.e. the audio from other sound areas or mainly containing interference noise) are identified and distinguished. Through the comparative analysis of the energy distribution, the system can accurately determine the source of each audio signal and judge whether it belongs to the target audio that needs to be processed preferentially. Such analysis can effectively exclude the interference of background noise and other irrelevant sound sources, so as to ensure that the system can focus on the wake-up recognition of the target audio, significantly reducing the false wake-up caused by environmental noise or non-target voice interference. In a complex multi-audio environment, this processing ensures that the system can accurately focus on the correct audio signal, greatly improving the wake-up accuracy and the stability of the system.
[0118] The approximate words of the target channel are normally absorbed and intercepted, and the wake-up result of the non-target channel audio is ignored.
[0119] For the approximate words of the target channel audio, the absorption and interception are still carried out to process the approximate words in the target sound area, so as to effectively prevent false wake-up. For the approximate words in the non-target channel audio, if a false wake-up signal appears, the wake-up response of the signal is selected to be ignored. This process accurately analyzes the audio channel of the signal source and adopts corresponding processing strategies according to the source of different audio signals. Through the targeted processing of target channel audio and non-target channel audio, the system significantly reduces the risk of false wake-up caused by non-target sound area (which may contain background noise or irrelevant signals), thereby optimizing the efficiency and accuracy of the wake-up engine. In addition, the standard absorption and interception mechanism for the target channel audio ensures that the device can efficiently and accurately respond when identifying voice commands, avoiding the negative impact of external interference or irrelevant audio signals on the wake-up effect.
[0120] When there is an overlap between the processing of the target channel audio and the non-target channel audio, the wake-up result of the target channel audio is used preferentially.
[0121] When the target path audio overlaps with the non-target path audio at a certain moment, the system will prioritize the wake-up result of the target path audio and ignore the influence of the non-target path audio on the wake-up decision. By focusing the processing on the target path audio source, ensuring that the wake-up decision is based on the most relevant and clear signal source, the strategy of prioritizing the target path audio effectively enhances the accuracy and consistency of the system. Especially in the case of overlapping multi-channel audio signals, it helps to improve the reliability of the wake-up response. By setting priorities, the system can effectively avoid conflicts or mutual interference between multi-channel audio, thereby enhancing stability in complex audio environments. This mechanism ensures that the system can accurately respond to the wake-up instructions of the tester and minimize the false wake-up interference caused by speech or background noise from non-target areas.
[0122] The following further explains and illustrates the method for intercepting false wake-up of intelligent devices according to the present invention through a specific implementation process.
[0123] Figure 5 Shows a score schematic diagram of different human voice audios for the target wake-up word "Hello, Dreame" according to an embodiment of the present invention. As Figure 5 shown, the black part represents the score change of the common prefix "Hello" of the human voice, and the blue part represents the score change of the "Dreame" part of the human voice in "Hello, Dreame". When the score reaches the threshold of 0.8, the device will perform a wake-up response. The green part represents the score trend of the "XGIMI" part of the human voice in "Hello, XGIMI". Although the score of this part lags behind that of the normal wake-up word, its final score may still exceed the threshold, resulting in a false wake-up of the system.
[0124] To avoid the above problems, the present invention is optimized by adding the known approximate word "Hello, XGIMI" to the absorption wake-up word list. Different from the target wake-up word, once this approximate word reaches the wake-up condition, the system does not perform a wake-up operation, thus preventing false wake-ups. At this time, the wake-up word list of the intelligent device contains two wake-up words: the target wake-up word "Hello, Dreame" and the absorption wake-up word "Hello, XGIMI". Whenever a human voice is received, the device will calculate the scores of the human voice for these two wake-up words simultaneously. If the target wake-up word "Hello, Dreame" reaches the wake-up condition first, the system will clear the previous audio signal and perform a wake-up; if the absorption wake-up word "Hello, XGIMI" reaches the wake-up condition first, the previous human voice signal will be cleared, and the system will not make any response, avoiding false wake-ups.
[0125] Figure 6 Shows a score schematic diagram of the human voice audio "Hello, Dreame" for the target wake-up word and the absorption wake-up word according to an embodiment of the present invention.
[0126] Refer to Figure 6In this solution, in order to effectively reduce false wake-up without affecting normal wake-up, the threshold of the absorption word needs to be finely adjusted. Specifically, first, set the initial threshold of the absorption wake-up word "Hello Ximi" according to the actual use. For example, when the threshold of the absorption wake-up word "Hello Ximi" is set to 0.7, the user shouts "Hello Zhui Mi", the wake-up point W of the target wake-up word will be earlier than the wake-up point A of the absorption wake-up word, and the system can normally wake up the device. At this time, the absorption wake-up word will not interfere with the normal wake-up.
[0127] However, if the threshold of the absorption wake-up word is adjusted to 0.6, the wake-up point W of the target wake-up word "Hello Zhui Mi" will lag behind the wake-up point B of the absorption wake-up word, and the normal wake-up signal will be mistakenly intercepted as false wake-up, affecting the response of the device. Therefore, in order to ensure that the device can normally respond to the wake-up instruction without being disturbed by false wake-up, the threshold of the absorption wake-up word should be gradually adjusted from high to low, and a critical value should be accurately set so that the wake-up point of the absorption word always lags behind the wake-up point of the target wake-up word.
[0128] Through this adjustment process, false wake-up phenomenon can be effectively avoided while maintaining a high wake-up rate, thereby optimizing the wake-up experience of the intelligent device. This method ensures that the false wake-up interception solution minimizes the impact on the wake-up rate, achieving the best balance between the performance of the intelligent device and the user experience.
[0129] In some embodiments, the present application can add an absorption wake-up word "speech nihao zhui mi" to the absorption wake-up word library, where "speech" is the score of all pinyin units. When the voice contains a non-sentence-initial target wake-up word, the system absorbs the preposed voice segment through approximate words, processes these audio and prevents the wake-up response. When the clear sentence-initial wake-up word "Hello Zhui Mi" is detected, the system will execute the wake-up command, which can effectively reduce the false wake-up caused by the context and improve the accuracy and reliability of the device wake-up.
[0130] According to the above embodiments, by dynamically adjusting the wake-up threshold of the absorption wake-up word, the accuracy and flexibility of the system in intercepting false wake-up are improved. By adjusting the wake-up threshold of the absorption wake-up word from high to low, the system can adapt to the characteristics of voice signals in different environmental conditions, effectively avoiding false interception. Under the premise of ensuring the recognition accuracy of the target wake-up word, the system can trigger an interception response in time when the absorption wake-up word reaches the preset wake-up threshold, preventing false wake-up. At the same time, based on the optimization of score calculation and processing logic, the system can accurately judge and eliminate the false recognition of the absorption wake-up word, avoiding false response. The technical solution of the present application enables the intelligent device to dynamically adjust and optimize the recognition strategy of the wake-up word, thereby improving the intelligence, refinement and user experience of the system. Especially in complex or changing use environments, the system can more stably and accurately execute the wake-up instruction.
[0131] The embodiment of the present application also provides an intelligent device false wake-up interception system. Figure 7 As shown in the figure, the intelligent device false wake-up interception system comprises:
[0132] An approximate word generation module 101 is configured to generate an approximate word list according to a voice sample of a target wake-up word.
[0133] A threshold setting module 102 is configured to test and set a wake-up threshold of the target wake-up word, and set a corresponding wake-up threshold for each approximate word based on the wake-up threshold of the target wake-up word.
[0134] An audio analysis module 103 is configured to acquire human voice audio and calculate scores of the human voice audio for the target wake-up word and each approximate word.
[0135] A signal processing module 104 is configured to clear a received human voice audio signal and wake up a device when a score rate of the human voice audio for the target wake-up word first reaches the wake-up threshold of the target wake-up word, and clear the received human voice audio signal and make no response when a score rate of the human voice audio for the approximate word first reaches the wake-up threshold of the approximate word.
[0136] The above is applicable to the intelligent device false wake-up interception system, and the specific implementation of setting the wake-up threshold and the rule for the target wake-up word and the approximate word refers to the related content of the embodiment of the above method, which will not be described in detail here.
[0137] The embodiment of the present application also provides an intelligent device false wake-up interception device, which comprises a processor 201, a memory 202, and a computer program stored in the memory 202 and configured to be executed by the processor 201, and the processor 201 implements the intelligent device false wake-up interception method of any of the above embodiments when executing the computer program.
[0138] The processor 201 implements the steps in the above intelligent device false wake-up interception method embodiment when executing the computer program, for example, all the steps of the intelligent device false wake-up interception method shown in the figure. Figure 1 Or, the processor 201 implements the functions of the modules / units in the above intelligent device false wake-up interception system when executing the computer program, for example, the functions of the modules in the intelligent device false wake-up interception system shown in the figure. Figure 7
[0139] For example, the computer program can be divided into one or more modules, and the one or more modules are stored in the memory 202 and executed by the processor 201 to complete the present application. The one or more modules can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the intelligent device false wake-up interception system.
[0140] The processor 201 can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The processor 201 is the control center of the smart device false wake-up interception system, and connects various parts of the smart device false wake-up interception system through various interfaces and lines.
[0141] The memory 202 can be used to store computer programs and / or modules. The processor 201 realizes various functions of the smart device false wake-up interception system by running or executing the computer programs and / or modules stored in the memory 202, and calling the data stored in the memory 202. The memory 202 can mainly include a program storage area and a data storage area. The program storage area can store an operating system, at least one application required by a function, etc. The data storage area can store data created according to the use of the smart device false wake-up interception system, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state memory device.
[0142] If the modules / units of the intelligent device false wake-up interception system are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer-readable storage medium. When the processor executes the computer program, the steps of the above-mentioned various method embodiments can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium can include any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0143] The embodiments of the present application also provide a computer-readable storage medium, and the method according to the embodiments of the present application can be implemented in hardware, firmware, or recorded in a storage medium, or stored in a remote storage medium or a non-transitory machine-readable storage medium and stored in a local storage medium through network downloading of computer code, so that the method described herein can be processed by such software on a storage medium using a general-purpose computer, a special-purpose processor, or programmable or special-purpose hardware. The storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state disk, etc. Further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that the computer, the processor, the microprocessor controller, or the programmable hardware includes a storage component that can store or receive software or computer code, which, when accessed and executed by the computer, the processor, or the hardware, implements the method shown in the above-mentioned embodiments.
[0144] The above-mentioned embodiments only express several embodiments of the present application, and the description is more specific and detailed, but it cannot be understood as a limitation on the scope of the patent. It should be noted that for ordinary skilled persons in the art, without departing from the concept of the present application, several modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.
Claims
1. An intelligent device false wake-up intercepting method, characterized in that, The method comprises the following steps: Obtain an approximate word list; Add the approximate words in the approximate word list to an absorption wake-up word list; Set a wake-up threshold of a target wake-up word through testing; Set a wake-up threshold corresponding to each approximate word in the absorption wake-up word list based on the wake-up threshold of the target wake-up word; Obtain human voice audio and score the human voice audio to obtain a score of the human voice audio corresponding to each approximate word and a score of the human voice audio corresponding to the target wake-up word; When the score of the human voice audio corresponding to the target wake-up word first reaches the wake-up threshold of the target wake-up word, clear the received human voice audio signal and wake up the device; when the score of the human voice audio corresponding to the approximate word first reaches the wake-up threshold of the approximate word, clear the received human voice audio signal and do not respond; The setting of the wake-up threshold of the target wake-up word through testing comprises: Perform speech recognition on the audio segment of each false wake-up point to generate corresponding false wake-up text and compare it with the approximate words in the absorption wake-up word list; When the false wake-up text belongs to the approximate words in the absorption wake-up word list, ignore this false wake-up; when the false wake-up text does not belong to the approximate words in the absorption wake-up word list, mark it as a valid false wake-up and record it; When the number of valid false wake-ups exceeds a preset false wake-up number, increase the wake-up threshold of the target wake-up word until the number of valid false wake-ups reaches a preset false wake-up number standard within a specified time. 2.The method of claim 1, wherein, The obtaining of the approximate word list comprises the following steps: Collect speech samples of the target wake-up word issued by multiple users to obtain corresponding audio signals, wherein the users include sample groups with different pronunciation habits and dialects; Perform speech recognition on the audio signals to generate text results corresponding to the audio signals, wherein the text results include correctly recognized target wake-up word texts and approximate word texts generated due to unclear pronunciation of the users or recognition deviation; According to the text results, analyze the phonetic feature similarity and the proximity of the pinyin unit sequence of the texts to screen out approximate words that are easily misrecognized as the target wake-up word and generate an approximate word list. 3.The method of claim 1, wherein, The setting of the wake-up threshold of the target wake-up word through testing comprises the following steps: Collect audio samples in multiple different scenes, including correct pronunciation samples of the target wake-up word and samples of background noise or other irrelevant speech; Perform acoustic feature analysis on the audio samples to calculate the score of the target wake-up word in each sample and the score of the background noise or other irrelevant speech; Based on the score analysis result, select an initial wake-up threshold of the target wake-up word and test the initial wake-up threshold of the target wake-up word through multiple audio samples in different environments; Observe the number of false wake-ups and adjust the initial wake-up threshold of the target wake-up word until a preset false wake-up number standard is reached within a specified time; Determine the final wake-up threshold of the target wake-up word. 4.The method of claim 1, wherein, The setting of the wake-up threshold corresponding to each approximate word in the absorption wake-up word list based on the wake-up threshold of the target wake-up word comprises the following steps: On the basis of setting the wake-up threshold of the target wake-up word, a human voice audio sample is collected; Scores of the human voice audio sample for the target wake-up word and each approximate word are calculated respectively, and score distribution curves of the target wake-up word and each approximate word are drawn; Based on the wake-up threshold of the target wake-up word, the wake-up threshold of each approximate word is adjusted from high to low to ensure that the wake-up point of the target wake-up word is not later than the wake-up point of each approximate word. 5.The method of claim 1, wherein, The steps of obtaining the human voice audio and scoring the human voice audio to obtain the score of the human voice audio for each approximate word and the score of the human voice audio for the target wake-up word include the following steps: Obtain the audio features of the target wake-up word and each approximate word, and calculate the probability distribution of all pinyin units in each frame of audio in combination with an acoustic model; Based on the probability of each frame of pinyin unit, a combination algorithm is used to calculate the comprehensive score of the target wake-up word and each approximate word at the word level. 6.The method of claim 1, wherein, Further comprising: Obtain multiple audio signals; Based on the energy distribution comparison analysis of the multiple audio signals, determine the target audio and the non-target audio; Normally absorb and intercept the approximate words of the target audio, and ignore the wake-up results of the non-target audio; When there is an overlap between the target audio and the non-target audio processing, the wake-up result of the target audio is used preferentially.
7. An intelligent device false wake-up interception system employing the intelligent device false wake-up interception method according to any one of claims 1-6, characterized in that, The system comprises: An approximate word generation module configured to generate a list of approximate words based on a voice sample of a target wake-up word; A threshold setting module configured to set a wake-up threshold of the target wake-up word through testing, and set a corresponding wake-up threshold for each approximate word based on the wake-up threshold of the target wake-up word; An audio analysis module configured to obtain human voice audio and calculate scores of the human voice audio for the target wake-up word and each approximate word; A signal processing module configured to clear the received human voice audio signal and wake up the device when the score of the human voice audio for the target wake-up word reaches the wake-up threshold of the target wake-up word first, and clear the received human voice audio signal and do not respond when the score of the human voice audio for the approximate word reaches the wake-up threshold of the approximate word first.
8. An apparatus, comprising: Comprise: A processor, a memory, a communication interface and a communication bus, the processor, the memory and the communication interface complete mutual communication through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction makes the processor execute the operation corresponding to the intelligent device false wake-up interception method in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer program is stored in the computer readable storage medium, and the computer program is executed by the processor to realize the steps of the intelligent device false wake-up interception method in any one of claims 1-6.
Citation Information
Patent Citations
Speech recognition model training method and device, equipment, storage medium and product
CN114299933A
Detection method and device for false identification of similar sounds and computer equipment
CN114333799A