False wakeup interception method and system for intelligent equipment and storage medium
By setting the wake-up threshold of approximate words in the intelligent voice control device and scoring the human voice audio, the false wake-up problem is solved, the accuracy and user experience of the device are improved, and the stability and recognition accuracy of the device are ensured in different environments.
Patent Information
- Application Number
- CN202510116565.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The problem of false wake-up is common in existing intelligent voice control devices. Existing solutions such as dynamic adjustment of wake-up word thresholds and evasion of wake-up during a specific time period cannot effectively solve the precise control of the false wake-up frequency and the stability of the device under different environmental conditions.
By obtaining the approximate word list, add the approximate word list to the absorbed wake word list, and set the wake threshold for each approximate word based on the wake threshold of the target wake word. Get and score the audio of the person, and wake up the device when the score of the target wake word first reaches the threshold; when the score of the approximate word reaches the threshold, the clear signal does not respond.
It effectively improves the accuracy of smart devices, avoids false wake-up caused by different voice pronunciations or environments, optimizes the user experience, reduces useless wake-up, and improves the stability and recognition accuracy of the system in different environments.
Smart Images

Figure CN120048257A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and in particular, to a method, a system, and a storage medium for intercepting accidental wake-up of intelligent devices. Background Art
[0002] At present, the problem of accidental wake-up in intelligent voice control devices widely exists. Existing solutions mainly reduce the probability of accidental wake-up by dynamically adjusting the threshold of the wake-up word. For example, some methods automatically adjust the wake-up threshold according to the user's wake-up habit, gradually reducing the threshold as the number of uses increases, or adjusting the threshold over time. However, these methods have defects. In the case of no wake-up for a long time, an increase in the wake-up word threshold may make the first wake-up difficult, while an overly low threshold after frequent wake-ups is prone to cause accidental wake-up, affecting the accuracy of device response. In addition, some technologies avoid accidental wake-up by detecting the do-not-disturb time period, such as not responding to wake-up instructions at night to avoid accidental wake-up during this time period. However, this method has a narrow scope of application, is limited to a specific time period, and cannot meet the need to normally start the device during the do-not-disturb period, affecting the user experience.
[0003] Therefore, the prior art still faces a series of challenges such as how to more precisely control and optimize the frequency of accidental wake-up, improve the stability of the device under all environmental conditions, and how to balance the sensitivity of wake-up and the prevention of accidental wake-up. Summary of the Invention
[0004] In view of this, the present invention is committed to providing a method, a system, and a storage medium for intercepting accidental wake-up of intelligent devices, which are used to solve technical problems such as too high wake-up probability, inaccurate threshold adjustment, and inability to balance sensitivity and prevention of accidental wake-up in the prior art.
[0005] In a first aspect, the present invention provides a method for intercepting accidental wake-up of an intelligent device, and the method includes the following steps:
[0006] Obtain a list of approximate words;
[0007] Add the approximate words in the list of approximate words to the absorption wake-up word list;
[0008] Set the wake-up threshold of the target wake-up word through testing;
[0009] Set the wake-up threshold corresponding to each approximate word in the absorption wake-up word list based on the wake-up threshold of the target wake-up word;
[0010] Obtain human voice audio, and score the human voice audio to obtain the score of the human voice audio corresponding to each approximate word and the score of the human voice audio corresponding to the target wake-up word;
[0011] When the score of the human voice audio for the target wake-up word reaches the wake-up threshold of the target wake-up word first, clear the received human voice audio signal and wake up the device; when the score of the human voice audio for the approximate word reaches the wake-up threshold of the approximate word first, clear the received human voice audio signal and do not make any response.
[0012] Optionally, the obtaining of the approximate word list includes the following steps:
[0013] Collect voice samples of the target wake-up word uttered by multiple users to obtain corresponding audio signals, where the users include sample groups with different vocalization habits and dialects;
[0014] Perform speech recognition on the audio signal to generate a text result corresponding to the audio signal, where the text result includes the correctly recognized target wake-up word text and approximate word texts generated due to unclear pronunciation or recognition deviation of the user;
[0015] According to the text result, by analyzing the speech feature similarity and pinyin unit sequence proximity of the text, screen out approximate words that are easily misrecognized as the target wake-up word and generate an approximate word list.
[0016] Optionally, the setting of the wake-up threshold of the target wake-up word through testing includes the following steps:
[0017] Collect audio samples in multiple different scenarios, including correct pronunciation samples of the target wake-up word and samples of background noise or other irrelevant speech;
[0018] Perform acoustic feature analysis on the audio samples to calculate the score of the target wake-up word and the score of the background noise or other irrelevant speech in each sample;
[0019] Based on the score analysis result, select an initial wake-up threshold of the target wake-up word and test the initial wake-up threshold of the target wake-up word with audio samples in multiple different environments;
[0020] Observe the number of false awakenings and adjust the initial wake-up threshold of the target wake-up word until the preset false awakening times standard is reached within the specified time;
[0021] Determine the final wake-up threshold of the target wake-up word.
[0022] Optionally, the setting of the wake-up threshold corresponding to each approximate word in the absorption wake-up word list based on the wake-up threshold of the target wake-up word includes the following steps:
[0023] On the basis of the set wake-up threshold of the target wake-up word, collect human voice audio samples;
[0024] Calculate the scores of the human voice audio sample for the target wake-up word and each approximate word respectively, and draw the score distribution curves of the target wake-up word and each approximate word;
[0025] Based on the wake-up threshold of the target wake-up word, adjust the wake-up thresholds of each approximate word from high to low to ensure that the wake-up point of the target wake-up word is not later than the wake-up points of each approximate word.
[0026] Optionally, the step of adjusting the wake-up threshold of the initial target wake-up word according to the number of false awakenings until the preset number of false awakening times standard is reached within a specified time includes the following steps:
[0027] Perform speech recognition on the audio segment of each false awakening point to generate the corresponding false awakening text, and compare it with the approximate words in the absorption wake-up word list;
[0028] When the false awakening text belongs to the approximate words in the absorption wake-up word list, ignore this false awakening; when the false awakening text does not belong to the approximate words in the absorption wake-up word list, mark it as a valid false awakening and record it;
[0029] When the number of valid false awakenings exceeds the preset number of false awakening times, increase the wake-up threshold of the target wake-up word until the number of valid false awakenings reaches the preset number of false awakening times standard within a specified time.
[0030] Optionally, the step of obtaining the human voice audio and scoring the human voice audio to obtain the scores of the human voice audio for each approximate word and the score of the human voice audio for the target wake-up word includes the following steps:
[0031] Obtain the audio features of the target wake-up word and each approximate word, and calculate the probability distribution of all phonetic units in each frame of audio in combination with the acoustic model;
[0032] Based on the probability of each frame of phonetic unit, use the combination algorithm to calculate the comprehensive scores of the target wake-up word and each approximate word at the word level.
[0033] Optionally, it further includes:
[0034] Obtain multiple audio signals;
[0035] Based on the comparative analysis of the energy distribution of multiple audio signals, determine the target path audio and non-target path audio;
[0036] Perform normal absorption and interception on the approximate words of the target path, and ignore the wake-up results of the non-target path audio;
[0037] When there is an overlap in the processing of the target path audio and the non-target path audio, give priority to using the wake-up result of the target path audio.
[0038] In a second aspect, the present invention further provides an intelligent device false wake-up interception system, which includes:
[0039] An approximate word generation module, configured to generate a list of approximate words according to the voice sample of the target wake-up word;
[0040] A threshold setting module, configured to set the wake-up threshold of the target wake-up word through testing, and set corresponding wake-up thresholds for each approximate word based on the wake-up threshold of the target wake-up word;
[0041] An audio analysis module, configured to obtain human voice audio and calculate the scores of the human voice audio for the target wake-up word and each approximate word;
[0042] A signal processing module, configured to clear the received human voice audio signal and wake up the device when the score of the human voice audio for the target wake-up word reaches the wake-up threshold of the target wake-up word first; when the score of the human voice audio for the approximate word reaches the wake-up threshold of the approximate word first, clear the received human voice audio signal and make no response.
[0043] In a third aspect, the present invention further provides a device, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete mutual communication through the communication bus;
[0044] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above intelligent device false wake-up interception method.
[0045] In a fourth aspect, the present invention further provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above intelligent device false wake-up interception method are implemented.
[0046] According to the first aspect of the present invention, by specially processing approximate words that are likely to cause false wake-up as absorption wake-up words, the accuracy of intelligent devices is effectively improved. By setting an independent wake-up threshold for the absorption wake-up word and clearing the false wake-up signal, when the absorption wake-up word reaches the wake-up condition before the target wake-up word, the device does not make any response, thus avoiding false wake-up caused by different voice pronunciations or voice environments. This solution ensures that the target wake-up word can trigger the device at the accurate timing, thereby optimizing the user experience and reducing useless wake-up.
[0047] Furthermore, the present invention dynamically adjusts the wake-up threshold for absorbing wake-up words, improving the accuracy and flexibility of the system in intercepting false wake-ups. By adjusting the wake-up threshold of the absorbing wake-up words from high to low, it can flexibly adapt to the speech signal characteristics under different environmental conditions, effectively avoiding the situation of false interception. On the premise of ensuring the recognition accuracy of the target wake-up word, it can trigger an interception reaction in a timely manner when the absorbing wake-up word reaches the preset wake-up threshold, preventing false wake-ups. At the same time, based on the optimization of the score calculation and processing logic, the system can accurately judge and eliminate misrecognized absorbing wake-up words, avoiding false responses. The technical solution of the present invention enables intelligent devices to dynamically adjust and optimize the recognition strategy of wake-up words, thereby enhancing the intelligence, refinement and user experience of the system. Especially in complex or changing usage environments, it can execute wake-up instructions more stably and accurately.
[0048] The above description is only an overview of the technical solution of the present invention. In order to understand the technical means of the present invention more clearly and implement it in accordance with the content of the specification, the following describes the preferred embodiments of the present invention in detail. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 Shows a schematic flowchart of a method for intercepting false wake-ups of an intelligent device according to an embodiment of the present invention;
[0050] Figure 2 Shows Figure 1 A schematic flowchart of the method for obtaining an approximate word list in step S100 shown;
[0051] Figure 3 Shows Figure 1 A schematic flowchart of the method for setting the wake-up threshold of the target wake-up word through testing in step S300 shown;
[0052] Figure 4 Shows Figure 1 A schematic flowchart of the method for setting the wake-up threshold of the target wake-up word through testing in step S400 shown;
[0053] Figure 5 Shows a score diagram of different human voice frequencies for the target wake-up word "Hello, Dreame" according to an embodiment of the present invention;
[0054] Figure 6 Shows a score diagram of the human voice frequency "Hello, Dreame" for the target wake-up word and the absorbing wake-up word according to an embodiment of the present invention;
[0055] Figure 7 Shows a structural block diagram of an intelligent device false wake-up interception system according to an embodiment of the present invention;
[0056] Figure 8The structural block diagram of the intelligent device wake-up false alarm interception device according to an embodiment of the present invention is shown. Detailed implementation manners
[0057] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the convenience of description, only the parts related to the present invention rather than all the structures are shown in the drawings. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0058] The terms "comprising" and "having" in the present invention and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices.
[0059] Referring to "embodiments" herein means that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0060] Based on the problems in the above background technology, during the design process, the inventors once considered setting a wake-up threshold for each character of the target wake-up word. This method aims to distinguish the target wake-up word from similar words by refining the requirements for speech features. However, due to the influence of users' pronunciation habits and environmental noise in actual use, each character of the target wake-up word may not always be clearly distinguishable. This method has limited effect in intercepting similar words, and at the same time, the overly strict character-level threshold setting significantly reduces the wake-up rate, making it difficult for the device to effectively respond to users' instructions. This limitation has prompted further exploration of innovative solutions to improve the false alarm interception effect while ensuring the wake-up rate.
[0061] Finally, after multiple research and experiments, the inventors proposed the intelligent device wake-up false alarm interception method of the present invention. Figure 1 The schematic flowchart of the intelligent device wake-up false alarm interception method according to an embodiment of the present invention is shown. As Figure 1 shown, the intelligent device wake-up false alarm interception method includes:
[0062] Step S100, obtain a list of approximate words.
[0063] Collect all approximate words that are similar to the target wake-up word in terms of speech features. These words are usually easily misrecognized as the target wake-up word during the speech recognition process.
[0064] Step S200, add the approximate words in the approximate word list to the absorption wake-up word list.
[0065] Add the recognized approximate words to the "absorption wake-up word list" of the device, that is, these words are filtered as non-wake-up instructions. The absorption wake-up word list is used to store all approximate words that may cause miswakening.
[0066] Step S300, set the wake-up threshold of the target wake-up word through testing.
[0067] Collect audio samples in different scenarios, analyze the acoustic features of the target wake-up word, and select a suitable wake-up threshold based on the analysis of background noise, environmental noise, etc. Determine the most suitable threshold through multi-scenario testing to enable the device to achieve accurate wake-up in different noise environments and improve the stability of the device.
[0068] Step S400, set the wake-up threshold corresponding to each approximate word in the absorption wake-up word list based on the wake-up threshold of the target wake-up word.
[0069] Use the wake-up threshold of the target wake-up word as a reference to set a corresponding threshold for each approximate word, which can ensure that approximate words that cause miswakening do not wake up the device prematurely or repeatedly. Ensure that each approximate word has a suitable threshold to avoid miswakening. This step further precisely manages the wake-up threshold, enabling the device to have higher accuracy when processing approximate words.
[0070] Step S500, obtain human voice audio, score the human voice audio, and obtain the score of the human voice audio corresponding to each approximate word and the score of the human voice audio corresponding to the target wake-up word.
[0071] After the device receives a human voice signal, perform speech recognition on it, and compare the scores of each approximate word and the target wake-up word to determine whether the human voice audio is a valid instruction. Precisely calculating the response score of each sound source to the wake-up word can effectively distinguish the difference between the target wake-up word and noise or approximate words, thereby reducing the possibility of miswakening.
[0072] Step S600, when the score of the human voice audio for the target wake-up word reaches the wake-up threshold of the target wake-up word first, clear the received human voice audio signal and wake up the device; when the score of the human voice audio for the approximate word reaches the wake-up threshold of the approximate word first, clear the received human voice audio signal and do not make any response.
[0073] Compare the scores of the identified target wake-up words with the scores of the approximate words. If the target wake-up words first meet the threshold condition, trigger device wake-up; if the scores of the approximate words exceed the threshold, ignore this instruction and do not respond. Ensure that the device responds to valid wake-up words and does not respond to received approximate words, minimizing the occurrence of false wake-ups. At the same time, the efficiency of the device and the user experience are maintained.
[0074] According to the above embodiments, by specially processing approximate words that are prone to false wake-ups as absorption wake-up words, the accuracy of intelligent devices is effectively improved. By setting an independent wake-up threshold for the absorption wake-up words and clearing false wake-up signals, when the absorption wake-up words reach the wake-up condition before the target wake-up words, the device does not make any response, thus avoiding false wake-ups caused by different voice pronunciations or voice environments. This solution ensures that the target wake-up words can trigger the device at the accurate time, thereby optimizing the user experience and reducing useless wake-ups.
[0075] Figure 2 Shows Figure 1 A schematic flowchart of the method for obtaining the approximate word list in step S100 shown. As Figure 2 shown, this step S100 includes:
[0076] Step S110, collect voice samples of the target wake-up words uttered by multiple users to obtain corresponding audio signals, where the users include sample groups with different vocalization habits and dialects.
[0077] In this step, in order to cover the diversity of different user groups, collect the target wake-up words uttered by multiple users in different environments. The user groups may include different accents, dialects, vocalization habits, and voice characteristics. An intelligent device or a voice collection platform can be used to set up a special voice collection task to ensure the diversity of the samples. By widely collecting different voice samples, it can be ensured that the device adapts to the vocalization methods, dialects, or accents of different users, enabling the wake-up system to better understand the voice inputs of different groups and improving the recognition accuracy of the device in diverse user environments.
[0078] Step S120, perform speech recognition on the audio signals to generate text results corresponding to the audio signals, where the text results include the correctly recognized target wake-up word text and approximate word text generated due to unclear enunciation or recognition deviation of the user.
[0079] Convert the collected audio signal into text through a speech recognition system (such as an ASR system). The key point of this step is whether each target wake-up word in the audio is correctly recognized during the speech recognition process. Sometimes, due to unclear pronunciation or different accents of users, the speech recognition system may mis-recognize the target wake-up word as a word with a similar pronunciation but not exactly the same. These mis-recognized words are called "approximate words". The generated text includes the correctly recognized target wake-up words and the approximate words mis-recognized by the system due to unclear enunciation or dialect differences. Converting the speech signal into text through speech recognition technology can reveal the possible biases in speech recognition, capture the differences between the target wake-up word and the approximate word, and effectively provide data support for further screening of approximate words.
[0080] Step S130, according to the text result, screen out the approximate words that are easily mis-recognized as the target wake-up word by analyzing the speech feature similarity and the proximity of the pinyin unit sequence of the text, and generate an approximate word list.
[0081] In this step S130, a detailed speech feature similarity analysis is performed on the text result, comparing the audio features of the target wake-up word with other words in the text, such as pitch, speech speed, speech rhythm, etc., so as to identify those words with similar pronunciations. In addition, through the analysis of the proximity of the pinyin unit sequence, words with similar pinyin unit sequences are detected. These words may cause errors during speech recognition and be misjudged as the target wake-up word. By combining the analysis of speech features and pinyin sequences, the system can more accurately identify potential approximate words and classify them into the approximate word list. This method can effectively reduce the occurrence probability of mis-recognitions and ensure a clearer distinction between the target wake-up word and the mis-wake-up word, thereby improving the accuracy of speech recognition, ensuring that the user's wake-up command is not interfered by other sounds or similar words, and optimizing the user experience.
[0082] Figure 3 Shows Figure 1 A schematic flowchart of the method for setting the wake-up threshold of the target wake-up word through testing in step S300 as shown. As Figure 3 shown, this step S300 includes:
[0083] Step S310, collect audio samples in multiple different scenarios, including correct pronunciation samples of the target wake-up word and samples of background noise or other irrelevant speech.
[0084] Step S310 collects audio samples in a variety of different scenarios, which may include noisy environments, quiet environments, and normal call environments, etc. In addition, the audio samples include correct pronunciation samples of the target wake-up word and background noise or other irrelevant speech samples. By covering a variety of environments and audio types, it can effectively simulate different situations in actual use and provide comprehensive basic data for subsequent acoustic feature analysis. The benefit is that through diversified sample collection, it helps to comprehensively evaluate the performance of the target wake-up word in different environments, ensuring that the device can be accurately awakened in a variety of environments and enhancing the applicability and anti-interference ability of the device.
[0085] In step S320, acoustic feature analysis is performed on the audio samples, and the scores of the target wake-up word and background noise or other irrelevant speech are calculated for each sample.
[0086] Acoustic feature analysis is performed on the collected audio samples, and the scores of the target wake-up word and background noise or irrelevant speech are calculated for each sample. By performing a detailed analysis of the audio samples, the acoustic features (such as pitch, duration, sound quality, etc.) of the target wake-up word and the feature differences of background noise or irrelevant speech are identified and quantified. The benefit of this step S320 is that it can accurately evaluate the distinguishability between the target wake-up word and background noise and provide data support for subsequent threshold setting, ensuring that the wake-up threshold can be reasonably adjusted under different environmental conditions and improving the accuracy of wake-up.
[0087] In step S330, based on the score analysis results, an initial wake-up threshold for the target wake-up word is selected, and the initial wake-up threshold for the target wake-up word is tested with audio samples in multiple different environments.
[0088] Based on the score analysis results in step S320, an initial wake-up threshold for the target wake-up word is selected, and the effectiveness of this threshold is tested with audio samples in multiple different environments. In this process, the initial selection of the threshold needs to consider the obtained score distribution to ensure that the selected threshold can accurately identify the target wake-up word in most environments and avoid false wake-up. The benefit is that this step selects a reasonable initial threshold by comprehensively considering the environment and sound scores and verifies its performance through multi-scenario tests, which provides strong data support for further adjusting and optimizing the wake-up sensitivity.
[0089] In step S340, observe the number of false wake-ups and adjust the initial wake-up threshold of the target wake-up word until the preset false wake-up number standard is reached within the specified time.
[0090] This step S340 mainly optimizes the threshold setting of the target wake-up word by pre-recording the audio data of the real environment. At this stage, the device does not have the actual wake-up function and is mainly used to collect audio samples for algorithm model analysis. By inputting this audio data into the algorithm model in the development environment, the system generates a corresponding score map, showing the matching degree of the audio at each time point with the target wake-up word. During the analysis process, the staff observes the possible false wake-up peaks that may occur during the recognition process of the model. By adjusting the score threshold, the staff can control the sensitivity of the system to false wake-ups and accurately calibrate the accuracy of the device's response. This adjustment process ensures that the model can identify and filter out potential false wake-up points in the actual environment without the target wake-up word, thereby optimizing the performance of the system.
[0091] This step can effectively reduce the probability of false wake-up by finely adjusting the threshold, while ensuring the reliability and stability of the system in actual use. Finally, only the most accurate matching results will trigger wake-up, thereby improving the overall recognition ability of the system and the user experience.
[0092] Step S350, determine the wake-up threshold of the final target wake-up word.
[0093] After multiple adjustments in step S340, when the number of false wake-ups of the target wake-up word meets the preset standard, the final threshold can be determined. This threshold has been fully tested in the actual environment to ensure that the device can be effectively woken up under accurate recognition while avoiding false wake-ups. Through adjustment and verification, the most suitable wake-up threshold is finally obtained to ensure the stable performance of the intelligent voice device in various environments, while improving the accuracy and reliability of the device.
[0094] In some embodiments, the above step S340 includes:
[0095] Step S341, perform speech recognition on the audio segment of each false wake-up point to generate a corresponding false wake-up text, and compare it with the approximate words in the absorption wake-up word list.
[0096] This step S341 performs speech recognition on the audio segment of each false wake-up point to generate a corresponding false wake-up text, and compares this text with the approximate words in the absorption wake-up word list. During this process, the false wake-up text will first be processed by the speech recognition system and converted into text form, and then compared with the approximate words in the preset absorption wake-up word list to determine whether this text belongs to the range of approximate words. The advantage is that this step can effectively distinguish whether the false wake-up is caused by approximate words, thereby providing a basis for subsequent judgment on whether to record the false wake-up. This refined comparison mechanism helps to avoid misjudging non-genuine false wake-ups due to reasons such as language accents and inaccurate pronunciations.
[0097] Step S342: When the mis-awakening text belongs to an approximate word in the absorption wake-up word list, this mis-awakening is ignored; when the mis-awakening text does not belong to an approximate word in the absorption wake-up word list, it is marked as a valid mis-awakening and recorded.
[0098] After speech recognition and comparison, it is determined whether the mis-awakening text belongs to an approximate word in the absorption wake-up word list. If the mis-awakening text belongs to one of these approximate words, the system will ignore this mis-awakening and not record it; if the mis-awakening text does not belong to an approximate word in the absorption wake-up word list, this mis-awakening is considered a valid mis-awakening and is recorded. The advantage of this step is that by processing approximate words, irrelevant errors can be reduced, and mis-awakenings caused by pronunciation or other non-standard speech features can be prevented from being counted as valid mis-awakenings, thereby improving the accuracy of mis-awakening statistics. This process provides an effective tolerance mechanism for the intelligent device to the interference sounds or non-target wake-up words in the environment.
[0099] Step S343: When the number of valid mis-awakenings exceeds the preset number of mis-awakenings, the wake-up threshold of the target wake-up word is increased until the number of valid mis-awakenings reaches the preset mis-awakening number standard within a specified time.
[0100] When the number of valid mis-awakenings exceeds the preset mis-awakening number standard, the system will increase the wake-up threshold of the target wake-up word. By increasing this threshold, the system will reduce the occurrence frequency of mis-awakenings until the number of valid mis-awakenings no longer exceeds the preset standard. This step S343 can dynamically adjust the wake-up threshold according to the actual situation and flexibly handle mis-awakening problems in different environments and usage scenarios. If the occurrence frequency of mis-awakenings is too high, the tolerance of the system to environmental noise and approximate words is increased by increasing the threshold, thereby reducing unnecessary mis-responses and ensuring the precise response of the device to the target wake-up word. This method provides a gradual optimization process for the device to ensure the long-term effectiveness and accuracy of the system.
[0101] Figure 4 shows Figure 1 The schematic flowchart of the method for setting the wake-up threshold of the target wake-up word through testing in step S400 as shown. As Figure 4 shown, this step S400 includes:
[0102] Step S410: On the basis of the set wake-up threshold of the target wake-up word, human voice audio samples are collected.
[0103] Based on the set wake-up threshold of the target wake-up word, human voice audio samples are collected. In this process, the system first sets a preliminary target wake-up word threshold, and then collects audio data containing the target wake-up word and other related language features. The advantage of this step is that it provides the real audio data required for subsequent testing and optimization processes, ensuring that the adjustment and optimization work are based on the actually collected data. Collecting audio samples can cover various situations, which helps to more comprehensively analyze the score distributions of the target wake-up word and similar words.
[0104] Step S420: Calculate the scores of the human voice audio samples for the target wake-up word and each similar word respectively, and draw the score distribution curves of the target wake-up word and each similar word.
[0105] In this step S420, the system will evaluate each collected audio sample, calculate the recognition scores of the audio for the target wake-up word and each similar word respectively, and draw the score distribution diagram. Through the score distribution curve, the system can analyze the performance of the wake-up word and similar words under different audio conditions, intuitively reflect their relative score situations, provide objective data support for subsequent threshold adjustment, clearly show the score distributions of different wake-up words, and thus evaluate the necessity and rationality of threshold adjustment.
[0106] Step S430: Based on the wake-up threshold of the target wake-up word, adjust the wake-up thresholds of each similar word from high to low to ensure that the wake-up point of the target wake-up word is not later than the wake-up points of each similar word.
[0107] Based on the set threshold of the target wake-up word, gradually adjust the wake-up thresholds of each similar word from high to low to ensure that the wake-up point of the target wake-up word is not later than the wake-up points of any similar word. That is, by adjusting the thresholds of each similar word, when a voice signal is received, the target wake-up word can be preferentially recognized as a wake-up signal, and it can be prevented from being mis-triggered as a certain similar word. The advantage of this is that this adjustment process can effectively prevent the problem of mis-awakening delay of the target wake-up word. By strictly controlling the wake-up timing of similar words and the target wake-up word, it is ensured that the target wake-up word can accurately and timely respond to the user's wake-up request, improve the user experience while reducing the phenomenon of mis-triggering, and enhance the stability and accuracy of the system.
[0108] In some embodiments, the obtaining of the human voice audio in step S500 and the scoring of the human voice audio to obtain the scores of the human voice audio for each corresponding similar word and the score of the human voice audio for the target wake-up word include:
[0109] Step S510: Obtain the audio features of the target wake-up word and each similar word, and calculate the probability distribution of all phoneme units in each frame of audio in combination with the acoustic model.
[0110] In this step, the system performs precise probability calculations on the pinyin units in each frame of the audio signal by obtaining the audio features of the target wake-up word and its approximate words and combining with the acoustic model. Specifically, this process involves the analysis of phonemes, syllables, and their pinyin compositions in the audio signal, and assigns corresponding probability values to each pinyin unit through the acoustic model. This basic audio recognition work ensures that the system can capture and quantify the audio characteristics in the speech in detail, including the pronunciation patterns of phonemes and pinyin. Its advantage is that by means of the probability analysis of pinyin units by the acoustic model, the system can accurately identify and process various speech signal features, and can effectively distinguish the target wake-up word from its approximate words in both clear and ambiguous speech environments. This not only improves the recognition accuracy of the wake-up word, but also provides reliable data support for subsequent score calculation, thus enhancing the robustness and flexibility of the speech recognition system, ensuring its stable operation and providing accurate wake-up responses in dynamic and complex environments.
[0111] Step S520, based on the probabilities of each frame of pinyin units, uses a combination algorithm to calculate the comprehensive scores of the target wake-up word and each approximate word at the word level.
[0112] Based on the probabilities of each frame of pinyin units, the system uses a combination algorithm (such as a weighted or sorting algorithm) to calculate the comprehensive scores of the target wake-up word and each approximate word. This algorithm not only considers the matching degree of pinyin units, but also combines speech features and semantic context to accurately integrate multi-dimensional information, so as to obtain more accurate scores. Through this calculation method of multi-dimensional feature fusion, the system can efficiently process the comprehensive scores of each word, significantly improving the reliability and accuracy of recognition. This algorithm effectively improves the discrimination between the target wake-up word and its approximate words, reduces the probability of misrecognition and false wake-up, thus enhancing the adaptability of the system in complex speech environments. Through this flexible and efficient algorithm, the device can provide more accurate and reliable wake-up responses under variable audio conditions.
[0113] In some embodiments, the above method for intercepting false wake-up of intelligent devices further includes the following steps:
[0114] Obtain multiple audio signals.
[0115] The system captures audio data from different positions or devices in parallel through multiple audio signal sources (such as multiple microphones or audio channels), so as to achieve multi-channel audio acquisition. This multi-channel audio recording provides richer and more comprehensive sound inputs, greatly enhancing the adaptability of the system in complex audio environments, especially in the case of noise interference, echo reflection, or coexistence of multiple sound sources. By comprehensively processing multiple audio signals, the system can significantly improve the recognition accuracy of the target speech, and at the same time effectively distinguish background noise from meaningful speech signals, thus optimizing the accuracy and reliability of wake-up word recognition.
[0116] Based on the comparative analysis of the energy distribution of multiple audio signals, the target-channel audio and non-target-channel audio are determined.
[0117] By analyzing the energy distribution of signals from multiple audio channels and based on the intensity differences of signals in each channel, the target-channel audio (i.e., the audio from the sound area where the tester is located) and the non-target-channel audio (i.e., the audio from other sound areas or mainly containing interference noise) are identified and distinguished. Through the comparative analysis of the energy distribution, the system can accurately determine the source of each audio signal and judge whether it belongs to the target audio that needs to be processed preferentially. Such analysis can effectively exclude the interference of background noise and other irrelevant sound sources, so as to ensure that the system can focus on the wake-up recognition of the target audio, significantly reducing the false wake-up caused by environmental noise or non-target speech interference. In a complex multi-audio environment, this processing ensures that the system can accurately focus on the correct audio signal, greatly improving the wake-up accuracy and the stability of the system.
[0118] The approximate words of the target-channel audio are normally absorbed and intercepted, and the wake-up results of the non-target-channel audio are ignored.
[0119] For the approximate words of the target-channel audio, absorption and interception are still carried out to handle the approximate words in the target sound area, so as to effectively prevent false wake-up. For the approximate words in the non-target-channel audio, if a false wake-up signal appears, the wake-up response of this signal is selected to be ignored. This process accurately analyzes the audio channels of the signal sources and adopts corresponding processing strategies according to the sources of different audio signals. Through the targeted processing of the target-channel audio and the non-target-channel audio, the system significantly reduces the risk of false wake-up caused by the non-target sound area (which may contain background noise or irrelevant signals), thus optimizing the efficiency and accuracy of the wake-up engine. In addition, the standard absorption and interception mechanism for the target-channel audio ensures that the device can respond efficiently and accurately when recognizing voice commands, avoiding the negative impact of external interference or non-related audio signals on the wake-up effect.
[0120] When there is an overlap in the processing of the target-channel audio and the non-target-channel audio, the wake-up result of the target-channel audio is preferentially used.
[0121] When the target path audio overlaps with the non-target path audio at a certain moment, the system will preferentially adopt the wake-up result of the target path audio and ignore the influence of the non-target path audio on the wake-up decision. By focusing the processing on the target path audio source, ensuring that the wake-up decision is based on the most relevant and clear signal source, the strategy of preferentially processing the target path audio effectively enhances the accuracy and consistency of the system. Especially in the case of overlapping multi-channel audio signals, it helps to improve the reliability of the wake-up response. By setting priorities, the system can effectively avoid conflicts or mutual interference between multi-channel audio, thereby enhancing stability in complex audio environments. This mechanism ensures that the system can accurately respond to the wake-up instructions of the tester and minimize false wake-up interference caused by speech or background noise from non-target areas.
[0122] The following further explains the intelligent device false wake-up interception method of the present invention through a specific implementation process.
[0123] Figure 5 The figure shows a score schematic diagram of different human voice audios for the target wake-up word "Hello, Dreame" according to an embodiment of the present invention. As Figure 5 shown, the black part represents the score change of the common prefix "Hello" of the human voice, and the blue part is the score change of the human voice of the "Dreame" part in "Hello, Dreame". When the score reaches the threshold of 0.8, the device will perform a wake-up response. The green part represents the score trend of the human voice of the "XGIMI" part in "Hello, XGIMI". Although the score of this part lags behind that of the normal wake-up word, its final score may still exceed the threshold, resulting in a false wake-up of the system.
[0124] To avoid the above problems, the present invention optimizes by adding the known approximate word "Hello, XGIMI" to the absorption wake-up word list. Different from the target wake-up word, once this approximate word reaches the wake-up condition, the system does not perform the wake-up operation, thereby preventing false wake-up. At this time, the wake-up word list of the intelligent device contains two wake-up words: the target wake-up word "Hello, Dreame" and the absorption wake-up word "Hello, XGIMI". Whenever a human voice is received, the device will calculate the scores of the human voice for these two wake-up words simultaneously. If the target wake-up word "Hello, Dreame" reaches the wake-up condition first, the system will clear the previous audio signal and perform the wake-up; if the absorption wake-up word "Hello, XGIMI" reaches the wake-up condition first, the previous human voice signal will be cleared, and the system will not make any response to avoid false wake-up.
[0125] Figure 6 The figure shows a score schematic diagram of the human voice audio "Hello, Dreame" for the target wake-up word and the absorption wake-up word according to an embodiment of the present invention.
[0126] Refer to Figure 6, in this solution, in order to effectively reduce false awakenings without affecting normal awakenings, it is necessary to finely adjust the threshold of the absorption word. Specifically, first set the initial threshold of the absorption wake-up word "Hello, XGIMI" according to the actual usage. For example, when the threshold of the absorption wake-up word "Hello, XGIMI" is set to 0.7, when the user shouts "Hello, Dreame", the wake-up point W of the target wake-up word will be ahead of the wake-up point A of the absorption wake-up word, and the system can normally wake up the device. At this time, the absorption wake-up word will not interfere with the normal wake-up.
[0127] However, if the threshold of the absorption wake-up word is adjusted to 0.6, the wake-up point W of the target wake-up word "Hello, Dreame" will lag behind the wake-up point B of the absorption replacement word, and the normal wake-up signal will be wrongly intercepted as a false awakening, affecting the response of the device. Therefore, in order to ensure that the device can normally respond to the wake-up instruction without being interfered by false awakenings, the threshold of the absorption wake-up word should be gradually adjusted from high to low to precisely set a critical value so that the wake-up point of the absorption word always lags behind the wake-up point of the target wake-up word.
[0128] Through this parameter adjustment process, it is possible to effectively avoid false awakening phenomena while maintaining a high wake-up rate, thereby optimizing the wake-up experience of intelligent devices. This method ensures that the false awakening interception scheme has the least impact on the wake-up rate, achieving the best balance between the performance of intelligent devices and user experience.
[0129] In some embodiments, the present invention can add the absorption wake-up word "speech nihao zhui mi" to the absorption wake-up word library, where "speech" is the score summary of all pinyin units. When the voice contains a target wake-up word that is not at the beginning of a sentence, the system absorbs the preceding voice segment through approximate words, processes these audio signals, and prevents the wake-up response. When a clear wake-up word "Hello, Dreame" at the beginning of a sentence is detected, the system will execute the wake-up command, which can effectively reduce false awakenings caused by context and improve the accuracy and reliability of device wake-up.
[0130] According to the above embodiments, by dynamically adjusting the wake-up threshold of the absorption wake-up word, the accuracy and flexibility of the system to intercept false awakenings are improved. By adjusting the wake-up threshold of the absorption wake-up word from high to low, it can flexibly adapt to the voice signal characteristics under different environmental conditions and effectively avoid the situation of false interception. On the premise of ensuring the recognition accuracy of the target wake-up word, it can trigger an interception reaction in time when the absorption wake-up word reaches the preset wake-up threshold to prevent false awakenings. At the same time, based on the optimization of the score calculation and processing logic, the system can accurately judge and eliminate misrecognized absorption wake-up words to avoid incorrect responses. The technical solution of the present invention enables intelligent devices to dynamically adjust and optimize the recognition strategy of wake-up words, thereby enhancing the intelligence, refinement, and user experience of the system. Especially in complex or changing usage environments, it can execute wake-up instructions more stably and accurately.
[0131] An embodiment of the present invention further provides a system for intercepting accidental wake-up of an intelligent device. As Figure 7 shown, the system for intercepting accidental wake-up of the intelligent device includes:
[0132] An approximate word generation module 101, configured to generate a list of approximate words according to the voice sample of the target wake-up word.
[0133] A threshold setting module 102, configured to set the wake-up threshold of the target wake-up word through testing, and set corresponding wake-up thresholds for each approximate word based on the wake-up threshold of the target wake-up word.
[0134] An audio analysis module 103, configured to obtain human voice audio and calculate the scores of the human voice audio for the target wake-up word and each approximate word.
[0135] A signal processing module 104, configured to clear the received human voice audio signal and wake up the device when the score of the human voice audio for the target wake-up word reaches the wake-up threshold of the target wake-up word first; when the score of the human voice audio for the approximate word reaches the wake-up threshold of the approximate word first, clear the received human voice audio signal and make no response.
[0136] Regarding the specific implementation manners of setting the wake-up thresholds and rules for the target wake-up word and the approximate words in the above-mentioned system for intercepting accidental wake-up of an intelligent device, refer to the relevant content of the embodiments in the above method, and details are not described herein again.
[0137] An embodiment of the present invention further provides an intelligent device for intercepting accidental wake-up, including: a processor 201, a memory 202, and a computer program stored in the memory 202 and configured to be executed by the processor 201. When the processor 201 executes the computer program, it implements the method for intercepting accidental wake-up of an intelligent device according to any of the above embodiments.
[0138] When the processor 201 executes the computer program, it implements the steps in the above embodiments of the method for intercepting accidental wake-up of an intelligent device, such as Figure 1 all the steps of the method for intercepting accidental wake-up of an intelligent device shown. Or, when the processor 201 executes the computer program, it implements the functions of each module / unit in the above-mentioned system for intercepting accidental wake-up of an intelligent device, such as Figure 7 the functions of each module of the system for intercepting accidental wake-up of an intelligent device shown.
[0139] Exemplarily, the computer program can be divided into one or more modules. One or more modules are stored in the memory 202 and executed by the processor 201 to complete the present invention. One or more modules can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the system for intercepting accidental wake-up of an intelligent device.
[0140] The so-called processor 201 can be a Central Processing Unit (CPU), or it can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The processor 201 is the control center of the intelligent device accidental wake-up interception system, and connects all parts of the intelligent device accidental wake-up interception system through various interfaces and lines.
[0141] The memory 202 can be used to store computer programs and / or modules. The processor 201 realizes various functions of the intelligent device accidental wake-up interception system by running or executing the computer programs and / or modules stored in the memory 202, and by calling the data stored in the memory 202. The memory 202 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the intelligent device accidental wake-up interception system, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.
[0142] Among them, if the modules / units of the intelligent device misawakening interception system are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0143] The embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded in a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above-mentioned embodiment is implemented.
[0144] The above-mentioned embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. A method for intercepting false wake-up of a smart device, characterized in that: The method comprises the following steps: Get a list of similar words; Adding the similar words in the similar word list to the absorption wake-up word list; Set the wake-up threshold of the target wake-up word through testing; Setting a wake-up threshold corresponding to each approximate word in the absorbed wake-up word list based on the wake-up threshold of the target wake-up word; Acquire human voice audio, and score the human voice audio to obtain the score of the human voice audio corresponding to each of the approximate words and the score of the human voice audio for the target wake-up word; When the score of the human voice audio for the target wake-up word first reaches the wake-up threshold of the target wake-up word, the received human voice audio signal will be cleared and the device will be woken up; when the score of the human voice audio for the approximate word first reaches the wake-up threshold of the approximate word, the received human voice audio signal will be cleared and no response will be made.
2. The method for intercepting false wakeup of a smart device according to claim 1, characterized in that: The method of obtaining a list of similar words comprises the following steps: Collecting speech samples of the target wake-up word uttered by multiple users to obtain corresponding audio signals, wherein the users include sample groups with different vocalization habits and dialects; Performing speech recognition on the audio signal to generate a text result corresponding to the audio signal, the text result including a correctly recognized target wake-up word text and an approximate word text generated due to unclear pronunciation of the user or recognition deviation; According to the text results, by analyzing the similarity of the text's speech features and the proximity of the phonetic unit sequences, similar words that are easily misidentified as target wake-up words are screened out, and a list of similar words is generated.
3. The method for intercepting false wakeup of a smart device according to claim 1, characterized in that: The step of setting the wake-up threshold of the target wake-up word through testing includes the following steps: Collect audio samples in multiple different scenarios, including samples of the correct pronunciation of the target wake-up word and samples of background noise or other irrelevant speech; Performing acoustic feature analysis on the audio samples to calculate the score of the target wake-up word and the score of the background noise or other irrelevant speech in each sample; Based on the score analysis result, an initial wake-up threshold of a target wake-up word is selected, and the initial wake-up threshold of the target wake-up word is tested using audio samples in multiple different environments; Observe the number of false awakenings and adjust the awakening threshold of the initial target wake-up word until a preset number of false awakenings is reached within a specified time; Determine the final wake-up threshold of the target wake-up word.
4. The method for intercepting false wakeup of a smart device according to claim 1, characterized in that: The step of setting the wake-up threshold corresponding to each approximate word in the absorbed wake-up word list based on the wake-up threshold of the target wake-up word comprises the following steps: Based on the set wake-up threshold of the target wake-up word, collect human voice audio samples; Calculate the scores of the human voice audio sample for the target wake-up word and each similar word respectively, and draw score distribution curves of the target wake-up word and each similar word; Based on the wake-up threshold of the target wake-up word, the wake-up threshold of each approximate word is adjusted from high to low to ensure that the wake-up point of the target wake-up word is not later than the wake-up point of each approximate word.
5. The method for intercepting false wakeup of a smart device according to claim 3, characterized in that: The step of adjusting the initial target wake-up word wake-up threshold according to the number of false awakenings until a preset number of false awakenings is reached within a specified time includes the following steps: Performing speech recognition on the audio clip of each false awakening point, generating corresponding false awakening text, and comparing it with similar words in the absorption awakening word list; When the false awakening text belongs to an approximate word in the list of absorbed awakening words, ignore the false awakening; when the false awakening text does not belong to an approximate word in the list of absorbed awakening words, mark it as a valid false awakening and record it; When the effective number of false awakenings exceeds the preset number of false awakenings, the awakening threshold of the target awakening word is increased until the effective number of false awakenings reaches the preset number of false awakenings within a specified time.
6. The method for intercepting false wakeup of a smart device according to claim 1, characterized in that: The obtaining of human voice audio, scoring the human voice audio, and obtaining the score of the human voice audio corresponding to each of the approximate words and the score of the human voice audio for the target wake-up word comprises the following steps: Obtain the audio features of the target wake-up word and each similar word, and calculate the probability distribution of all phonetic units in each frame of audio in combination with the acoustic model; Based on the probability of phonetic units in each frame, a combined algorithm is used to calculate the comprehensive scores of the target wake-up word and each similar word at the word level.
7. The method for intercepting false wakeup of a smart device according to claim 1, characterized in that: Also includes: Get multi-channel audio signals; Based on the comparative analysis of the energy distribution of multiple audio signals, the target audio channel and the non-target audio channel are determined; Normally absorb and intercept the similar words of the target channel, and ignore the wake-up results of the non-target channel audio; When the target channel audio and non-target channel audio processing overlap, the wake-up result of the target channel audio is used first.
8. A smart device false wakeup interception system, characterized in that: The system comprises: An approximate word generation module, configured to generate an approximate word list according to a speech sample of a target wake-up word; A threshold setting module, configured to set a wake-up threshold of a target wake-up word through testing, and to set a corresponding wake-up threshold for each similar word based on the wake-up threshold of the target wake-up word; An audio analysis module, configured to obtain human voice audio and calculate a score of the human voice audio for the target wake-up word and each similar word; The signal processing module is configured to clear the received human voice audio signal and wake up the device when the score of the human voice audio for the target wake-up word first reaches the wake-up threshold of the target wake-up word; and clear the received human voice audio signal without making any response when the score of the human voice audio for the approximate word first reaches the wake-up threshold of the approximate word.
9. A device, characterized in that: include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform an operation corresponding to the method for intercepting false wake-up of a smart device as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for intercepting false wake-up of a smart device according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Speech recognition model training method and device, equipment, storage medium and product
CN114299933A
Detection method and device for false identification of similar sounds and computer equipment
CN114333799A
Voice wake-up method and device
CN115966199A
Keyword detection model training method, electronic equipment and storage medium
CN116110376A
Method and device for waking up via speech based on artificial intelligence
US20180158449A1