Intelligent household appliance awakening system and method based on voice control
By combining ultra-low power acoustic event detection with an adaptive noise-robust wake-up word recognition unit, the problem of balancing low power consumption and high robustness in existing technologies is solved, and a voice wake-up system with high accuracy and low false wake-up rate in complex environments is realized.
Patent Information
- Application Number
- CN202511201210.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-11-04
AI Technical Summary
Existing local voice wake-up technology is prone to high false alarm and missed alarm rates in complex home environments, making it difficult to balance low power consumption and high robustness, resulting in a poor user experience.
A two-level wake-up architecture is adopted, consisting of an ultra-low power acoustic event detection unit and an adaptive noise-robust wake-up word recognition unit. Combined with an environmental context awareness unit and a system control and power management unit, accurate wake-up word recognition is achieved through multi-level noise suppression and adaptive threshold adjustment.
While maintaining low power consumption, it significantly reduces false wake-up rate and missed wake-up rate, improves the accuracy of wake word recognition and user experience, adapts to different environments and user behaviors, and achieves personalized response.
Smart Images

Figure CN120895038A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence technology, specifically a voice-controlled smart wake-up system and method for home appliances. Background Technology
[0002] With the continuous evolution of the Internet of Things, artificial intelligence, and human-computer interaction technologies, smart homes have become an irreversible trend in modern life. Among these technologies, voice control, due to its intuitive and convenient characteristics, has gradually become one of the core ways for users to interact with smart home appliances. Technological advancements in this field have greatly improved the user experience, making daily chores, information access, and entertainment more intelligent and efficient. Specifically, voice control systems collect users' voice commands and, through a series of complex signal processing, feature extraction, and pattern recognition processes, ultimately convert acoustic signals into machine-understandable instructions, thereby driving smart home appliances to perform corresponding operations.
[0003] In the existing technological system, realizing voice wake-up functionality in home appliances typically relies on two main technical approaches. One is a cloud-based speech recognition solution. This solution uploads user-collected voice data to a remote server and utilizes high-performance cloud computing resources to perform large-scale acoustic and language model calculations, thereby achieving high-accuracy, large-vocabulary speech recognition. The advantage of this approach lies in its powerful computing capabilities and continuous iterative model optimization, enabling the system to handle complex voice commands and adapt to diverse languages and accents. However, its inherent drawbacks are also significant, primarily its strong dependence on network connectivity and the resulting communication latency. Especially when network conditions are poor or server response is slow, the user experience will be greatly diminished. Furthermore, the cloud uploading and processing of user voice data has raised growing privacy and security concerns, which are particularly sensitive in smart home environments.
[0004] Secondly, there is Keyword Spotting (KWS) technology based on the local device. The core of this solution lies in the device's built-in low-power voice processing module, which continuously monitors preset wake words. Once a matching wake word is detected, it triggers a subsequent system response. The significant advantages of this solution are its low latency, lack of network connection requirement, and good user privacy protection. Because voice data is processed locally, the risk of sensitive information leakage is avoided, while also enabling rapid response to user wake-up commands. Therefore, KWS technology has been widely used in smart home appliances such as smart speakers and smart TVs, which have high requirements for power consumption and response speed. Specifically, traditional KWS systems typically employ lightweight acoustic models, such as those based on Dynamic Time Warping (DTW) or Shallow Neural Networks, combined with pre-trained wake word templates for matching. These models are designed to achieve 24 / 7 monitoring with extremely low computational resource consumption, thereby maintaining the device's low-power standby state to extend battery life or reduce energy costs.
[0005] However, with the continuous development of related technologies and users' increasingly stringent requirements for smart home experiences, some inherent characteristics of the aforementioned local KWS technology solutions at the principle level have gradually revealed their limitations in addressing new challenges, evolving into a deep-seated technical contradiction. The reason for this is that traditional local KWS solutions, in order to achieve extremely low power consumption, often have to sacrifice model complexity and robustness. This means they face severe performance challenges in real, complex home environments. For example, in the presence of interference factors such as background music, television playback, multiple conversations, or environmental noise (such as air conditioning or fan noise), traditional KWS models are prone to high false alarm rates (FAR) and false rejection rates (FRR). High false alarm rates lead to frequent false wake-ups of home appliances without explicit wake-up commands, not only causing unnecessary power consumption but also potentially disrupting users' normal lives or work, causing unnecessary inconvenience and privacy concerns. In contrast, a high false alarm rate manifests as the device failing to respond after the user repeatedly calls the wake word, severely damaging the smoothness and convenience of the user experience, rendering the advantages of "intelligent wake-up" meaningless, and even causing user frustration.
[0006] Furthermore, the limited vocabulary and relatively fixed acoustic models of existing KWS technologies make it difficult to adapt to users' diverse pronunciation habits, speech rates, and accents, thus exacerbating false alarms and missed alarms. To improve the accuracy and robustness of KWS, researchers have attempted to introduce more complex deep learning models, but this often means higher computational resource requirements and power consumption, creating a sharp conflict with the strict constraints of low power consumption and low cost on local devices, making such performance improvements difficult to widely apply in practical products. In other words, in the pursuit of low power consumption, local processing to protect privacy, and ensuring low latency, existing KWS technologies face an inherent contradiction in achieving a "smart wake-up" function with high robustness, low false alarm rate, and low missed alarm rate. Users expect a wake-up mechanism that can respond accurately and quickly in any environment, but current technologies, under limited resource constraints, cannot effectively balance high wake-up accuracy and low false wake-up rate while maintaining extremely low power consumption, resulting in a generally poor user experience.
[0007] Therefore, the present invention provides a smart wake-up system and method for home appliances based on voice control. Summary of the Invention
[0008] In order to overcome the shortcomings of the prior art, at least one technical problem raised in the background art is solved.
[0009] The technical solution adopted by this invention to solve its technical problem is: a voice-controlled smart wake-up system for home appliances, comprising: The voice acquisition unit (MAU) is used to continuously acquire acoustic signals in the environment and convert them into digital audio data; An ultra-low power acoustic event detection unit (UAED) is electrically connected to the voice acquisition unit. It is used to perform preliminary processing on the digital audio data to detect whether there are potential acoustic events or wake-up word fragments. When the potential acoustic events or wake-up word fragments are detected, an activation signal is sent to the system control and power management unit. An Adaptive Noise Robust Wake-up Word Recognition Unit (ANR-KWS) is electrically connected to the ultra-low power acoustic event detection unit and the system control and power management unit. After receiving a wake-up command from the system control and power management unit, it performs in-depth analysis and accurate recognition of the potential acoustic events or wake-up word fragments contained in the digital audio data collected by the voice acquisition unit to determine whether they are preset wake-up words. An Environmental Context Awareness Unit (CAM), electrically connected to the system control and power management unit, is used to monitor the ambient noise level, user activity status, and other relevant context information in real time, and transmit the context information to the system control and power management unit. The System Control and Power Management Unit (SCU&PMU) is electrically connected to the Ultra-Low Power Acoustic Event Detection Unit, the Adaptive Noise Robust Wake-up Word Recognition Unit, and the Environmental Context Awareness Unit. It is used to receive the activation signal sent by the Ultra-Low Power Acoustic Event Detection Unit, dynamically adjust the recognition threshold of the Adaptive Noise Robust Wake-up Word Recognition Unit based on the context information provided by the Environmental Context Awareness Unit, and coordinately control the power status and operating mode of each unit.
[0010] In a preferred embodiment of the present invention, the voice acquisition unit includes at least two high-performance microelectromechanical systems (MEMS) microphones arranged in an array for multi-channel acoustic signal acquisition. The microphone array converts the acquired analog acoustic signals into a high-resolution digital audio stream in real time via digital pulse density modulation (PDM) or an inter-integrated circuit audio bus (I2S) interface. The digital audio stream has a sampling frequency of 16 kHz and a quantization bit depth of 16 bits. The voice acquisition unit further includes a low-latency digital signal processor (DSP) core equipped with a preprocessing module for preliminary array signal processing of the digital audio stream, including but not limited to delay-and-sum-based beamforming algorithms to enhance speech signals in specific directions and suppress ambient noise. The preprocessing module can also implement adaptive beamforming algorithms, such as minimum variance distortionless response (MVDR) beamformers, to achieve better noise suppression and signal separation, thereby improving the signal input quality of subsequent units.
[0011] In a preferred embodiment of the present invention, the ultra-low power acoustic event detection unit (UAED) includes: Ultra-low power digital signal processing core, which has microwatt-level power consumption characteristics for continuous operation around the clock; The Voice Activity Detection (VAD) module, configured within the ultra-low power digital signal processing core, is used to perform real-time analysis of the digital audio data based on features such as short-time energy, zero-crossing rate, spectral entropy, or fundamental frequency change, combined with a Gaussian mixture model (GMM) or a shallow neural network model, in order to distinguish between speech segments and non-speech segments. The pre-wake word feature extraction module is configured within the ultra-low power digital signal processing core and is used to extract a small number of acoustic features from the speech segment when the speech activity detection module recognizes the speech segment, such as low-dimensional mezzanine frequency cepstral coefficients (MFCC, e.g., 4th-6th order) or energy envelope. A lightweight matching module, configured within the ultra-low power digital signal processing core, performs pattern matching between the pre-wake word features and a pre-trained, highly simplified wake word template. The template can be implemented based on a Dynamic Time Warping (DTW) algorithm or a Shallow Neural Network to quickly determine the presence of potential wake word fragments. The recognition threshold of the lightweight matching module is set to a relatively lenient value to minimize the false alarm rate and ensure that all potential wake word fragments are captured. When the matching score of the lightweight matching module exceeds a preset initial threshold, the ultra-low power acoustic event detection unit generates a hardware interrupt signal or an event trigger signal and sends it to the system control and power management unit.
[0012] In a preferred embodiment of the present invention, the adaptive noise robust wake-up word recognition unit (ANR-KWS) includes: A high-performance, low-power embedded digital signal processor or a dedicated artificial intelligence (AI) acceleration core, which has milliwatt-level operating power consumption and is in a deep sleep state most of the time, only starting up rapidly when it receives a wake-up command from the system control and power management unit; An advanced noise suppression (ANR) module, configured within the embedded digital signal processor or AI acceleration core, is used to perform refined noise removal on wake-word segments in the digital audio data provided by the speech acquisition unit. The ANR module employs spectral subtraction, Wiener filtering, or deep learning-based speech enhancement algorithms (such as lightweight temporal convolutional networks or residual networks) to effectively suppress background noise and reverberation, significantly improving the signal-to-noise ratio of the speech signal. A robust feature extraction module, configured within the embedded digital signal processor or AI acceleration core, is used to extract richer and more robust acoustic features from noise-suppressed speech segments, such as full-order Mel-frequency cepstral coefficients (MFCC, e.g., 13-26th order), fundamental frequency (Pitch), formants, linear predictive cepstral coefficients (LPCC), or deep learning-based acoustic embeddings (such as x-vectors or d-vectors). A deep learning wake word recognition model is configured within the embedded digital signal processor or AI acceleration core. The model employs a quantized (e.g., 8-bit integer quantization) convolutional neural network (CNN) architecture, a recurrent neural network (RNN, such as a gated recurrent unit (GRU) or a long short-term memory (LSTM) network), or a lightweight Transformer architecture based on an attention mechanism. The model is trained on a large-scale wake word dataset and a diverse set of background noise datasets, exhibiting high accuracy and strong robustness. The model takes the robust acoustic features as input and outputs a confidence score representing the probability of the wake word's presence. An adaptive decision logic module, configured within the embedded digital signal processor or AI acceleration core, is used to compare the confidence score output by the deep learning wake word recognition model with the recognition threshold dynamically set by the system control and power management unit to make a final wake word recognition decision.
[0013] In a preferred embodiment of the present invention, the environment context-aware unit (CAM) includes: An ambient noise sensor, such as the digital microphone itself integrated into the voice acquisition unit, quantifies the ambient noise level in real time by calculating the short-time root mean square energy of the audio stream, the signal-to-noise ratio estimate, or the energy distribution of a specific frequency band, and outputs noise level data in decibels (dB).
[0014] Ambient light sensors, such as silicon photodiodes or ambient light sensors (ALS), are used to measure the light intensity (Lux value) in the environment to help determine the usage status of the environment, such as day or night, and whether the user may be asleep.
[0015] One or more user activity status sensors, such as passive infrared (PIR) sensors, millimeter-wave radar sensors, or accelerometers, are used to detect the physical movement or presence of a user near the device to generate a user activity index. An internal status recording module, stored in the non-volatile memory of the system control and power management unit, is used to record historical wake-up success rate, number of false wake-ups, frequency of user manual interaction, and current operating status of home appliances (e.g., whether music or TV programs are being played).
[0016] The environmental context awareness unit integrates the aforementioned environmental noise level, light intensity, user activity index, and internal state data into a multi-dimensional context vector, and transmits it to the system control and power management unit in real time.
[0017] In a preferred embodiment of the present invention, the system control and power management unit (SCU&PMU) includes: The master microcontroller, such as a microcontroller based on the ARM Cortex-M4 or Cortex-M7 architecture, is used to coordinate the work of each unit and execute complex control logic; The power management module, integrated into the main control microcontroller or existing as an independent chip, is used to finely control the power supply voltage and current of each functional unit and switch the power state of each unit in different working modes, including but not limited to: ultra-low power standby mode, partial wake-up mode and full-function operation mode. An adaptive threshold adjustment module, configured within the main control microcontroller, receives a multi-dimensional context vector provided by the environmental context awareness unit and dynamically calculates and updates the recognition threshold of the adaptive noise-robust wake-word recognition unit based on a preset adaptive strategy or machine learning model. The adaptive strategy includes: increasing the recognition threshold to reduce the false alarm rate when the environmental noise level exceeds a preset threshold; appropriately decreasing the recognition threshold to improve response sensitivity when the user activity index is high or there is frequent manual interaction; and further adjusting the threshold based on the voice characteristics of the playback content when the home appliance is in media playback mode. The adaptive threshold adjustment module can employ fuzzy logic reasoning, decision tree models, or reinforcement learning-based strategies for threshold optimization. The system state machine management module, configured within the main control microcontroller, is used to drive the system to switch between different power states and processing modes based on the activation signal of the ultra-low power acoustic event detection unit and the recognition result of the adaptive noise robust wake word recognition unit. After successful wake word recognition, it sends a wake-up command to the main processor of the home appliance, causing it to switch from a low-power state to an active state.
[0018] The present invention provides a voice-controlled smart wake-up method for home appliances, applicable to the aforementioned system, comprising the following steps: S1. Ultra-low power standby monitoring: The voice acquisition unit continuously acquires ambient acoustic signals at a preset low sampling rate and low bit depth, and transmits the digital audio stream to the ultra-low power acoustic event detection unit. S2. Preliminary Acoustic Event Detection: The ultra-low power acoustic event detection unit continuously runs the Voice Activity Detection (VAD) module and the lightweight matching module. The VAD module determines the presence of speech based on features such as short-time energy and zero-crossing rate. Once a speech segment is detected, the pre-wake word feature extraction module extracts a small number of acoustic features from the speech segment, and the lightweight matching module determines whether there is a potential wake word fragment by performing pattern matching with a simplified wake word template. If the matching score exceeds a preliminary threshold, the ultra-low power acoustic event detection unit sends an activation signal to the system control and power management unit and caches the digital audio data of the potential wake word fragment. S3. Context information perception: The environmental context perception unit collects and integrates environmental noise level, light intensity, user activity index and historical interaction data in real time, generates a multi-dimensional context vector, and transmits it to the system control and power management unit. S4. Adaptive Threshold Calculation and Wake-up Unit Activation: After receiving the activation signal, the system control and power management unit first wakes up the power supply of the adaptive noise robust wake-up word recognition unit from deep sleep state, and obtains the latest multi-dimensional context vector from the environmental context awareness unit. Based on the context vector and the preset adaptive strategy, the system control and power management unit dynamically calculates and sets the precise recognition threshold required for the current wake-up cycle, and transmits the threshold to the adaptive noise robust wake-up word recognition unit. S5. Precise Wake-up Word Recognition: The adaptive noise-robust wake-up word recognition unit receives digital audio data of cached potential wake-up word segments. The advanced noise suppression module performs refined noise removal on the segments. Subsequently, the robust feature extraction module extracts multi-dimensional robust acoustic features from the denoised speech segments. The deep learning wake-up word recognition model uses the robust acoustic features as input to calculate a confidence score representing the probability of the wake-up word's presence. S6. Wake-up Decision and Appliance Control: The adaptive decision logic module of the adaptive noise robust wake-up word recognition unit compares the confidence score with the precise recognition threshold set by the system control and power management unit. If the confidence score is higher than the threshold, the wake-up word recognition is confirmed to be successful. The system control and power management unit sends a wake-up command to the main processor of the appliance, causing it to switch from a low-power state to a working state and start the subsequent voice command recognition module. If the confidence score is not higher than the threshold, it is determined to be a non-wake-up word or a false trigger. The adaptive noise robust wake-up word recognition unit returns to a deep sleep state, and the ultra-low power acoustic event detection unit continues to execute step S1.
[0019] In a preferred embodiment of the present invention, the voice acquisition unit acquires audio at a lower sampling rate (e.g., 8 kHz) and a lower bit depth (e.g., 12 bits) in S1 to reduce the computational load on the ultra-low power acoustic event detection unit. In S1, the adaptive noise-robust wake-word recognition unit can instruct the voice acquisition unit to resample or recapture the cached wake-word fragments at a higher sampling rate (e.g., 16 kHz or 24 kHz) and a higher bit depth (e.g., 16 bits or 24 bits) to provide higher quality audio data for accurate recognition. This approach further balances power consumption and recognition accuracy.
[0020] In a preferred embodiment of the present invention, the lightweight matching module of the ultra-low power acoustic event detection unit is implemented using a mini model based on a recurrent neural network (RNN) or a gated recurrent unit (GRU). The mini model has a very small number of parameters and undergoes deep quantization processing. Its operation is mainly performed on a dedicated computing unit of the ultra-low power digital signal processing core to ensure that the initial voice matching function is achieved at sub-milliwatt power consumption.
[0021] In a preferred embodiment of the present invention, the deep learning wake-up word recognition model of the adaptive noise-robust wake-up word recognition unit adopts a convolutional neural network architecture based on depthwise separable convolution. The model parameters have been pruned and quantized to adapt to the computational limitations of the embedded hardware platform. The model is trained to withstand various common background noises (such as white noise, pink noise, music, television audio, keyboard clicks, etc.) to improve robustness in complex acoustic environments.
[0022] In a preferred embodiment of the present invention, in S4, the adaptive threshold adjustment module uses a pre-trained decision model (e.g., Support Vector Machine (SVM) or a small neural network) to map the context vector to a dynamic adjustment factor. This factor is used to multiplicatively or additively adjust the baseline recognition threshold. For example, when the ambient noise level is higher than 60 dB, the recognition threshold is increased by 0.05; when the user activity index is at rest and the ambient light intensity is lower than 10 Lux, the recognition threshold is decreased by 0.02 to improve response sensitivity in quiet nighttime environments.
[0023] In a preferred embodiment of the present invention, the system further includes a user feedback and model optimization module configured in the system control and power management unit. This module is used for: Record wake-up success and failure events, including: the validity of the wake word command, the recognition result (success / failure), the current context information, and the confidence score of the recognition model; Collect user feedback data when users perform manual operations or express their satisfaction with the wake-up result through other means; Based on accumulated recorded data and user feedback, the adaptive strategy of the adaptive threshold adjustment module is periodically iteratively optimized through offline or online methods. When the device connects to the network, it anonymizes and uploads a portion of the recorded data to the cloud server for large-scale model retraining. After the model is updated in the cloud, the updated deep learning wake word recognition model parameters or quantized weights are downloaded to the adaptive noise-robust wake word recognition unit via a secure channel, enabling online model updates and continuous performance improvement. The uploaded data strictly adheres to user privacy protection protocols, containing only non-sensitive acoustic features, recognition results, and contextual metadata, and does not include raw speech data or any information that could identify the user.
[0024] The voice-controlled smart wake-up system and method for home appliances provided by this invention achieves a balance between power consumption and performance by constructing a two-level wake-up architecture consisting of an ultra-low-power acoustic event detection unit and an adaptive noise-robust wake-up word recognition unit. The ultra-low-power acoustic event detection unit performs all-weather, low-cost monitoring and initial screening, significantly reducing overall average power consumption. The adaptive noise-robust wake-up word recognition unit, after precise activation, utilizes more powerful computing resources for high-precision, interference-resistant wake-up word recognition. The linkage between the environmental context awareness unit and the system control and power management unit allows the wake-up word recognition threshold to be dynamically adjusted based on real-time environmental conditions and user behavior, thereby effectively reducing the false wake-up rate (false alarm rate) and missed wake-up rate (missed alarm rate). Specifically, in noisy environments… By increasing the recognition threshold and utilizing advanced noise suppression techniques, false activations caused by background noise can be reduced. In quiet or active user environments, appropriately lowering the recognition threshold can improve wake-up sensitivity and response speed, meeting users' personalized needs in different scenarios. At the hardware level, the system adopts a hierarchical heterogeneous computing architecture, allocating power-sensitive tasks to ultra-low-power processors and computationally intensive tasks to high-performance processors, and optimizing energy consumption through refined power management. At the software level, multimodal context fusion and adaptive decision-making algorithms enhance the system's intelligence and robustness. Therefore, this invention significantly improves the accuracy and user experience of voice wake-up while maintaining low power consumption and real-time performance, effectively solving the technical challenge of balancing high accuracy and low power consumption in existing technologies.
[0025] The present invention provides a voice-controlled smart wake-up system and method for home appliances. At the hardware level, the technical solution is specifically deployed as follows: the MEMS microphone array of the voice acquisition unit is directly connected via an I2S interface to the ultra-low-power digital signal processing core (e.g., a customized RISC-V processor or ARM Cortex-M0+ core with low-power DSP instruction set extensions) of the ultra-low-power acoustic event detection unit. This core integrates a low-power analog front-end (AFE) and an analog-to-digital converter (ADC). The ultra-low-power acoustic event detection unit is connected to the main control microcontroller (e.g., an ARM Cortex-M4 core) of the system control and power management unit via a GPIO interrupt line and an SPI or I2C bus. The high-performance, low-power embedded digital signal processor (e.g., an ARM Cortex-M7 core with an AI acceleration unit) of the adaptive noise-robust wake-up word recognition unit receives the digital audio data buffered by the ultra-low-power acoustic event detection unit via a high-speed SPI or QSPI bus, and interacts with the main control microcontroller for control commands and recognition results via a UART or I2C bus. The various sensors of the environmental context awareness unit (e.g., environmental noise measurement circuit, photoresistor or digital light sensor, PIR sensor) are connected to the main control microcontroller via analog input, I2C, or SPI interface. The power management module integrated into the main control microcontroller or a separate power management integrated circuit (PMIC) manages the power distribution and mode switching of the entire system, including enabling / disabling the power supply to the ultra-low power acoustic event detection unit, the adaptive noise robust wake-word recognition unit, and other peripherals, as well as setting low-power modes. The PMIC is configured with multiple programmable voltage output channels, enabling it to independently control the power rails of each unit according to the instructions of the main control microcontroller, thereby achieving fine-grained power management.
[0026] Furthermore, the voice acquisition unit comprises at least two MEMS microphones configured in a linear or circular array, with their spacing optimized based on the microphone's directivity and the target frequency band characteristics. For example, for the human voice frequency band from 2kHz to 8kHz, the microphone spacing can be set to 5cm to 10cm. The microphones have a sensitivity of -26dBFS±1dB, a signal-to-noise ratio (SNR) of 64dB, and an acoustic overload point (AOP) of 120dBSPL, enabling high-fidelity capture of voice signals with a wide dynamic range. The beamforming algorithm in the preprocessing module is implemented using an 8th-order FIR filter to effectively filter out interference from non-target directions.
[0027] Furthermore, the ultra-low power acoustic event detection unit's ultra-low power digital signal processing core operates at a clock frequency of 10MHz, with a memory configuration of 64KBSRAM and 256KB flash memory. The VAD module employs an acoustic classifier based on a Gaussian mixture model (GMM), which has four mixture components, each containing 12-dimensional MFCC features, with only a few kilobytes of parameters. The lightweight matching module uses a single-layer long short-term memory (LSTM) network with 16 hidden layer nodes, input features of 4-dimensional MFCC, and undergoes 8-bit integer quantization. It processes approximately 100 frames of audio data per second, with a total power consumption of less than 500 microwatts.
[0028] Furthermore, the high-performance, low-power embedded digital signal processor of the adaptive noise-robust wake-word recognition unit employs a Cortex-M7 core running at a clock speed of 150MHz and integrates a dedicated neural network accelerator (NPU) core, achieving a computing power of several TOPs / W. The advanced noise suppression (ANR) module uses a lightweight speech enhancement model based on a temporal convolutional network (TCN). This model has four convolutional layers and two gated activation units, with approximately 100KB of parameters, and can process audio streams with a 16kHz sampling rate in real time. The robust feature extraction module extracts 26th-order MFCC features and combines them with the fundamental frequency (pitch) and zero-crossing rate as supplementary features, resulting in a total of 30-dimensional feature vectors. The deep learning wake word recognition model uses a deep separable convolutional neural network based on the MobileNetV2 architecture. This network has 5 convolutional blocks, input features are 96x30 MFCC feature maps, output layer is a Sigmoid activation function, the number of parameters is about 300KB and is quantized by 8 bits, the inference latency is less than 50ms, the peak power consumption of a single wake-up processing is 20mW, and the average power consumption is less than 1mW.
[0029] Furthermore, the noise sensor of the environmental context awareness unit performs energy calculation and spectrum analysis on the digital audio stream of the voice acquisition unit, and runs a real-time noise level estimation algorithm on the main control microcontroller, updating the environmental noise level data once per second. The light sensor is a digital ambient light sensor with an I2C interface, and its measurement range is from 0.01 Lux to 65535 Lux. The user activity status sensor includes a millimeter-wave radar sensor that detects micro-movements within a 3-meter range in front of the device in a non-contact manner, and outputs a user activity index of 0-100, where 0 represents no activity and 100 represents intense activity.
[0030] Furthermore, the adaptive threshold adjustment module of the system control and power management unit employs a fuzzy logic inference system. This system takes ambient noise level, light intensity, and user activity index as inputs and outputs a dynamic threshold adjustment factor. The fuzzy logic system includes three fuzzy variables, each with three fuzzy sets (e.g., "low," "medium," and "high" noise levels), and 27 fuzzy rules, which are defuzzified using the centroid method. The adjustment factor multiplicatively adjusts the baseline recognition threshold (e.g., 0.85) within the range of [0.9, 1.1]. For example, under the fuzzy rule of "high noise, medium light intensity, low activity," the adjustment factor might be set to 1.08, thereby raising the recognition threshold to 0.918 to reduce the false alarm rate.
[0031] The beneficial effects of this invention are as follows: This invention discloses a voice-controlled smart wake-up system and method for home appliances. Through the coordinated operation and refined design of the aforementioned components, it achieves an effective balance between accuracy, robustness, and low power consumption in wake-up word recognition. In extremely noisy environments, the system effectively filters out interference and reduces false wake-up rates by employing beamforming with a multi-channel microphone array, advanced noise suppression algorithms, and a dynamic high-threshold strategy based on contextual information. In quiet environments, the success rate of user wake-up is improved by lowering the recognition threshold and optimizing model sensitivity. Furthermore, the phased wake-up mechanism ensures that only ultra-low-power units operate during most of the standby time, keeping average power consumption at an extremely low level, thereby extending the device's battery life or reducing normal energy consumption. By introducing user feedback and model optimization mechanisms, the system possesses the ability to continuously learn and self-evolve, continuously improving wake-up performance and user satisfaction over time and with the accumulation of user habits. Attached Figure Description
[0032] The invention will now be further described with reference to the accompanying drawings.
[0033] Figure 1 This is a system structure framework diagram of the present invention; Figure 2 This is a schematic diagram of the method flow of the present invention.
[0034] In the diagram: 1. Voice acquisition unit; 2. Ultra-low power acoustic event detection unit; 3. Adaptive noise-robust wake-up word recognition unit; 4. Environmental context awareness unit; 5. System control and power management unit. Detailed Implementation
[0035] This invention provides a voice-controlled smart wake-up system and method for home appliances, aiming to solve the technical challenge of balancing low power consumption and high wake-up accuracy in existing technologies. Through refined component design and collaborative control, it effectively improves the accuracy and robustness of wake-up word recognition while significantly reducing false alarm and missed alarm rates. This embodiment will elaborate on the specific hardware deployment of the system, the technical implementation details of each functional unit, and the workflow and key parameters of the method employed.
[0036] Reference Figure 1 As shown, the present invention provides a voice-controlled smart wake-up system for home appliances, the core structure of which includes a voice acquisition unit 1, an ultra-low power acoustic event detection unit 2, an adaptive noise-robust wake-up word recognition unit 3, an environmental context awareness unit 4, and a system control and power management unit 5. These units are interconnected through a sophisticated electronic interface to form a highly efficient, intelligent, and low-power integrated system.
[0037] First, at the hardware deployment level, the core component of the voice acquisition unit 1 is at least two high-performance microelectromechanical systems (MEMS) microphones arranged in an array. In a specific embodiment, these MEMS microphones can be configured as a linear array or a circular array, with their spacing optimized according to the wavelength characteristics of the target human voice frequency band. For example, for the human voice frequency band of 2kHz to 8kHz, the center-to-center spacing of the microphones is set within the range of 5cm to 10cm to effectively achieve sound source localization and beamforming effects. The microphones are rigorously selected, with a sensitivity calibrated to -26dBFS±1dB, a signal-to-noise ratio (SNR) of 64dB, and an acoustic overload point (AOP) as high as 120dBSPL, ensuring high-fidelity capture of wide dynamic range voice signals even in high-noise and high-sound-pressure environments. These microphones are directly connected to the ultra-low-power digital signal processing core of the ultra-low-power acoustic event detection unit 2 via a digital pulse density modulation (PDM) or inter-integrated circuit audio bus (I2S) interface. Specifically, the digital audio stream is transmitted in real time at a sampling frequency of 16 kHz and a quantization depth of 16 bits, ensuring the integrity of the voice information. The voice acquisition unit 1 further includes a low-latency digital signal processor (DSP) core equipped with a preprocessing module responsible for preliminary array signal processing of the digital audio stream. This preprocessing module integrates a delay-and-sum-based beamforming algorithm, which enhances the voice signal from a specific direction and suppresses ambient noise from other directions by adjusting and summing the delays of the signals from each microphone. In a more advanced implementation, the preprocessing module can also implement an adaptive beamforming algorithm, such as a minimum variance distortionless response (MVDR) beamformer, which adaptively adjusts the beamforming weights by estimating the noise covariance matrix in real time to achieve better noise suppression and signal separation, thereby providing higher quality signal input for subsequent processing units. The beamforming algorithm in the preprocessing module is specifically implemented using an 8th-order FIR filter, which effectively filters out interference from non-target directions and further optimizes signal quality.
[0038] Next, the ultra-low power acoustic event detection unit 2 is electrically connected to the voice acquisition unit 1 and communicates with the system control and power management unit 5. Its core component is an ultra-low power digital signal processing core, which features microwatt-level power consumption and is designed for continuous operation around the clock. In a specific embodiment, this core can be implemented based on a customized RISC-V processor or ARM Cortex-M0+ core, and integrates a low-power analog front-end (AFE) and an analog-to-digital converter (ADC) to directly process digital audio signals from the microphone array. The core operates at a clock frequency of 10MHz and is equipped with 64KB of SRAM and 256KB of flash memory to support its lightweight algorithm. The ultra-low power acoustic event detection unit 2 includes a voice activity detection (VAD) module, configured within the aforementioned ultra-low power digital signal processing core. This VAD module analyzes the digital audio data in real time based on features such as short-time energy, zero-crossing rate, spectral entropy, or fundamental frequency changes, combined with a Gaussian mixture model (GMM) or a shallow neural network model, to accurately distinguish between speech segments and non-speech segments. In one specific embodiment, the VAD module employs an acoustic classifier based on a Gaussian Mixture Model (GMM). This GMM model has four mixture components, each containing 12-dimensional MFCC features, with a parameter count of only a few kilobytes, ensuring efficient operation in ultra-low power environments. When the VAD module recognizes a speech segment, a pre-wake word feature extraction module is activated. This module extracts a small number of acoustic features from the speech segment, such as low-dimensional Mel-frequency cepstral coefficients (MFCCs, typically 4th-6th order) or energy envelopes. These features form the basis for subsequent lightweight matching. Subsequently, a lightweight matching module, also configured within the ultra-low power digital signal processing core, performs pattern matching between the pre-wake word features and a pre-trained, highly simplified wake word template. This template can be implemented based on a Dynamic Time Warping (DTW) algorithm or a Shallow Neural Network to quickly determine the presence of potential wake word segments. In one specific embodiment, the lightweight matching module is implemented using a single-layer Long Short-Term Memory (LSTM) network with 16 hidden layer nodes and 4-dimensional MFCC input features. This LSTM network undergoes 8-bit integer quantization, processing approximately 100 frames of audio data per second with a total power consumption of less than 500 microwatts, thus ensuring preliminary voice matching functionality at sub-milliwatt power levels. The recognition threshold of the lightweight matching module is set to a relatively lenient value to minimize the false alarm rate and ensure that all potential wake-word fragments are captured.When the matching score of the lightweight matching module exceeds a preset initial threshold, the ultra-low power acoustic event detection unit 2 generates a hardware interrupt signal or event trigger signal and sends it to the system control and power management unit 5. Simultaneously, it caches the digital audio data of the potential wake-up word fragment to prepare for subsequent accurate recognition. The ultra-low power acoustic event detection unit 2 is connected to the main microcontroller of the system control and power management unit 5 via a GPIO interrupt line and an SPI or I2C bus.
[0039] Next, the adaptive noise-robust wake-word recognition unit 3 is electrically connected to the ultra-low-power acoustic event detection unit 2 and the system control and power management unit 5. The core of this unit is a high-performance, low-power embedded digital signal processor or a dedicated artificial intelligence (AI) acceleration core. This core has milliwatt-level operating power consumption and remains in deep sleep mode most of the time, only rapidly starting up when it receives a wake-up command from the system control and power management unit 5. In a specific embodiment, the processor uses an ARM Cortex-M7 core running at a clock speed of 150MHz and integrates a dedicated neural network accelerator (NPU) core, with a computing power of several TOPs / W, capable of efficiently executing complex deep learning inference tasks. The unit includes an advanced noise suppression (ANR) module configured within the embedded digital signal processor or AI acceleration core. This module is used to perform refined noise removal on wake-word segments in the digital audio data provided by the voice acquisition unit 1. The advanced noise suppression module employs spectral subtraction, Wiener filtering, or deep learning-based speech enhancement algorithms, such as lightweight temporal convolutional networks (TCNs) or residual networks, to effectively suppress background noise and reverberation, significantly improving the signal-to-noise ratio of the speech signal. In one specific embodiment, the ANR module uses a lightweight speech enhancement model based on a temporal convolutional network (TCN). This model has four convolutional layers and two gated activation units, with approximately 100KB of parameters, capable of processing audio streams with a 16kHz sampling rate in real time, and exhibiting extremely low latency. Subsequently, a robust feature extraction module, configured within the embedded digital signal processor or AI acceleration core, extracts richer and more robust acoustic features from the noise-suppressed speech segments. These features include full-order Mel-frequency cepstral coefficients (MFCCs, e.g., orders 13-26), fundamental frequency (Pitch), formants, linear predictive cepstral coefficients (LPCCs), or deep learning-based acoustic embeddings (such as x-vectors or d-vectors). In one specific embodiment, the robust feature extraction module extracts 26th-order MFCC features and combines them with the fundamental frequency (pitch) and zero-crossing rate as supplementary features to form a 30-dimensional feature vector for a more comprehensive representation of the speech signal. A deep learning wake word recognition model, configured within the embedded digital signal processor or AI acceleration core, takes the robust acoustic features as input and outputs a confidence score representing the probability of the wake word's presence. The model employs a quantized (e.g., 8-bit integer quantization) convolutional neural network (CNN) architecture, a recurrent neural network (RNN, such as a gated recurrent unit GRU or a long short-term memory network LSTM), or a lightweight Transformer architecture based on an attention mechanism. This model, trained on a large-scale wake word dataset and diverse background noise datasets, exhibits high accuracy and strong robustness.In one specific embodiment, the deep learning wake word recognition model employs a deep separable convolutional neural network based on the MobileNetV2 architecture. This network has 5 convolutional blocks, input features are 96x30 MFCC feature maps, and the output layer uses a Sigmoid activation function with approximately 300KB of parameters, processed by 8-bit integer quantization. Its inference latency is less than 50ms, the peak power consumption for a single wake-up process is 20mW, and the average power consumption is less than 1mW, achieving efficient and low-power wake word recognition. Finally, an adaptive decision logic module, configured within the embedded digital signal processor or AI acceleration core, compares the confidence score output by the deep learning wake word recognition model with the recognition threshold dynamically set by the system control and power management unit 5 to make the final wake word recognition decision. The adaptive noise-robust wake word recognition unit 3 receives digital audio data buffered by the ultra-low-power acoustic event detection unit 2 via a high-speed SPI or QSPI bus, and interacts with the main control microcontroller of the system control and power management unit 5 via a UART or I2C bus to exchange control commands and recognition results.
[0040] Furthermore, the environmental context awareness unit 4 is electrically connected to the system control and power management unit 5, and is used to monitor the environmental noise level, user activity status, and other relevant context information in real time. This unit includes an environmental noise sensor, such as the digital microphone integrated into the voice acquisition unit 1. By calculating the short-time root-mean-square energy, signal-to-noise ratio estimate, or energy distribution of a specific frequency band of the audio stream, and running a real-time noise level estimation algorithm on the main control microcontroller of the system control and power management unit 5, the environmental noise level data is updated once per second and output in decibels (dB). This unit also includes an ambient light sensor, such as a silicon photodiode or a digital ambient light sensor (ALS), used to measure the ambient light intensity (Lux value) to help determine the usage status of the environment, such as day or night, and whether the user may be asleep. In a specific embodiment, the light sensor is an I2C interface digital ambient light sensor with a measurement range of 0.01 Lux to 65535 Lux, capable of covering various environments from extremely dark to extremely bright. In addition, the environmental context awareness unit 4 includes one or more user activity state sensors, such as passive infrared (PIR) sensors, millimeter-wave radar sensors, or accelerometers, to detect the physical movement or presence of the user near the device, thereby generating a user activity index. In one specific embodiment, a millimeter-wave radar sensor is used to detect micro-movements within a 3-meter radius in front of the device in a non-contact manner and outputs a user activity index of 0-100, where 0 represents no activity and 100 represents intense activity. Finally, an internal status recording module, stored in the non-volatile memory of the system control and power management unit 5, records historical wake-up success rates, false wake-up counts, user manual interaction frequency, and the current operating status of home appliances (e.g., whether music or television programs are playing). The environmental context awareness unit 4 integrates the above-mentioned environmental noise level, light intensity, user activity index, and internal status data into a multi-dimensional context vector and transmits it to the system control and power management unit 5 in real time via analog input, I2C, or SPI interface.
[0041] The core system control and power management unit 5 is electrically connected to the ultra-low power acoustic event detection unit 2, the adaptive noise robust wake-word recognition unit 3, and the environmental context awareness unit 4. This unit includes a master microcontroller, such as one based on an ARM Cortex-M4 or Cortex-M7 architecture, used to coordinate the work of each unit and execute complex control logic. The power management module integrated into the master microcontroller or a separate power management integrated circuit (PMIC) is responsible for managing the power distribution and mode switching of the entire system, including enabling, disabling, and setting low-power modes for the ultra-low power acoustic event detection unit 2, the adaptive noise robust wake-word recognition unit 3, and other peripherals. The PMIC is configured with multiple programmable voltage output channels, capable of independently controlling the power rails of each unit according to the instructions of the master microcontroller to achieve fine-grained power management. This unit also includes an adaptive threshold adjustment module, configured within the master microcontroller, used to receive multi-dimensional context vectors provided by the environmental context awareness unit 4 and dynamically calculate and update the recognition threshold of the adaptive noise robust wake-word recognition unit 3 based on a preset adaptive strategy or machine learning model. The adaptive strategy includes: increasing the recognition threshold to reduce the false alarm rate when the ambient noise level exceeds a preset threshold; appropriately decreasing the recognition threshold to improve response sensitivity when the user activity index is high or there is frequent manual interaction; and further adjusting the threshold based on the voice characteristics of the playback content when the home appliance is playing media. The adaptive threshold adjustment module can use fuzzy logic reasoning, decision tree models, or reinforcement learning-based strategies for threshold optimization. In a specific embodiment, the adaptive threshold adjustment module employs a fuzzy logic reasoning system. This system takes ambient noise level (e.g., 0-90dB), light intensity (e.g., 0-1000Lux), and user activity index (e.g., 0-100) as inputs and outputs a dynamic threshold adjustment factor. The fuzzy logic system contains three fuzzy variables, each with three fuzzy sets (e.g., "low," "medium," and "high" for noise level; "dark," "medium," and "bright" for light intensity; "still," "medium," and "active" for user activity), and 27 fuzzy rules, which are defuzzified using the centroid method. The adjustment factor multiplicatively adjusts the baseline recognition threshold (e.g., 0.85) within the range of [0.9, 1.1]. For example, under a fuzzy rule of "high noise, medium light, low activity," the adjustment factor might be set to 1.08, thereby raising the recognition threshold to 0.918 to reduce the false alarm rate. Conversely, under a fuzzy rule of "low noise, darkness, stillness," the adjustment factor might be set to 0.95, lowering the recognition threshold to 0.8075 to improve response sensitivity in quiet nighttime environments.Finally, a system state machine management module is configured in the main control microcontroller. It is used to drive the system to switch between different power states and processing modes based on the activation signal of the ultra-low power acoustic event detection unit 2 and the recognition result of the adaptive noise robust wake word recognition unit 3. After the wake word is successfully recognized, it sends a wake-up command to the main processor of the home appliance to switch it from the low power state to the active state.
[0042] The present invention provides a method for intelligent wake-up of home appliances based on voice control, referring to... Figure 2 As shown, its execution process includes the following key steps: S1. Ultra-low power standby monitoring: In this initial stage, the voice acquisition unit 1 continuously acquires ambient acoustic signals at a preset low sampling rate and low bit depth. In a specific embodiment, the audio acquisition uses an 8kHz sampling rate and a 12-bit quantization bit depth to minimize the computational load and power consumption of the ultra-low power acoustic event detection unit 2. The acquired digital audio stream is transmitted in real time to the ultra-low power acoustic event detection unit 2 for processing; S2. Preliminary Acoustic Event Detection: In this step, the ultra-low power acoustic event detection unit 2 continuously runs the speech activity detection (VAD) module and the lightweight matching module. The VAD module determines the presence of speech based on features such as short-time energy and zero-crossing rate. Once a speech segment is detected, the pre-wake word feature extraction module extracts a small number of acoustic features from the speech segment, such as 4th order MFCC. Subsequently, the lightweight matching module performs pattern matching with a pre-trained, simplified wake word template to determine whether there are potential wake word segments. If the matching score exceeds a preset preliminary threshold (e.g., 0.7), it indicates the presence of a potential wake word. The ultra-low power acoustic event detection unit 2 immediately sends an activation signal (e.g., a hardware interrupt) to the system control and power management unit 5 and simultaneously caches the digital audio data of the potential wake word segment in its internal SRAM for subsequent high-precision recognition. S3. Context Information Awareness: During this process, the environmental context awareness unit 4 collects and integrates real-time data on ambient noise levels, light intensity, user activity index, and historical interaction data from the internal state recording module to generate a multi-dimensional context vector, which is then transmitted to the system control and power management unit 5. For example, the system may perceive that the current ambient noise is 55dB, the light intensity is 300Lux, and the user activity index is 20 (indicating slight activity). S4. Adaptive Threshold Calculation and Wake-up Unit Activation: After receiving the activation signal from the ultra-low power acoustic event detection unit 2, the system control and power management unit 5 first quickly wakes up the power supply of the adaptive noise robust wake-up word recognition unit 3 from deep sleep state, putting it into working state. Simultaneously, it obtains the latest multi-dimensional context vector from the environmental context awareness unit 4. Based on the context vector and a preset adaptive strategy (e.g., a fuzzy logic inference system), the system control and power management unit 5 dynamically calculates and sets the precise recognition threshold required for the current wake-up cycle. For example, if the context vector indicates high environmental noise and low user activity, the system may adjust the baseline threshold of 0.85 to 0.918 using an adjustment factor (e.g., 1.08) calculated by the fuzzy logic system. The calculated precise recognition threshold is then transmitted to the adaptive noise robust wake-up word recognition unit 3. In this step, to improve recognition accuracy, the system control and power management unit 5 may instruct the voice acquisition unit 1 to resample or recapture the cached wake-up word fragments at a higher sampling rate (e.g., 16kHz) and a higher bit depth (e.g., 16-bit) to provide higher quality audio data for accurate recognition. S5. Precise Wake-up Word Recognition: The adaptive noise-robust wake-up word recognition unit 3 receives digital audio data of potential wake-up word segments cached (or recaptured) by the ultra-low-power acoustic event detection unit 2. First, the advanced noise suppression module performs refined noise removal on the segment, significantly improving the signal-to-noise ratio. Subsequently, the robust feature extraction module extracts multi-dimensional robust acoustic features from the denoised speech segment, such as 30-dimensional MFCC, fundamental frequency, and zero-crossing rate combination features. The deep learning wake-up word recognition model (e.g., a CNN based on MobileNetV2) takes the robust acoustic features as input and performs high-speed inference with the assistance of the NPU accelerator to calculate a confidence score representing the probability of the wake-up word's presence. S6. Wake-up Decision and Appliance Control: The adaptive decision logic module of the adaptive noise robust wake-up word recognition unit 3 compares the confidence score output by the deep learning wake-up word recognition model with the precise recognition threshold set by the system control and power management unit 5. If the confidence score is higher than the threshold, for example, a confidence score of 0.93 and a threshold of 0.918, the wake-up word recognition is confirmed to be successful. At this time, the system control and power management unit 5 sends a wake-up command to the main processor of the appliance (e.g., via UART or SPI) to switch it from a low-power state to an active state and start the subsequent voice command recognition module. If the confidence score is not higher than the threshold, it is determined to be a non-wake-up word or a false trigger, and the adaptive noise robust wake-up word recognition unit 3 quickly returns to a deep sleep state to save energy. At the same time, the ultra-low power acoustic event detection unit 2 continues to execute step S1 and returns to the ultra-low power standby listening state.
[0043] The system provided by this invention further includes a user feedback and model optimization module, configured in the system control and power management unit 5. This module has the following functions: First, it records each successful and failed wake-up event, including the validity of the wake-up word command (whether it's a genuine wake-up or a false wake-up), the recognition result (success / failure), the current context information (such as noise level, illumination), and the confidence score of the recognition model. Second, when the user performs manual operations or explicitly expresses satisfaction with the wake-up result through other means, this module collects user feedback data. For example, when a user turns off a falsely woken device using a physical button, the system records a negative feedback event. Third, based on the accumulated recorded data and user feedback, the adaptive strategy of the adaptive threshold adjustment module is periodically iteratively optimized offline or online. This may include fine-tuning the rule weights of the fuzzy logic system, adjusting the splitting conditions of the decision tree, or updating the parameters of the support vector machine / small neural network model to improve the accuracy and adaptability of the threshold adjustment. Finally, when the device connects to the network, the module strictly adheres to the user privacy protection protocol, uploading anonymized portions of the recorded data (containing only non-sensitive acoustic features, recognition results, and contextual metadata, excluding raw speech data or any information that could identify the user) to the cloud server for large-scale model retraining in the cloud. After the model is updated in the cloud, the updated deep learning wake word recognition model parameters or quantized weights are downloaded to the adaptive noise-robust wake word recognition unit 3 via a secure channel, enabling online model updates and continuous performance improvement.
[0044] Example This embodiment aims to verify the actual performance of the voice-controlled smart wake-up system and method for home appliances provided by the present invention. We constructed a smart speaker prototype device that conforms to the above technical solution, with its wake-up phrase set to "Hello, Xiao Zhi".
[0045] System Configuration: Voice acquisition unit 1: Employs a linear array of two InvenSense ICS-43432 digital MEMS microphones, spaced 8cm apart. Outputs a 16kHz / 16bit digital audio stream via an I2S interface.
[0046] Ultra-low power acoustic event detection unit 2: Based on the NXP Pinetis KL27 microcontroller (Cortex-M0+ core, 10MHz), the VAD uses a 4-Gaussian mixture model, and the lightweight matching module uses a single-layer LSTM (16 hidden layer nodes, 8-bit quantization). The initial threshold is set to 0.7. The average power consumption is 500μW.
[0047] Adaptive noise robust wake word recognition unit 3: Based on STM32H7A3 microcontroller (Cortex-M7 core, 150MHz, integrated neural network accelerator), ANR adopts a 4-layer TCN model, and the wake word model adopts MobileNetV2 deep separable convolutional network (8-bit quantization), with inference latency <50ms.
[0048] Environmental Context Awareness Unit 4: integrates a digital microphone for noise assessment, a BH1750FVI digital light sensor, and an Infineon BGT60TR13C millimeter-wave radar sensor.
[0049] System Control and Power Management Unit 5: Based on the STM32L4P5 microcontroller (Cortex-M4 core), it uses a fuzzy logic system for threshold adjustment, with a base threshold of 0.85.
[0050] Experimental Scenario Design: The experiment was conducted in a 20-square-meter room, simulating a home environment. We set up three typical scenarios to evaluate system performance: Quiet environment: The background noise level is kept at 30-40 dBSPL, with no other voice or music interference.
[0051] Moderate noise environment: Playing TV programs or soft music, with a background noise level of 50-60 dBSPL.
[0052] High-noise environment: Simulates a family gathering or vacuum cleaner operation, with a background noise level of 70-80 dBSPL.
[0053] In each scenario, we conducted 1000 wake-word ("Hello, Xiaozhi") call tests and 1000 non-wake-word (such as "Okay," "Time to eat," "How's the weather?") or background noise interference tests. The test personnel maintained a distance of 1.5 meters from the device.
[0054] Data Collection and Analysis: We recorded the wake-up success rate (Rec.Rate), false wake-up rate (FalseAlarmRate, FAR), and average system power consumption under different scenarios. Among these: Wake-up success rate = (Number of correctly recognized wake-up words / Total number of wake-up word calls) * 100% False wake-up rate = (Non-wake words incorrectly identified as wake words / Number of background noise occurrences / Total number of non-wake words / Number of background noise tests) * 100% Average system power consumption: Calculated by combining the current measurement function of the integrated power management module with the operating time ratio of each unit.
[0055] Experimental Results (Example): Comparative Example To further highlight the superiority of the technical solution of this invention, we have set up a comparative system. This comparative system represents a traditional intelligent wake-up system that does not employ a hierarchical wake-up architecture and context-adaptive threshold adjustment.
[0056] Comparative system configuration: Single-level wake-up architecture: The ultra-low power acoustic event detection unit 2 is eliminated, and instead a high-performance wake word recognition module (similar in configuration to the adaptive noise robust wake word recognition unit 3 in this invention, but continuously running) directly receives data from the voice acquisition unit 1.
[0057] Fixed recognition threshold: The recognition threshold is fixed at 0.85 and is not dynamically adjusted.
[0058] Noise suppression: Noise suppression is performed using only basic spectral subtraction, without employing deep learning speech enhancement.
[0059] Power Management: The high-performance wake word recognition module remains continuously active and cannot enter deep sleep mode.
[0060] We used the same wake word "Hello, Xiaozhi" to evaluate the performance of the comparative system under three experimental scenarios and testing methods that were exactly the same as those in the example.
[0061] Experimental results (comparative example): Comparative analysis and performance verification: By comparing the data from the above embodiments with those from comparative examples, the performance advantages of the voice-controlled smart wake-up system and method for home appliances provided by this invention are clearly demonstrated.
[0062] Firstly, regarding power consumption, the system of this invention consumes an average of only 0.6mW in a quiet environment, 1.8mW in a moderately noisy environment, and 3.2mW in a high-noise environment. This contrasts sharply with the comparative system's continuous power consumption of up to 150mW. This is mainly due to the invention's unique hierarchical wake-up architecture: the ultra-low-power acoustic event detection unit 2 performs initial listening and screening with microwatt-level power consumption around the clock, only waking up the higher-power adaptive noise-robust wake-up word recognition unit 3 when a potential wake-up word fragment is detected, thereby significantly reducing the overall average power consumption of the system. This refined power management strategy enables the device to remain in standby mode for extended periods, greatly extending battery life or reducing routine power consumption.
[0063] Secondly, regarding the wake-up success rate, the system of this invention achieved a wake-up success rate of 98.7% in a quiet environment, slightly higher than the 97.8% of the comparative example. However, in moderate and high noise environments, the advantages of this invention are more pronounced, reaching 96.2% and 92.5% respectively, far exceeding the 88.5% and 72.3% of the comparative example. This is mainly attributed to the advanced noise suppression (ANR) module and robust feature extraction used in the adaptive noise robust wake-up word recognition unit 3, as well as the powerful noise resistance of the deep learning wake-up word recognition model. In noisy environments, the system of this invention can more effectively separate the target speech from complex acoustic backgrounds and perform accurate recognition.
[0064] Furthermore, the system of this invention exhibits significant superiority in terms of false wake-up rate. In a quiet environment, its false wake-up rate is only 0.1%, in a moderately noisy environment it is 0.3%, and in a high-noise environment it is 0.8%. In contrast, the false wake-up rates of the comparative system are 0.6%, 2.5%, and 6.8%, respectively. This significant reduction in false wake-up rate is directly attributed to the adaptive threshold adjustment module of the system control and power management unit 5. This module dynamically adjusts the recognition threshold based on environmental context information (such as noise level), increasing the threshold when noise is high. This effectively avoids background noise or non-wake words being incorrectly identified as wake words, thereby greatly improving the stability and reliability of the user experience. The comparative system, due to its fixed threshold, is more prone to misidentifying noise as wake words in high-noise environments, leading to frequent false activations.
[0065] In summary, the voice-controlled smart wake-up system and method for home appliances provided by this invention achieves a balance between power consumption and performance in the wake-up function by constructing a two-level wake-up architecture consisting of an ultra-low-power acoustic event detection unit and an adaptive noise-robust wake-up word recognition unit. The ultra-low-power acoustic event detection unit performs all-weather, low-cost monitoring and initial screening, significantly reducing overall average power consumption. The adaptive noise-robust wake-up word recognition unit, after being accurately activated, utilizes more powerful computing resources for high-precision, interference-resistant wake-up word recognition. The linkage between the environmental context awareness unit and the system control and power management unit allows the wake-up word recognition threshold to be dynamically adjusted based on real-time environmental conditions and user behavior, thereby effectively reducing the false wake-up rate (false alarm rate) and missed wake-up rate (missed alarm rate). Specifically, in noisy environments, increasing the recognition threshold and utilizing advanced noise suppression technology can reduce false activations caused by background noise; in quiet or active user environments, appropriately lowering the recognition threshold can improve wake-up sensitivity and response speed, meeting users' personalized needs in different scenarios. The system employs a hierarchical heterogeneous computing architecture at the hardware level, allocating power-sensitive tasks to ultra-low-power processors and computationally intensive tasks to high-performance processors, while optimizing energy consumption through refined power management. At the software level, multimodal context fusion and adaptive decision-making algorithms enhance the system's intelligence and robustness. Through user feedback and model optimization modules, the system possesses the ability to continuously learn and self-evolve, constantly improving wake-up performance and user satisfaction over time and with the accumulation of user habits. Therefore, this invention significantly improves the accuracy of voice wake-up and user experience while maintaining low power consumption and real-time performance, effectively solving the technical challenge of balancing high accuracy and low power consumption in existing technologies.
[0066] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A voice-controlled smart wake-up system for home appliances, characterized in that, include: The voice acquisition unit (1) is used to continuously acquire acoustic signals in the environment and convert them into digital audio data; The ultra-low power acoustic event detection unit (2) is electrically connected to the voice acquisition unit (1) and is used to perform preliminary processing on the digital audio data to detect whether there are potential acoustic events or wake word fragments. When the potential acoustic events or wake word fragments are detected, an activation signal is sent to the system control and power management unit (5), and the digital audio data of the potential wake word fragments is cached. An adaptive noise robust wake word recognition unit (3) is electrically connected to the ultra-low power acoustic event detection unit (2) and the system control and power management unit (5). After receiving the wake-up command from the system control and power management unit (5), it performs in-depth analysis and accurate recognition of the potential wake word fragments cached by the ultra-low power acoustic event detection unit (2) to determine whether it is a preset wake word. The environmental context awareness unit (4) is electrically connected to the system control and power management unit (5). The system control and power management unit (5) is electrically connected to the ultra-low power acoustic event detection unit (2), the adaptive noise robust wake word recognition unit (3), and the environmental context perception unit (4). It is used to receive the activation signal sent by the ultra-low power acoustic event detection unit (2), dynamically adjust the recognition threshold of the adaptive noise robust wake word recognition unit (3) according to the context information provided by the environmental context perception unit (4), and coordinately control the power status and working mode of each unit.
2. The smart home appliance wake-up system based on voice control according to claim 1, characterized in that, The voice acquisition unit (1) includes: At least two high-performance microelectromechanical system microphones are arranged in an array for multi-channel acoustic signal acquisition; The digital signal processor core is equipped with a preprocessing module for performing preliminary array signal processing on the digital audio stream.
3. The smart home appliance wake-up system based on voice control according to claim 2, characterized in that, The voice acquisition unit (1) has at least two MEMS microphones configured in either a linear array or a circular array.
4. The smart home appliance wake-up system based on voice control according to claim 1, characterized in that, The ultra-low power acoustic event detection unit (2) includes: Ultra-low power digital signal processing core; The speech activity detection module is configured within the ultra-low power digital signal processing core. It combines either a Gaussian mixture model or a shallow neural network model to perform real-time analysis on the digital audio data to distinguish between speech segments and non-speech segments. The speech activity detection module uses an acoustic classifier based on a Gaussian mixture model. A pre-wake word feature extraction module is configured within the ultra-low power digital signal processing core and is used to extract low-Viemel frequency cepstral coefficients from the speech segment when the speech activity detection module identifies the speech segment. The lightweight matching module is configured in the ultra-low power digital signal processing core and is used to perform pattern matching between the pre-wake word features and the pre-trained, highly simplified wake word template. When the matching score of the lightweight matching module exceeds the preset initial threshold, the ultra-low power acoustic event detection unit (2) will generate a hardware interrupt signal or event trigger signal and send it to the system control and power management unit (5).
5. A smart home appliance wake-up system based on voice control according to claim 1, characterized in that, The adaptive noise-robust wake word recognition unit (3) includes: A high-performance, low-power embedded digital signal processor or a dedicated artificial intelligence acceleration core, which is rapidly activated only upon receiving a wake-up command from the system control and power management unit (5), and the processor integrates a neural network accelerator core; An advanced noise suppression module is used to perform refined noise removal on wake word segments in the digital audio data provided by the speech acquisition unit (1). The advanced noise suppression module adopts a lightweight speech enhancement model based on a temporal convolutional network. The adaptive decision logic module is used to compare the confidence score output by the deep learning wake word recognition model with the recognition threshold dynamically set by the system control and power management unit (5) to make the final wake word recognition decision.
6. A smart home appliance wake-up system based on voice control according to claim 1, characterized in that, The environment context awareness unit (4) includes: An ambient noise sensor is integrated into the digital microphone of the voice acquisition unit (1). The noise sensor runs a real-time noise level estimation algorithm on the main control microcontroller of the system control and power management unit (5) and updates the ambient noise level data once per second. Ambient light sensor; Multiple user activity status sensors are used to detect the physical movement or presence of a user near the device to generate a user activity index, said user activity status sensors including a millimeter-wave radar sensor; The internal status recording module is stored in the non-volatile memory of the system control and power management unit (5) and is used to record the historical wake-up success rate, the number of false wake-ups, the frequency of user manual interaction and the current operating status of the home appliances. The environmental context awareness unit (4) integrates the above-mentioned environmental noise level, light intensity, user activity index and internal state data into a multi-dimensional context vector and transmits it to the system control and power management unit (5) in real time.
7. A voice-controlled smart wake-up system for home appliances according to claim 1, characterized in that, The system control and power management unit (5) includes: The main microcontroller is used to coordinate the work of each unit and execute complex control logic; The power management module is used to control the power supply voltage and current of each functional unit and switch the power state of each unit in different working modes. An adaptive threshold adjustment module is configured in the main control microcontroller to receive the multi-dimensional context vector provided by the environment context awareness unit (4) and dynamically calculate and update the recognition threshold of the adaptive noise robust wake word recognition unit (3) according to a preset adaptive strategy or machine learning model. The system state machine management module is configured in the main control microcontroller and is used to drive the system to switch between different power states and processing modes according to the activation signal of the ultra-low power acoustic event detection unit (2) and the recognition result of the adaptive noise robust wake word recognition unit (3). After the wake word is successfully recognized, the module sends a wake-up command to the main processor of the home appliance.
8. A smart home appliance wake-up system based on voice control according to claim 1, characterized in that, The system further includes a user feedback and model optimization module, configured in the system control and power management unit (5), which is used for: Record wake-up success and failure events, including: the validity of the wake-up word command, the recognition result, the current context information, and the confidence score of the recognition model; Collect user feedback data when users perform manual operations or express their satisfaction with the wake-up result through other means; Based on accumulated recorded data and user feedback, the adaptive strategy of the adaptive threshold adjustment module is periodically iteratively optimized through offline or online methods.
9. A voice-controlled smart wake-up method for home appliances, applicable to the voice-controlled smart wake-up system for home appliances as described in any one of claims 1-8, characterized in that, Includes the following steps: S1, Ultra-low power standby monitoring: The voice acquisition unit (1) continuously acquires ambient acoustic signals at a preset low sampling rate and low bit depth, and transmits the digital audio stream to the ultra-low power acoustic event detection unit (2). S2, Preliminary acoustic event detection; S3, Context Information Awareness: The environmental context awareness unit (4) collects and integrates environmental noise level, light intensity, user activity index and historical interaction data in real time, generates a multi-dimensional context vector, and transmits it to the system control and power management unit (5). S4. Adaptive threshold calculation and wake-up unit activation: After receiving the activation signal, the system control and power management unit (5) first wakes up the power supply of the adaptive noise robust wake word recognition unit (3) from the deep sleep state, and obtains the latest multi-dimensional context vector from the environmental context perception unit (4). Based on the context vector and the preset adaptive strategy, the system control and power management unit (5) dynamically calculates and sets the accurate recognition threshold required for the current wake-up cycle, and transmits the threshold to the adaptive noise robust wake word recognition unit (3). S5. Precise wake word recognition: The adaptive noise robust wake word recognition unit (3) receives digital audio data of potential wake word segments cached by the ultra-low power acoustic event detection unit (2). The advanced noise suppression module performs refined noise removal on the segment. Subsequently, the robust feature extraction module extracts multi-dimensional robust acoustic features from the denoised speech segment. The deep learning wake word recognition model takes the robust acoustic features as input and calculates a confidence score representing the probability of the wake word's existence. S6. Wake-up Decision and Home Appliance Control: The adaptive decision logic module of the adaptive noise robust wake word recognition unit (3) compares the confidence score with the precise recognition threshold set by the system control and power management unit (5). If the confidence score is higher than the threshold, the wake word recognition is confirmed to be successful. The system control and power management unit (5) sends a wake-up command to the main processor of the home appliance to switch it from a low-power state to a working state and start the subsequent voice command recognition module. If the confidence score is not higher than the threshold, it is judged to be a non-wake word or a false touch. The adaptive noise robust wake word recognition unit (3) returns to a deep sleep state, and the ultra-low power acoustic event detection unit (2) continues to execute step S1.
10. A method for intelligent wake-up of home appliances based on voice control according to claim 9, characterized in that, The voice acquisition unit (1) acquires audio in S1 at a lower sampling rate and lower bit depth to reduce the computational load of the ultra-low power acoustic event detection unit (2).
Citation Information
Patent Citations
Voice wake-up method and apparatus, terminal, and processing method thereof
CN105575395A
Voice wake-up method, system and intelligent terminal
CN107767863A
Wake-up word preset confidence threshold adjustment method and system
CN108847219A
Control method and system capable of automatically adjusting voice recognition sensitivity, air conditioner and readable storage medium
CN110556107A
Control method, intelligent terminal and readable storage medium
CN114093357A