Smart speaker wake-up method and smart speaker

By integrating vibration sensors and deep learning models into smart speakers, vibration signals are used to wake up the smart speakers and adjust linked devices, solving the problem of low wake-up rate in high-noise environments and achieving reliable voice interaction and device control.

CN121922131BActive Publication Date: 2026-06-09HANGZHOU ROBAM APPLIANCES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU ROBAM APPLIANCES CO LTD
Filing Date
2026-03-26
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Smart speakers have difficulty accurately recognizing wake words in high-noise environments, resulting in low wake-up success rates or failure to wake up properly, which affects the user's interactive experience.

Method used

By adding a vibration sensor to the smart speaker, the vibration signal generated by physical tapping is collected as the wake-up trigger condition. Combined with a deep learning model, the vibration characteristics are extracted and compared, and the system actively detects environmental noise and adjusts the working level of the linked devices to create a low-noise interactive environment.

Benefits of technology

In high-noise environments, reliable wake-up and accurate interaction of smart speakers were achieved, improving the wake-up success rate and ensuring the device's response reliability and interaction stability in complex acoustic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121922131B_ABST
    Figure CN121922131B_ABST
Patent Text Reader

Abstract

The application provides a wake-up method of a smart speaker and the smart speaker, the smart speaker is provided with a vibration sensor, and the method comprises the following steps: in response to a vibration signal collected by the vibration sensor, waking up the smart speaker. In the method, the vibration sensor is additionally arranged in the smart speaker to collect the vibration signal generated by physical tapping, so that the interactive intention of a user can be accurately perceived without relying on acoustic speech recognition, thereby effectively avoiding the technical defect of low speech wake-up rate in a high-noise environment such as a kitchen frying scene, and further ensuring the wake-up success rate and reliability of the smart speaker in an extreme acoustic interference scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart home technology, and in particular to a method for waking up a smart speaker and a smart speaker. Background Technology

[0002] Smart speakers, serving as the control hub for smart homes, are increasingly being used in complex acoustic environments such as kitchens to enable convenient voice control of home appliances. For example, in a kitchen setting, when a range hood operates in a high-power mode such as the stir-fry setting, it generates significant ambient noise.

[0003] Currently, smart speakers primarily rely on users uttering specific voice wake-up words to activate them. The microphone array of a smart speaker picks up the user's voice and uses voice recognition technology to identify the wake-up command. However, in noisy environments, the noise generated by home appliances can severely interfere with the microphone's effective reception of the user's voice. This interference makes it difficult for existing voice signal processing technologies to accurately identify the wake-up word, resulting in a low wake-up success rate for smart speakers, or even complete failure to wake them up. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a wake-up method for a smart speaker and a smart speaker. By adding a vibration sensor to the smart speaker to collect vibration signals generated by physical tapping, the user's interaction intention can be accurately perceived without relying on acoustic voice recognition. This effectively avoids the technical defect of low voice wake-up rate in high-noise environments such as kitchen stir-fry stalls, and thus ensures the wake-up success rate and reliability of the smart speaker in extreme acoustic interference scenarios.

[0005] In a first aspect, the present invention provides a method for waking up a smart speaker, wherein the smart speaker is equipped with a vibration sensor, and the method includes:

[0006] The smart speaker is activated in response to vibration signals collected by the vibration sensor.

[0007] In an optional implementation, the smart speaker is further equipped with a microphone and a wireless communication unit. The smart speaker establishes a communication connection with a corresponding linked device through the wireless communication unit. The linked device is a noise source. After the step of waking up the smart speaker in response to the vibration signal collected by the vibration sensor, the method further includes:

[0008] In response to the noise signal picked up by the microphone exceeding the corresponding noise threshold, the control linkage device is reduced to lower its operating level in order to pick up the corresponding voice signal for dialogue.

[0009] In an optional implementation, after the step of controlling the linkage device to reduce its operating level to pick up the corresponding voice signal for dialogue in response to the noise signal picked up by the microphone exceeding a corresponding noise threshold, the method further includes:

[0010] In response to the end of the dialogue, the control linkage equipment resumes its original working position.

[0011] In an optional implementation, the step of waking up the smart speaker in response to a vibration signal collected by a vibration sensor includes:

[0012] The smart speaker is activated when the signal characteristics of the vibration signal meet the corresponding feature matching conditions.

[0013] In an optional implementation, the step of waking up the smart speaker in response to the signal characteristics of the vibration signal meeting the corresponding feature matching conditions includes:

[0014] The vibration signal is input into a pre-trained vibration signal feature processing model to determine whether the signal features of the vibration signal meet the corresponding feature matching conditions.

[0015] In an optional implementation, the vibration signal feature processing model includes a preprocessing module and a feature comparison module. The preprocessing module performs pre-emphasis processing, framing processing, and windowing processing on the vibration signal to obtain FBank features corresponding to the vibration signal. The feature comparison module extracts features from the FBank features to obtain corresponding vibration signal feature values ​​and calculates the similarity between the vibration signal feature values ​​and the corresponding sample signal feature values. The step of inputting the vibration signal into the pre-trained vibration signal feature processing model to determine whether the signal features of the vibration signal meet the corresponding feature matching conditions includes:

[0016] If the similarity exceeds the corresponding similarity threshold, then the signal characteristics of the vibration signal are determined to meet the corresponding feature matching conditions.

[0017] In an optional implementation, the feature comparison module includes an x-vector model, which is used to perform frame-level processing, statistical pooling processing, and segment-level processing on FBank features to output vibration signal feature values.

[0018] In an optional implementation, the vibration signal is generated by at least two consecutive taps on the smart speaker.

[0019] In an alternative implementation, the vibration sensor is located on the back of the top panel of the smart speaker.

[0020] In a second aspect, the present invention provides a smart speaker, comprising:

[0021] The smart speaker itself.

[0022] The controller is located inside the smart speaker itself.

[0023] A vibration sensor, installed inside the smart speaker and connected to the controller, is used to collect vibration signals from the panel.

[0024] The controller is configured to perform a wake-up method for a smart speaker as described in any of the foregoing embodiments.

[0025] This application provides a method for waking up a smart speaker and a smart speaker in its embodiments. By integrating a vibration sensor into the smart speaker and using the physical vibration signal it collects as the wake-up trigger condition, it provides users with a non-voice interaction method independent of the acoustic dimension. It can effectively utilize the characteristic that solid vibration signals are not affected by aerodynamic noise, fundamentally solving the technical pain point of traditional voice wake-up failing due to signal submersion in high background noise environments such as stir-frying in the kitchen. This significantly improves the wake-up success rate of the smart speaker in extreme noise scenarios, thereby ensuring the response reliability and interaction stability of the device in complex acoustic environments.

[0026] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the application. The objectives and other advantages of this application are realized and obtained through the structures particularly pointed out in the description, claims and drawings.

[0027] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0029] Figure 1 A schematic diagram of the module architecture of a smart speaker provided in an embodiment of this application;

[0030] Figure 2 A schematic diagram of the module architecture of another smart speaker provided in an embodiment of this application;

[0031] Figure 3 This is a schematic diagram of the structure of a smart speaker provided in an embodiment of this application;

[0032] Figure 4This is a schematic diagram of the model architecture provided for an embodiment of this application.

[0033] Icons: 1-Smart speaker body; 2-Controller; 3-Vibration sensor; 4-Microphone; 5-Wireless communication unit; 6-Memory; 7-Panel. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0035] To help those skilled in the art better understand this application, a brief introduction to its application scenarios and design concepts is provided.

[0036] While voice recognition is widely used in smart homes, it still has significant limitations in certain scenarios.

[0037] In real-world applications, voice-interactive devices such as smart speakers often face complex acoustic interference. For example, in a kitchen environment, when linked devices (such as range hoods) are operating at high fan speed or on the stir-fry setting, they generate extremely high decibels of aerodynamic noise. This high-intensity background noise can severely overwhelm the user's voice signal, making it difficult for the device's microphone to pick up clear wake-up word features, resulting in a very low wake-up success rate, or even situations where the device cannot be woken up at all or is triggered falsely. Furthermore, even if the device is managed to wake up, the continuous high noise can interfere with the recognition of subsequent voice commands, causing frequent interruptions or recognition errors in the interaction process, severely impacting the user experience in specific high-noise interference scenarios.

[0038] Based on this, this application provides a wake-up method for multiple smart speakers and a smart speaker. This application adds a vibration sensor to the device, using solid vibration signals generated by physical tapping as the wake-up trigger condition. Since vibration signals belong to the non-acoustic dimension, their propagation is not affected by air noise, and even when the linked devices generate significant background noise, the user's wake-up intention can still be accurately perceived. Furthermore, after determining a manual wake-up action, this application actively detects ambient noise and automatically lowers the operating level of the linked devices (such as range hoods, high-power fans, etc.), thereby artificially creating a low-noise voice interaction environment. Simultaneously, this application uses a pre-trained deep learning model to extract and compare vibration features. By comparing the similarity between the real-time extracted vibration feature values ​​and pre-stored samples, interference from non-human tapping (such as environmental vibration or device self-vibration) can be effectively eliminated, ensuring the accuracy of wake-up action recognition. Moreover, this application has a state memory and recovery function, automatically restoring the linked devices to their original operating levels before the dialogue ends. This application ensures a high success rate for interaction while minimizing the impact on the original functions of the linked equipment (such as smoke extraction and ventilation), thus achieving a balance between the convenience of interaction and the reliability of functions.

[0039] To facilitate understanding of this embodiment, the embodiments of this application will be described in detail below.

[0040] This application provides a smart speaker, referring to... Figure 1 and Figure 3 The smart speaker provided in this application embodiment includes:

[0041] Smart speaker body 1.

[0042] Controller 2 is located in the smart speaker body 1.

[0043] Vibration sensor 3 is installed in the smart speaker body 1 and connected to controller 2 to collect vibration signals from panel 7.

[0044] The controller 2 is configured to execute the wake-up method for the smart speaker as described above.

[0045] Here, refer to Figure 2 The smart speaker includes a smart speaker body 1, and a controller 2, a vibration sensor 3, a microphone 4, a wireless communication unit 5, and a memory 6 disposed inside the body.

[0046] The smart speaker body 1 serves as the base for all components, and a panel 7 is located on the top of the smart speaker body 1. A controller 2 is located inside the smart speaker body 1. The controller 2 is the control center of the smart speaker, responsible for processing various sensor signals and performing logical judgments.

[0047] Vibration sensor 3 is electrically connected to controller 2. Vibration sensor 3 is disposed on the back of the top panel 7 of the smart speaker body 1. In one specific embodiment of this application, vibration sensor 3 is disposed on the back of panel 7 and slightly below the microphone array 4. Vibration sensor 3 is a high-sensitivity accelerometer sensor, which can capture vibration signals generated by a finger tapping panel 7 and mechanical vibrations generated by the operation of linked devices.

[0048] Microphone 4 is connected to controller 2 and is used to collect ambient sound. In this embodiment, there can be four microphones 4. The four microphones 4 are distributed on the front panel 7 of the smart speaker body 1 and are used to collect external ambient sound signals and user voice commands.

[0049] The wireless communication unit 5 and the memory 6 are respectively connected to the controller 2. The wireless communication unit 5 is used to establish a communication connection with external linked devices, including home appliances such as range hoods, dishwashers, or steam ovens. In this embodiment, the wireless communication unit 5 can be a WiFi module. The memory 6 is used to store preset tapping action commands, pre-trained vibration signal feature processing models, and pre-extracted vibration feature value samples.

[0050] Controller 2 is configured to execute the wake-up and interaction control methods of the smart speaker, and the specific functional logic is as follows:

[0051] Controller 2 controls vibration sensor 3 to monitor the physical vibration state of panel 7 in real time. When the user taps the smart speaker panel 7 at least twice consecutively, the vibration signal generated by the consecutive taps is transmitted to controller 2 through vibration sensor 3. Controller 2 inputs the vibration signal into the vibration signal feature processing model in memory 6. Controller 2 uses this model to pre-emphasize, frame, and window the vibration signal to obtain FBank (Filter Bank features), and uses an x-vector model to extract the corresponding vibration signal feature values. Controller 2 calculates the similarity between the extracted vibration signal feature values ​​and the sample signal feature values ​​in memory 6. When the similarity exceeds 75%, controller 2 determines that the smart speaker has been successfully woken up.

[0052] After the smart speaker is activated, controller 2 controls microphone 4 to monitor the current ambient sound intensity. When the noise signal picked up by microphone 4 exceeds 65 dB, controller 2 sends a status query request to the linked device via wireless communication unit 5. Controller 2 obtains the current operating level of the linked device through wireless communication unit 5. When the linked device is at a preset high-noise level (such as the high setting or stir-fry setting on a range hood), controller 2 sends a down-level command to the linked device through wireless communication unit 5. Upon receiving the down-level command, the linked device automatically reduces its operating power, thereby reducing background noise interference with voice interaction.

[0053] After the linked device lowers its operating level, controller 2 controls microphone 4 to enter dialogue mode to pick up the user's voice signal. During the dialogue, controller 2 records the original operating level of the linked device before the adjustment. When controller 2 determines that the voice dialogue has ended, controller 2 sends a recovery command to the linked device through wireless communication unit 5. Upon receiving the recovery command, the linked device automatically returns to its original operating level, allowing it to continue performing the functions and tasks performed before the dialogue.

[0054] Through the combination of the aforementioned hardware structure and functional logic, the smart speaker can accurately recognize the user's double-tap to wake it up in high-noise scenarios such as stir-frying in the kitchen, and ensure smooth voice dialogue by automatically adjusting the status of linked devices.

[0055] Based on the above embodiments, this application provides a method for waking up a smart speaker, wherein the smart speaker includes a vibration sensor. The method for waking up a smart speaker provided in this application includes:

[0056] The smart speaker is activated in response to vibration signals collected by the vibration sensor.

[0057] Here, a vibration sensor is installed inside the smart speaker. The smart speaker's controller monitors the output status of the vibration sensor in real time. When a user wants to wake up the smart speaker, they can tap it with their finger, tap it repeatedly, or strike the panel of the smart speaker with another object, thereby generating a physical vibration signal. The vibration sensor captures this physical vibration signal and converts it into an electrical signal, which is then transmitted to the controller.

[0058] In one alternative implementation, the frequency at which the smart speaker collects vibration signals can be adjusted based on the actual hardware performance. The controller performs feature analysis on the vibration signals using a built-in signal processing algorithm. To broaden the protection range, the vibration signal is not limited to a specific number of taps; it can also be continuous vibration generated by pressing and holding the panel, or a specific rhythmic tapping sequence.

[0059] To further improve the accuracy of recognition, the smart speaker uses a pre-trained vibration signal feature processing model to analyze the vibration signal. This model can extract time-domain, frequency-domain, or spatial-domain features from the vibration signal.

[0060] Specifically, the vibration signal feature processing model preprocesses the original vibration signal, including but not limited to pre-emphasis processing, framing processing, and windowing processing. The preprocessed signal is converted into an FBank sequence. Next, the controller uses a deep learning model, such as an x-vector model or TDNN, to process the filter bank feature sequence. This model includes frame-level processing, statistical pooling processing, and segment-level processing, ultimately outputting a fixed-dimensional vibration signal feature value.

[0061] The controller compares the real-time extracted vibration signal feature values ​​with the sample signal feature values ​​pre-stored in the memory. If the similarity between the real-time extracted vibration signal feature values ​​and the sample signal feature values ​​exceeds a preset similarity threshold, the controller determines that the current vibration signal meets the feature matching conditions and classifies it as a valid user wake-up action.

[0062] In the process of waking up the smart speaker in response to vibration signals, the method also incorporates environmental parameters for comprehensive judgment to avoid accidental triggering by the environment.

[0063] The smart speaker uses a built-in microphone to collect ambient sound signals. The controller calculates the sound pressure level of the ambient sound signal. If the noise signal picked up by the microphone exceeds a preset noise threshold, it indicates that the current environment is in a high-noise state.

[0064] In one implementation, a high noise state is one of the prerequisites for triggering the vibration wake-up logic. That is, in a low noise environment, the smart speaker prioritizes voice wake-up, while in a high noise environment, the smart speaker automatically enables or prioritizes the vibration wake-up logic.

[0065] Meanwhile, the smart speaker can establish communication connections with nearby linked devices via wireless communication units (such as Wi-Fi or Bluetooth). Linked devices can be range hoods, dishwashers, air purifiers, or other motor-driven home appliances. The controller acquires the current operating status information of the linked devices. When a linked device is operating at a preset high-power setting (such as the high setting or stir-fry setting of a range hood), the controller confirms that there is indeed severe acoustic interference in the current environment.

[0066] Once the controller determines that the speaker has been successfully woken up via vibration, the smart speaker enters dialogue mode. To ensure the success rate of subsequent voice command recognition, the smart speaker sends a noise reduction command to the linked devices via its wireless communication unit. Upon receiving the noise reduction command, the linked devices automatically reduce their operating power or switch to a low-noise level.

[0067] During voice conversations, the smart speaker picks up the user's voice signal through its microphone and performs local or cloud-based recognition. Because the linked devices have reduced noise output, users can effectively control kitchen appliances or other smart home devices without having to shout.

[0068] When the controller determines that the voice conversation has ended, or the preset interaction time has elapsed, the smart speaker sends a recovery command to the linked device again via the wireless communication unit. Upon receiving the recovery command, the linked device automatically returns to its original operating state before being woken up, thus ensuring that the original working efficiency of the linked device is not continuously affected.

[0069] Through the above steps, the smart speaker achieves reliable wake-up by using vibration signals generated by physical contact without requiring the user to repeat voice commands multiple times, and further optimizes the overall interactive environment through device linkage.

[0070] In an optional implementation, the smart speaker is further equipped with a microphone and a wireless communication unit. The smart speaker establishes a communication connection with a corresponding linked device through the wireless communication unit. The linked device is a noise source. After the step of waking up the smart speaker in response to the vibration signal collected by the vibration sensor, the method further includes:

[0071] In response to the noise signal picked up by the microphone exceeding the corresponding noise threshold, the control linkage device is reduced to lower its operating level in order to pick up the corresponding voice signal for dialogue.

[0072] Here, the smart speaker itself is equipped with a microphone and a wireless communication unit. The smart speaker establishes a communication connection with corresponding linked devices through the wireless communication unit. The linked devices are noise sources, including but not limited to kitchen appliances with power devices such as range hoods, steam ovens, or air purifiers.

[0073] After the smart speaker responds to the vibration signal collected by the vibration sensor and successfully wakes up, it enters the environmental noise monitoring phase. The smart speaker uses four microphones to pick up environmental noise signals in real time. The controller performs real-time sound pressure level analysis on the noise signals picked up by the microphones and compares the intensity of the noise signals with a preset noise threshold. In this embodiment, the preset noise threshold is 65dB.

[0074] When the noise signal picked up by the microphone exceeds 65dB, the controller sends a status query request to the linked equipment through the wireless communication unit to obtain the current working level information of the linked equipment.

[0075] The controller determines whether the linked device is in a preset high-noise operating mode. For example, when the linked device is a range hood, the controller determines whether the range hood is currently on a high, strong, or stir-fry setting. Only when the noise signal exceeds the noise threshold and the linked device is in a preset high-noise setting will the controller determine that the current environmental interference is caused by the operation of the linked device and trigger subsequent noise reduction control logic.

[0076] In response to a noise signal exceeding a corresponding noise threshold, the controller sends a downshift command to the linked equipment via the wireless communication unit. Upon receiving the downshift command, the linked equipment automatically adjusts its operating speed from a high-noise level (such as the stir-fry setting) to a low-noise level (such as the low-noise setting).

[0077] By controlling the linked device to lower its operating level, the background noise level around the smart speaker is significantly reduced, allowing the microphone to clearly pick up the corresponding voice signal input by the user for subsequent voice recognition and dialogue interaction. While lowering the linked device's operating level, the controller records the original operating level of the linked device before the adjustment, so that the state can be restored after the interaction ends.

[0078] During a voice conversation, the smart speaker interacts with the user. Upon the end of the conversation, the smart speaker determines that the user no longer needs to perform voice control operations. At this point, the controller sends a recovery command to the linked device again via the wireless communication unit. Upon receiving the recovery command, the linked device automatically resumes from the low-noise mode to the previously recorded original operating mode, ensuring that the linked device continues to operate according to the performance parameters initially set by the user.

[0079] Through the above steps, this application not only solves the wake-up problem in high-noise environments by using vibration sensors, but also actively creates good voice interaction conditions through deep coupling with linkage devices, thereby improving the overall success rate of voice control of kitchen appliances.

[0080] In an optional implementation, after the step of controlling the linkage device to reduce its operating level to pick up the corresponding voice signal for dialogue in response to the noise signal picked up by the microphone exceeding a corresponding noise threshold, the method further includes:

[0081] In response to the end of the dialogue, the control linkage equipment resumes its original working position.

[0082] Here, as the controller lowers the operating level of the linked equipment and picks up the corresponding voice signals for dialogue, the controller monitors the execution status of the voice interaction in real time. The controller monitors whether the user continues to input voice commands through the microphone, and combines the semantic recognition results returned by the cloud server or the local natural language processing module to determine whether the current interaction logic has been completed.

[0083] In response to the end of the dialogue, the controller confirms that the user has completed the voice control request. The conditions for determining the end of the dialogue include, but are not limited to: the controller detecting that the user has uttered a preset closing phrase; the controller not detecting any new voice input within a preset time interval; or the controller has completed the control feedback corresponding to the user's command.

[0084] Before controlling the linked device to lower its operating level, the controller has already acquired and recorded the real-time operating status of the linked device via the wireless communication unit. The controller stores this real-time operating status as the original operating level in the smart speaker's memory.

[0085] When the controller determines that the dialogue has ended, it retrieves the original operating level value from the memory. This original operating level represents the user's performance requirements for the linked devices (such as range hoods, high-power fans, etc.) before initiating the interaction, for example, the high-power or stir-fry level that the range hood was originally in during a stir-fry scenario.

[0086] In response to the end of the dialogue, the controller sends a recovery command to the linked device via the wireless communication unit. This recovery command contains the parameter information of the original operating mode. Upon receiving the recovery command, the linked device automatically switches its operating state back from the low-noise, low-power mode to the original operating mode.

[0087] By controlling the linked devices to restore their original operating levels, the smart speaker can immediately restore the linked devices to their expected working efficiency (such as restoring the high-volume smoke extraction capacity of the range hood) after completing the voice interaction task, thereby ensuring the air quality of the kitchen environment or the continuity of the original functions of the linked devices.

[0088] In an optional implementation, the step of waking up the smart speaker in response to a vibration signal collected by a vibration sensor includes:

[0089] The smart speaker is activated when the signal characteristics of the vibration signal meet the corresponding feature matching conditions.

[0090] Here, the smart speaker's controller monitors the mechanical vibration state of the smart speaker body in real time using a vibration sensor. Once the vibration sensor detects a vibration signal, the controller performs a digital conversion. In one specific implementation, the smart speaker's central processing unit converts the voltage signal output from the vibration sensor from analog to digital at a sampling frequency of 8 kHz (kilohertz), thereby obtaining a continuous sequence of original vibration signals.

[0091] In response to the vibration signal collected by the vibration sensor, the controller further determines whether the signal characteristics of the vibration signal meet the corresponding feature matching conditions. Feature matching conditions refer to a series of algorithmic criteria used to distinguish between intentional human wake-up and environmental interference noise.

[0092] Feature matching conditions include signal similarity criteria. The controller inputs the currently acquired real-time vibration signal into a pre-trained feature extraction model to extract vibration signal feature values ​​that reflect the current vibration waveform structure. The controller then compares these vibration signal feature values ​​with pre-stored feature value samples in the memory. The feature value samples are standard feature vectors generated in advance by collecting a large amount of real striking action data and performing model inference.

[0093] When the similarity between the real-time extracted vibration signal feature value and a certain set of feature values ​​in the feature value sample exceeds a preset similarity threshold (e.g., 75%), the controller determines that the current vibration signal meets the feature matching condition.

[0094] The feature matching criteria also incorporate limitations on the tapping pattern. The controller determines whether the vibration signal is a double-tap action resulting from at least two consecutive taps on the smart speaker. During the dataset preparation phase, a large amount of vibration signal data from single taps and double-tap actions was pre-collected as the basis for model training, enabling the feature extraction model to accurately identify wake-up behaviors with specific rhythms.

[0095] If the signal characteristics of the vibration signal fully match the above feature matching conditions, the controller determines that the current vibration signal is not background vibration generated by the linked device, but a wake-up behavior actively triggered by the user. Subsequently, the controller executes the wake-up command, switching the smart speaker from standby mode to dialogue mode, thereby allowing the user to perform subsequent voice interaction operations.

[0096] Through the above steps, this application effectively eliminates the interference of external mechanical vibration on the wake-up logic, ensuring the interaction accuracy of the smart speaker in complex physical environments.

[0097] In an optional implementation, the step of waking up the smart speaker in response to the signal characteristics of the vibration signal meeting the corresponding feature matching conditions includes:

[0098] The vibration signal is input into a pre-trained vibration signal feature processing model to determine whether the signal features of the vibration signal meet the corresponding feature matching conditions.

[0099] Here, the training process of the vibration signal feature processing model is completed in advance on the server. To ensure the generalization ability of the vibration signal feature processing model, the training dataset includes normal samples and abnormal samples. Normal samples are obtained by collecting vibration signals from ten typical models of range hoods running at four speeds: low, medium, high, and stir-fry. A total of forty datasets, each sixty seconds long, are obtained. During the data processing phase, the server slices these range hood vibration data into 400-millisecond time slices, resulting in 6,000 slices. Abnormal samples are obtained by collecting vibration signal data from 200 people, each performing 20 single taps and 20 double taps, resulting in 4,000 sets of tapping vibration data. The final training dataset contains a total of 10,000 data sets.

[0100] The vibration signal feature processing model uses the x-vector deep learning model, whose internal structure is based on the time-delay neural network (TDNN).

[0101] During the manufacturing or initialization phase of the smart speaker, a total of one hundred inference samples, generated by twenty people each performing five habitual double-click actions, are input into a pre-trained vibration signal feature processing model for inference. Through inference, one hundred sets of vibration features for the double-click actions are obtained. These one hundred sets of vibration features are pre-stored as feature value samples in the smart speaker's memory for subsequent real-time comparison.

[0102] After acquiring real-time vibration signals, the smart speaker's controller first performs front-end processing on these signals. This front-end processing includes pre-emphasis processing, frame segmentation, windowing, and FBank feature calculation. Through this front-end processing, the controller obtains a variable-length feature sequence reflecting the vibration frequency distribution.

[0103] The controller inputs this feature sequence into the pre-trained x-vector model. The data propagates forward layer by layer within the model, reaching the statistical pooling layer after frame layer processing. The controller extracts the activation values ​​of the second layer in the segment layer after the statistical pooling layer as the feature values ​​of the input signal. This vibration signal feature value is a 512-dimensional vector, which can highly abstractly represent the biometric features of the user's tapping action.

[0104] The controller compares the 512-dimensional vibration signal feature values ​​extracted in real time with the 100 pre-stored feature value samples in the memory for similarity.

[0105] If the similarity between the real-time extracted vibration signal feature value and a certain set of feature values ​​in the feature value sample is greater than 75%, the controller determines that the vibration signal feature meets the corresponding feature matching condition. In this case, the controller determines that a valid double-tap wake-up action has occurred and triggers the smart speaker to enter the wake-up state. If the similarity does not meet the above requirements, the controller determines that the current vibration is environmental interference or an invalid action and returns to the vibration signal acquisition step.

[0106] Through the feature processing based on the deep learning model described above, the smart speaker can accurately extract the user's tapping signal from the complex vibrations of the kitchen environment, greatly reducing the false trigger rate and improving the wake-up sensitivity in high-noise environments.

[0107] In an optional implementation, the vibration signal feature processing model includes a preprocessing module and a feature comparison module. The preprocessing module performs pre-emphasis processing, framing processing, and windowing processing on the vibration signal to obtain FBank features corresponding to the vibration signal. The feature comparison module extracts features from the FBank features to obtain corresponding vibration signal feature values ​​and calculates the similarity between the vibration signal feature values ​​and the corresponding sample signal feature values. The step of inputting the vibration signal into the pre-trained vibration signal feature processing model to determine whether the signal features of the vibration signal meet the corresponding feature matching conditions includes:

[0108] If the similarity exceeds the corresponding similarity threshold, then the signal characteristics of the vibration signal are determined to meet the corresponding feature matching conditions.

[0109] In an optional implementation, the feature comparison module includes an x-vector model, which is used to perform frame-level processing, statistical pooling processing, and segment-level processing on FBank features to output vibration signal feature values.

[0110] Here, the vibration signal feature processing model includes a preprocessing module and a feature comparison module. The preprocessing module performs preliminary digital analysis and feature transformation on the acquired raw vibration signals, providing standard feature input for subsequent deep learning models. The feature comparison module extracts deep biological features from the preprocessed data and performs consistency evaluation with pre-stored samples.

[0111] Reference Figure 4 The smart speaker's controller invokes the preprocessing module to perform a series of time-domain and frequency-domain transformations on the original vibration signal. The specific process is as follows:

[0112] First, the preprocessing module pre-emphasizes the vibration signal to compensate for the loss of the vibration signal in the high-frequency part.

[0113] Next, the preprocessing module performs frame segmentation on the pre-emphasized signal. In this embodiment, the preprocessing module slices the vibration signal according to a time span of 400ms (milliseconds).

[0114] Subsequently, the preprocessing module performs windowing on each time slice to reduce spectral leakage caused by frame splitting.

[0115] Finally, the preprocessing module calculates the FBank for each frame of the signal, thereby obtaining the FBank feature sequence corresponding to the vibration signal. This FBank feature sequence completely preserves the energy distribution information of the user's tapping action within a specific frequency range.

[0116] The feature comparison module performs deep feature extraction on FBank features using the built-in x-vector model.

[0117] During feature extraction, the x-vector model processes the input FBank feature sequence sequentially as follows:

[0118] First, frame-level processing. The x-vector model extracts local features from the input sequence at the frame level, capturing the dynamic changes of the vibration signal over a short period of time.

[0119] Second, statistical pooling layer processing. The statistical pooling layer summarizes the outputs of all frame layers, calculating the mean and standard deviation of all frames. Through statistical pooling, the variable-length frame layer sequence is converted into a fixed-length statistical feature vector, thereby eliminating the influence of the user's tapping speed on the recognition result.

[0120] Third, segment-level processing. The x-vector model performs further nonlinear transformation on the output of the statistical pooling layer. In this embodiment, the controller extracts the activation values ​​of the second layer in the segment layer, obtaining a 512-dimensional vector. This 512-dimensional vector is the characteristic value of the output vibration signal.

[0121] After obtaining the 512-dimensional vibration signal feature values, the feature comparison module calculates the similarity between these vibration signal feature values ​​and the sample signal feature values ​​pre-stored in the memory. The sample signal feature values ​​are feature benchmarks obtained by pre-collecting habitual double-click actions from multiple testers and inferring them using the same model.

[0122] The controller determines whether the calculated similarity exceeds a corresponding similarity threshold. In this embodiment, the similarity threshold is set to 75%. If the similarity between the real-time extracted vibration signal feature values ​​and any set of feature values ​​in the sample signal feature value set exceeds 75%, the controller determines that the vibration signal signal features meet the corresponding feature matching conditions. At this time, the controller determines that a valid user double-click action has been detected and triggers the smart speaker's wake-up logic.

[0123] In this way, this application can accurately pinpoint the user's tapping intention from the chaotic vibrations of kitchen machinery, ensuring the professionalism and reliability of the wake-up determination process.

[0124] In an optional implementation, the vibration signal is generated by at least two consecutive taps on the smart speaker.

[0125] Here, in the wake-up method of the smart speaker, the vibration signal is generated by a specific physical interaction action performed by the user on the smart speaker. In this embodiment, the vibration signal is generated by at least two consecutive taps on the smart speaker. This consecutive tapping action is usually manifested as a double-tap action in practical applications.

[0126] By setting at least two consecutive taps as the wake-up trigger, the smart speaker's controller can effectively distinguish between a user's conscious wake-up behavior and random vibrations caused by accidental single touches, object collisions, or environmental factors. To ensure the robustness of this specific action recognition, during the system's development and model training phases, engineers collected vibration signal data from 200 people, each recording 20 single taps and 20 double taps. This data, containing taps of varying force and frequency, was used as sample input for model training. This allows the controller to accurately identify the specific vibration characteristics generated by continuous taps, thereby significantly reducing the false trigger rate.

[0127] In an alternative implementation, the vibration sensor is located on the back of the top panel of the smart speaker.

[0128] Here, the physical mounting position of the vibration sensor on the smart speaker body has a decisive impact on the sensitivity and accuracy of signal acquisition. In this embodiment, the vibration sensor is located on the back of the top panel of the smart speaker.

[0129] Specifically, the vibration sensor is mounted in the center of the back of the top panel, slightly below the microphone array. The vibration sensor is a high-sensitivity accelerometer. Because the vibration sensor is mounted close to the back of the panel, the accelerometer can directly sense the minute mechanical waves generated when the panel is subjected to force.

[0130] This layout allows the vibration sensor to effectively capture background vibration signals transmitted to the smart speaker through a solid medium from external devices such as range hoods during operation. It also enables the sensor to sensitively detect pulsed vibration signals generated when a user taps the front of the panel with their finger. This layout minimizes vibration signal attenuation during transmission, providing high-quality raw electrical signals for the controller to perform subsequent feature extraction, model inference, and similarity comparison.

[0131] The wake-up method for smart speakers provided in this application introduces a vibration sensor into the smart speaker and combines it with a feature value extraction model to recognize tapping actions. This can accurately distinguish the user's wake-up intention in high-noise environments, thereby avoiding the inability to correctly recognize voice commands due to environmental noise interference. This improves the wake-up success rate and voice interaction experience of smart speakers in complex scenarios such as kitchens.

[0132] The computer program product provided in this application includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.

[0133] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and apparatus described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0134] Furthermore, in the description of the embodiments of this application, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0135] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0136] In the description of this application, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0137] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of protection of the claims.

Claims

1. A method for waking up a smart speaker, characterized in that, The smart speaker is equipped with a vibration sensor, and the method includes: The smart speaker is woken up in response to the vibration signal collected by the vibration sensor. The smart speaker is also equipped with a microphone and a wireless communication unit. The smart speaker establishes a communication connection with a corresponding linked device through this wireless communication unit. The linked device is a noise source. In response to the vibration signal collected by the vibration sensor, after the step of waking up the smart speaker, the method further includes: In response to the noise signal picked up by the microphone exceeding the corresponding noise threshold, the linkage device is controlled to reduce its working level in order to pick up the corresponding voice signal for dialogue. After the step of controlling the linkage device to reduce its operating level in response to the noise signal picked up by the microphone exceeding a corresponding noise threshold, so as to pick up the corresponding voice signal for dialogue, the method further includes: In response to the end of the dialogue, the control device is restored to its original working position.

2. The method according to claim 1, characterized in that, The step of waking up the smart speaker in response to the vibration signal collected by the vibration sensor includes: The smart speaker is activated when the signal characteristics of the vibration signal meet the corresponding feature matching conditions.

3. The method according to claim 2, characterized in that, The steps for waking up the smart speaker in response to the vibration signal's signal characteristics matching the corresponding feature matching conditions include: The vibration signal is input into a pre-trained vibration signal feature processing model to determine whether the signal features of the vibration signal meet the corresponding feature matching conditions.

4. The method according to claim 3, characterized in that, The vibration signal feature processing model includes a preprocessing module and a feature comparison module. The preprocessing module performs pre-emphasis processing, framing processing, and windowing processing on the vibration signal to obtain FBank features corresponding to the vibration signal. The feature comparison module extracts features from the FBank features to obtain corresponding vibration signal feature values ​​and calculates the similarity between the vibration signal feature values ​​and the corresponding sample signal feature values. The step of inputting the vibration signal into the pre-trained vibration signal feature processing model to determine whether the signal features of the vibration signal meet the corresponding feature matching conditions includes: If the similarity exceeds the corresponding similarity threshold, then the signal characteristics of the vibration signal are determined to meet the corresponding feature matching conditions.

5. The method according to claim 4, characterized in that, The feature comparison module includes an x-vector model, which is used to perform frame-level processing, statistical pooling processing, and segment-level processing on the FBank features to output the vibration signal feature values.

6. The method according to any one of claims 1-5, characterized in that, The vibration signal is generated by at least two consecutive taps on the smart speaker.

7. The method according to claim 6, characterized in that, The vibration sensor is located on the back of the top panel of the smart speaker.

8. A smart speaker, characterized in that, include: The smart speaker itself; The controller is located in the main body of the smart speaker; A vibration sensor is installed in the smart speaker body and connected to the controller to collect vibration signals from the panel. The controller is configured to perform the wake-up method for the smart speaker as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Voice system suitable for portable intelligent device and application method thereof

    CN109346077A

  • Smart home awakening interaction method and device

    CN115314334A