Wake-up method and apparatus, electronic device, and storage medium

By using bone conduction sensors and microphone signal data padding technology, the problem of reduced wake-up performance caused by microphone delay was solved, achieving efficient wake-up word recognition and low power consumption in high-noise environments.

CN122135703APending Publication Date: 2026-06-02GEER TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GEER TECH CO LTD
Filing Date
2026-02-27
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In traditional voice wake-up systems, there is a delay between when the microphone is activated and when it outputs a stable signal, which leads to the loss of information in the beginning part of the voice and affects wake-up performance, especially in high-noise or far-field scenarios.

Method used

The system continuously monitors speech activity using a bone voiceprint sensor. When the pre-wake-up condition is met, the microphone is activated, and the microphone signal is filled with data based on the bone voiceprint signal to generate a dual-channel input that is time-synchronized with the bone voiceprint signal for wake-up model recognition.

Benefits of technology

It improves wake-up performance, especially in high-noise environments, effectively recognizing the start information of the wake word while maintaining low power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135703A_ABST
    Figure CN122135703A_ABST
Patent Text Reader

Abstract

This application discloses a wake-up method, apparatus, electronic device, and storage medium, relating to the field of speech recognition technology. The wake-up method includes: acquiring a bone voiceprint signal collected by a bone voiceprint sensor; activating a microphone and acquiring a first microphone signal collected by the microphone when the bone voiceprint signal meets preset pre-wake-up conditions; filling the first microphone signal with data based on the bone voiceprint signal to obtain a second microphone signal, wherein the second microphone signal is synchronized with the bone voiceprint signal in time; inputting the second microphone signal and the bone voiceprint signal into a preset wake-up model to obtain a wake-up recognition result; and executing a wake-up operation corresponding to the wake-up recognition result. This application effectively improves wake-up performance while maintaining low power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and in particular to wake-up methods, devices, electronic devices, and storage media. Background Technology

[0002] In traditional voice wake-up systems, a tiered wake-up architecture is typically used to reduce standby power consumption: the device shuts down or puts the microphone (MIC) to sleep in standby mode, relying on low-power sensors (e.g., bone conduction sensors) to continuously monitor ambient voice activity. Once a potential voice signal is detected, the microphone is then powered on, and the wake-up word recognition process begins. However, due to the tens of milliseconds delay between microphone activation and the output of a stable signal, the initial part of the voice is difficult to capture effectively during this period. This results in the audio signal used for wake-up recognition lacking crucial initial information, leading to a decline in wake-up performance. Summary of the Invention

[0003] The main purpose of this application is to provide a wake-up method, device, electronic device and storage medium. The embodiments of this application effectively improve wake-up performance while taking into account low power consumption.

[0004] To achieve the above objectives, this application proposes a wake-up method, the method comprising:

[0005] Acquire bone voiceprint signals collected by a bone voiceprint sensor; If the bone voiceprint signal meets the preset pre-wake-up conditions, the microphone is activated and the first microphone signal collected by the microphone is acquired. Based on the bone voiceprint signal, the first microphone signal is filled with data to obtain the second microphone signal, wherein the second microphone signal is synchronized with the bone voiceprint signal in time; The second microphone signal and the bone conduction signal are input into a preset wake-up model to obtain the wake-up recognition result; Execute the wake-up operation corresponding to the wake-up recognition result.

[0006] In one embodiment, the step of filling the first microphone signal with data based on the bone conduction signal to obtain the second microphone signal includes: Identify the start time of the appearance of the speech signal in the bone voiceprint signal; Based on the start time and the acquisition time of the first microphone signal, determine the missing time period that needs to be filled with data; Generate filler data that matches the missing time period; The padding data is concatenated with the first microphone signal to obtain the second microphone signal.

[0007] In one embodiment, the filling data includes at least one of: silent data, ambient noise data, and preset audio template data.

[0008] In one embodiment, prior to the step of enabling the microphone, the method further includes: Detect whether there is a speech signal in the bone voiceprint signal; If it is determined that there is a speech signal in the bone voiceprint signal, the bone voiceprint signal is determined to meet the preset pre-wake-up conditions, and the step of enabling the microphone and subsequent steps are executed.

[0009] In one embodiment, the step of detecting whether a speech signal exists in the bone voiceprint signal includes: Determine whether the signal energy of the bone voiceprint signal within a preset time window is greater than a preset energy threshold; If the signal energy of the bone voiceprint signal within a preset time window is greater than a preset energy threshold, then it is determined that there is a speech signal in the bone voiceprint signal. If the signal energy of the bone voiceprint signal within a preset time window is less than or equal to a preset energy threshold, then it is determined that there is no speech signal in the bone voiceprint signal.

[0010] In one embodiment, prior to the step of inputting the second microphone signal and the bone conduction signal to a preset wake-up model, the method further includes: Acquire bone voiceprint training signals collected by a bone voiceprint sensor and microphone first training signals collected by a microphone, wherein both the bone voiceprint training signals and the microphone first training signals contain a preset wake-up word; The data in the first training signal of the microphone that is located in the preset simulated activation delay period is replaced with preset training padding data to obtain the second training signal of the microphone; Obtain the wake-up tags corresponding to the bone voiceprint training signal and the second microphone training signal; The bone voiceprint training signal, the microphone second training signal, and the wake-up tag are used as a training sample, and a wake-up training sample set is obtained based on the acquired training samples. The wake-up training sample set is used to train the preset wake-up model to obtain the wake-up model.

[0011] In one embodiment, the step of acquiring the wake-up tag corresponding to the bone conduction training signal and the microphone second training signal includes: The wake-up recognition result corresponding to the first training signal of the microphone is determined as the wake-up tag.

[0012] Furthermore, to achieve the above objectives, embodiments of this application also propose a wake-up device, the device comprising: The acquisition module is used to acquire bone voiceprint signals collected by the bone voiceprint sensor; The activation module is used to activate the microphone and acquire the first microphone signal collected by the microphone when it is determined that the bone voiceprint signal meets the preset pre-wake-up conditions. A filling module is used to fill the first microphone signal with data according to the bone voiceprint signal to obtain a second microphone signal, wherein the second microphone signal is synchronized with the bone voiceprint signal in time; The input module is used to input the second microphone signal and the bone conduction signal into a preset wake-up model to obtain a wake-up recognition result; The execution module is used to perform the wake-up operation corresponding to the wake-up recognition result.

[0013] Furthermore, to achieve the above objectives, this application also proposes an electronic device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the wake-up method described above.

[0014] Furthermore, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the wake-up method described above.

[0015] In addition, to achieve the above objectives, this application also proposes a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the wake-up method described above.

[0016] One or more technical solutions proposed in this application have at least the following technical effects: A wake-up method is provided that continuously monitors speech activity using a bone conduction sensor, and activates the microphone only when pre-wake-up conditions are met to maintain low power consumption. However, since there is a certain delay in microphone activation, the first microphone signal collected may lack the initial speech segment. Therefore, this application embodiment fills the first microphone signal with data based on the bone conduction signal to generate a second microphone signal that is time-synchronized with the bone conduction signal. Both are then synchronously input into a wake-up model, enabling the model to effectively identify the start information of the wake-up word by combining the noise immunity of the bone conduction signal with the complete temporal structure of the filled microphone signal, thereby improving wake-up performance while maintaining low power consumption. Attached Figure Description

[0017] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments in line with this application, and are used together with the specification to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 It is a schematic diagram of the scenario provided for the first embodiment of this application; Figure 2 It is a flow schematic provided for the first embodiment of this application Figure 1 ; Figure 3 It is a flow schematic provided for the first embodiment of this application Figure 2 ; Figure 4 It is a schematic diagram of the frame structure of the wake-up device involved in the embodiments of this application; Figure 5 It is a schematic diagram of the structure of the hardware operating environment of the electronic device involved in the embodiments of this application.

[0020] The realization of the purpose of this application, functional features and advantages will be further described in combination with the embodiments and the accompanying drawings. Specific Embodiments

[0021] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.

[0022] To better understand the technical solutions of this application, the following will be described in detail in combination with the drawings of the specification and specific embodiments.

[0023] In traditional voice wake-up systems, to reduce the power consumption of the device in the standby state, a hierarchical wake-up architecture is generally adopted: the microphone is in a sleep or off state during standby, and only a low-power sensor (such as a bone voiceprint sensor) continuously monitors the voice activities in the environment. When the bone voiceprint sensor detects a possible voice signal, it triggers the microphone to power on and start the subsequent wake-up word recognition process. However, the signal energy collected by the bone voiceprint sensor is mainly concentrated in the low-frequency band and is less sensitive to high-frequency components. Mainstream Chinese wake-up words (such as "Xiaoyi Xiaoyi", "Xiaoi Tongxue", etc.) usually start with voiceless consonants (such as the "x" sound in the word "xiao"), and such phonemes have the characteristics of high frequency, weak energy, and strong transients, and are difficult to effectively represent in the bone voiceprint sensor signal, resulting in the loss of information in the initial part of the voice during the pre-wake-up stage, with a time span of about 100 milliseconds (corresponding Figure 1The region from time A to time B). Secondly, from the bone conduction sensor detecting speech activity and issuing a wake-up command to the microphone powering on and outputting a stable audio signal, there is an inherent hardware response delay, typically about 30 milliseconds (corresponding to...). Figure 1 The region from time B to time C. During this period, actual airborne speech cannot be captured by the microphone. That is to say, in traditional voice wake-up systems, the microphone signal used for wake-up word recognition only includes the part after time C. For a typical wake-up word, the key acoustic information at the beginning of its first 130 milliseconds (especially the initial consonant that determines the semantic recognition) is completely missing. This directly leads to a decrease in wake-up performance, which is more obvious in high-noise or far-field scenarios.

[0024] In view of this, this application provides a wake-up method that continuously monitors speech activity using a bone conduction sensor and activates the microphone only when pre-wake-up conditions are met to maintain low power consumption. However, since there is a certain delay in microphone activation, the first microphone signal collected may lack the initial speech segment. Therefore, this application fills the first microphone signal with data based on the bone conduction signal to generate a second microphone signal that is time-synchronized with the bone conduction signal. Both signals are then synchronously input into a wake-up model, enabling the model to effectively identify the start information of the wake-up word by combining the noise immunity of the bone conduction signal with the complete temporal structure of the filled microphone signal, thus improving wake-up performance while maintaining low power consumption.

[0025] It should be noted that the executing entity of the wake-up method can be an electronic device, which can be a local device, such as a television, computer, laptop, mobile phone, smartwatch, etc., or a virtual device. This application embodiment does not impose any limitations on this. For ease of description, the execution entity is omitted from the following description of each embodiment.

[0026] Reference Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the wake-up method of this application. In this embodiment, the wake-up method includes the following steps: Step S10: Acquire bone voiceprint signals collected by the bone voiceprint sensor; In one feasible embodiment, bone voiceprint signals collected by a bone voiceprint sensor are acquired to enable continuous monitoring of speech activity via the bone voiceprint sensor.

[0027] A bone conduction voiceprint sensor (VPU) is a vibration sensing device based on bone conduction principles, used to collect vocal cord vibration signals (i.e., bone voiceprint signals) transmitted through the skull when a user speaks. Unlike traditional air conduction microphones, bone voiceprint sensors are directly attached to or close to the user's temporal bone, mandible, etc., acquiring voice information by detecting the mechanical vibration of the bones. Bone voiceprint sensors have strong resistance to environmental noise, maintaining a high signal-to-noise ratio even in high-noise environments. Furthermore, their simple structure and low operating current allow them to remain active for extended periods without significantly increasing system standby power consumption, making them suitable as a voice activity detection unit in the pre-wake-up phase. Although the frequency response of bone voiceprint sensors is biased towards low frequencies (typically with an effective bandwidth concentrated between 100 Hz and 1 kHz), and their response to high-frequency components such as voiceless consonants is weak, they have excellent perception capabilities for vowels and the starting point of speech energy, reliably triggering subsequent wake-up processes.

[0028] Optionally, bone voiceprint signals can reflect the vibration characteristics of the skull when a user speaks. Since the bone conduction path is not sensitive to environmental noise, bone voiceprint signals can still effectively characterize the existence of speech activity in high-noise scenarios, thus providing a reliable basis for pre-wake-up judgment.

[0029] Step S20: If the bone voiceprint signal meets the preset pre-wake-up conditions, the microphone is activated and the first microphone signal collected by the microphone is acquired. In one feasible embodiment, it is dynamically determined whether the bone voiceprint signal meets preset pre-wake-up conditions. These pre-wake-up conditions refer to a set of criteria used to determine whether there is potentially valid speech activity in the current bone voiceprint signal. By using pre-wake-up conditions, environmental sounds can be initially screened using a low-power bone voiceprint sensor without enabling a high-power microphone. The microphone is only powered on when there is a high probability of speech activity, thus maximizing system low-power operation while ensuring wake-up sensitivity.

[0030] Optionally, the pre-wake condition may include at least one of the following: The signal energy of the bone voiceprint signal is greater than the preset energy threshold within a preset time window (e.g., 200 ms); Bone voiceprint signals contain continuous human voice features; The similarity between the bone voiceprint signal and the preset wake word start template (such as the rising edge of the vibration energy corresponding to the word "small") exceeds the preset similarity threshold.

[0031] Step S30: Based on the bone voiceprint signal, the first microphone signal is filled with data to obtain the second microphone signal, wherein the second microphone signal is synchronized with the bone voiceprint signal in time. In one feasible embodiment, due to the inherent hardware response delay (e.g., 20-50 ms) from when the microphone is activated to when it outputs a stable and usable audio signal, the first microphone signal it acquires only contains the portion after the start of the speech, which may result in the missing key phonemes at the beginning of the wake word (such as "Xiaoyi Xiaoyi"). To solve this problem, embodiments of this application introduce a time alignment and data padding mechanism based on bone voiceprint signals. According to the time structure of the bone voiceprint signal, the specific time periods in the first microphone signal are filled, so that the second microphone signal and the bone voiceprint signal are strictly aligned on the time axis, and both are based on the same start time, forming a synchronous dual-channel input.

[0032] Step S40: Input the second microphone signal and bone conduction signal into the preset wake-up model to obtain the wake-up recognition result; In one feasible embodiment, since the second microphone signal and the bone voiceprint signal are strictly synchronized on the time axis, they can form a time-aligned dual-channel input. The first channel is the bone voiceprint signal, which has the advantages of stable low-frequency vibration energy and strong resistance to environmental noise, especially providing reliable speech presence cues in the speech initiation segment. The second channel is the filled microphone signal (i.e., the second microphone signal). Although it does not have real high-frequency data in the initial interval, it contains complete full-band speech information after the initial interval, and can accurately restore key phoneme features such as voiceless consonants. Then, the second microphone signal and the bone voiceprint signal are input into a preset wake-up model to obtain the wake-up recognition result. Through this dual-channel collaborative recognition mechanism, the model can effectively overcome the misjudgment problem caused by the lack of initial signal in a single microphone signal, significantly improve the robustness of recognition of wake words containing weak-energy voiceless consonants, and avoid the defect of insufficient high-frequency information introduced by over-reliance on bone voiceprint signals.

[0033] Optionally, the wake-up model can be a deep neural network structure, such as a dual-branch convolutional neural network, a temporally fused Transformer, or a lightweight TDNN (temporally delayed neural network) architecture, which uses dual-channel samples constructed with the same temporal alignment during the training phase for learning.

[0034] Optionally, the wake-up model can primarily rely on the low-frequency context of the bone voiceprint signal to determine whether a valid speech start exists within the initial interval; in the interval after the initial interval, it can fuse the features of the two channels and use the high-frequency details of the microphone signal and the noise-resistant characteristics of the bone voiceprint signal for joint discrimination.

[0035] Optionally, the output of the wake-up model may include the confidence score and / or classification probability of the wake-up word. When the score exceeds a preset threshold, it is determined that the wake-up word has been "hit", and the corresponding wake-up recognition result is generated.

[0036] Optionally, the wake-up recognition result may include at least one of the following: wake-up successful (target wake-up word detected), wake-up failed (target wake-up word not detected), and multiple wake-up word category identifier (such as distinguishing between "Xiaoyi" and "Xiaoai").

[0037] Step S50: Execute the wake-up operation corresponding to the wake-up recognition result.

[0038] In one feasible embodiment, the wake-up operation corresponding to the wake-up recognition result is performed, such as starting the main processor or activating the voice assistant.

[0039] Optionally, refer to Figure 3 During the wake-up process, the VPU (Voice Processing Unit, i.e., bone conduction sensor) continuously collects bone conduction signals and sends them to the VAD (Voice Activity Detection) module. The VAD module can analyze the bone conduction signals in real time to determine whether a voice signal exists. If the determination is "NO," meaning no voice signal, the system maintains its current state, and the MIC (microphone) remains in sleep mode to reduce standby power consumption. If the determination is "YES," meaning a voice signal exists, the MIC is powered on, and the microphone begins to collect audio signals. After the MIC is powered on, the system acquires its output first microphone signal. Subsequently, based on the bone conduction signal, the first microphone signal is padded with data to obtain a second microphone signal that is synchronized with the bone conduction signal in time. The bone conduction signal and the padded second microphone signal are synchronously input into a preset wake-up model. This model, based on a trained neural network structure, integrates the low-frequency noise reduction characteristics of the VPU with the high-frequency detail information of the MIC to complete the recognition of the wake-up word and output the recognition result.

[0040] In this embodiment, a bone conduction sensor continuously monitors speech activity, and the microphone is activated only when the pre-wake-up conditions are met to maintain low power consumption. However, due to the slight delay in microphone activation, the first microphone signal may lack the initial speech segment. Therefore, this embodiment fills in the first microphone signal with data based on the bone conduction signal to generate a second microphone signal that is time-synchronized with the bone conduction signal. Both signals are then synchronously input into the wake-up model, enabling the model to effectively identify the start information of the wake-up word by combining the noise immunity of the bone conduction signal with the complete temporal structure of the filled microphone signal, thus improving wake-up performance while maintaining low power consumption.

[0041] Based on the first embodiment described above, a second embodiment of the wake-up method of this application is proposed. In this embodiment, before step S20, the step of enabling the microphone, the method further includes: Step S11: Detect whether there is a speech signal in the bone voiceprint signal; In one feasible embodiment, the presence of a speech signal in the bone voiceprint signal is detected to determine whether there is valid speech activity in the bone voiceprint signal.

[0042] In one feasible implementation, step S11, the step of detecting whether a speech signal exists in the bone voiceprint signal, includes: Step S111: Determine whether the signal energy of the bone voiceprint signal within a preset time window is greater than a preset energy threshold. In one feasible embodiment, the presence of a speech signal in the bone voiceprint signal is determined by judging whether the signal energy of the bone voiceprint signal within a preset time window is greater than a preset energy threshold.

[0043] Optionally, the preset time window can be, for example, a sliding window of 20ms to 100ms, and the signal energy of the current frame or multiple consecutive frames within the preset time window can be calculated.

[0044] Optionally, the signal energy can be the sum of squares or the root mean square value of the signal amplitude.

[0045] Optionally, the energy threshold can be dynamically adjusted according to the background vibration level of the environment in which the device is located, or it can be determined by calibration at the factory to distinguish between the silent state and the user's voice state.

[0046] Step S112: If the signal energy of the bone voiceprint signal within the preset time window is greater than the preset energy threshold, then it is determined that there is a speech signal in the bone voiceprint signal. Step S113: If the signal energy of the bone voiceprint signal within the preset time window is less than or equal to the preset energy threshold, then it is determined that there is no speech signal in the bone voiceprint signal.

[0047] In one feasible embodiment, if the signal energy of the bone voiceprint signal within a preset time window is greater than a preset energy threshold, it is determined that a speech signal exists in the bone voiceprint signal, i.e., the pre-wake-up condition is met. If the signal energy of the bone voiceprint signal within the preset time window is less than or equal to the preset energy threshold, it is determined that no speech signal exists in the bone voiceprint signal, i.e., the pre-wake-up condition is not met.

[0048] In this embodiment, the energy threshold detection method has advantages such as simple calculation, low resource consumption, and fast response speed. Meanwhile, since the bone conduction signal itself is minimally affected by ambient air noise, its energy changes mainly reflect real speech vibrations. Therefore, this criterion still possesses high reliability in high-noise environments, effectively avoiding the false triggering problem caused by background noise in traditional air microphones.

[0049] Step S12: If it is determined that there is a speech signal in the bone voiceprint signal, determine that the bone voiceprint signal meets the preset pre-wake-up conditions, and execute the step of enabling the microphone, as well as subsequent steps.

[0050] In one feasible embodiment, if it is determined that there is a speech signal in the bone voiceprint signal, then it is determined that the bone voiceprint signal meets the preset pre-wake-up conditions, and step S20, the step of enabling the microphone, and subsequent steps are executed.

[0051] In this embodiment, rapid coarse screening for pre-wake-up is achieved by detecting the speech signal in the bone voiceprint signal, maintaining low power consumption, and improving trigger accuracy.

[0052] In one feasible implementation, step S30, which involves filling the first microphone signal with data based on the bone conduction signal to obtain the second microphone signal, includes: Step S31: Identify the start time of the appearance of the speech signal in the bone voiceprint signal; In one feasible embodiment, the start time of the appearance of the speech signal in the bone voiceprint signal is identified (e.g., denoted as time S).

[0053] Optionally, the start time of the voice signal can be obtained by the timestamp when the pre-wake condition is triggered or by signal energy detection.

[0054] Step S32: Determine the missing time period that needs to be filled with data based on the start time and the acquisition time of the first microphone signal; In one feasible embodiment, the moment when the microphone completes power-on and outputs a stable signal is recorded, i.e., the acquisition time of the first microphone signal (e.g., denoted as time C). Since the first microphone signal is only valid from time C, its time axis is [C, E], where E is the end time of the speech signal, while the time axis of the bone voiceprint signal is a complete [S, E]. Therefore, based on the time structure of the bone voiceprint signal, the missing time period that needs to be filled with data can be determined, i.e., the time difference ΔT = C between the start time S and the stable time C. S.

[0055] Step S33: Generate filler data that matches the missing time period; In one feasible embodiment, filler data matching the missing time period is generated, wherein the filler data includes at least one of: silent data, ambient noise data, and preset audio template data.

[0056] Optionally, silent data refers to audio frames with amplitudes close to zero or at the level of system noise floor, typically representing a state of no effective speech input. In digital audio, this can be directly represented by all-zero sequences (e.g., PCM value of 0) or extremely low-energy random noise. This clearly conveys the semantic information that there is no reliable microphone input during this period to the wake-up model. If the same zero-padding strategy is used during model training, the network can learn to primarily rely on bone conduction signals for judgment in this interval, avoiding the introduction of false high-frequency features that interfere with recognition.

[0057] Optionally, ambient noise data refers to background noise samples collected before the microphone is activated or during device standby, such as wind noise, traffic noise, and office noise. This data can be cached before the microphone goes into sleep mode or obtained through brief sampling at the beginning of each wake-up process. If ambient noise data is used, the padded second microphone signal can more closely resemble the real physical scene, avoiding spectral distortion caused by abrupt "zero jumps," and helping the wake-up model better distinguish speech from background in high-noise environments.

[0058] Optionally, the preset audio template data refers to pre-stored synthetic or real speech segments that match the statistical characteristics of the starting part of the target wake word. For example, for "Xiao Yi Xiao Yi", multiple voiceless consonant + vowel transition templates starting with the word "Xiao" can be pre-stored (such as / (Typical waveform or Mel spectrum of i / ). In the absence of key starting phonemes (such as "x"), the pre-set audio template data can provide reasonable high-frequency prior information to assist the model in completing word beginning reconstruction.

[0059] Step S34: The padding data is concatenated with the first microphone signal to obtain the second microphone signal.

[0060] In one feasible embodiment, the padding data is concatenated with the first microphone signal, that is, within the time interval [S, C), the microphone signal value is set as the padding data, thereby generating a complete signal sequence covering the time range [S, E], which is the second microphone signal.

[0061] In this embodiment, through the above processing, the second microphone signal and the bone conduction signal are strictly aligned on the time axis, both using the same start time S as a reference, forming a synchronous dual-channel input. Although the second microphone signal has no real speech data in the interval [S, C), it retains the complete temporal context structure of the wake-up word, enabling the subsequent wake-up model to accurately perceive the speech start position and, combined with the low-frequency vibration information of the bone conduction signal in this interval, achieve semantic compensation for the missing high-frequency components.

[0062] Based on any of the above embodiments, a third embodiment of the wake-up method of this application is proposed. In this embodiment, before step S40, which inputs the second microphone signal and the bone conduction signal to the preset wake-up model, the method further includes: Step A10: Obtain the bone voiceprint training signal collected by the bone voiceprint sensor and the microphone first training signal collected by the microphone, wherein both the bone voiceprint training signal and the microphone first training signal contain a preset wake-up word. In one feasible embodiment, during the training data preparation phase, two signals from the same user's pronunciation are collected simultaneously. One signal is the bone voiceprint training signal collected from the bone voiceprint sensor, and the other signal is the microphone first training signal collected from the microphone. All training signals contain a preset wake-up word (such as "Xiaoyi Xiaoyi"), and the two signals are ensured to be originally aligned in time.

[0063] Step A20: Replace the data in the first training signal of the microphone that is located in the preset analog activation delay period with preset training padding data to obtain the second training signal of the microphone; In one feasible embodiment, to simulate the situation where the microphone lacks a start segment due to power-on delay in actual deployment, a data segment of a preset duration (e.g., ΔT) is artificially truncated at the beginning of the microphone's first training signal (i.e., data located within the simulated activation delay period). The simulated activation delay period can be 30 ms, 50 ms, 130 ms, etc., covering a typical hardware delay range. Then, the data located within the preset simulated activation delay period replaces the preset training filler data, wherein the training filler data may include at least one of: silent data, ambient noise data, and preset audio template data, to obtain the microphone's second training signal.

[0064] Step A30: Obtain the wake-up tags corresponding to the bone voiceprint training signal and the microphone second training signal; In one feasible embodiment, the wake-up recognition result corresponding to the first training signal of the microphone is obtained as the wake-up tag corresponding to the bone voiceprint training signal and the second training signal of the microphone.

[0065] Step A40: Use the bone voiceprint training signal, the microphone second training signal, and the wake-up tag as a training sample, and obtain the wake-up training sample set based on the obtained training samples. In one feasible embodiment, the bone voiceprint training signal, the microphone second training signal, and the wake-up tag are used as a training sample, and a wake-up training sample set is obtained based on the acquired training samples.

[0066] Optionally, a large-scale wake-up training sample set can be constructed by traversing a large number of recordings from different speakers, in different noise environments, and at different speaking speeds. The key feature of this sample set is that the microphone channel always contains an initial missing + padding structure consistent with the actual deployment, thereby enabling the model to learn how to use bone conduction signals to compensate for missing microphone information during training.

[0067] Step A50: Train the preset wake-up model using the wake-up training sample set to obtain the wake-up model.

[0068] In one feasible embodiment, a wake-up training sample set is used to train a wake-up model (such as a convolutional neural network CNN, a time-delay neural network TDNN, a lightweight Transformer, etc.) to obtain a wake-up model.

[0069] Alternatively, the optimization objective during model training can be to minimize the loss between the predicted label and the actual wake-up label (such as cross-entropy loss).

[0070] Optionally, the wake-up method of this application embodiment is compared with the conventional wake-up method, and the results are shown in Table 1 below: Table 1

[0071] As can be seen, the advantages of the solution in this application embodiment become more and more obvious as the signal-to-noise ratio decreases.

[0072] In this embodiment, since the wake-up model has seen a large number of simulated missing scenarios during training, when facing real delays in actual deployment, it can mainly rely on bone voiceprint signals to determine the presence of speech during the missing period, and then fuse dual-channel features in the effective interval after the missing period to improve high-frequency recognition, thereby achieving a balance between high wake-up rate and low false wake-up rate.

[0073] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the activation method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0074] This application provides a wake-up device, referring to... Figure 4 The device includes: Acquisition module 10 is used to acquire bone voiceprint signals collected by the bone voiceprint sensor; The activation module 20 is used to activate the microphone and acquire the first microphone signal collected by the microphone when it is determined that the bone voiceprint signal meets the preset pre-wake-up conditions. The filling module 30 is used to fill the first microphone signal with data according to the bone voiceprint signal to obtain a second microphone signal, wherein the second microphone signal is synchronized with the bone voiceprint signal in time; The input module 40 is used to input the second microphone signal and the bone conduction signal into a preset wake-up model to obtain a wake-up recognition result; The execution module 50 is used to execute the wake-up operation corresponding to the wake-up recognition result.

[0075] The wake-up device provided in this application, employing the wake-up method described in the above embodiments, can effectively improve wake-up performance while maintaining low power consumption. Compared with the prior art, the beneficial effects of the wake-up device provided in this application are the same as those of the wake-up method described in the above embodiments, and other technical features in the wake-up device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0076] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the wake-up method in the first embodiment described above.

[0077] The following is for reference. Figure 5 The diagrams show structural schematics of electronic devices suitable for implementing the embodiments of this application. The electronic devices in the embodiments of this application may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0078] like Figure 5As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. While electronic devices with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0079] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0080] The electronic device provided in this application, employing the wake-up method described in the above embodiments, can effectively improve wake-up performance while maintaining low power consumption. Compared with the prior art, the beneficial effects of the electronic device provided in this application are the same as those of the wake-up method provided in the above embodiments, and other technical features of the electronic device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0081] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0082] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0083] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the wake-up method in the above embodiments.

[0084] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0085] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.

[0086] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to: acquire bone voiceprint signals collected by a bone voiceprint sensor; activate a microphone and acquire a first microphone signal collected by the microphone when the bone voiceprint signal meets preset pre-wake-up conditions; fill the first microphone signal with data based on the bone voiceprint signal to obtain a second microphone signal, wherein the second microphone signal is synchronized with the bone voiceprint signal in time; input the second microphone signal and the bone voiceprint signal into a preset wake-up model to obtain a wake-up recognition result; and execute a wake-up operation corresponding to the wake-up recognition result.

[0087] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0088] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0089] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0090] The readable storage medium provided in this application embodiment is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described wake-up method, which can effectively improve wake-up performance while maintaining low power consumption. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application embodiment are the same as the beneficial effects of the wake-up method provided in the above embodiments, and will not be repeated here.

[0091] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A wake-up method, characterized in that, The method includes: Acquire bone voiceprint signals collected by a bone voiceprint sensor; If the bone voiceprint signal meets the preset pre-wake-up conditions, the microphone is activated and the first microphone signal collected by the microphone is acquired. Based on the bone voiceprint signal, the first microphone signal is filled with data to obtain the second microphone signal, wherein the second microphone signal is synchronized with the bone voiceprint signal in time; The second microphone signal and the bone conduction signal are input into a preset wake-up model to obtain the wake-up recognition result; Execute the wake-up operation corresponding to the wake-up recognition result.

2. The method as described in claim 1, characterized in that, The step of filling the first microphone signal with data based on the bone conduction signal to obtain the second microphone signal includes: Identify the start time of the appearance of the speech signal in the bone voiceprint signal; Based on the start time and the acquisition time of the first microphone signal, determine the missing time period that needs to be filled with data; Generate filler data that matches the missing time period; The padding data is concatenated with the first microphone signal to obtain the second microphone signal.

3. The method as described in claim 2, characterized in that, The filling data includes at least one of the following: silent data, ambient noise data, and preset audio template data.

4. The method as described in claim 1, characterized in that, Prior to the step of enabling the microphone, the following is also included: Detect whether there is a speech signal in the bone voiceprint signal; If it is determined that there is a speech signal in the bone voiceprint signal, the bone voiceprint signal is determined to meet the preset pre-wake-up conditions, and the step of enabling the microphone and subsequent steps are executed.

5. The method as described in claim 4, characterized in that, The step of detecting whether a speech signal exists in the bone voiceprint signal includes: Determine whether the signal energy of the bone voiceprint signal within a preset time window is greater than a preset energy threshold; If the signal energy of the bone voiceprint signal within a preset time window is greater than a preset energy threshold, then it is determined that there is a speech signal in the bone voiceprint signal. If the signal energy of the bone voiceprint signal within a preset time window is less than or equal to a preset energy threshold, then it is determined that there is no speech signal in the bone voiceprint signal.

6. The method as described in claim 1, characterized in that, Before the step of inputting the second microphone signal and the bone conduction signal to the preset wake-up model, the method further includes: Acquire bone voiceprint training signals collected by a bone voiceprint sensor and microphone first training signals collected by a microphone, wherein both the bone voiceprint training signals and the microphone first training signals contain a preset wake-up word; The data in the first training signal of the microphone that is located in the preset simulated activation delay period is replaced with preset training padding data to obtain the second training signal of the microphone; Obtain the wake-up tags corresponding to the bone voiceprint training signal and the second microphone training signal; The bone voiceprint training signal, the microphone second training signal, and the wake-up tag are used as a training sample, and a wake-up training sample set is obtained based on the acquired training samples. The wake-up training sample set is used to train the preset wake-up model to obtain the wake-up model.

7. The method as described in claim 6, characterized in that, The step of acquiring the wake-up tag corresponding to the bone conduction training signal and the second microphone training signal includes: The wake-up recognition result corresponding to the first training signal of the microphone is determined as the wake-up tag.

8. A wake-up module, characterized in that, The module includes: The acquisition module is used to acquire bone voiceprint signals collected by the bone voiceprint sensor; The activation module is used to activate the microphone and acquire the first microphone signal collected by the microphone when it is determined that the bone voiceprint signal meets the preset pre-wake-up conditions. A filling module is used to fill the first microphone signal with data according to the bone voiceprint signal to obtain a second microphone signal, wherein the second microphone signal is synchronized with the bone voiceprint signal in time; The input module is used to input the second microphone signal and the bone conduction signal into a preset wake-up model to obtain a wake-up recognition result; The execution module is used to perform the wake-up operation corresponding to the wake-up recognition result.

9. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the wake-up method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the wake-up method as described in any one of claims 1 to 7.