Voice wake-up method and apparatus, and electronic device and storage medium

By adjusting the self-supervised training speech feature extraction sub-model, its noise resistance is enhanced, solving the problem of low wake-up rate in high-noise environments. This enables accurate device wake-up in noisy environments, improving user experience and device applicability.

WO2026007709A1PCT designated stage Publication Date: 2026-01-08BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/102042
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-02
Filing Date
2025-06-19
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

In high-noise environments, existing voice wake-up solutions struggle to effectively detect wake-up keywords, resulting in low wake-up rates.

Method used

A reference speech feature extraction sub-model is obtained by adjusting the candidate speech feature extraction sub-model based on pre-self-supervised training, thereby enhancing its anti-noise interference performance. Feature information is extracted from the speech signal to be processed through this model to achieve wake-up control.

Benefits of technology

Significantly improved wake-up rate in high-noise environments, ensuring accurate and reliable wake-up of devices in complex noisy scenarios, improving user experience and enhancing device applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025102042_08012026_PF_FP_ABST
    Figure CN2025102042_08012026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a voice wake-up method and apparatus, and an electronic device and a storage medium. The voice wake-up method comprises: determining a voice signal to be processed which is acquired by a target device; determining a reference voice feature extraction sub-model associated with the target device, and extracting, from said voice signal by means of the reference voice feature extraction sub-model, voice feature information to be processed, wherein the reference voice feature extraction sub-model is obtained by means of performing adjustment on the basis of a candidate voice feature extraction sub-model that has been subjected to self-supervised training in advance, and the reference voice feature extraction sub-model has inherited the noise interference resistance performance of said candidate voice feature extraction sub-model during voice feature extraction; and on the basis of the voice feature information to be processed, performing wake-up control on the target device or a target application. The technical solution of the present disclosure improves the capability of detecting a wake-up keyword in a high-noise environment, thereby greatly increasing the wake-up rate.
Need to check novelty before this filing date? Find Prior Art

Description

Voice wake-up method and device, electronic device, and storage medium

[0001] This application claims priority to Chinese Patent Application No. 202410883020.1, filed on July 2, 2024, the disclosure of which is incorporated herein in its entirety as part of the present application. TECHNICAL FIELD

[0002] Embodiments of the present disclosure relate to a voice wake-up method, device, electronic device, and storage medium. BACKGROUND

[0003] Generally, a device needs to be woken up from a sleep state to a working state to process instructions normally, such as wake-up methods including touch wake-up (such as a lock screen key), timing wake-up (such as an alarm clock), passive wake-up (such as a phone), and the like. In order to make device wake-up more convenient, voice wake-up technology is gradually used to wake up a device by voice to switch the device from a sleep state to a working state. However, the related voice wake-up scheme has difficulty in effectively solving the wake-up problem in a high-noise environment, which specifically manifests in that when in a larger external noise environment, such as a subway, an airport, a shopping mall, and the like, it is difficult to detect a wake-up keyword, thereby resulting in a low wake-up rate. SUMMARY

[0004] The present disclosure provides a voice wake-up method, device, electronic device, and storage medium to solve the problem of low wake-up rate caused by difficulty in detecting a wake-up keyword in a high-noise environment.

[0005] In a first aspect, the embodiments of the present disclosure provide a voice wake-up method, which comprises:

[0006] determining a to-be-processed voice signal obtained by a target device, the to-be-processed voice signal supporting carrying a preset wake-up keyword used for waking up the target device;

[0007] determining a reference voice feature extraction sub-model associated with the target device, and extracting to-be-processed voice feature information from the to-be-processed voice signal through the reference voice feature extraction sub-model, the reference voice feature extraction sub-model being obtained by adjusting a pre-self-supervised training candidate voice feature extraction sub-model, and the reference voice feature extraction sub-model having inherited the anti-noise interference performance of the pre-self-supervised training candidate voice feature extraction sub-model when performing voice feature extraction;

[0008] controlling the target device or the target application to wake up based on the to-be-processed voice feature information.

[0009] In a second aspect, the embodiments of the present disclosure further provide a voice wake-up device, which comprises:

[0010] a first determining module configured to determine a to-be-processed voice signal obtained by a target device, the to-be-processed voice signal supporting carrying of a preset wake-up keyword for waking up the target device or a target application associated with the target device;

[0011] a second determining module configured to determine a reference voice feature extraction sub-model associated with the target device, and extract to-be-processed voice feature information from the to-be-processed voice signal through the reference voice feature extraction sub-model, the reference voice feature extraction sub-model being obtained by adjusting a pre-self-supervised training candidate voice feature extraction sub-model, and the reference voice feature extraction sub-model having inherited an anti-noise interference performance of the pre-self-supervised training candidate voice feature extraction sub-model when performing voice feature extraction;

[0012] a control module configured to perform wake-up control on the target device or the target application based on the to-be-processed voice feature information.

[0013] In a third aspect, an electronic device is provided, and the electronic device includes:

[0014] one or more processors;

[0015] a storage device configured to store one or more programs,

[0016] when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement a voice wake-up method according to any of the embodiments of the present disclosure.

[0017] In a fourth aspect, a storage medium containing computer executable instructions is provided, and the computer executable instructions, when executed by a computer processor, are used to perform a voice wake-up method according to any of the embodiments of the present disclosure.

[0018] It should be understood that the contents described in this section are not intended to identify key or important features of the embodiments of the present disclosure, nor are they used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0019] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent as various embodiments of the present disclosure are described in conjunction with the following drawings, in which like reference numbers represent like elements throughout the drawings. It should be understood that the drawings are schematic and elements are not necessarily drawn to scale.

[0020] FIG. 1 is a flow diagram of a voice wake-up method according to an embodiment of the present disclosure;

[0021] Fig. 2 is a structural schematic diagram of a voice wake-up device according to an embodiment of the present disclosure; and

[0022] Fig. 3 is a structural schematic diagram of an electronic device implementing a voice wake-up method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] Embodiments of the present disclosure will be described in more detail with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be construed as being limited to the embodiments set forth herein, but rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are merely for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.

[0024] It should be understood that the various steps recited in the method embodiments of the present disclosure can be performed in different orders, and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this respect.

[0025] The term "comprising" and variations thereof as used herein are used inclusively, i.e., "comprising, but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions will be given in the description below.

[0026] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.

[0027] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise explicitly stated in the context, it should be understood as "one or more".

[0028] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are merely for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0029] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained in accordance with relevant laws and regulations.

[0030] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will require obtaining and using personal information of the user. Thus, the user can autonomously select whether to provide the personal information to the software or hardware, such as an electronic device, an application, a server or a storage medium, etc. performing the operation of the technical solution of the present disclosure according to the prompt information.

[0031] As an optional but non-limiting implementation, in response to receiving an active request of a user, the prompt information can be sent to the user in the form of a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0032] It can be understood that the above notification and obtaining user authorization process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0033] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the obtaining or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.

[0034] FIG. 1 is a flow diagram of a voice wake-up method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the case of waking up a device from a sleep state to a working state by a voice. The voice wake-up method can be executed by a voice wake-up apparatus, which can be implemented in the form of software and / or hardware, and is generally integrated on any electronic device with network communication function, such as a mobile terminal, a PC terminal or a server.

[0035] As shown in FIG. 1, the voice wake-up method of the embodiment of the present disclosure can include the following processes:

[0036] S110, determining a target device to obtain a to-be-processed voice signal, and the to-be-processed voice signal supports carrying a preset wake-up keyword for waking up the target device or a target application associated with the target device.

[0037] The target device can refer to a device that is switched from a sleep state to a working state through a specific voice instruction including a preset wake-up keyword. The target device needs to receive a specific voice instruction including a preset wake-up keyword to start or activate its main function when it is in a standby, sleep or inactivated state. For example, a smart speaker is in a low-power standby state when it is not woken up. Only when the user speaks the preset wake-up word, the smart speaker will start responding and executing the user's subsequent voice instructions, such as playing music, querying the weather, etc. For another example, some smart home appliances, such as smart TVs or smart air conditioners, can also be set to a voice wake-up mode. Only after receiving the correct voice wake-up instruction including the preset wake-up keyword, the smart home appliances will switch from the sleep state to the working state and be ready to receive further operation instructions.

[0038] For another example, the target device can be a terminal device such as a mobile phone, which can be installed with a target application program (such as a voice assistant) that can be started and controlled through voice instructions. The target application program can be woken up by a wake-up word received by the terminal device. For another example, the target device can be a headset connected to a smart device such as a mobile phone, which can be installed with a target application program (such as a voice assistant) that can be started through voice instructions. The user can input a wake-up word to the voice assistant through the headset to wake up the voice assistant, input various voice instructions to the voice assistant through the headset, and obtain voice responses provided by the voice assistant through the headset.

[0039] The voice signal to be processed can include sound information obtained within a preset distance range around the location of the target device. The sound information corresponding to the voice signal to be processed contains various voice contents, which can be speech uttered by a sound source close to the target device, a mixture of noise and voice in the environment of the location where the target device is located, or a preset keyword related to waking up the target device or a target application associated with the target device. Therefore, the voice signal to be processed supports carrying a preset wake-up keyword for waking up the target device. The preset wake-up keyword can be a specific voice instruction or phrase, which is used to trigger the device to switch from a sleep or standby state to a working state and be ready to receive and process subsequent voice commands or perform related operations. When the target device or the target application associated with the target device receives the preset wake-up keyword through voice, it can switch from the sleep state to the working state.

[0040] Optionally, the voice signal to be processed can cover the sound information collected within a preset distance range around the target device in a reference noise environment. The reference noise environment can be an environment where the noise intensity is relatively high, such as a noise loudness greater than a preset noise intensity. In such an environment with high noise intensity, the voice is often significantly disturbed, and the overall acoustic environment is relatively complex and noisy, which will inevitably make the key detection of the voice signal in the noise environment more difficult, and thus adversely affect the success rate of voice wake-up and reduce the voice wake-up rate.

[0041] As an optional but non-limiting implementation, the voice signal to be processed obtained by the target device includes the following steps A1-A2:

[0042] Step A1, using the sound collector configured on the target device to obtain the voice signal in the preset distance range of the target device in real time.

[0043] Step A2, obtaining the voice signal to be processed after preprocessing the voice signal obtained in real time, the preprocessing including at least one of noise removal and filtering processing.

[0044] The sound collector configured on the target device can be a device that can convert the sound in the surrounding environment of the target device into an electrical signal. The sound collector configured on the target device has a specific sensitivity and receiving range, and can collect the sound information within the preset distance range around the target device by installing the sound collector on the target device. Then, at least one preprocessing method such as noise removal and filtering processing can be performed on the voice signal collected in real time to obtain the voice signal to be processed.

[0045] The preset distance range can be a pre-set distance area. For example, if the preset distance range is 3 meters, the sound collector will collect the sound within a radius of 3 meters centered on the target device. The sound collector configured on the target device will continuously and immediately convert the sound information generated within the preset distance range of the target device into an electrical signal and transmit it to the target device or the target application associated with the target device, so as to perform subsequent analysis, processing or corresponding operations, such as waking up the device, to obtain the voice information within a specific distance around the target device in time and provide data support for various voice-related functions of the device.

[0046] S120, determine the reference speech feature extraction sub-model associated with the target device, and extract the to-be-processed speech feature information from the to-be-processed speech signal through the reference speech feature extraction sub-model, the reference speech feature extraction sub-model is obtained by adjusting the candidate speech feature extraction sub-model based on the pre-self-supervised training, and the reference speech feature extraction sub-model has inherited the anti-noise interference performance of the candidate speech feature extraction sub-model in the pre-self-supervised training when performing speech feature extraction.

[0047] The reference speech feature extraction sub-model can be a model for extracting speech feature information from a to-be-processed speech signal. The reference speech feature extraction sub-model is obtained by adjusting the candidate speech feature extraction sub-model based on pre-self-supervised training, and is ensured to be adjusted through the candidate speech feature extraction sub-model during the process of adjusting to form the reference speech feature extraction sub-model. The reference speech feature extraction sub-model automatically captures the speech features in the speech signal, and gradually inherits the anti-noise interference performance of the candidate speech feature extraction sub-model in the pre-self-supervised training when performing speech feature extraction. The speech feature information extracted by the reference speech feature extraction sub-model includes but is not limited to the frequency spectrum, time domain feature, and acoustic feature of the speech.

[0048] For the to-be-processed speech signal obtained by the target device, a specific reference speech feature extraction sub-model associated with the target device is obtained, and the to-be-processed speech signal is processed using this reference speech feature extraction sub-model to extract relevant to-be-processed speech feature information. The reference speech feature extraction sub-model is not produced out of thin air, but is obtained by adjusting the candidate speech feature extraction sub-model that is pre-trained by self-supervision.

[0049] Moreover, the reference speech feature extraction sub-model is obtained by adjusting the candidate speech feature extraction sub-model based on pre-self-supervised training. By adjusting the candidate speech feature extraction sub-model, it can better adapt to the environment and noise situation of the target device. For example, the parameters of the candidate speech feature extraction sub-model can be fine-tuned according to the type and intensity of the noise to improve its anti-interference ability in a noisy environment, so that the reference speech feature extraction sub-model inherits the ability of the candidate speech feature extraction sub-model in pre-self-supervised training to resist noise interference when extracting speech features, that is, the candidate speech feature extraction sub-model can accurately extract speech features in a noisy environment.

[0050] The reason for selecting the reference speech feature extraction sub-model is that the reference speech feature extraction sub-model inherits the ability of the pre-self-supervised trained candidate speech feature extraction sub-model to resist noise interference in speech feature extraction, which means that even if the speech signal to be processed is in a noisy environment, the reference speech feature extraction sub-model can relatively accurately extract useful speech features, reduce the adverse effects of noise on feature extraction, and thus provide a more reliable and effective basis for subsequent speech processing tasks (such as speech recognition, speech understanding, etc.).

[0051] As an optional but non-limiting implementation, the determination process of the reference speech feature extraction sub-model in the embodiments of the present disclosure includes the following steps B1-B2:

[0052] Step B1, determine the candidate speech feature extraction sub-model corresponding to the target device, which is obtained by training and adjusting the speech feature extraction sub-model based on the mask prediction self-supervised learning method on the candidate speech data set, the data amount of the candidate speech data set is greater than the preset data amount, and the candidate speech feature extraction sub-model has stable robustness to noise.

[0053] Optionally, a large amount of speech data is collected, including speech signals in various noisy environments and sound information within a preset distance range around the location of the target device, such as a million-hour-level speech data set for training to improve the generalization ability and robustness to noise of the model. Then, based on the mask prediction self-supervised learning model (such as the pre-training method of HuBERT), a part of the input speech data is randomly masked, and then the features of the masked frames are predicted, so as to learn the intrinsic representation of the speech, and fine-tuning training is performed on a large-scale speech data set, such as using a stochastic gradient descent optimization algorithm for training fine-tuning, and the anti-noise interference performance of the adjusted reference speech feature extraction sub-model is evaluated to ensure that it can effectively extract speech features in a noisy environment. After multiple adjustments and optimizations, a candidate speech feature extraction model with relatively good robustness to noise is obtained, which can be used for actual speech processing tasks such as speech recognition, speech enhancement, and speech wake-up.

[0054] For the candidate speech feature extraction sub-model corresponding to the target device, the candidate speech feature extraction sub-model is trained based on the mask prediction self-supervised learning method on a large amount of candidate speech data set, the data amount of the candidate speech data set needs to be greater than the preset data amount, the model data amount of the candidate speech feature extraction sub-model is relatively large, which can ensure that the model can learn enough speech features, so that the candidate speech feature extraction sub-model has stable robustness to noise, and has a certain ability to resist noise interference.

[0055] Step B2, model compression based on the pre-self-supervised training candidate speech feature extraction sub-model, to obtain a reference speech feature extraction sub-model associated with the target device, so that the reference speech feature extraction sub-model requires less computing resources than the candidate speech feature extraction sub-model when performing speech feature extraction.

[0056] After determining the candidate speech feature extraction sub-model, model compression needs to be performed on the candidate speech feature extraction sub-model. The purpose of model compression is to reduce the computing resource requirements of the candidate speech feature extraction sub-model while maintaining the performance of the model. Through model compression, a reference speech feature extraction sub-model associated with the target device can be obtained. This reference speech feature extraction sub-model not only inherits the anti-noise interference performance of the candidate speech feature extraction sub-model when performing speech feature extraction, but also requires less computing resources than the candidate speech feature extraction sub-model when performing speech feature extraction.

[0057] Optionally, model compression based on the pre-self-supervised training candidate speech feature extraction sub-model can include using pruning, quantization, low-rank decomposition, knowledge distillation, etc. Model compression methods to perform model compression based on the pre-self-supervised training candidate speech feature extraction sub-model.

[0058] Among them, parameter quantization can be to quantize the model's weight parameters from high precision (such as 32-bit floating point numbers) to low precision (such as 8-bit integers) to reduce the storage space and calculation of parameters. Pruning can be to reduce the number of parameters in the model by deleting unimportant connections or neurons in the model. Knowledge distillation can be to transfer the knowledge of a complex large model (teacher model) to a smaller model (student model) so that the student model can achieve performance close to the teacher model at a smaller scale. Low-rank decomposition can be to decompose the weight matrix of the model into the product of low-rank matrices to reduce the number of parameters.

[0059] With the above scheme, the reference speech feature extraction sub-model can run in an extremely efficient manner on the target device due to careful compression processing, and still accurately extract accurate speech feature information while significantly reducing the occupation of computing resources. Compared with the original model without compression, the required memory space is greatly reduced, and the calculation amount is also significantly reduced, so that the speech feature extraction task can be smoothly and quickly executed on the target device with limited resources, such as mobile terminals or embedded devices with relatively weak processing capabilities, while reducing the consumption of computing resources, without affecting the accuracy and reliability of extracting speech feature information.

[0060] As an optional but non-limiting implementation manner, the reference speech feature extraction sub-model supports speech feature extraction under reference noise, the reference noise being ambient noise in which the target device or the target application associated with the target device is located when the target device or the target application associated with the target device is woken up by speech, and the reference noise having a noise level greater than a preset noise level.

[0061] The reference speech feature extraction sub-model has excellent performance in speech feature extraction under a reference noise environment, the reference noise being ambient noise in a preset distance range around the target device or the target application associated with the target device when the target device or the target application associated with the target device is woken up by speech, and the reference noise having a noise level greater than a preset noise level threshold. The reference speech feature extraction sub-model supports speech feature extraction under reference noise, which fully indicates that the reference speech feature extraction sub-model has strong ability to cope with harsh noise conditions, and can still effectively extract valuable speech features even in an environment with high noise intensity. Thus, a solid and reliable foundation is built for subsequent speech processing and analysis, and the smoothness and accuracy of the entire speech processing process are effectively ensured.

[0062] As an optional but non-limiting implementation manner, the candidate speech feature extraction sub-model based on pre-self-supervised training is compressed to obtain a reference speech feature extraction sub-model associated with the target device, including the following steps C1-C2.

[0063] Step C1, determining a second speech feature extraction sub-model corresponding to a first speech feature extraction sub-model, the first speech feature extraction sub-model being a candidate speech feature extraction sub-model pre-trained by self-supervision, and the second speech feature extraction sub-model being a speech feature extraction sub-model having a model size smaller than that of the first speech feature extraction sub-model.

[0064] Step C2, guiding the second speech feature extraction sub-model to learn in a direction of reducing a reference loss in a speech feature extraction training task by the first speech feature extraction sub-model, so as to transfer the anti-noise interference performance of the first speech feature extraction sub-model in speech feature extraction to the second speech feature extraction sub-model, the reference loss including a prediction loss of the second speech feature extraction sub-model in the speech feature extraction training task and an output difference loss between the first speech feature extraction sub-model and the second speech feature extraction sub-model.

[0065] The first speech feature extraction sub-model is a candidate speech feature extraction sub-model obtained in advance through self-supervised training, and a second speech feature extraction sub-model with a smaller model size than the first speech feature extraction sub-model is found. The purpose of this is to pass the anti-noise interference performance of the first speech feature extraction sub-model to the second speech feature extraction sub-model in subsequent training, while reducing the consumption of computing resources.

[0066] The reference loss includes a prediction loss of the second speech feature extraction sub-model in the speech feature extraction training task and an output difference loss between the first speech feature extraction sub-model and the second speech feature extraction sub-model. The first speech feature extraction sub-model is used to guide the training learning of the second speech feature extraction sub-model to learn the anti-noise performance of the first speech feature extraction sub-model in speech feature extraction, and the second speech feature extraction sub-model can learn how to extract speech features with anti-noise interference performance by minimizing the reference loss, thereby improving the speech recognition accuracy in a noisy environment.

[0067] The anti-interference ability of the first speech feature extraction sub-model with a larger size in speech feature extraction is transferred to the second speech feature extraction sub-model with a smaller size, so that the second speech feature extraction sub-model with a smaller size and lower computing resource demand can learn the knowledge and patterns contained in the first speech feature extraction sub-model, thereby approaching or even reaching the level of the first speech feature extraction sub-model in performance, while having lower computing cost and smaller model size, facilitating deployment and application on resource-constrained devices.

[0068] In the process of transferring the anti-interference ability of the first speech feature extraction sub-model with a larger size in speech feature extraction to the second speech feature extraction sub-model with a smaller size, the output (such as probability distribution, feature representation, etc.) of the first speech feature extraction sub-model is usually used as a soft target to guide the training of the second speech feature extraction sub-model, and real labels can also be combined as hard targets to achieve more effective transfer of anti-interference ability in speech feature extraction. By continuously adjusting the parameters of the second speech feature extraction sub-model, it can fit the output of the first speech feature extraction sub-model and the distribution of real data, thereby realizing knowledge transfer and model compression optimization.

[0069] If the first speech feature extraction sub-model has rich and accurate knowledge, and the distillation process can effectively transfer this knowledge to the second speech feature extraction sub-model, and the second speech feature extraction sub-model has sufficient learning ability to absorb this knowledge, the second speech feature extraction sub-model may approach or even exceed the first speech feature extraction sub-model in accuracy, especially in the case of relatively small amount of data or limited computing resources.

[0070] Exemplarily, the first and second speech feature extraction sub-models are selected: first, a first speech feature extraction sub-model with excellent speech feature extraction performance, large scale and complexity, and a first speech feature extraction sub-model with small scale and low demand for computing resources are determined. Then, the first speech feature extraction sub-model is fully trained using a large amount of training data to achieve good performance. The reference loss in knowledge distillation is defined, including the prediction loss of the second speech feature extraction sub-model when performing speech feature extraction on the speech signal, and the difference loss between the output of the second speech feature extraction sub-model and the output of the first speech feature extraction sub-model (for example, the difference loss can be based on the difference of probability distribution). The second speech feature extraction sub-model is guided by the first speech feature extraction sub-model in the speech feature extraction training task, and the parameters of the second speech feature extraction sub-model are adjusted so that the prediction loss and the difference loss between the outputs of the two models are minimized at the same time. According to the training effect, the second speech feature extraction sub-model is fine-tuned, for example, the learning rate, the number of training rounds, the hyperparameters, etc., to further improve the performance of the second speech feature extraction sub-model.

[0071] By using the above scheme, the second speech feature extraction sub-model is continuously guided by the first speech feature extraction sub-model, and continuously tries and adjusts to find the most suitable parameter settings, so that the second speech feature extraction sub-model can effectively inherit the anti-interference ability in speech feature extraction from the first speech feature extraction sub-model, and achieve a good balance between performance and resource utilization.

[0072] Optionally, the model size of the speech feature extraction sub-model can be described in detail using at least one of the following dimension indicators: training data size, model parameter number, model calculation amount, model memory occupation amount, and model complexity, etc.

[0073] The training data size can include the number of samples, the feature dimension, and the total amount of data (in bytes) of the training data. For example, if the training data contains tens of billions or even hundreds of billions of samples, and each sample has thousands of features, the data size is considered large. The number of model parameters can include the number of learnable parameters in a statistical model. When the number of parameters reaches tens of millions, hundreds of millions or even billions, it is usually considered a large-scale model. The model calculation amount can be the number of floating point operations required by the calculation model during training or inference. If the calculation amount is measured in the order of tens of billions, hundreds of billions or even higher, it indicates that the model has a large scale. The model memory occupation can consider the memory space required by the model during training or inference. When the model needs a large amount of memory to store parameters, intermediate results and data, it indicates that it has a large scale. The model complexity can be the number of layers, the number of neurons, and the complexity of the network structure, etc. A complex network structure is usually related to a large-scale model.

[0074] S130, wake up control is performed on the target device or the target application based on the to-be-processed voice feature information.

[0075] The to-be-processed voice feature information not only includes acoustic features such as prosody and timbre, but also includes acoustic features for describing a preset wake-up keyword. Whether the preset wake-up keyword exists in the to-be-processed voice signal is determined by recognizing the to-be-processed voice feature information, and then whether the preset wake-up keyword is recognized from the to-be-processed voice signal is used to realize the wake-up control on the target device or the target application.

[0076] The reference voice feature extraction sub-model inherits the anti-noise interference performance of the candidate voice feature extraction sub-model pre-trained in a self-supervised manner. The to-be-processed voice feature information extracted from the to-be-processed voice signal is more accurate and reliable. Even in a noisy environment, key voice features can be effectively captured, which enables the target device or the target application to accurately recognize the wake-up instruction when facing various complex noise environments, greatly improving the success rate and stability of device wake-up. The accurately extracted to-be-processed voice feature information can reduce the occurrence of false wake-up, avoid unnecessary energy consumption and system burden caused by mistakenly recognizing environmental noise or irrelevant voice as a wake-up instruction, and improve the energy efficiency and operation efficiency of the device.

[0077] As an optional but non-limiting implementation manner, the wake-up control is performed on the target device or the target application based on the to-be-processed voice feature information, including the following steps D1-D2:

[0078] Step D1, input the to-be-processed voice feature information into a reference voice feature detection sub-model associated with the target device, and connect the output of the reference voice feature extraction sub-model to the input of the reference voice feature detection sub-model in series. The reference voice feature detection sub-model can support recognizing the voice feature of the output of the reference voice feature extraction sub-model and detecting the possibility of the existence of the preset wake-up keyword in the voice signal based on the voice feature.

[0079] Step D2, according to the possibility of the existence of the preset wake-up keyword in the to-be-processed voice signal output by the reference voice feature detection sub-model, the wake-up control is performed on the target device or the target application.

[0080] After the reference speech feature extraction sub-model, a speech feature detection sub-model is connected in series, and the classification loss can be used to train on the woken-up speech data set. The specific steps are as follows: collect and organize the training data set containing the preset wake-up keyword. Then, a neural network architecture suitable for speech processing can be selected as the speech feature extraction sub-model, such as a recurrent neural network (RNN), a long short-term memory network (LSTM), or a convolutional neural network (CNN), etc., and a classification layer is added behind the Hubert model. After the reference speech feature extraction sub-model, a speech feature detection sub-model is connected in series, such as using a model suitable for classification tasks, using the prepared data set to train the model, continuously adjusting the parameters of the model through the back propagation algorithm to minimize the loss function, and obtaining the reference speech feature detection sub-model.

[0081] The reference speech feature detection sub-model can be used to detect whether the preset wake-up keyword exists in the speech signal. Its core function is to analyze and process the input speech features to determine whether they contain information related to the wake-up keyword. The reference speech feature detection sub-model usually uses a series of algorithms and techniques to perform feature extraction, pattern matching, and classification on the input speech features. By inputting the to-be-processed speech feature information into the reference speech feature detection sub-model, the reference speech feature detection sub-model can analyze and process the to-be-processed speech feature information and output a result representing the possibility of the presence of the preset wake-up keyword in the to-be-processed speech signal. This result can be used to implement wake-up control of the target device or target application, for example, when the possibility exceeds a certain threshold, the wake-up operation of the target device or target application is triggered.

[0082] As an optional but non-limiting implementation, according to the possibility of the presence of the preset wake-up keyword in the to-be-processed speech signal output by the reference speech feature detection sub-model, the wake-up control of the target device or target application includes the following steps E1-E2:

[0083] Step E1, if the possibility of the presence of the preset wake-up keyword in the to-be-processed speech signal is greater than a preset probability threshold, the target device or target application is switched from the sleep state to the working state.

[0084] Step E2, if the possibility of the presence of the preset wake-up keyword in the to-be-processed speech signal is not greater than the preset probability threshold, the target device or target application is not switched from the sleep state to the working state.

[0085] By setting a preset probability threshold to determine whether to start the state switching of the device, the truly effective wake-up instruction can be accurately identified, the false wake-up can be avoided, and the accuracy of the wake-up is improved. When the possibility is greater than the threshold, the switching is started, so that only the second voice signal containing the preset wake-up keyword with high possibility can wake up the device, and unnecessary state switching caused by misjudgment is reduced. When the possibility is not greater than the threshold, the switching is not started, and the target device or the target application continues to wait for an effective wake-up instruction input, so that the device is prevented from being incorrectly woken up from the sleep state due to uncertain or incorrect judgment, and device resources and energy are saved. According to the possibility of the preset wake-up keyword in the voice signal, whether to switch the state of the device is determined, so that the resources and energy consumption of the device can be reasonably allocated. Through the above scheme, the effective wake-up instruction of the user can be responded in time, and the disturbance caused by false wake-up can be avoided, so that the user feels convenient and comfortable when using the voice wake-up function.

[0086] The technical scheme of the embodiment of the present disclosure determines the target device corresponding to the to-be-processed voice signal carrying the preset wake-up keyword, provides a data basis for subsequent wake-up operation of the target device or the target application, and extracts feature information from the to-be-processed voice signal by using a reference voice feature extraction sub-model adjusted by using a candidate voice feature extraction sub-model based on pre-self-supervised training. Since the reference voice feature extraction sub-model inherits the anti-noise interference performance of the candidate voice feature extraction sub-model in voice feature extraction, the key features can be effectively extracted from complex voice signals, and the interference of noise is reduced. Therefore, the technical scheme significantly improves the ability to detect the wake-up keyword in a high-noise environment by using the feature extraction sub-model with excellent anti-noise performance, thereby greatly improving the wake-up rate, so that the target device or the target application can be more accurately and reliably woken up in various noisy environments, greatly improving the user experience in a complex noise scene, and enhancing the applicability and practicality of the device.

[0087] FIG. 2 is a structural schematic diagram of a voice wake-up device provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the case of waking up a device from a sleep state to a working state by a voice mode. The voice wake-up device can be implemented in the form of software and / or hardware, and is generally integrated on any electronic device with network communication function, which can be a mobile terminal, a PC terminal, or a server, etc.

[0088] As shown in FIG. 2, the voice wake-up device of the embodiment of the present disclosure can include the following:

[0089] The first determination module 210 is configured to determine a to-be-processed voice signal obtained by a target device, wherein the to-be-processed voice signal supports carrying a preset wake-up keyword used for waking up the target device or a target application associated with the target device.

[0090] The second determination module 220 is configured to determine a reference speech feature extraction sub-model associated with the target device, and extract the to-be-processed speech feature information from the to-be-processed speech signal through the reference speech feature extraction sub-model, wherein the reference speech feature extraction sub-model is obtained by adjusting a pre-self-supervised trained candidate speech feature extraction sub-model, and the reference speech feature extraction sub-model has inherited the anti-noise interference performance of the pre-self-supervised trained candidate speech feature extraction sub-model when performing speech feature extraction.

[0091] The control module 230 is configured to perform wake-up control on the target device or the target application based on the to-be-processed speech feature information.

[0092] On the basis of the above-mentioned embodiments, optionally, the to-be-processed speech signal obtained by the target device comprises:

[0093] Real-time acquisition of a voice signal within a preset distance range of the target device by using a sound pickup device configured on the target device;

[0094] The to-be-processed speech signal is obtained by pre-processing the real-time acquired voice signal, and the pre-processing comprises noise removal and filtering processing.

[0095] On the basis of the above-mentioned embodiments, optionally, the determination process of the reference speech feature extraction sub-model comprises:

[0096] A candidate speech feature extraction sub-model corresponding to the target device is determined, wherein the candidate speech feature extraction sub-model is obtained by training and adjusting a speech feature extraction sub-model based on a mask prediction self-supervised learning manner on a candidate speech data set, the data amount of the candidate speech data set is greater than a preset data amount, the candidate speech feature extraction sub-model has stable robustness to noise, and model compression is performed based on the pre-self-supervised trained candidate speech feature extraction sub-model to obtain the reference speech feature extraction sub-model associated with the target device, so that the reference speech feature extraction sub-model requires less computing resources than the candidate speech feature extraction sub-model when performing speech feature extraction.

[0097] On the basis of the above-mentioned embodiments, optionally, the model compression based on the pre-self-supervised trained candidate speech feature extraction sub-model to obtain the reference speech feature extraction sub-model associated with the target device comprises:

[0098] determining a second speech feature extraction sub-model corresponding to the first speech feature extraction sub-model, the first speech feature extraction sub-model being a pre-self-supervised trained candidate speech feature extraction sub-model, and the second speech feature extraction sub-model being a speech feature extraction sub-model for performing speech feature extraction and having a model scale smaller than that of the first speech feature extraction sub-model;

[0099] guiding the second speech feature extraction sub-model to learn in a direction of reducing a reference loss in a speech feature extraction training task by the first speech feature extraction sub-model, so as to transfer the anti-noise interference performance of the first speech feature extraction sub-model in speech feature extraction to the second speech feature extraction sub-model, the reference loss including a prediction loss of the second speech feature extraction sub-model in the speech feature extraction training task and an output difference loss between the first speech feature extraction sub-model and the second speech feature extraction sub-model.

[0100] On the basis of the above-mentioned embodiments, optionally, the reference speech feature extraction sub-model supports speech feature extraction under reference noise, the reference noise being environmental noise when the target device or the target application is woken up by speech, and the noise size of the reference noise being greater than a preset noise size.

[0101] On the basis of the above-mentioned embodiments, optionally, the wake-up control of the target device or the target application based on the to-be-processed speech feature information includes:

[0102] inputting the to-be-processed speech feature information into a reference speech feature detection sub-model associated with the target device, the output of the reference speech feature extraction sub-model being connected in series to the input of the reference speech feature detection sub-model, and the reference speech feature detection sub-model being capable of identifying the speech feature of the output of the reference speech feature extraction sub-model and detecting the possibility of the presence of a preset wake-up keyword in a speech signal based on the speech feature;

[0103] controlling the wake-up of the target device or the target application according to the possibility of the presence of the preset wake-up keyword in the to-be-processed speech signal output by the reference speech feature detection sub-model.

[0104] On the basis of the above-mentioned embodiments, optionally, the wake-up control of the target device or the target application according to the possibility of the presence of the preset wake-up keyword in the to-be-processed speech signal output by the reference speech feature detection sub-model includes:

[0105] if the possibility of the presence of the preset wake-up keyword in the to-be-processed speech signal is greater than a preset probability threshold, starting to switch the target device or the target application from a sleep state to a working state;

[0106] If the possibility that the to-be-processed voice signal contains the preset wake-up keyword is less than the preset probability threshold, the target device or the target application is not switched from the sleep state to the working state.

[0107] The technical solution of the embodiments of the present disclosure determines the to-be-processed voice signal carrying the preset wake-up keyword corresponding to the target device, thereby providing a data basis for subsequent wake-up operation of the target device or the target application; and the reference voice feature extraction sub-model obtained by adjusting the candidate voice feature extraction sub-model based on the pre-self-supervised training is used to extract feature information from the to-be-processed voice signal. Since the reference voice feature extraction sub-model inherits the anti-noise interference performance of the candidate voice feature extraction sub-model in voice feature extraction, it can effectively extract key features from complex voice signals and reduce the interference of noise. Therefore, the technical solution significantly improves the ability to detect the wake-up keyword in a high-noise environment by using the feature extraction sub-model with excellent anti-noise performance, thereby greatly improving the wake-up rate. As a result, the target device or the target application can be more accurately and reliably woken up in various noisy environments, greatly improving the user's experience in complex noise scenarios and enhancing the applicability and practicality of the device.

[0108] The voice wake-up apparatus provided by the embodiments of the present disclosure can perform the voice wake-up method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of the execution method.

[0109] It should be noted that each unit and module included in the above apparatus is only divided according to the function logic, but is not limited to the above division, as long as the corresponding function can be realized; in addition, the specific names of each functional unit are only for convenient mutual distinction, and do not serve to limit the protection scope of the embodiments of the present disclosure.

[0110] FIG. 3 is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. Referring to FIG. 3, a structural schematic diagram of an electronic device (for example, a terminal device or a server in FIG. 3) 300 suitable for implementing the embodiments of the present disclosure is shown. The terminal device in the embodiments of the present disclosure can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablets), PMPs (portable multimedia players), vehicle-mounted terminals (for example, vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. The electronic device shown in FIG. 3 is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.

[0111] As shown in FIG. 3, the electronic device 300 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 302 or loaded into a random access memory (RAM) 303 from a storage device 308. Various programs and data required for the operation of the electronic device 300 are also stored in the RAM 303. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0112] Generally, the following devices can be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 can allow the electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although FIG. 3 shows the electronic device 300 with various devices, it should be understood that all of the illustrated devices are not required to be implemented or possessed. More or fewer devices can be alternatively implemented or possessed.

[0113] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product including a computer program carried on a non-transitory computer readable medium, the computer program containing program code for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.

[0114] The names of messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0115] The electronic device provided by the embodiments of the present disclosure and the voice wake-up method provided by the above-described embodiments belong to the same inventive concept, and technical details not described in detail in the present embodiments can be referred to the above-described embodiments, and the present embodiments have the same beneficial effects as the above-described embodiments.

[0116] The embodiments of the present disclosure provide a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the voice wake-up method provided by the above-described embodiments.

[0117] Note that the computer readable medium described above in the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination thereof. The computer readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination thereof. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer readable program code is contained. Such a propagated data signal can take any of a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to wire, cable, RF, etc., or any suitable combination thereof.

[0118] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.

[0119] The computer readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and can be accessed via the electronic device.

[0120] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: determine a to-be-processed voice signal obtained by a target device, the to-be-processed voice signal supporting carrying of a preset wake-up keyword for waking up the target device or a target application associated with the target device; determine a reference voice feature extraction sub-model associated with the target device, and extract to-be-processed voice feature information from the to-be-processed voice signal through the reference voice feature extraction sub-model, the reference voice feature extraction sub-model being obtained by adjusting a pre-self-supervised trained candidate voice feature extraction sub-model, and the reference voice feature extraction sub-model having inherited an anti-noise interference performance of the pre-self-supervised trained candidate voice feature extraction sub-model when performing voice feature extraction; and perform wake-up control on the target device or the target application based on the to-be-processed voice feature information.

[0121] Computer program code for carrying out operations of the present disclosure can be written in any of one or more programming languages or combinations of languages including object or visual programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0122] The flow and block diagrams in the drawings show architectural, functional, and operational representations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may

[0123] The units described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the names of the units do not constitute a limitation on the units themselves. For example, the first obtaining unit can also be described as a unit that obtains at least two Internet protocol addresses.

[0124] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), etc.

[0125] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0126] The above description is merely that of the preferred embodiments of the present disclosure and the description of the technical principles of the application. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by the combinations of the above technical features or equivalent features without departing from the above disclosed concept. For example, the above technical features can be replaced with the technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.

[0127] Moreover, while operations are depicted in a particular order, this should not be understood as requiring such an order nor infringing on the scope of the disclosure. Certain of the operations described in the discussion are combinable into a single operation, and certain operations can be separated into several operations. In some embodiments, the operations described in the discussion can be performed in an order different than presented in the discussion. In some embodiments, the operations described in the discussion can be performed concurrently. Also, while several specific implementation details are discussed in the discussion, these should not be interpreted as limiting the scope of the disclosure. Rather, certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.

[0128] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. A voice wake-up method, comprising: determining a target device to obtain a to-be-processed voice signal, wherein the to-be-processed voice signal supports carrying a preset wake-up keyword for waking up the target device or a target application associated with the target device; determining a reference voice feature extraction sub-model associated with the target device, and extracting to-be-processed voice feature information from the to-be-processed voice signal by using the reference voice feature extraction sub-model, wherein the reference voice feature extraction sub-model is obtained by adjusting a pre-self-supervised training candidate voice feature extraction sub-model, and the reference voice feature extraction sub-model has inherited the anti-noise interference performance of the pre-self-supervised training candidate voice feature extraction sub-model when performing voice feature extraction; controlling the target device or the target application to wake up based on the to-be-processed voice feature information.

2. The method of claim 1, wherein, The determination of the to-be-processed voice signal obtained by the target device comprises: using a sound collector configured on the target device to obtain a voice signal in a preset distance range of the target device in real time; obtaining the to-be-processed voice signal by preprocessing the real-time obtained voice signal, wherein the preprocessing comprises at least one of noise removal and filtering processing.

3. The method of claim 1 or 2, wherein, The determination process of the reference voice feature extraction sub-model comprises: determining a candidate voice feature extraction sub-model corresponding to the target device, wherein the candidate voice feature extraction sub-model is obtained by training and adjusting the voice feature extraction sub-model based on a mask prediction self-supervised learning manner on a candidate voice data set, the data amount of the candidate voice data set is greater than a preset data amount, and the candidate voice feature extraction sub-model has stable robustness to noise; performing model compression based on the pre-self-supervised training candidate voice feature extraction sub-model to obtain the reference voice feature extraction sub-model associated with the target device, so that the reference voice feature extraction sub-model requires less computing resources than the candidate voice feature extraction sub-model when performing voice feature extraction.

4. The method of claim 3, wherein, The model compression based on the pre-self-supervised training candidate voice feature extraction sub-model to obtain the reference voice feature extraction sub-model associated with the target device comprises: determining a second voice feature extraction sub-model corresponding to a first voice feature extraction sub-model, wherein the first voice feature extraction sub-model is a pre-self-supervised training candidate voice feature extraction sub-model, and the second voice feature extraction sub-model is a voice feature extraction sub-model used for voice feature extraction and having a model size smaller than that of the first voice feature extraction sub-model. The first speech feature extraction sub-model guides the second speech feature extraction sub-model to learn in a direction of reducing a reference loss in a speech feature extraction training task, so as to transfer the anti-noise interference performance of the first speech feature extraction sub-model in speech feature extraction to the second speech feature extraction sub-model, and the reference loss includes a prediction loss of the second speech feature extraction sub-model in the speech feature extraction training task and an output difference loss between the first speech feature extraction sub-model and the second speech feature extraction sub-model.

5. The method of claim 3 or 4, wherein, The reference speech feature extraction sub-model supports speech feature extraction under reference noise, the reference noise is environmental noise when the target device or the target application is woken up by speech, and the noise size of the reference noise is greater than a preset noise size.

6. The method of any one of claims 1 to 5, wherein, The wake-up control on the target device or the target application based on the to-be-processed speech feature information includes: inputting the to-be-processed speech feature information into a reference speech feature detection sub-model associated with the target device, the output of the reference speech feature extraction sub-model being connected in series with the input of the reference speech feature detection sub-model, the reference speech feature detection sub-model being capable of identifying the speech feature of the output of the reference speech feature extraction sub-model and detecting the possibility of the presence of a preset wake-up keyword in a speech signal based on the speech feature; controlling the wake-up of the target device or the target application according to the possibility of the presence of the preset wake-up keyword in the to-be-processed speech signal output by the reference speech feature detection sub-model.

7. The method of claim 6, wherein, The controlling the wake-up of the target device or the target application according to the possibility of the presence of the preset wake-up keyword in the to-be-processed speech signal output by the reference speech feature detection sub-model includes: if the possibility of the presence of the preset wake-up keyword in the to-be-processed speech signal is greater than a preset probability threshold, switching the target device or the target application from a sleep state to an active state; if the possibility of the presence of the preset wake-up keyword in the to-be-processed speech signal is not greater than the preset probability threshold, not switching the target device or the target application from the sleep state to the active state.

8. A speech wake-up apparatus, comprising: a first determination module configured to determine a to-be-processed speech signal obtained by a target device, the to-be-processed speech signal supporting carrying a preset wake-up keyword for waking up the target device or a target application associated with the target device; a second determination module configured to determine a reference speech feature extraction sub-model associated with the target device, and extract to-be-processed speech feature information from the to-be-processed speech signal through the reference speech feature extraction sub-model, the reference speech feature extraction sub-model being obtained by adjusting a candidate speech feature extraction sub-model based on pre-self-supervised training, and the reference speech feature extraction sub-model having inherited the anti-noise interference performance of the pre-self-supervised training candidate speech feature extraction sub-model in speech feature extraction; The control module is configured to perform wake-up control on the target device or the target application based on the to-be-processed voice feature information. 9.An electronic device, comprising: one or more processors; a storage device configured to store one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the voice wake-up method according to any one of claims 1-7.

10. A storage medium containing computer-executable instructions, wherein, The computer executable instructions, when executed by a computer processor, are used to perform the voice wake-up method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Voice awakening method and device, equipment and computer readable storage medium

    CN110634468A

  • Method for voice wake-up, electronic equipment, storage medium, and program

    CN114038457A

  • Voice wake-up method and system based on voice similarity matching

    CN115223551A

  • Speech recognition model training method and device, storage medium and electronic equipment

    CN116343781A

  • Unified speech representation learning

    US20220366898A1